Prefetching of Weights Used in Neural Network Processor

The systolic array with a weight fetcher unit in a special-purpose hardware circuit addresses the inefficiencies of existing neural network computations by prefetching weights, improving computational efficiency and reducing execution time in neural network processing.

JP7710018B2Active Publication Date: 2025-07-17GOOGLE LLC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023190654
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-09-03
Filing Date
2023-11-08
Publication Date
2025-07-17
Estimated Expiration
2036-04-29

AI Technical Summary

Technical Problem

Existing neural network computations, particularly in convolutional layers, are computationally intensive and time-consuming due to the need for numerous matrix multiplications, with limited parallelization capabilities in existing processor architectures.

Method used

A special-purpose hardware circuit with a systolic array and weight fetcher unit that prefetched weight inputs to cells, allowing efficient execution of neural network calculations by eliminating direct memory access and synchronizing weight shifts for multiple convolution operations.

Benefits of technology

This approach enables more efficient neural network processing by reducing the need for external memory connections and synchronizing weight inputs, thereby enhancing computational efficiency and reducing execution time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710018000001
    Figure 0007710018000001
  • Figure 0007710018000002
    Figure 0007710018000002
  • Figure 0007710018000003
    Figure 0007710018000003
Patent Text Reader

Abstract

To provide a circuit and a method for computing a neural network inference in hardware.SOLUTION: An architecture 300 including a matrix computation unit, includes a two-dimensional systolic array 306 including a plurality of cells, and a weight fetcher interface. For each of a plurality of neural network layers, the weight fetcher interface sends a plurality of weight inputs to cells along a first dimension of the two-dimensional systolic array for the neural network layer. For each of the plurality of neural network layers, a plurality of weight sequencers, each of which is coupled to a distinct cell along the first dimension of the two-dimensional systolic array, shift, for the neural network layer, the plurality of weight inputs to cells along the second dimension of the two-dimensional systolic array over a plurality of clock cycles. The cells in the two-dimensional systolic array each compute a product of an activation input and a respective weight input using a multiplication circuit.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Background This specification relates to calculating neural network inference values in hardware.

Background Art

[0002] A neural network is a machine learning model that uses one or more layers of the model to generate an output, such as a classification, for the received input. Some neural networks include one or more hidden layers in addition to the output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer of the network. Each layer of the network generates an output from the received input according to the current values of each set of parameters.

[0003] Some neural networks include one or more convolutional neural network layers. Each convolutional neural network layer has an associated set of kernels. Each kernel contains values constructed by a neural network model created by a user. In some implementations, the kernels identify specific image contours, shapes, or colors. A kernel can be represented as a matrix structure of weighted inputs. Each convolutional layer can also process a set of activation inputs. The set of activation inputs can also be represented as a matrix structure.

[0004] Some existing systems perform calculations in software for a given convolutional layer. For example, the software can apply each kernel of the layer to a set of activation inputs. That is, for each kernel, the software can overlay a kernel, which can be represented multidimensionally, over a first portion of the activation inputs, which can be represented multidimensionally. The software can then calculate a dot product from the overlapping elements. The dot product can correspond to a single activation input, for example, an activation input element having an upper left position within the overlapping multidimensional space. Then, for example, using a sliding window, the software can shift the kernel to overlay a second portion of the activation inputs and calculate another dot product corresponding to another activation input. The software can repeatedly execute this process until each activation input has a corresponding dot product. In some implementations, the dot product is an input to an activation function that generates activation values, which can be combined, for example, pooled, before being sent to a subsequent layer of the neural network.

[0005] One way to calculate the convolution operation requires a large number of matrix multiplications in a large dimensional space. The processor can calculate the matrix multiplications by brute force. For example, although computationally intensive and time intensive, the processor can repeatedly calculate individual sums and products for the convolution operation. The degree to which the processor can parallelize the calculations is limited by its architecture. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION

[0006] Overview Overall, this document describes a special-purpose hardware circuit for calculating neural network inference values. MEANS FOR SOLVING THE PROBLEMS

[0007] As a whole, one innovative aspect of the subject matter described herein can be implemented in a circuit for performing neural network computations on a neural network having multiple layers, the circuit comprising a systolic array having a plurality of cells and a weight fetcher unit, the weight fetcher unit configured to send a plurality of weight inputs to cells along a first dimension of the systolic array for the neural network layer for each of the plurality of neural network layers, the circuit further comprising a plurality of weight sequencer units, each weight sequencer unit coupled to an individual cell along the first dimension of the systolic array, the plurality of weight sequencer units configured to shift the plurality of weight inputs to cells along a second dimension of the systolic array for the neural network layer over a plurality of clock cycles for each of the plurality of neural network layers, each weight input being stored within a respective cell along the second dimension, and each cell being configured to calculate a product of an activation input and a respective weight input using a multiplication circuit.

[0008] The implementation example may include one or more of the following. The value sequencer unit is configured to send a plurality of activation inputs to cells along the second dimension of the systolic array for the neural network layer for each of the plurality of neural network layers. The first dimension of the systolic array corresponds to the rows of the systolic array, and the second dimension of the systolic array corresponds to the columns of the systolic array. Each cell is configured to pass a weight control signal to an adjacent cell, and the weight control signal causes a circuit in the adjacent cell to shift or load a weight input for the adjacent cell. The weight path register is configured to store the weight input shifted to the cell, the weight register is coupled to the weight path register, the weight control register is configured to determine whether to store the weight input in the weight register, the activation register is configured to store an activation input, and is configured to send the activation input to another activation register in a first adjacent cell along the first dimension. The multiplication circuit is coupled to the weight register and the activation register, and the multiplication circuit is configured to output a product of the weight input and the activation input. The summation circuit is coupled to the multiplication circuit and is configured to receive the product and a first partial sum from a second adjacent cell along the second dimension. The summation circuit is configured to output a second partial sum of the product and the first partial sum. The partial sum register is coupled to the summation circuit and is configured to store the second partial sum, and the partial sum register is configured to send the second partial sum to another summation circuit in a third adjacent cell along the second dimension. Each weight sequencer unit includes a pause counter corresponding to the weight control register in the corresponding cell coupled to the weight sequencer unit, and a decrement circuit. The decrement circuit is configured to decrement an input to the weight sequencer unit to generate a decremented output, and send the decremented output to the pause counter.The values in each pause counter are the same, and each weight sequencer unit is configured to load the corresponding weight input into the corresponding individual cell of the systolic array, and the loading comprises sending the weight input to the multiplication circuit. The values in each pause counter reach a predetermined value to pause the shift of the plurality of weight inputs along the second dimension in the plurality of weight sequencer units. The systolic array is configured to generate a cumulative output for the neural network layer from each product for each of the plurality of neural network layers.

[0009] Certain embodiments of the subject matter described herein can be realized to achieve one or more of the following advantages. By prefetching weights, it becomes possible for the neural network processor to execute calculations more efficiently. The processor is associated with loading weight inputs into the systolic array using a weight fetcher unit and a weight sequencer unit, thereby eliminating the wires that couple an external memory unit to each cell within the systolic array. The processor can pause, i.e., "freeze", the shift of weight inputs to synchronize the execution of multiple convolution calculations.

[0010] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

[0012] Like reference numerals and names in the various figures indicate like elements. Detailed Description A neural network having multiple layers can be used to calculate an estimate. For example, given an input, the neural network can calculate an estimate for the input. The neural network calculates this estimate by processing the input through each layer of the neural network. In particular, the layers of the neural network are arranged in a certain order with each having a respective set of weights. Each layer receives an input, processes the input according to the set of weights of that layer, and generates an output.

[0013] Thus, to calculate an estimate from the received input, the neural network receives the input, processes it through each of the neural network layers in that order to generate an estimate, and the output from one neural network layer is provided as the input to the next neural network layer. The data input to a neural network layer, such as the input to the neural network or the output of a layer below that layer in the order, to the neural network layer, can be referred to as the activation input to that layer.

[0014] In some implementations, the layers of the neural network are arranged in a directed graph. That is, any given layer can receive multiple inputs, multiple outputs, or both. Also, the layers of the neural network can be arranged so that the output of a layer can be fed back as an input to a previous layer.

[0015] FIG. 1 is a flowchart of an exemplary process 100 for performing calculations for a given layer of a neural network using a special-purpose hardware circuit. For convenience, method 100 will be described in the context of a system having one or more circuits that execute method 100. Method 100 can be performed for each layer of the neural network to calculate an estimate from received inputs.

[0016] The system receives a plurality of sets of weight inputs for a given layer (step 102) and receives a plurality of sets of activation inputs for the given layer (step 104). The plurality of sets of weight inputs and the plurality of sets of activation inputs can each be received from the dynamic memory and integrated buffer of the special-purpose hardware circuit. In some implementations, both the plurality of sets of weight inputs and the plurality of sets of activation inputs may be received from the integrated buffer.

[0017] The system uses a matrix multiplication unit of the special-purpose hardware circuit to generate a cumulative value from the weight inputs and the activation inputs (step 106). In some implementations, the cumulative value is the dot product of the plurality of sets of weight inputs and the plurality of sets of activation inputs. That is, for one set of weights, the system can multiply each weight input by each activation input and sum the products to form a cumulative value. The system can then calculate the dot product of other sets of weights and other pluralities of activation inputs.

[0018] The system can generate a layer output from the cumulative value using a vector calculation unit of a specific purpose hardware circuit (step 108). In some implementations, the vector calculation unit applies an activation function to the cumulative value. The output of the layer may be stored in an integration buffer so as to be used as an input to a subsequent layer in the neural network, or may be used to determine a speculation value. When the system processes the received input through each layer of the neural network to generate a speculation value for the received input, the system ends processing the neural network.

[0019] FIG. 2 shows an exemplary application specific integrated circuit 200 for performing neural network calculations. System 200 includes a host interface 202. The host interface 202 can receive instructions including parameters for neural network calculations. The parameters can include at least one or more of how many layers are to be processed, the corresponding multiple sets of weighted inputs for each layer, the first set of activation inputs, i.e., the input to the neural network for calculating a speculation value, the corresponding input and output sizes for each layer, the stride value for neural network calculations, and the type of layer to be processed, such as a convolutional layer or a fully connected layer.

[0020] The host interface 202 can send instructions to the sequencer 206, and the sequencer 206 converts the instructions into low-level control signals for controlling the circuit to execute neural network calculations. In some embodiments, the control signals adjust the data flow within the circuit, such as how multiple sets of weight inputs and multiple sets of activation inputs flow through the circuit. The sequencer 206 can send the control signals to the integrated buffer 208, the matrix calculation unit 212, and the vector calculation unit 214. In some embodiments, the sequencer 206 also sends control signals to the direct memory access engine 204 and the dynamic memory 210. In some embodiments, the sequencer 206 is a processor that generates a clock signal. The sequencer 206 can use the timing of the clock signal to send the control signals to each component of the circuit 200 at the appropriate time. In some other embodiments, the host interface 202 passes a clock signal from an external processor.

[0021] The host interface 202 can send multiple sets of weight inputs and the first set of activation inputs to the direct memory access engine 204. The direct memory access engine 204 can store multiple sets of activation inputs in the integrated buffer 208. In some embodiments, the direct memory access stores multiple sets of weights in the dynamic memory 210, which can be a memory unit. In some embodiments the dynamic memory is located away from the circuit.

[0022] The integrated buffer 208 is a memory buffer. The integrated buffer 208 can be used to store the set of activation inputs from the direct memory access engine 204 and the output of the vector calculation unit 214. The direct memory access engine 204 can also read the output of the vector calculation unit 214 from the integrated buffer 208.

[0023] The dynamic memory 210 and the integrated buffer 208 can each send a plurality of sets of weight inputs and a plurality of sets of activation inputs to the matrix calculation unit 212. In some implementations, the matrix calculation unit 212 is a two-dimensional systolic array. The matrix calculation unit 212 may be a one-dimensional systolic array or other circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the matrix calculation unit 212 is a general matrix processor.

[0024] The matrix calculation unit 212 can process the weight inputs and the activation inputs to provide an output vector to the vector calculation unit 214. In some implementations, the matrix calculation unit sends the output vector to the integrated buffer 208, and the integrated buffer 208 sends the output vector to the vector calculation unit 214. The vector calculation unit can process the output vector and store the processed output vector in the integrated buffer 208. For example, the vector calculation unit 214 can apply a non-linear function to the output of the matrix calculation unit, such as a vector of cumulative values, to generate activation values. In some implementations, the vector calculation unit 214 generates normalization values, pooling values, or both. The processed output vector can be used as an activation input to the matrix calculation unit 212, for example, for use in subsequent layers within a neural network. The matrix calculation unit 212 will be described in further detail below with reference to FIGS. 3 and 4.

[0025] FIG. 3 shows an exemplary architecture 300 including a matrix calculation unit. The matrix calculation unit is a two-dimensional systolic array 306. The array 306 includes a plurality of cells 304. In some implementations, the first dimension 320 of the systolic array 306 corresponds to columns of cells, and the second dimension 322 of the systolic array 306 corresponds to rows of cells. The systolic array may have more rows than columns, more columns than rows, or an equal number of columns and rows.

[0026] In the example shown, value loader 302 sends start-up inputs to the rows of array 306, and weight fetcher interface 308 sends weight inputs to the columns of array 306. However, in some other implementations, the start-up inputs are sent to the columns of array 306 and the weight inputs are sent to the rows of array 306.

[0027] Value loader 302 can receive start-up inputs from an integrated buffer, such as integrated buffer 208 of FIG. 2. Each value loader can send the corresponding start-up input to an individual leftmost cell of array 306. The leftmost cell can be a cell along the leftmost column of array 306. For example, value loader 312 can send a start-up input to cell 314. The value loader can also send start-up inputs to adjacent value loaders, and the start-up inputs can be used in another leftmost cell of array 306. Thereby, the start-up inputs can be shifted so that they can be used in another specific cell of array 306.

[0028] Weight fetcher interface 308 can receive weight inputs from a memory unit, such as dynamic memory 210 of FIG. 2. Weight fetcher interface 308 can send the corresponding weight input to an individual topmost cell of array 306 . The topmost cell can be a cell along the topmost row of array 306. For example, weight fetcher interface 308 can send weight inputs to cells 314 and 316.

[0029] In some implementations, a host interface, such as host interface 202 of FIG. 2, shifts start-up inputs along one dimension across the entire array 306, for example to the right, and shifts weight inputs along another dimension across the entire array 306, for example downwards. For example, in one clock cycle, the start-up input at cell 314 can be shifted to the start-up register at cell 316 to the right of cell 314. Similarly, the weight input at cell 316 can be shifted to the weight register at cell 318 below cell 314.

[0030] In each clock cycle, each cell can process a given weight input and a given activation input to generate an accumulated output. The accumulated output can also be passed to adjacent cells along the same dimension as the given weight input. Individual cells will be further described below with reference to FIG. 4.

[0031] In some implementations, the weights and activations are shifted across two or more cells during a given clock cycle to transition from one convolution operation to another.

[0032] The accumulated output can be passed along the same column as the weight input, for example, to the lower part of the column within the array 306. In some implementations, the array 306 can include an accumulator unit 310 at the lower part of each column that stores and accumulates each accumulated output when performing calculations in a layer with more weight inputs than columns or a layer with more activation inputs than rows. In some implementations, each accumulator unit stores a plurality of parallel accumulated values. This will be further described below with reference to FIG. 6. The accumulator unit 310 can accumulate each accumulated output to generate a final accumulated value. The final accumulated value can be sent to the vector calculation unit. In some other implementations, the accumulator unit 310 passes the accumulated value to the vector calculation unit without performing any accumulation when processing a layer with fewer weight inputs than columns or a layer with fewer activation inputs than rows.

[0033] When the activation input and the weight input flow through the circuit, the circuit can "freeze" or temporarily stop the flow of the set of weight inputs to accurately calculate the accumulated value. That is, the circuit can temporarily stop the set of weight inputs, so that a particular set of weight inputs can be applied to a particular set of activation inputs.

[0034] In some implementations, the weight sequencer 324 configures whether the weight input shifts to an adjacent cell. The weight sequencer 326 can receive a control value from a host, such as the host interface 202 in FIG. 2, or an external processor. Each weight sequencer can pass the control value to the corresponding cell in the array 306. In particular, the control value can be stored in a weight control register within the cell, such as the weight control register 414 in FIG. 4. The control value can determine whether the weight input is shifted or loaded along the dimensions of the array, which will be described below with reference to FIG. 8. The weight sequencer can also send the control value to an adjacent weight sequencer, and the adjacent weight sequencer can adjust the shift or load of the corresponding weight input for the corresponding cell.

[0035] In some implementations, the control value is represented as an integer. Each weight sequencer may include a pause counter register that stores the integer. Also, the weight sequencer can decrement the integer before storing the control value in the pause counter register. After storing the control value in the pause counter register, the weight sequencer can send the integer to an adjacent weight sequencer and to the corresponding cell. For example, each weight sequencer may have a decrement circuit configured to generate the decremented integer from the control value. The decremented integer can be stored in the pause counter register. The stored control value can be used to associate a simultaneous pause of the shift across the columns of the array, which will be further described below with reference to FIG. 8.

[0036] In some implementations, pausing the weights in the circuit enables the developer to debug the circuit.

[0037] Other ways to pause the weights are also possible. For example, instead of passing the value in the pause counter register to an adjacent pause counter register, a tree may be used to pass the control value. That is, in a given cell, a signal can be passed not only to one adjacent cell but to all adjacent cells, thereby quickly dispersing the signal throughout the systolic array.

[0038] Figure 4 shows an exemplary architecture 400 of a cell within a systolic array, such as systolic array 306 of FIG. 3.

[0039] The cell may include an activation register 406 for storing an activation input. The activation register can receive the activation input from an adjacent cell on the left, i.e., an adjacent cell located to the left of a given cell, or from an integration buffer, depending on the position of the cell within the systolic array. The cell may include a weight register 402 for storing a weight input. The weight input can be transmitted from an adjacent cell above or from a weight fetcher interface, depending on the position of the cell within the systolic array. The cell may also include a sum register 404. The sum register 404 can store the cumulative value from the adjacent cell above. The multiplication circuit 408 can be used to multiply the weight input from the weight register 402 and the activation input from the activation register 406. The multiplication circuit 408 can output the product to the summation circuit 410.

[0040] The summation circuit can sum the product and the cumulative value from the sum register 404 to generate a new cumulative value. Then, the summation circuit 410 can send the new cumulative value to another sum register located in the adjacent cell below. The new cumulative value can be used as an operand for summation in the adjacent cell below.

[0041] In some implementations, a cell also includes a general control register. The control register can store a control signal that determines whether the cell should shift a weight input or a start-up input to an adjacent cell. In some implementations, shifting the weight input or the start-up input requires two or more clock cycles. The control signal can also determine whether to send the start-up input to the multiplication circuit 408 or send the weight input to the multiplication circuit 408, or determine whether the multiplication circuit 408 operates on the start-up input and the weight input. The control signal can also be passed to one or more adjacent cells using, for example, wires.

[0042] In some implementations, the weight is pre-shifted to the weight path register 412. The weight path register 412 can receive a weight input from, for example, an upper adjacent cell and send the weight input to the weight register 402 based on a control signal. The weight register 402 can statically store the weight input so that, for example, when the start-up input is sent to the cell in multiple clock cycles via the start-up register 406, the weight input remains within the cell and is not sent to an adjacent cell. Thus, the weight input can be applied to multiple start-up inputs using, for example, the multiplication circuit 408, and each accumulated value can be sent to an adjacent cell. Accordingly, the weight input can be applied to multiple start-up inputs using, for example, the multiplication circuit 408, and each accumulated value can be sent to an adjacent cell.

[0043] In some implementation examples, the weight control register 414 controls whether the weight input is stored in the weight register 402. For example, when the weight control register 414 stores a control value of 0, the weight register 402 can store the weight input sent by the weight path register 412. In some implementation examples, storing the weight input in the weight register 402 is referred to as loading the weight input. When the weight input is loaded, the weight input can be sent to the multiplication circuit 408 for processing. When the weight control register 414 stores a non-zero control value, the weight register 402 can ignore the weight input sent by the weight path register 412. The control value stored in the weight control register 414 can be passed to one or more adjacent cells of a given cell, for example, and the control value can be sent to the weight control register in the cell located to the right of the given cell.

[0044] Also, the cell can shift the weight input and the activation input to adjacent cells. For example, the weight path register 412 can send the weight input to another weight path register in the adjacent cell below. The activation register 406 can send the activation input to another activation register in the adjacent cell to the right. Thus, both the weight input and the activation input can be reused by other cells in the array in subsequent clock cycles.

[0045] FIG. 5 shows an exemplary matrix structure 500 having a spatial dimension and a feature dimension. The matrix structure 500 can represent either a set of activation inputs or a set of weight inputs. The matrix structure for the set of activation inputs is referred to herein as the activation matrix structure, and the matrix structure for the set of weight inputs is referred to herein as the kernel matrix structure. The matrix structure 500 has three dimensions, namely two spatial dimensions and one feature dimension.

[0046] In some implementations, the spatial dimensions correspond to the space or location of a set of activation inputs. For example, when a neural network is processing an image having two dimensions, the matrix structure may have two spatial dimensions corresponding to the spatial coordinates of the image, i.e., XY coordinates.

[0047] The feature dimensions correspond to features from the activation inputs. Each feature dimension may have a depth level. For example, matrix structure 500 has depth levels 502, 504, and 506. By way of example, if matrix structure 500 represents a 3×3×3 image sent as a set of activation inputs to a first layer, the X and Y dimensions (3×3) of the image may be spatial dimensions, and the Z dimension (3) may be a feature dimension corresponding to R, G, and B values. That is, depth level 502 may correspond to nine “1” activation inputs, e.g., features of red values, depth level 504 may correspond to nine “2” activation inputs, e.g., features of green values, and depth level 506 may correspond to nine “3” activation inputs, e.g., features of blue values.

[0048] In the example of FIG. 5, only three depth levels of the feature dimensions are shown, but a given feature dimension may have a large number of feature dimensions, e.g., hundreds of feature dimensions. Similarly, although only one feature dimension is shown, a given matrix structure may have multiple feature dimensions.

[0049] To perform calculations for a convolutional layer using matrix structure 500, the system must convert the convolutional calculations into two-dimensional matrix multiplications.

[0050] FIG. 6 shows an exemplary diagram of how matrix structure 500 of FIG. 5 is processed by systolic array 606 in a given convolutional layer. Matrix structure 600 is a set of activation inputs. It can be a dot. Generally, a neural network processor can send an activation input, for example, elements within a matrix structure 600, and a weight input, for example, kernels A - D 610, to the rows and columns of an array respectively. The activation input and the weight input can be shifted to the right side and the lower part of the systolic array respectively and must reach a specific location, for example, a specific register in a specific cell. For example, when it is determined that the input has reached a predetermined position by verifying a control signal, the processor can use the input stored in the cell to execute calculations to generate the output of a given layer.

[0051] As described above, the neural network processor "flattens" the matrix structure 600 before sending a part of the structure 600 to the rows of the systolic array. That is, the neural network processor can divide the depth layers 602 of the matrix structure 600, for example, the depth layers 602, 604, and 606 in FIG. 6, and send each depth layer to individual cells. In some implementations, each depth layer is sent to cells in different rows of the systolic array 606. For example, the processor can send the activation input from the first depth layer, for example, a matrix of 9 "1" activation inputs, to the leftmost cell in the first row of the systolic array 606, the activation input from the second depth layer, for example, a matrix of 9 "2" activation inputs, to the leftmost cell in the second row of the systolic array 606, the activation input from the third depth layer, for example, a matrix of 9 "3" activation inputs, to the leftmost cell in the third row of the systolic array 606, and so on.

[0052] A given layer can have multiple kernels, for example, kernels A - D 610. Kernels A - D 610 can have a matrix structure of dimension 3×3×10. The processor can send each kernel matrix structure to cells in an individual column of the systolic array 606. For example, kernel A can be sent to the upper cell in the first column, kernel B can be sent to the upper cell in the second column, and so on.

[0053] When the row-column structure is sent to the cells, the first element of the row-column can be stored in the cells within one clock cycle. In the next clock cycle, the next element can be stored in the cells. As described above with reference to FIG. 4, the first stored element can be shifted to adjacent cells. The shifting of the input can continue until all elements of the row-column structure are stored in the systolic array 606. Both the activation input and the weight input can be shifted across each cell after one or more clock cycles. The shifting of the input within the systolic array will be further described below with reference to FIG. 7.

[0054] FIG. 7 shows an exemplary diagram 700 of the weight input within the cells of an exemplary 3×3 systolic array after three clock cycles. As described above with reference to FIG. 5, each cell can store a weight input and an activation input. As described above with reference to FIG. 7, the weight input can be sent to the cells in individual columns of the systolic array for convolution operations. By way of example, the system sends a first kernel row-column structure having weight inputs of 1, 2, and 4 to the first column of the systolic array. The system sends a second kernel structure having weight inputs of 3, 5, and 7 to the second column. The system sends a third kernel structure having weights 6, 8, and 10 to the third column. After any clock cycle, the weight input can be shifted one-dimensionally, for example, from top to bottom, while the activation input can be shifted in a different dimension, for example, from left to right (not shown).

[0055] The weight inputs can be stored in the cells in an alternating pattern. That is, the state of the systolic array after the first clock cycle 702 shows a "1" in the top left cell. The "1" represents that the weight input of "1" is stored in the cell. In the next clock cycle 704, the "1" is shifted to the cell below the top left cell, and another weight input from the kernel, namely "2", is stored in the top left cell, and similarly the weight input of "3" is stored in the top cell in the second column. 2 of the columns.

[0056] In the third clock cycle 706, each weight is shifted again. In the first column, the bottom cell stores a weight input of "1", a weight input of "2" is stored where a weight input of "1" was stored in the previous cycle, and a weight input of "4" is stored in the top leftmost cell. Similarly, in the second column, "3" is shifted down and a weight input of "5" is stored in the upper middle cell. In the third column, a weight input of "6" is stored in the top rightmost cell.

[0057] In some embodiments, a control signal for the weight input that determines whether the weight input should be shifted is also shifted together with the weight input.

[0058] The start input can be shifted in the other dimension, for example, from left to right, in a similar manner.

[0059] When the start input and the weight input reach a predetermined position, the processor can execute a convolution operation by using, for example, a multiplication circuit and a summation circuit in the cell to generate a set of accumulated values used in the vector calculation unit.

[0060] Although the system has been described with the weight input being sent to the columns of the array and the start input being sent to the rows of the array, in some embodiments, the weight input is sent to the rows of the array and the start input is sent to the columns of the array.

[0061] FIG. 8 is an exemplary diagram of how control values can shift or load weight inputs. As described above with reference to FIG. 3, the control value 806 can be sent by the host and stored by the weight sequencer. The values in the graph represent, on a clock cycle 802 basis, the control values stored in the weight sequencers 808 - 814 corresponding to rows 1 - 4 804 of the systolic array, respectively.

[0062] In some implementations, if the control value in a given weight sequencer is non-zero, the weight input in the corresponding cell of the systolic array will shift to the adjacent cell. If the control value in a given weight sequencer is zero, the weight input can be loaded into the corresponding cell and used to calculate the product with the activation input within the cell.

[0063] As an example, the host can determine that it should shift before loading four weight inputs. At clock cycle 0, the host can send a control value of 5 to weight sequencer 808, which corresponds to row 1, i.e., the weight sequencer. Weight sequencer 808 includes a decrement circuit that takes one clock cycle to output a control value of 4 based on the control value of 5. Thus, the control value of 4 is stored in weight sequencer 808 in the subsequent clock cycle, i.e., clock cycle 1.

[0064] At clock cycle 1, the host sends a control value of 4 to weight sequencer 808. Thus, at clock cycle 2, weight sequencer 808 stores a control value of 3, for example, using a decrement circuit. At clock cycle 1, weight sequencer 808 can send a control value of 4 to weight sequencer 810. Thus, at clock cycle 2, after the control value of 4 is processed by the decrement circuit of weight sequencer 810, weight sequencer 810 can store a control value of 3.

[0065] Similarly, the host can send a control value of 3, a control value of 2, and a control value of 1 at clock cycles 2, 3, and 4, respectively. Since the decrement circuits in each of the weight sequencers 808 - 814 introduce a delay when decrementing the control value, by decrementing the control value 806 in each clock cycle, ultimately, the same control value can be stored in each weight sequencer, i.e., a control value of 1 at clock cycle 4 and a control value of 0 at clock cycle 5. ​

[0066] In some embodiments, when each weight sequencer outputs a control value of 0, the systolic array pauses the shift of the weight input and loads the weight input into each cell. That is, by loading the weight input, the systolic array enables the use of the weight input as an operand in the dot product calculation, thereby starting to process the layers within the neural network.

[0067] In some embodiments, after the calculation is completed, in order to resume the shift of the weights, the host changes the control value to a non-zero number, for example, sends a control value of 5 during clock cycle 7. The process of shifting can be repeated as described above with reference to clock cycle 0.

[0068] In some embodiments, the control value starts from another offset, for example, 1. Embodiments of the subject matter and the functional operations described in this specification may be implemented in digital electronic circuitry, or in tangibly implemented computer software or firmware, or in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory program carrier for execution by, or to control the operation of, a data processing apparatus. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information and transmit it to a suitable receiver apparatus for execution by a data processing apparatus. A computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.

[0069] The term "data processing apparatus" encompasses, by way of example, all kinds of apparatus, devices and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus may include dedicated logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits). The apparatus may also include, in addition to hardware, code for creating an execution environment for the computer program in question, such as processor firmware, protocol stack, database management system, operating system, or code constituting one or more combinations thereof.

[0070] (which may also be referred to as a program, software, software application, module, software module, script, or code, or described as such) A computer program may be written in any form of programming language, including compiler-type or interpreter-type languages, or functionally pure or declarative or procedural languages, and may be deployed in any form, including as a stand-alone program, or as a module, component, subroutine or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program may be stored as part of a file that holds other programs or data, such as one or more scripts stored in a markup language document, may be stored in a single file dedicated to the program in question, or may be stored in multiple cooperating files, such as files storing one or more modules, subprograms or portions of code. A computer program may be deployed to be executed on one computer, or may be deployed to be executed on multiple computers located at one location or distributed at multiple locations and interconnected by a communication network.

[0071] The processes and logical flows described in this specification may be executed by one or more programmable computers that execute one or more computer programs to function by operating on input data to generate output. Also, the processes and logical flows may be executed by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit), and the apparatus may be implemented as special-purpose logic circuitry, such as an FPGA or ASIC.

[0072] Computers suitable for the execution of a computer program may be, by way of example, general purpose or special purpose microprocessors, or any other kind of central processing unit. In general, a central processing unit receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks or optical disks, or is operatively coupled to receive data from, or transmit data to, or both receive and transmit data to and from, one or more mass storage devices. However, a computer need not have such devices. Furthermore, a computer may be coupled to another device, such as, by way of example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System :It may be implemented on a GPS receiver or on a portable memory device, such as a universal serial bus (USB) flash drive.

[0073] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0074] To require interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and a pointing device by which the user can send input to the computer, such as a mouse or a trackball. Other types of devices may also be used to require interaction with the user. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user may be received in any form including acoustic input, voice input, or tactile input. Also, the computer may interact with the user by sending documents to or receiving documents from the device used by the user, for example, by sending a web page to the web browser of the user's client device in response to a request received from the web browser. It may interact with the user by sending documents to or receiving documents from the device used by the user, for example, by sending a web page to the web browser of the user's client device in response to a request received from the web browser.

[0075] Embodiments of the subject matter described herein may be implemented in a computing system that includes back-end components, such as a data server, or in a computing system that includes middleware components, such as an application server, or in a computing system that includes front-end components, such as a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the subject matter described herein, or in a computing system that includes any combination of one or more such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0076] The computing system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship between the client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other.

[0077] This specification includes details of many specific implementation examples, but these should not be construed as limiting the scope of the invention or what may be claimed, but rather as describing features that may be specific to particular embodiments of a particular invention. Specific features described herein in the context of separate embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately in multiple embodiments or in any suitable partial combination. Further, features may be described as operating in a particular combination and may even initially be claimed as such, but one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a partial combination or a variation of a partial combination.

[0078] Similarly, operations are shown in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in the particular order shown or in a sequential order to achieve the desired result, nor as requiring that all of the shown operations be performed. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of the various system modules and components in the above embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be implemented in a single software product or packaged into multiple software products.

[0079] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims may be performed in a different order and still achieve desirable results. As one example, the processes illustrated in the accompanying figures do not necessarily require the particular order or sequential order shown to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

**Claim 1** A system for performing neural network calculations on a neural network having a plurality of neural network layers, comprising: a matrix calculation unit including an array of circuits, each circuit in the array: obtains weight inputs for a neural network layer among the plurality of neural network layers; receives a control signal; and is configured to determine whether to shift the weight inputs to another circuit based on the control signal for reuse of the weight inputs by the other circuit in a subsequent clock cycle. **Claim 2** The system according to claim 1, wherein the circuit is further configured to shift the weight inputs in response to a determination to shift the weight inputs to the other circuit. **Claim 3** The circuit is further configured to: obtain activation inputs for the neural network layer; and determine whether to shift the activation inputs to the other circuit based on the control signal for reuse of the activation inputs by the other circuit. **Claim 4** The system according to claim 3, wherein obtaining the activation inputs includes obtaining the activation inputs from a value loader. **Claim 5** Obtaining the weight inputs includes obtaining shifted weight inputs, and obtaining the activation inputs includes obtaining shifted activation inputs. **Claim 6** The system according to claim 3, further comprising: a first memory configured to provide activation inputs to the plurality of neural network layers; and a second memory configured to provide weight inputs to the plurality of neural network layers. **Claim 7** The system according to claim 6, further comprising a vector calculation unit including a circuit, the circuit: receives one or more accumulated values from the matrix calculation unit; determines a vector based on the one or more accumulated values; and is configured to provide the vector to the first memory. **Claim 8** The system according to claim 7, further comprising a sequencer circuit configured to provide one or more control signals to at least one of the first memory, the second memory, the vector calculation unit, or the matrix calculation unit. **Claim 9** The system according to any one of claims 1 to 4, wherein obtaining the weight input includes obtaining the weight input from a weight fetcher interface.

10. The system according to any one of claims 1 to 9, wherein determining whether to shift the weight input based on the control signal includes determining whether the control signal satisfies a predetermined value.

11. The circuit includes one or more weight control registers configured to store the control signal, and one or more weight registers configured to load the weight input. The system according to any one of claims 1 to 10.

12. The system according to claim 1, further comprising a sequencer including a circuit configured to provide the control signal to the matrix calculation unit.

13. The system according to claim 12, wherein the circuit of the sequencer includes a decrement circuit configured to decrement the value of the control signal in each clock cycle.

14. A method implemented by a computing unit comprising an array of circuits for performing neural network calculations on a neural network having a plurality of neural network layers, the method comprising: each circuit of the array obtaining a weight input for a neural network layer of the plurality of neural network layers; receiving a control signal; determining, based on the control signal, whether to shift the weight input to another circuit for reuse by the other circuit in a subsequent clock cycle.

15. The method according to claim 14, further comprising shifting the weight input in response to a determination that the circuit shifts the weight input to the other circuit.

16. The method according to claim 14 or 15, further comprising: the circuit obtaining an activation input for the neural network layer; and determining, based on the control signal, whether the circuit shifts the activation input to the other circuit for reuse by the other circuit.

17. The method according to any one of claims 14 to 16, wherein the step of determining whether to shift the weight input based on the control signal includes a step of determining that the control signal satisfies a predetermined value.

18. A matrix calculation unit for performing neural network calculations on a neural network having a plurality of neural network layers, the matrix calculation unit comprising an array of circuits, each of the circuits of the array obtains a weight input for a neural network layer among the plurality of neural network layers, receives a control signal, A matrix calculation unit configured to determine whether to shift the weight input to another circuit based on the control signal so that the other circuit can reuse the weight input in a subsequent clock cycle.

19. The matrix calculation unit according to claim 18, wherein the circuit is further configured to shift the weight input in response to a determination to shift the weight input to the other circuit.

20. The circuit is further configured to obtain an activation input for the neural network layer, The matrix calculation unit according to claim 18 or 19, configured to determine whether to shift the activation input to another circuit based on the control signal for reusing the activation input.

21. A program for causing a computer to execute the method according to any one of claims 14 to 17.

Citation Information

Patent Citations

  • Controllable shift matrix

    JP1983039106A

  • Communication system in parallel computer

    JP1985028345A

  • Communication method for parallel computer

    JP1988293668A

  • Two-dimensional contraction array and method for neural network

    JP1991131965A

  • Learning machine

    JP1992229362A