Pipeline operation in neural networks
The integrated circuit efficiently performs matrix operations in neural networks by using a large scale multiplier and combiner circuits to perform parallel multiplications, addressing the inefficiencies of sequential methods and reducing processing time and power consumption.
Patent Information
- Application Number
- JP2023563294
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-04-15
- Filing Date
- 2022-04-05
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2042-04-05
AI Technical Summary
Existing technologies for matrix operations in neural networks are inefficient due to sequential multiplication of weights with inputs, leading to high time and cost requirements.
An integrated circuit (IC) implementing an M×N aperture function on an R×C source array, using a large scale multiplier circuit to multiply input values by all required weights in parallel, and an array of combiner circuits to combine these products with initial values and delayed values, producing a stream of output values.
This approach significantly reduces the time and cost of matrix operations by performing all multiplications in parallel, allowing for faster processing and lower power consumption compared to traditional sequential methods.
Smart Images

Figure 0007681348000018 
Figure 0007681348000019 
Figure 0007681348000020
Abstract
Description
[Technical field]
[0001] This application is a continuation-in-part of co-pending application Ser. No. 17 / 071,875, filed Oct. 15, 2020. The entire disclosure of the parent application is incorporated at least by reference.
[0002] The present invention belongs to the technical field of computer operations involving input and output of matrices, and more specifically to circuits designed for large scale multiplication in matrix operations. [Background technology]
[0003] The use of computers in matrix operations is widely known in the art, with specific examples being image processing and the development and use of neural networks. Neural networks are an important part of artificial intelligence, and as such are a very popular subject in intellectual property development at the time of filing this patent application. Generally speaking, in this type of computer operation, a significant number of input values are processed in a regular pattern, which is most often a matrix. The processing of the input values may include biasing and applying weights by which individual input values may be multiplied.
[0004] The inventors believe that the sophisticated and computationally intensive operations in neural network technology, where an incoming value is multiplied by each of a plurality of weight values, are a step open to innovation to provide distinct advantages in the art. The inventors also believe that there are advantages to be gained by modifying the order in which the mathematical operations are applied.
[0005] The inventors believe that they have determined general changes in the order and manner of the mathematical operations implemented in such applications, which changes can result in very significant reductions in the time and cost of such operations. Summary of the Invention [Means for solving the problem]
[0006] In one embodiment of the invention, an integrated circuit (IC) is provided that implements an M×N aperture function on an R×C source array to generate an R×C destination array, the IC having an input port that receives an ordered stream of independent input values from the source array, an output port that generates an ordered output stream of output values into the destination array, and a large scale multiplier circuit coupled to the input port that multiplies, in parallel, each input value in turn by all weights required by the aperture function to generate a stream of products on a set of parallel conductive product paths on the IC, each product path being dedicated to a single multiplication by an input weight value. The system includes a multiplier circuit, an M×N array of combiner circuits on the IC, each combiner circuit associated with a subfunction of the aperture function at an (m,n) location and coupled by a dedicated path to each of a set of product paths carrying products generated from weight values associated with the subfunctions, a single dedicated path between the combiners, a delay circuit on the IC that receives values on the dedicated path from the combiners and provides a delayed value on the dedicated path to another downstream combiner at a later time, a finalization circuit, and control circuitry that implements a counter and generates control signals that are coupled to the combiners, the delay circuits, and the finalization circuit. At each source interval, a combiner combines values received from the dedicated connections into parallel conductive paths, and further combines the result with an initial value for that combiner, or a value on a dedicated path from an adjacent upstream combiner, or a value received from a delay circuit, and posts the combined result to a register coupled to a dedicated path to an adjacent downstream combiner, or a delay circuit, or both, and when the last downstream combiner has produced a complete combination of values for the output of the aperture function at a particular location of the input R×C array, the combined value is passed to a finalization circuit, which processes the value and posts the result to an output port as one value in the output stream.
[0007] In one embodiment, the aperture function is for a convolutional neural node, and at each source interval, the combiner adds the products of the weights with the inputs, adds the sum of the products to an initial bias, or a value on a dedicated path from an adjacent upstream combiner, or a value received from a delay circuit, and posts the sum to an output register. Also, in one embodiment, the aperture function produces truncated results for aperture positions where the M×N input patch overlaps the left and right edges of the R×C input array, and for certain source intervals where the source input position represents the first or last column of the R×C input array, the truncated patch results are delayed and accessed by the combiner and integrated with the complete interior patch flow. And, in one embodiment, the aperture function produces truncated results for those certain positions where the M×N input patch overlaps the top edge of the R×C input array, and for certain source intervals where the source input position represents the first row of the R×C input array, the truncated patch results are delayed and accessed by the combiner and integrated with the complete interior patch flow.
[0008] In one embodiment, the aperture function produces truncated results for those particular locations where the M×N input patch overlaps the bottom edge of the R×C input array, and for particular source intervals where the source input location represents the first row of the R×C input array, the truncated patch results are delayed and merged with the flow of full interior patches. And, in one embodiment of this IC, certain outputs of the aperture function are excluded from the output stream in a fixed or variable stepping pattern.
[0009] In another aspect of the invention, a method is provided for implementing an M×N aperture function on an R×C source array to generate an R×C destination array, the method comprising the steps of: providing an ordered stream of independent input values from the source array to an input port of an integrated circuit (IC); multiplying, in parallel, each input value in turn by all weight values required by the aperture function by a massive multiplier circuit on the IC coupled to the input port; generating, by the massive multiplier, a stream of products on a set of parallel conductive product paths on the IC, each product path being dedicated to a single product by an input weight value; and an M×N array of combiner circuits on the IC, each associated with a sub-function of the aperture function. providing, by a dedicated connection from the stream of products to each of the combiner circuits, products generated from the sub-functions and associated weight values; providing, by a control circuit implementing a counter and generating control signals, control signals to the combiner, the plurality of delay circuits, and the finalization circuit; combining, by the combiner, at each source cycle, values received in the stream of products from the dedicated connection with an initial value for that combiner, or with a value on a dedicated path to an adjacent upstream combiner, or with a value received from one of the plurality of delay circuits, and posting the result to a register coupled to a dedicated path to an adjacent downstream combiner, or to one of the plurality of delay circuits. When the last downstream combiner has generated a complete combination of values for the output of the aperture function at a particular location in the R×C array of inputs, providing the complete combination to the finalization circuit; processing the complete combination by the finalization circuit and posting the result to an output port as one value in an ordered output stream; and continuing operation of the IC until all input elements have been received and a final output value has been generated in the output stream.
[0010] In one embodiment of the method, the aperture function is for a convolutional neural node, and at each source interval, the combiner adds the products of the weights with the inputs, adds the sum of the products to an initial bias, or a value on a dedicated path from an adjacent upstream combiner, or a value received from a delay circuit, and posts the sum to an output register. Also, in one embodiment, the aperture function produces truncated results for aperture positions that overlap an M×N input patch with the left or right edge of an R×C input array, and for certain source intervals where the source input position represents the first or last column of the R×C input array, the truncated patch results are delayed and accessed by the combiner and merged with the full internal patch flow.
[0011] In one embodiment of the method, the aperture function produces truncated results for certain locations where the M×N input patch overlaps the top edge of the R×C input array, and for certain source intervals where the source input location represents the first row of the R×C input array, the truncated patch results are delayed, accessed by the combiner, and merged with the full interior patch flow. In one embodiment, the aperture function produces truncated results for certain locations where the M×N input patch overlaps the bottom edge of the R×C input array, and for certain source intervals where the source input location represents the first row of the R×C input array, the truncated patch results are delayed and merged with the full interior patch flow. And in one embodiment, certain outputs of the aperture function are excluded from the output stream in a fixed or variable stepping pattern. [Brief description of the drawings]
[0012] [Figure 1] 1 is an illustration of an embodiment in which the large multipliers applied to each common source are fixed and wired directly to the processing circuitry. [Diagram 2] 13 is an illustration of an embodiment in which the large multipliers applied to each common source are dynamic and routed through multiplexers to processing circuits. [Diagram 3]13 illustrates a simple embodiment in which shifted terms corresponding to set bits in each large multiplier are summed to form a product. [Figure 4] 13 is an illustration of an enhanced embodiment in which additions of shifted terms and subtractions from each other are mixed to form an equivalent solution of lower complexity. [Figure 5A] 1 is an illustration of a pipelined embodiment that maximizes clock frequency by building sub-synthesis from pairwise operations only. [Figure 5B] 1 is an illustration of an embodiment in which multiples are formed directly by a fixed set of cases, without reference to standard arithmetic operations. [Figure 6] 1 is an illustration of a pipelined embodiment that maximizes circuit density by building sub-compositions out of up to every fourth operation. [Figure 7] FIG. 1 illustrates the structure and connectivity in one embodiment of the present invention that receives an input stream, pre-processes the input stream, and feeds the results through a unique digital device to generate an output stream. [Figure 8A] FIG. 1 illustrates the structure and connectivity for generating a source-channel product. [Figure 8B] FIG. 2 illustrates additional details of the controllers and functions in one embodiment of the present invention. [Figure 9A] 1 is a partial diagram of a general case of pipelined operation in one embodiment of the present invention. [Figure 9B] 4 is another partial diagram of a general case of pipelined operation in one embodiment of the present invention. [Figure 9C] 4 is another partial diagram of a general case of pipelined operation in one embodiment of the present invention. [Figure 10A] FIG. 9C illustrates the internal structure of combiners 905a, 905b, and 905c of FIGS. 9A and 9B in one embodiment of the present invention. [Figure 10B]FIG. 9C illustrates the internal structure of combiners 902a, 902b, and 902c of FIGS. 9A and 9B in one embodiment of the present invention. [Figure 10C] FIG. 9B illustrates the internal structure of the combiner 904 of FIG. 9A in one embodiment of the present invention. [Figure 10D] FIG. 9B illustrates the internal structure of the combiner 901 of FIG. 9A in one embodiment of the present invention. [Figure 10E] FIG. 9C illustrates the internal structure of combiners 903a, 903b, and 903c of FIGS. 9B and 9C in one embodiment of the present invention. [Figure 10F] FIG. 9C illustrates the internal structure of combiners 907a, 907b, and 907c of FIGS. 9A and 9B in one embodiment of the present invention. [Figure 10G] FIG. 9B illustrates the internal structure of the combiner 906 of FIG. 9A in one embodiment of the present invention. [Figure 11] FIG. 9D illustrates the internal structure and operation of delay stages 908a, 908b, 908c, 908d, 908e, and 908f of FIG. 9C in one embodiment of the present invention. [Figure 12] 9D illustrates the operation of delay stage 909 of FIG. 9C in one embodiment of the present invention. [Figure 13] 9D illustrates the operation of delay stages 910a and 910b of FIG. 9C in one embodiment of the present invention. [Figure 14] FIG. 9D illustrates the operation of finalization step 911 in FIG. 9C. [Figure 15] FIG. 1 illustrates a specific case of pipelined operations in one implementation of the invention implementing a 5×5 convolution node. [Figure 16] 1 is an illustration of an IC for a 4×4 aperture function in one embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] A variety of image and data algorithms make extensive use of the matrix form of linear algebra, both to prove propositions and to compute results. In this application, by "algorithm" is meant a process or set of rules to be followed, especially in a computational or other problem-solving operation. An algorithm, in this application, should not be construed as software without exception. The algorithms described in this application may typically and preferably be implemented in hardware.
[0014] Matrix operations are defined as orthogonal collections of one or more dimensions, and are commonly thought of as having the same number of elements in all iterations of each given dimension. As an example, an M×N matrix is often written as:
number
[0015] Conceptually, a matrix can have any number of dimensions, and a matrix can be represented as a set of tables that indicate the values for each dimension.
[0016] A subset of matrices of the form M×1 or 1×N are sometimes called vectors; vectors have their own defined properties and operations and are used extensively in 2D and 3D graphic simulations.
[0017] Degenerate subsets of matrices of the form 1×1 are sometimes called scalars, and constitute numbers that are very familiar to those skilled in the art.
[0018] When matrix values are constants and matrices are of suitable dimensions, some operations such as multiplication are well defined. A 3x4 matrix A can be multiplied with a 4x5 matrix B, often as follows: A×B=C
number
[0019] However, the operation B×A is not well defined because the inner dimensions do not match (5≠3) and k cannot have a single range that is compatible with the indices of B and A.
[0020] Matrices whose elements are vectors or other matrices are known as tensors (hence the name TensorFlow). A familiar form of a tensor can be an RGB image. One form of an RGB image is an HDMI frame as a 1080x1920 matrix of RGB values, where each pixel is a 3x1 vector of color components. A pixel is considered a true vector, because a linear operation on the red component does not affect the green or blue, and vice versa.
[0021] An HDMI frame is not generally considered to be a 5-dimensional matrix, because the processing of pixel positions in an image is not related to the processing of color. Cropping an image by discarding parts of the image that are not of interest is valid and quite meaningful, but there is no corresponding operation for cropping color components. Similarly, there may be many operations on color with easily understandable effects that would not make sense if applied to the elements of the containing array. Thus, an HDMI frame is clearly a few tensors, not a 5D array.
[0022] Many image processing algorithms are known that can be expressed as matrix operations, which are a concise way of expressing repetitive operations, and the rules of matrix mathematics are useful in proving certain propositions.
[0023] Execution of matrix-based algorithms in general-purpose computer processors is commonly accomplished by looping mechanisms, and both computer languages and hardware CPUs may have features that make such loops efficient, but there is nothing inherent in the mathematics of matrix definition that requires operations to be performed in a particular way or scheme in order to compute a correct result.
[0024] A modern hybrid of image processing and recognition is the Convolutional Neural Network (CNN). Training such networks has been extremely challenging for many years, but actually running a trained network is relatively trivial.
[0025] In a CNN, the output element of each convolution operates by passing an independent kernel on the input tensor to generate each component of the output tensor. Typically, when a neural network is used to process an image, the first layer of the network operates on an input array of RGB pixels of the image and generates an output array of related size that contains an arbitrary vector of output components that are structurally unrelated to the RGB vector of the input components. The components of the output vector are generally described as features or activations, and represent the response strength (degree of recognition) of each kernel. Subsequent layers in a CNN take as their input the output from the previous layer, so only the very first layer operates on pixel values; all the rest operate on features to generate more features. Each output feature of a convolution is independent and distinct from any other feature, just as color components are distinct from each other.
[0026] A common form of a CNN layer is a 3x3 convolution. In operation, a 3x3 kernel of constant weights is applied element-wise to each particular location of the input tensor (i.e., the image); that is, each of the weights is multiplied by the pixel component at the same relative location in the image, and the products are summed to produce a single component of the output for that location. A bias constant (which can be zero) provides an initial value to make it easier to solve the model to arrive at optimal weight values.
[0027] If there are three input components (which in the case of the first layer are colors), as there are in an RGB image, then there are three separate sets of 3x3 weights applied to each component value, but only one initial bias. Adding the bias to each convolution of the 3x3x3 weights forms a single output component value that corresponds to the pixel's position in the center of the 3x3 patch. Each output channel then applies its own 27 weight values until all output components for a given patch (a subset of input components at the same location as the output location, and corresponding to the relative location of the kernel weights) have been calculated. It is typical for a convolution to have between 64 and 256 output components, each of which has a unique specific set of 27 weights plus a bias.
[0028] In this example, each kernel multiplies its 27 weights with the same patch of 9 pixels in the 3 RGB components. For a relatively small set of 64 output components, each input component is multiplied with 64 arbitrary and unrelated weights. After the output components for each patch are calculated, the neighboring patch is loaded from the image and the full set of kernel weights is applied again. This process continues until the right edge of the image is reached, then the patch drops down one row and starts over with the left edge.
[0029] After the first layer is processed, the next convolutional layer processes the output of the first layer as input to the second layer. Thus, a 3×3 convolution now has 3×3×64 weights applied to the 3×3×64 input components of the patch. If this layer has 256 outputs, then 3×3×64×256=147,456 multiplications must be performed for each output position. Those skilled in the art will understand that this refers to a single layer in a deep neural network that may contain more than 40 layers.
[0030] The number of multiplications applied to each element of a patch is equal to the number of channels in the layer. In a standard CPU, these must always be done in some sequence. Many modern CPUs have the ability to perform a set of multiplications simultaneously, especially when the data format is small (i.e., 8-bit). In a GPU or TPU, the number of multipliers available is much larger, but each multiplier is designed to generate a product from two separate and unlimited factors.
[0031] State-of-the-art processors, CPUs, TPUs, or GPUs do not take advantage of the simple fact that in a CNN implementation, one of the factors for multiplication is common to all weights applied to an input channel during processing for a patch.
[0032] The inventors of this application propose a large scale multiplier that performs all multiplications in a single step, whereas conventionally all multiplications are instead performed sequentially. When the weights of the set of multiplications are all in a small precision (typically 8 bits for the TPU), the number of iterations is limited (2 8(This can be of any size; no matter what the precision of the common factor is, there are still only 256 possible multiples when 8-bit weights are applied.) In this case, there is a clear advantage to implementing a circuit that produces all the required outputs at once using many fewer elements than the same number of unlimited multipliers.
[0033] In one embodiment of the present invention, an equivalent large multiplier is dedicated to a single input channel and is not always shared, so operations have the option of using several clock cycles and multiple register stages, allowing operations to take a very simple and efficient form without impacting the overall throughput of the system.
[0034] In the general case where a single dynamic value is multiplied by many constants, using a single, multi-stage large scale multiplier circuit, as in embodiments of the present invention, instead of an equivalent set of independent single stage multiplier circuits results in a system that performs the same calculation with substantially higher throughput and substantially lower power and footprint. Even if the set of outputs is smaller than the number of actual multiples used, significant savings in power and space can still be possible.
[0035] Having established a clear advantage of the unique large scale multiplier in one implementation of the present invention over an independent multiplier, this advantage can be further increased by changing the order of the sequence of operations.
[0036] There is nothing in the mathematics of algorithms in neural networks (or other similar image processing) that requires any particular sequence of operations. The same operations can be performed in any order and will produce the same correct calculation. The inventors have observed that the usual order for software running on a CPU, GPU, or TPU based design is to generate all output channels for a given position simultaneously by multiplying the weights by the inputs and adding them immediately. Generating all output channels for a given position simultaneously by multiplying the weights by the inputs and adding them immediately minimizes the number of times the inputs must be read from RAM, and also limits the number of times the weights must be read from RAM. It does not preclude reading the inputs multiple times, since there is no place other than RAM to keep them when processing the next row below.
[0037] However, in one embodiment of the present invention, if the order of operation of a kernel or other aperture function defined to operate on an M×N patch of array inputs is reversed, i.e., effectively flipped, then each input value is used only once and no RAM buffer is required. Instead of generating outputs one at a time by redundantly reading the inputs as the aperture function passes over each row, this unique operation processes the inputs one at a time only when they are first presented, and keeps partial sums for all incomplete outputs. The partial sums can be kept in hardware shift registers or standard hardware first-in-first-out registers (FIFOs), with the number of registers required to hold the held values proportional to the height of the kernel and the width of the input row.
[0038] The function implementing the aperture function can be decomposed into a series of subfunctions, each of which operates on the results of the previous subfunction, so that the implementation of the kernel can be achieved by sequentially composing the subfunctions over time, resulting in a series of operations each of which operates immediately on the received data and theoretically is identical to applying the kernel. We refer to this recomposed function, including any initialization, as the aperture function, and the individual steps as subfunctions. An aperture function, as used herein, refers to any M×N computation implemented at multiple locations over a sliding window or patch of M×N inputs of a larger R×C array of inputs. The aperture function may also include initialization and finalization operations, as in the case of a full CNN kernel implementation. In the case of a CNN, initialization preloads bias values into an accumulator, and finalization transforms the raw output of the kernel through any activation function.
[0039] In this example of the invention, given the components of each new input position, the components at that position represent the first elements of the patches below and to the right, as well as the last elements of the patches above and to the left, and the middle elements of all other patches that intersect with the current position. This allows a computational circuit to be developed, as one embodiment of the invention, that always has a fixed number of elements in progress (with some possible exceptions near the edges of the input), and produces output as fast as it accepts inputs.
[0040] When the guiding algorithm requires evaluation of the aperture function on patches that extend beyond the edges of the input array, many special cases and challenges arise, but they are not insurmountable. Special case logic can be added so that partial results for overlapping patches match the general case without affecting overall throughput.
[0041] In the implementation of the present invention, the operation of this inverted form of the aperture function accepts inputs as a stream and produces outputs as a stream. The inputs do not need to be buffered in RAM because they are each referenced only once. Because the outputs are also in a stream, they can be processed by subsequent layers without buffering by RAM, a result attributable to the present invention that substantially increases processing speed over many other things that require necessary read and write operations to and from RAM.
[0042] In one embodiment of the present invention, instead of many layers sharing a single set of independent multipliers that operate, store, and then read the results back out to process the next layer sequentially, a pipeline can be created using dedicated large multipliers that process all layers simultaneously without waiting for any layer to be complete and feed the output stream of each layer to the input of the next layer.
[0043] A fully implemented pipeline in one embodiment of the present invention can thus reach an effective throughput measured two orders of magnitude better than traditional output-centric ordering processes, eliminating the contention for RAM (since it does not use RAM), which forms the main bottleneck in the case of GPU and TPU-based processing.
[0044] The latency of such a system in one embodiment of the present invention is reduced to the time from the input of the last pixel to the output of the final result. Since the last pixel of the image must, by definition of the algorithm, be the last data required to complete all of the final computations for all layers, the latency of the system is exactly the clock rate times the number of distinct clock stages in the pipeline that contain the final output.
[0045] The use of a single dedicated massive multiplier for each input channel throughout the neural network in one embodiment of the present invention (instead of a limited set of independent multipliers that must be reused and dynamically allocated) makes it possible to build a pixel-synchronous pipeline in which all multiplications are performed in parallel since only one massive multiplier is needed to handle any number of weights that are applied.
[0046] Having described the essential features of the large scale multiplier innovation, and also the advantages of inversion, the inventors present a specific example below: FIG. 1 illustrates one implementation of the present invention, in which each of a plurality of one or more source channels 1 to N are labeled 101a to 101d and assigned a dedicated large-scale multiplier 102a to 102d. Since each source channel in this example has a dedicated large-scale multiplier circuit that produces a set of multiples of the channel's value, the format of the source channels can vary between signed, unsigned, fixed, or floating point, in any precision convenient for the processing algorithm implemented in the hardware. The specific output of each large-scale multiplier circuit, such as large-scale multiplier circuit 102c, can be fed directly into one or more computation units 103a to 103d that can perform calculations requiring multiples of any or all of the source channels. These computation units can be used to implement independent output channels of a single algorithm or unrelated algorithms that are calculated on the same source channel. The output of the computation can be forwarded for further processing, shown at 104, that may be required by one or more algorithms implemented in the hardware. This situation arises, for example, when implementing neural networks in field programmable gate arrays (FPGAs), where the weight values applied as multiplicands do not change.
[0047] Figure 2 illustrates one embodiment of the present invention, where the output of each massive multiplier, such as massive multiplier 102a of Figure 1, is fed into computation units 203a-203d through a set of multiplexers 201a-201d such that the selected multiplier is chosen at system initialization or dynamically chosen as the system operates. The output of the computation may then be forwarded for further processing in 204, as previously described. This situation arises when implementing a neural network as an application specific integrated circuit (ASIC), where the structure of the computation is committed but the weight values used need to be changed.
[0048] FIG. 3 illustrates the internal structure of the large-scale multiplier 102a of FIG. 1 and FIG. 2 in one embodiment. This structure may be common to the large-scale multipliers 102b, 102c, and 102d, as well as to other large-scale multipliers in other embodiments of the present invention. In this structure, products 303a to 303f of the source channel multiplicand 101a, which is A bits, and all possible multipliers, which are B bits, are generated in parallel and delivered to a multiple 304. In this example, the A bits of the source multiplicand 101a are duplicated, shifted up by appending a 0 bit to the bottom position, and padded by prepending a 0 bit to the top position, so that a full set of all requested shifted values from 0 to B-1 is available in the form of a vector of A+B bit terms 302a to 302d. These terms may be formed simply by routing circuit connections, no registers or logic circuits are required. If the clock period is sufficient to allow the maximum value of B terms of A+B bits to be combined in a single cycle, registers or sub-combinations may not be required. The individual products of the added terms 303a through 303f may be stored locally or forwarded for further processing as combinatorial logic. Whenever a bit occurs in each multiplier, 1 through 2 of the source multiplicand 101a are BEach product of -1 may be formed by adding any or all of the corresponding terms 302a through 302d of B. Any source multiple 0 is an all-0 bit constant that may be included in multiple 304 for completeness when using multiplexers, but does not otherwise require circuitry. Any unused products 303a through 303f may be excluded, either by removing them from the circuit specification, allowing the synthesis tool to delete them, or by any other method. Unused terms 302a through 302d may also be excluded, but this generally has no effect since they do not occupy logic. In this way, all required multiples 304 of source multiplicand 101 may be formed as a single stage pipeline or as combinational logic.
[0049] 4 shows an optimized embodiment in which the set of terms 401 is composed of all required individual terms 302a to 302e, including 0 to B formed by A+B+1 bits. This allows the products 402a to 402f to include subtractions from larger terms instead of additions of smaller terms, which can be used to reduce the overall size of the circuitry, which may also increase the maximum possible clock frequency. For example, for any given input a and multiplier 15, 8a+4a+2a+1a=15a combines four components, while 16a-1a=15a combines only two, which is generally expected to be more compact and efficient. Each product 402a to 402f may be composed of additions and subtractions of terms 302a to 302e that produce the correct result, and each particular variant may be chosen based on the optimal tradeoff for a particular implementation technology. For example, subtraction of two N-bit quantities may require more logic than addition of two N-bit quantities, but in general, addition of three N-bit quantities will always require more logic than subtraction of two. The processing of required multiples 304 is not altered by the details of combining the individual products 402a through 402f.
[0050] FIG. 5A illustrates one embodiment of a large multiplier, in which the clock period is such that only a single addition of the A + B bit value (or A + B + 1 if subtraction is used) can be performed per cycle. In this case, in order to accommodate multiples where more than two terms are utilized, it is necessary to configure the required elements into a multi-stage pipeline. Term 401 is formed from each source channel 101 as described above, but is held in pipeline registers 501a and 501b one or more times for later reference. A pair 502 of two added terms is calculated and recorded, and then saved 503 as needed. Triple 504 is formed as the sum of pair 502 and term 501 being held. The value quad 505 of the terms is formed as the sum of pair 502. All unused elements may be excluded, and only the sequence of decreasing addends may be specified to increase overlap. This ensures that both redundant sums such as a + b and b + a are not utilized and are not held in the final circuit. Products 506a through 506f can utilize any addition or subtraction operation of any pair of the recorded sub-combinations that satisfy the timing constraints. By consistently using the maximum available elements, the overall size, and thus the power, can be reduced, but any combination of operations that produce the correct result is acceptable.
[0051] The embodiment of FIG. 5A is sufficient to generate all the required multiples for B = 8. For larger sets of multiples, the sub-combinations shown can be recombined again in additional pipeline stages such that all the required multiples 506a - 506f for any value of B are formed from a single clock operation of an extended set of sub-combinations that includes the previously disclosed and held term 501b, the held pair 503, the triple 504, the quad 505, and the other sub-combinations required to form a set of terms sufficient to form the multiples 506a - 506f by a single clock operation.
[0052] FIG. 5B illustrates an embodiment where multiples are formed directly by a fixed set of cases without reference to standard arithmetic operations. For each required multiple, a set of output values a*b is enumerated for each source channel value a. This allows a hardware circuit synthesis tool to determine the optimal logic circuit 507 to generate the full set of required multiples. The specification of the required output values for any given input value is typically created by enumeration in a Verilog "case" or "casex" statement. This is quite distinct from a lookup table where the output values are stored and accessed via an index formed from the inputs, because logic gates are used to implement a minimal subset of the operations required to generate the full set of output values, and redundant logic used to generate the relevant sub-expressions is combined.
[0053] Whether method 5A or 5B is the most efficient in terms of space, frequency, and power also depends on the particular values of A and B, as well as the core efficiency between arithmetic operations and arbitrary logic. The choice of which method to use may be based on direct observation, simulation, or other criteria.
[0054] 6 illustrates an embodiment in which the clock cycle allows, with sufficient levels of logic, the composition of additions and / or subtractions of four elements during each single clock cycle. By selecting from a set of sub-combinations, each product 605a-605f may be generated by combining four or fewer stored elements. As before, terms are stored in registers 501a and 501b, but triple 601 stored in 602 is composed directly from term 401, no pairs are used. Septet 603 and octet 604 are formed from triple 601 and stored term 501a.
[0055] The exemplary embodiment of Figure 6 is sufficient to generate all required multiples for B = 32. For larger multipliers, the sub-combinations shown are recombined four at a time in further pipeline stages to generate all required multiples for any value of B. While the elemental sub-combinations shown are necessary and sufficient to generate all products for B = 32, other sub-combinations (possibly chosen for consistency across different values of B) are acceptable.
[0056] When the set of multipliers is fixed, as is typical for FPGA applications, even a large coarse set of multipliers can be implemented efficiently because common elements can be merged and unused elements can be omitted. When a synthesis tool performs this function automatically, the representation of the circuit can include all possible elements without explicitly declaring which multipliers will be used.
[0057] If operations on A+B or A+B+1 bit values cannot be completed in a single clock cycle, a multi-stage pipeline adder can be inserted for any single stage of synthesis logic, with additional pipeline registers inserted as necessary so that all paths have the same number of clock periods. A pipeline stage period can be an instance of a single edge-to-edge clock transition, or a multi-cycle clock if throughput constraints permit. Neither the use of multiple clock stages per operation nor multi-cycle clock operations requires structural changes to any of the embodiments, other than the issues just mentioned.
[0058] An important objective of the present invention is to provide the industry with a large scale multiplier implemented in an integrated circuit for use in a variety of applications. To this end, the inventors provide, in one embodiment, a large scale multiplier implemented in an integrated circuit having a port for receiving a stream of discrete values, circuitry for simultaneously multiplying each value received at the port by a number of weight values, and an output channel for providing the generated large scale multiplier product.
[0059] In one version, the received discrete values may be unsigned binary values with a fixed width, the weight values may be unsigned binary with a fixed width of 2 or more bits, and each multiple may be synthesized as a sum of bit-shifted replicas of the input. In another version, the set of shifted replicas may be increased to allow the use of subtraction operations to reduce or otherwise optimize the circuit. Unused outputs of the set may be excluded, either explicitly or implicitly.
[0060] In one embodiment, the set of output products may be generated by combinational logic. In another embodiment, the set of output sets may be generated by a single stage pipeline using a single or multiple clock cycles. In another embodiment, the set of output multiples may be generated by a multi-stage pipeline by combining no more than two addends per stage. Unused elements of intermediate sub-compositions may be explicitly or implicitly excluded from the circuit.
[0061] In one embodiment, the set of output products may be generated by a multi-stage pipeline combining more than two addends per stage, and the sub-compositions are adjusted accordingly. Unused elements of intermediate sub-compositions may be explicitly or implicitly excluded from the circuit.
[0062] Another object of the present invention is to provide large scale multiplication in an integrated circuit for implementing substantially improved convolutional neural networks in the ongoing advancement of deep learning and artificial intelligence. In this end, the inventors provide a first convolutional neural network (CNN) node implemented as an integrated circuit having a first input channel defined as a stream of discrete values of a first component of an array of elements.
[0063] In this description, the inventor intends the nomenclature of an element of an array to mean an element that may have a single component or multiple components. A suitable example is an image, where an image may have pixels as elements, and each pixel may have a single element if the image is monochrome, or may have three color values in one example if the image is RGB color. Each color value in this example is a component of the element that is a pixel.
[0064] Continuing with the above description of a first convolutional neural network (CNN) node implemented in an integrated circuit with a first input channel defined as a stream of discrete values of a first component of an array of elements, the CNN further includes a first large scale multiplier circuit that simultaneously multiplies the received first component discrete value by a plurality of weight values, and an output channel providing an output stream of discrete values.
[0065] In one embodiment of the CNN node, the first output stream is formed from the product of the first large scale multiplier circuit, in some circumstances by combining the product with a constant, and in some circumstances by applying an activation function.
[0066] In another embodiment, the CNN node further comprises a second input channel defined as a stream of discrete values of the second component of the elements of the array, and a second massive multiplier circuit that simultaneously multiplies the received discrete values of the second component by multiple weight values. In another embodiment, there may be a third input channel defined as a stream of discrete values of a third component of the elements of the array, and a third massive multiplier circuit that simultaneously multiplies the received discrete values of the third component by multiple weight values.
[0067] Although a CNN node having one, two or three input component streams and a dedicated large-scale multiplier has been described, the inventor further provides a convolutional neural network (CNN) having a first convolutional neural network (CNN) node implemented as an integrated circuit with input channels defined as streams of discrete values of the components of the elements of the array, large-scale multiplier circuits dedicated to each input channel and multiplying the discrete values of the received components by multiple weight values simultaneously, and output channels providing an output stream of discrete values, and a second CNN node having an input that depends at least in part on the output of the first node. This CNN may have successive nodes and may operate as a deep neural network (DNN). It is not required that the successive nodes after the first node are CNN nodes.
[0068] Pipelined aperture function operation Referring back now to the earlier description herein, which discussed the order of operations in processing a CNN or other similarly selected aperture function that passes an array of computational subfunctions over an array of inputs to generate a net result, a specific description is now provided regarding the inverted form of operation of the aperture function in one embodiment of the present invention, which accepts inputs as a stream and generates outputs as a stream. In this embodiment of the present invention, the inputs are not and do not need to be buffered in RAM, since each input is referenced only once. Since the outputs are also generated as a stream, the output stream can be processed by subsequent layers without buffering by RAM. The inventors believe that this innovation substantially increases the processing speed in comparison to many other processing systems that require read and write operations to RAM.
[0069] In one embodiment of the present invention, an apparatus and method is provided in which the operation of passing a two-dimensional aperture function over a two-dimensional array operates on the incoming stream of inputs such that all inputs are processed immediately, partially completed calculations are held until all required inputs have been accepted and processed, and outputs are generated as a coherent stream, typically having the same or a lower data rate than the input stream. All inputs are accepted and processed at the rate at which they are provided, and are not required to be stored or accessed in any order other than the order given. Even if the application of the aperture function is defined such that more outputs are generated than inputs, the circuitry can still operate at the speed of the incoming data by selecting the processing clock rate with a sufficient increment so that the system never does not accept and process a given input.
[0070] The traditional way to implement the convolution of a kernel, or a more general aperture function, over a larger input array is to collect the required input patches, apply the function to the input, and output the result. As the aperture is passed over the input array, each successive patch overlaps with the one just processed, so some input can be retained and reused. Various mechanisms, such as FIFOs, can be used to avoid redundantly reading inputs from source storage as the patch advances to each new row, but the source data will still be applied to each location in the kernel to generate each output in turn where the input patch overlaps each particular data input location.
[0071] If there are many output channels and many independent aperture functions to be calculated, then a large scale multiplier can be used to provide all of the aperture functions with products of the patch of input values under consideration in parallel. However, with this configuration and order of operations, each location of the source data will require a set of products for each location in the kernel as it is combined into various overlapping output locations.
[0072] The mechanism of the present invention reverses or flips the order of operations, for the special advantage of using a single large multiplier applied only once per input channel to a given input value. Rather than retaining or re-reading source values for later use in the form of later product calculations, the process in one embodiment of the present invention calculates all required products of each input as they are given, and keeps running sums for each element of the aperture function that are complete up to the point in time when the current input appears.
[0073] Any aperture function that can be mathematically decomposed into a series of sub-functions applied sequentially can be implemented in this way. This mechanism can be easily applied because the CNN kernel is nothing more than a series of sums of weight-multiplied inputs, and the order of operations matches the order of the source inputs taken from left to right and top to bottom.
[0074] In one embodiment of the present invention, an array of combiners corresponding to sub-function elements of the aperture function is implemented on the IC, each keeping a running total of the aperture function's values as it moves along the input stream. The final combiner in the array outputs the complete value of the function, and all other combiners output partial values of the function.
[0075] In the simple case of applying a 3x3 kernel, the output of the top-left compositor reflects the first element of the kernel applied to the current input plus any initialization constants, the output of the top-middle compositor reflects the first two steps, and the output of the top-right compositor reflects the first three steps. The output of the top-right compositor needs to be delayed until it can be used again by the next line. The next line of compositors continues the pattern of accepting the partially completed function value, adding the contribution of each new input, and passing it forward. The last line of compositors completes the last step of the function and outputs the completed value for any further processing.
[0076] Noting that the progression of partial values of a function between compositors is generally left-to-right in the first row, then in subsequent rows, and finally to the last compositor in the last row, the flow of partial values can be thought of as a stream, and the compositors and flows can be referred to as upstream or downstream.
[0077] At all times, each combiner maintains a partial sum of the aperture functions up to and including the current source input. Each combiner is always operating on a different patch position of the output, specifically, on that patch position where the current input appears relative to the combiner's position in the aperture subfunction array.
[0078] The 3×3 kernel W is a function of the input A.
number
[0079] The circuitry required to compute these subfunctions then consists of a corresponding array of combiners:
number
number
[0080] Here, a i is the current value from the input stream, and a i-1 From a i-8 Until, in each case, a i is the previously processed input for a particular patch that appears at a position relative to the output of each individual combiner. Each combiner will calculate the value of the aperture function up to and including its corresponding position in the aperture array. Each combiner takes the current value of the input stream and combines it with the previous value to generate a different partial sum that corresponds to a partially processed patch in the input array, where the current input value appears at the relative position of that patch that corresponds to each combiner's position in the aperture function.
[0081] In this way, partial values of the aperture function, calculated in standard order and precision, will be kept in the input stream over time until the complete value is ready to be output.
[0082] While this technique is quite straightforward within the input array, complications arise when applied to patches that overlap the edges of the input array, because the aperture function is defined differently when not all inputs are available. In the case of the CNN kernel, an additional operation is omitted, which is equivalent to using zero as the input. The present invention is interested in maintaining a steady flow of partial sums through the combiner while handling such exceptions, as will be described below.
[0083] FIG. 7 is a diagram illustrating the structure and connections in one embodiment of the present invention that receives an input stream, pre-processes the input stream, and applies the results through a unique digital device to generate an output stream.
[0084] The set of input channels 701 and associated control signals 702 are used by a common circuit 703 to generate any products of the input channel set with weights for subsequent sub-functions. The products of the source channels are then distributed to a bank of sub-function computation circuits 704a, 704b, and 704c, each of which generates a single channel of the set of output channels 705. Any number of independent output channels can be supported by the common circuit 703.
[0085] FIG. 8A is a diagram illustrating the large-scale multipliers 801a, 801b, and 801c in the common circuit 703 of FIG. 7, which take each channel of the input channel set 701 and generate either a loose or complete set of multiples as required by the sub-functions defined. It should be noted that this illustration assumes three channels in the input channel set, such as for red, green, and blue pixel values when processing an RGB image. In other embodiments, there may be one, two, or more than three channels. Any or all of the products 802 (multiples of source input array values constructed by the large-scale multipliers) may be made available to the combiners shown in FIGS. 9A, 9B, and 9C, which are described in possible detail below. The combiners are examples of hardwired circuits in the unique device of the present invention that perform sub-functions on the products of source channels generated by the large-scale multipliers of FIG. 8A.
[0086] FIG. 8B is a diagram illustrating the structure of a synchronization circuit that provides normal and exception handling signals to all combiners of all output channels.
[0087] Control circuit 803 synchronizes all output and control counters to the source input stream and ensures that the output and control counters are set to their initial state whenever RST or INIT is asserted.
[0088] The colSrc counter 805 in this example counts up the interior dimensions of the array by columns across the rows, and advances as each set of source channel products is processed. At the end of each row, in this example, the colSrc counter returns to the leftmost position (0) and the rowSrc counter 804 advances by 1. At the end of the source array stream, the rowSrc and colSrc counters return to their initial states, ready to receive a new array of inputs.
[0089] In this example, the colDst counter 807 and the rowDst counter 806 act together in a manner similar to these counters for all output channels. The colDst counter and the rowDst counter are enabled by the output enable signal (DSTEN) 813 and determine when the post - processing enable signal (POSTEN) 812 is asserted.
[0090] It should be noted that the system shown in this example generates a single output of the aperture function, but will typically be used to generate a set of stream outputs of channel outputs that match the dimensions of the source input stream. Each independent output channel will share at least some of the computational circuits via a large multiplier and common control logic.
[0091] The output enable (DSTEN) signal 813 controls when the finalization function accepts and processes the results from the synthesizer. The first few rows are accepted from the source input array, but valid results are not given to the finalization function (see Figure 9C). The output enable signal 813 (DSTEN) is asserted either when the rowDst and colDst counters indicate that a valid result is available or, alternatively, when the delayed processing discards the result. The POSTEN signal 812 is asserted continuously or periodically to match the timing of the SRCEN signal 801. These signals are required to sequence all discarded final outputs of the synthesizer when processing the last row of the source input stream array.
Number
[0092] In this example, the POSTEN and DSTEN signals and colDst and rowDst counter values are independent of the SRCEN signal and colSrc and rowSrc counter values, and continue to process delayed results until all delayed results are finalized and sent to the output stream. The system can accept new inputs until the previous output is completed, allowing the system to process multiple frames of the source input stream without pausing between frames. POSTEN is not asserted and final results are captured from the compositor while the source stream data has not yet reached the end of the array. Immediately after the end of the source array is reached, the POSTEN signal is asserted for each additional output and final results are captured from the truncated delay lines 909, 910a, and 910b until the rowDst counter reaches the full number of output rows, as shown in FIG. 9C below, at which point rowDst and colDst are reset to their initial states in preparation for the next frame of data.
[0093] A first row signal 808 (ROWFST) is asserted when the rowSrc counter indicates that the source data set from the stream represents the first row of the array.
[0094] The last row signal 809 (ROWLST) is asserted when the rowSrc counter indicates that the source data set from the stream represents the last row of the array.
[0095] The first column signal 810 (COLFST) is asserted when the colSrc counter indicates that the source data set from the stream represents the first column of each row of the array.
[0096] The last column signal 811 (COLLST) is asserted when the colSrc counter indicates that the source data set from the stream represents the last column of each row of the array.
[0097] Figures 9A, 9B, and 9C illustrate the unique device described above in the general case where MxN subfunction elements of an aperture function are applied to each overlapping MxN patch of an array of RxC inputs, including those that overlap the edges, which inputs are presented as streams of related components at regular or irregular time intervals to generate a corresponding stream of RxC outputs, each output being the collective effect of the MxN function elements applied to the input patches as specified by the aperture function rules. The function elements applied to each location of the array are, in this device, hardwired combiners for each of the MxN subfunctions, as shown in the composite of Figures 9A, 9B, and 9C.
[0098] The effect of this circuit is to calculate a resynthesized value of the aperture function at each location of the RxC input array, using the same sequence of operations that would be used to calculate the aperture function on each patch individually. If any locations are not desired in the output stream, circuitry can be added to filter them out, producing a tiled or spaced output rather than perfectly overlapping.
[0099] The products 802 of the source channels and the source control signals 814 are made available to each of the combiners 901, 902a, 902b, 902c, 903a, 903b, 903c, 904, 905a, 905b, 905c, 906, 907a, 907b, and 907c. The source control signals are also connected to delays 908a, 908b, 908c, 908d, 908e, and 908f. The output channel control and counters 815 are made available to delays 909, 910a, and 910b, as well as to a finalization function 911. Additional pipeline stages may be inserted, either manually or by automated tools, to make the circuit routing appropriate for a given clock frequency, if and only if the order of operations is not changed. The timing control and counter signals are available to all elements of the circuit and are not shown individually.
[0100] Each combiner has a dedicated direct connection to either a specific input product or, alternatively, a programmable multiplexer that selects one of the products for each input value in the set and is preconfigured prior to execution of the circuit. Each dedicated connection is a parallel path with enough wires to carry the bits representing the product required for a single input interval. Optional preconfigured multiplexers that select which product is sent to each combiner for each set of elements allow for in-field upgrade of weight values. Fixed connections are used when weights are not upgraded and remain fixed throughout the life of the device. The choice of fixed or variable product selection does not affect the operation of the circuit, since the weight selection does not change during operation.
[0101] Each combiner receives a set of products corresponding to the sub-function weights from the large scale multiplier, one for each input channel, and performs sub-function calculations, typically simply adding them all together, to form that combiner's contribution to the overall aperture function value. Each combiner also receives partially completed results from the combiner immediately to its left, except for the one corresponding to the left column of the aperture function. Each combiner may also receive delayed partially completed results from the combiner in the row above, except for the one corresponding to the top row of the aperture function. Each combiner has at most one connection from the left and one delayed connection from above, but each of these connections is a parallel path with enough conductors to carry a bit representing a partially completed result as input to that combiner. Depending on the definition of the subfunctions with respect to the position of the current input patch relative to the edge of the input array, each combiner performs one of three operations: combine its partial result with an initial value if present, or combine its partial result with a partial result from the left combiner, or combine its partial result with a delayed partial result. The corrected result is placed in an output register of multiple bits sufficient to contain it and make it available to the right combiner and / or the delay and finalization circuitry in the following input interval. This corrected result can be either a partial, complete, or truncated result, depending on the position of the combiner in the aperture function and the state of the input stream position.
[0102] The combiner (0,0) is unique in that there are no combiners to its left or above it in the aperture function, and therefore always initializes the calculations with each set of inputs received.
[0103] The combiner (M-1,N-1) is unique in that the result it produces is always the final result, but is structurally identical to all other combiners 903a, 903b, or 903c.
[0104] The outputs of some combiners are tapped for delay or post-processing, in which case the width of the path through such delay or post-processing is sufficient to carry the bits representing the partial, truncated, or complete result. The outputs of some combiners are used only by the right-hand combiner. Calculations and output data formats internal to the combiners do not require modification depending on the use of the outputs.
[0105] The finalization circuitry takes results from several possible sources and multiplexes them to select which one to process in any given interval. After applying a finalization function, if present, the width of the final output may be reduced to form the output stream of this embodiment, which may be either the next input stream, the final output of a system including the present invention, or an output that may be used for further processing.
[0106] The data paths in the unique device in the implementation of the present invention are shown in Figures 9A, 9B, and 9C by bold lines with the direction indicated by the arrows, and ellipses indicate where the last column or row in the range is repeated in its entirety. The data path (a) from the products 802 of the source channels is a set of parallel conductive paths, one path dedicated to each product of input components, each product being the input component multiplied by one of the weight values of the aperture function. It should be clear that a 5x5 aperture function has 25 weight values for each input component. In the context of an aperture function for an RxC input array of R, G, and B color pixels, there are 75 weight values. Thus, line (a) in this context has 75 parallel paths, each path being a set of parallel conductors of a width that accommodates the number of bits desired for precision. Line (a) is referred to in the art as a set of point-to-point connections, as opposed to a bus.
[0107] Data paths (b) in Figures 9A, 9B, and 9C are not extensions of lines (a), but are dedicated connections to a particular subset of paths in lines (a). Lines (b) are not necessarily marked in every example in Figures 9A, 9B, and 9C, but every connection from line (a) directly to an individual one of the constituent circuits is a dedicated line (b). Dedication means that each combiner is connected to its subset of paths that carry the products of each input component and the weight value required by that combiner.
[0108] Data paths (c) in Figures 9A, 9B, and 9C are point-to-point paths between the output registers in each combiner and the next combiner to its right. These are dedicated paths of precise width that typically carry partial sums, as explained in possible detail elsewhere in this specification. Not every path (c) is marked in the drawing, but in this example, every direct connection from one combiner to another can be assumed to be a path (c). Note that there are cases where the output path (c) branches off to alternative circuits.
[0109] Another unique data path in one embodiment of the invention is marked (d) in Figures 9A, 9B, and 9C. These are dedicated data paths from delay circuits such as circuits 908A to 908f, either back to the combiner in the row below or to the left, or directly to another delay circuit. The delay circuits are designed to accept partial sums at the right end of the combiner row, delay the passage for the partial sums by a specific number of source intervals, and then pass those partial sums on to other combiners and / or other processing at the appropriate time. The overall functionality is described in possible detail elsewhere in this specification. The paths (d) between the delay circuits are likewise dedicated paths for the partial sums typically passing at specific source intervals.
[0110] If either M or N is reduced such that the last row or column of a range is not required, the last element is dropped and the implementation of the first row or column in the range is retained. In the degenerate case where either M or N or both are reduced to 2, the first and last row or column are retained and the intermediate rows and columns are dropped. In the degenerate case where either M or N is reduced to 1, the implementations of the first and last combiner are combined and no special initialization is required. In the particular example where both M and N are 1, inversion of the aperture function is not required, but the use of a large scale multiplier still provides a clear advantage.
[0111] The source channel product 802 can be any set of binary values given simultaneously in relation to a particular position of the R×C array and in some predefined sequence. The source channels of the input stream can be any combination of integer or fractional values in any format defined by any nature for the input of the aperture function. One example is pixel values from one or more video frames and / or any other sensor values scaled to match the array size R×C, as well as feature component values generated as the output of a CNN layer. It is emphasized that each node embodying the invention can accept outputs from other nodes in addition to or instead of the main source input. In one embodiment of the present invention, there is no restriction on the nature of the data to be processed, as long as it can be formatted into a stream representing the R×C array, although it is common for the first node or nodes to accept image pixels as the main input of the system.
[0112] In some embodiments of the invention, the source stream element sets may be presented in row-first order, with each subsequent column presented in strictly ascending order. In some embodiments of the invention, the rows and columns need not correspond to horizontal or vertical axes, but may be arbitrary in scanning the columns up or down and right to left. Rows R and columns C here simply refer to the major and minor axes of the stream format. The circuitry does not need to be adjusted for input signals that produce input streams in an orientation other than the standard video left-to-right, top-to-bottom order. The orientation of the aperture subfunctions can be made to match to produce identical outputs for each input array position.
[0113] In this example, the source inputs, which are the products of the source values and weights required by the aperture function, are presented by a signal (SRCEN, see FIG. 8B) that indicates when each new set of elements is valid. The inputs can be paused and resumed at any time. In some examples, a minimum spacing between inputs can be defined, and the circuit can use a multi-cycle or higher speed clock to reduce size and power or otherwise gain benefits, and the output channel set can use the same minimum spacing.
[0114] A common control and synchronization circuit 803 (FIG. 8B) provides counters and control signals that describe the current input position in the R×C array. The counters may continue to run for additional rows and columns after the last input, helping the finalization function 911 (FIG. 9C) to output the accumulated output generated beyond the input column by the last row of inputs. (See FIGS. 12, 13, and 14 and the discussion below.) The control signals are available to all other elements and are not shown in FIGS. 9A, 9B, and 9C.
[0115] The combiner circuits 901, 902a, 902b, 902c, 903a, 903b, 903c, 904, 905a, 905b, 905c, 906, 907a, 907b, and 907c each calculate that portion of the aperture function assigned to their position in the M×N function. All combiners operate on the same set of source channels and row and column counter states provided by control 803. Details of the data handling of the aperture functions are described below with reference to additional figures.
[0116] As a set of source inputs is received from the input stream, partially completed computations of the aperture function to be applied to all patches that overlap the current position in the input stream are passed from left to right and top to bottom inside the M×N array of the synthesizer. This operation accumulates the complete computation of the aperture function over time and outputs the correct implementation of the aperture function on each patch of the input array, producing the same result through the same sequence of operations as if the aperture function was implemented by reading the input values directly from the array. The replacement of random access with an array with stream access is a key feature of the present invention, eliminating the need for redundant accesses to random access memory.
[0117] Right column of the synthesizer
number
[0118] When processing the last column C-1 of each input row,
number
[0119] In this example, the combiner 903c in the (M-1,N-1) position always produces a completed accumulation of M×N subfunction elements, but is otherwise indistinguishable from the other combiners in its configuration 903c. As mentioned above, when processing the last column C-1 of each input row, the column
number
[0120] In this example, while processing the last row of the input, R-1,
number
number
[0121] FIG. 15 is a diagram illustrating a specific case of pipelined operation in one embodiment of the present invention implementing a 5×5 convolution node.
[0122] Source channel products 802 and source control signals (not shown here) are made available to each of the combiners 901, 902a, 902b, 903a, 903b, 904, 905a, 905b, 906, 907a, and 907b. The source control signals are also connected to delays 908a, 908b, 908c, and 908d. Output channel controls and counters are made available to delays 909, 910a, as well as to finalization 911. If, and only if, the order of operations is not changed, additional pipeline stages may be inserted, either manually or by automated tools, to make the circuit routing appropriate for a given clock frequency. Timing control and counter signals are available to all elements of the circuit and are not shown individually.
[0123] Given each set of source channel products in turn, each combiner selects the appropriate product to compute the sub-function that corresponds to its position in the aperture function. Each 5x5 patch that intersects with the current position in the input array is corrected to include a computation based on the product for that position. The net effect is that a single source stream of input is transformed into a parallel set of 5x5 streams of partial computations that are passed between combiners, until each time all operations for a patch are completed, this occurs typically in combiner (4,4), and occasionally in other combiners when processing the right or lower edge of the input array.
[0124] Note that only the width of the input array affects the size of the delay elements, because each must delay the partial result for a number of source input intervals corresponding to receiving an input in one column and an input in the same column in the next row.
[0125] Figure 16 illustrates a 4x4 embodiment of the IC of the present invention. It is noted that the kernel can have an odd number of subfunctions or an even number of subfunctions in a row or column. This even version is degenerate in the sense that the element 910* shown in the general case of Figure 9C and shown in Figure 15 for the specific case of a 5x5 aperture function (odd number of rows and columns) does not occur at all, since an additional line of output processing is omitted.
[0126] Odd sizes of kernels are symmetric around the center in both directions, but for even sizes the center is offset. In the implementation of the present invention, the IC calculates the center as
number
[0127] Other than these comments, the operation of the particular IC of FIG. 16 is as described for the other versions described.
[0128] Figure 10A is a diagram illustrating the internal structure and operation of combiners 905a, 905b, and 905c of Figures 9A and 9B or 15 in one embodiment of the present invention. A source input set 1001 of stream values in a channel set, either single or a mix of data types as required by the aperture function, is used to calculate the individual combiner contributions by circuit 1004.
[0129] Circuit 1005 uses the output of 1004 to calculate an initial value of the sub-function. Circuit 1006 uses the output of 1004 and a previously calculated partial value 1002 by the combiner immediately to the left to calculate an ongoing partial value of the sub-function. Circuit 1007 uses the output of 1004 and a previously calculated delayed partial value 1003 from one of 908a, 908b, 908c, 908d, 908e, and 908f in the row of combiners immediately above to calculate an ongoing partial value of the sub-function.
[0130] The operations of circuits 1005, 1006, and 1007 may occur simultaneously (same clock cycle) with the operations of circuit 1004, with their outputs shared, or may be implemented by a series of pipeline stages synchronized by the same clock.
[0131] A multiplexer 1008 selects which variant of the partial result is forwarded as the partial value of the sub-function as the output of the combiner 1009. If COLFST 811 is not asserted then the output of 1006 is selected, otherwise if ROWFST 808 is not asserted then the output of 1007 is selected, otherwise the output of 1005 is selected.
[0132] This conditional processing is a natural consequence of allowing the M×N aperture function to extend beyond the edges of the source input stream, which represents a set of values in an R×C array. A single location at the leftmost or topmost edge becomes the first computable element of the aperture function for some patches that touch or overlap with these edges. Thus, each and every combiner at the first computable location of the overlapping patch is required to be initialized with the base value of the aperture function. Furthermore, each and every combiner at the first computable location of the subsequent rows of the patch must be combined with the previous value of the computed partial value of the same patch from the previous row. In this way, the correct computation of all patches that overlap, touch, and are within the topmost and leftmost edges is guaranteed using a single circuit.
[0133] In Figures 10B to 10G, all elements introduced in Figure 10A and using the same reference numbers are functionally identical to those described with reference to Figure 10A.
[0134] Figure 10B is a diagram illustrating the internal structure and operation of combiners 902a, 902b, and 902c of Figures 9A and 9B or 15 in one embodiment of the present invention. A source input set 1001 of stream values is used by circuit 1004 to calculate the combiner contribution to the aperture function.
[0135] Circuit 1005 uses the output of 1004 to calculate an initial value of the sub-function, and circuit 1006 uses the output of 1004 and the partial value 1002 previously calculated by the combiner immediately to the left to calculate the ongoing partial value of the sub-function.
[0136] Multiplexer 1010 selects which variant of the partial result is forwarded as the partial value of the sub-function as the output of combiner 1009. If COLFST 811 is not asserted then the output of 1006 is selected, otherwise the output of 1005 is selected.
[0137] Figure 10C is a diagram illustrating the internal structure and operation of the combiner 904 of Figure 9A or Figure 15 in one embodiment of the present invention. A source input set 1001 of stream values is used by the circuit 1004 to calculate the contributions of the individual combiners.
[0138] Circuit 1005 uses the output of 1004 to calculate an initial value of the sub-function, and circuit 1007 uses the output of 1004 and a previously calculated and delayed partial value 1003 from one of 908a, 908b, 908c, 908d, 908e, and 908f in the combiner row immediately above to calculate an ongoing partial value of the sub-function.
[0139] Multiplexer 1011 selects which variant of the partial result is forwarded as the partial value of the sub-function as the output of combiner 1009. If ROWFST 808 is not asserted then the output of 1007 is selected, otherwise the output of 1005 is selected.
[0140] Figure 10D is a diagram illustrating the internal structure and operation of the combiner 901 of Figure 9A or Figure 15 in one embodiment of the present invention. A source input set 1001 of stream values is used by a circuit 1004 to calculate the contributions of the individual combiners.
[0141] Circuit 1005 uses the output of 1004 to calculate the initial value of the sub-function, which is forwarded as the partial value of the sub-function as the output of combiner 1009 .
[0142] Cell 901 (FIGS. 9A, 15), if used, is always the first value in any complete or truncated patch and therefore always generates the initialization value for that patch.
[0143] Figure 10E is a diagram illustrating the internal structure and operation of combiners 903a, 903b, and 903c of Figures 9B and 9C or 15 in one embodiment of the present invention. A source input set 1001 of stream values is used by circuit 1004 to calculate the contributions of the individual combiners.
[0144] Circuit 1006 uses the output of circuit 1004 and the partial value 1002 previously calculated by the adjacent combiner to the left to calculate the ongoing partial value of the sub-function, which is forwarded as the partial value of the sub-function as the output of combiner 1009.
[0145] Figure 10F is a diagram illustrating the internal structure and operation of combiners 907a, 907b, and 907c of Figures 9A and 9B or 15 in one embodiment of the present invention. A source input set 1001 of stream values is used to calculate the contributions of each combiner 1004.
[0146] Circuit 1006 calculates the ongoing partial value of the sub-function using the output of circuit 1004 and a previously calculated partial value 1002 from the adjacent combiner to the left. Circuit 1007 calculates the ongoing partial value of the sub-function using the output of 1004 and a previously calculated delayed partial value 1003 from one of 908a, 908b, 908c, 908d, 908e, and 908f in the adjacent row of combiners above.
[0147] Multiplexer 1012 selects which variant of the partial result is forwarded as the partial value of the sub-function as the output of combiner 1009. If COLFST 811 is not asserted then the output of 1006 is selected, otherwise the output of 1007 is selected.
[0148] Figure 10G is a diagram illustrating the internal structure and operation of the combiner 906 of Figure 9A or Figure 15 in one embodiment of the present invention. A source input set 1001 of stream values is used by a circuit 1004 to calculate the contributions of the individual combiners.
[0149] Circuit 1007 uses the output of 1004 and a previously calculated delayed partial value 1003 from one of 908a, 908b, 908c, 908d, 908e, and 908f in the row of combiners immediately above to calculate the ongoing partial value of the sub-function. The output of circuit 1007 is forwarded as the partial value of the sub-function as the output of combiner 1009.
[0150] 11 is a diagram illustrating the internal structure and operation of internal row delay lines 908a, 908b, 908c, 908d, 908e, and 908f (FIG. 9C). The delay lines are used to hold partially computed results from each row of the combiner for use in the next row.
[0151] When COLLST is asserted, the current position of the source input stream is at the rightmost edge and
number
[0152] The current position of the source input stream, colSrc, is
number
[0153] The column position of the source input stream is
number
[0154] The partial output selected by multiplexer 1106 is fed to a first-in-first-out (FIFO) circuit 1107 with CN positions which are arranged such that the source input stream positions are processed such that exactly one value is inserted, and the one values are removed in the same order that they were inserted. Because partially completed results from one position are not required until the source input stream returns to the same patch position in the next row, this introduces a delay such that the partial results computed by one row are given to the next row exactly when they are needed.
[0155] The partial output selected by multiplexer 1106 also provides the same value (1114) into delay lines 909, 910a, and 910b of the final result.
[0156] The partial output taken from FIFO 1107 is routed at 1108 to both the leftmost combiner in the next row (1111) and to a series of parallel access registers 1109 to 1110 which further delay the partial output by one source input stream interval as the data is passed through this register chain.
[0157] When the current position of the source input stream is at the leftmost edge, the FIFO directs the output data at 1108 and the delayed results 1109 to 1110 are made available to the next row of cells at 1111, 1112 to 1113 respectively.
[0158] Note that when the source input array stream position is near the right edge, the additional values from the right side of the source input array stream inserted into FIFO 1107 by multiplexer 1106 are accessed only via path 1111, while when the source input array stream is in the leftmost position to access data inserted from path 1103 as usual, only the additional parallel paths 1112 to 1113 are used. The apparent similarity in structure and requirements between processing the right edge and the left edge is a natural result of the symmetry of the overlap of subfunctions with the right and left edges of the source input stream array. When the value for N is even, the number of additional cells processed to support the right and left edges is not the same.
[0159] FIG. 12 is a diagram illustrating the internal structure and operation of the final truncated result delay line 909 (FIG. 9C).
[0160] When processing the last row of the source input stream array, the partial result from the auxiliary output 1201 of the internal row delay line 908d is considered to be the final result of the last row of the truncated patch and is held in a FIFO 1202 whose number of elements C is equal to the width of the source input stream array.
[0161] Immediately after recording the final result of the truncated patch, the output of FIFO 1202 is transmitted to a further delay line 910a via 1203, or directly to final processing 911 if the value of M is such that no other delay lines are involved.
[0162] FIG. 13 is a diagram illustrating the internal structure and operation of the final truncated resultant delay lines 910a and 910b.
[0163] When processing the last row of the source input stream array, the partial result 1301 from the auxiliary outputs of the internal row delay lines 908e to 908f is considered to be the final result for the last row of the truncated patch and is held in a FIFO 1304 whose number of elements C is equal to the width of the source input stream array.
[0164] When POSTEN is asserted, multiplexer 1303 switches from taking values from 1302 to taking values from the final truncated delay line of the row above, which has the effect of giving the final truncated result in row-original order that matches the ordering of all preceding output results.
[0165] Note that during the cycle of the input frame in which POSTEN is first asserted, the contents of FIFOs 1202 and 1304 are the final values of the truncated patch that overlaps the last row of the source input stream array. Any inhibition of execution in not processing the last row of the source input stream array is optional, since any data contained in FIFOs 1202 and 1304 will not be processed prior to that cycle.
[0166] Immediately after recording the final result of the truncated patch, the output of FIFO 1304 is sent to a further delay line via 1305, or directly to final processing 911 if the value of M is such that no other delay lines are involved.
[0167] FIG. 14 is a diagram illustrating the internal structure and operation of the final processing of all complete and truncated results.
[0168] As in FIG. 11, and using the same structure and functionality, if the current position of the source input stream is at the rightmost edge,
number
[0169] The current position of the source input stream is
number
[0170] While processing the source input stream array, multiplexer 1402 feeds the result selected by multiplexer 1106 directly to finalization (1403). When at the output of the post-processing phase of the truncated result, delay line 1401 is selected instead of finalization (1403).
[0171] The finalization circuit 1403 performs all additional computations, if any, to generate the final form of the output stream (1404) from the combined patch results. This may typically take the form of a Rectified Linear Activation (RELU) function, whereby negative values are set to zero and out-of-bounds values are set to the maximum acceptable value, or any other desired conditioning function, such as sigmoid or tanh. The post-processing function is not required to complete within a single source input stream cycle, but is required to accept each final result at the rate of the source input stream array.
[0172] When DSTEN is asserted, the finalization circuit 1403 provides the final result as one value in the destination output stream. Any partial or inaccurate values produced by the finalization circuit 1403 are ignored whenever DSTEN is not asserted, so any suppression of operations when the result is not used is optional.
[0173] In one implementation, the destination output stream array is processed by circuitry similar to that described above. In that case, it is advantageous for the timing of the final truncated result to be identical to all previous final results. To that end, the control of FIFOs 1202 and 1304 is adjusted by control circuit 702 to maintain an output rate identical to the primary output rate.
[0174] In other implementations, the destination output stream array is the final stage of the system and no further processing is required. In that case, it is advantageous to have the timing of the final truncated result completed as quickly as possible. To that end, the control of FIFOs 1202 and 1304 is adjusted by control circuit 702 to output their results at the maximum frequency supported.
[0175] Note that the implementation described above generates a single output element from a full set of input elements: in a complete system generating a larger set of output elements from a set of inputs, the entire mechanism described would be replicated once for all output channels, with the notable exception of the control circuitry 702, which may be shared by the output channels, since the timing of all the individual sub-functions is identical for the entire output set.
[0176] The inventors have constructed a working prototype of an IC in accordance with one embodiment of the present invention to test and verify the details and features of the present invention, the operation of which confirms the above description, and the inventors have developed a software supported simulator which has been used up until the time of filing this application to test and verify the above details and descriptions.
[0177] In another aspect of the invention, a system is provided for accepting an input stream of three-dimensional data as typically provided in medical imaging, where additional circuitry and buffering is included to allow a three-dimensional aperture function to be passed over the three-dimensional input array with corresponding calculations that correctly implement both the interior and edge cases for the first and last planes.
[0178] In yet another aspect of the present invention, for the complex process of training a deep neural network (DNN), a hardware-assisted neural network training system is provided in which the training algorithm simply needs to use statistics gathered from the forward inference, with the bulk of the effort put in by the forward inference engine to periodically adjust the weights and biases for the entire network to converge the model to the desired state. With the addition of appropriate accumulators that sum the input states as the forward inference process is computed, the present invention forms a hardware-assisted neural network training system.
[0179] In yet another aspect of the present invention, regarding the well-known problem that floating-point precision limitations impede the convergence of DNN models (known in the art as the "vanishing gradient problem"), a single large multiplier with limited bit-width precision is provided, which can be cascaded with additional adders to generate floating-point products of arbitrarily large precision. While this innovation is not typically required for forward inference calculations, it is crucial in DNN trainers to avoid problems that arise when the computed gradient becomes too small to measure.
[0180] Those skilled in the art will appreciate that all of the embodiments illustrated in the drawings and described above are exemplary and do not detail all possible forms the invention may take, and there may be various other forms which may be realized within the scope of the invention.
[0181] The present invention is limited only by the claims.
Claims
1. an integrated circuit (IC) for implementing an M×N aperture function on an R×C source array to generate an R×C destination array, an input port that receives an ordered stream of independent input values from a source array; an output port that produces an ordered output stream of output values into the destination array; a large scale multiplier circuit coupled to the input port, the large scale multiplier circuit multiplying each input value in turn in parallel by all weights required by the aperture function to generate a stream of products on a set of parallel conductive product paths on the IC, each product path being dedicated to a single multiplication by an input weight value; an M×N array of combiner circuits on the IC, each combiner circuit associated with a subfunction of the aperture function at an (m,n) location and coupled by a dedicated path to each of a set of product paths carrying products generated from weight values associated with the subfunctions; a single dedicated path between the combiners; a delay circuit on the IC that receives a value on a dedicated path from the combiner and provides a delayed value on a dedicated path to another downstream combiner at a later time; A finalization circuit; a control circuit for operating the counter and generating control signals coupled to the synthesizer, the delay circuit, and the finalization circuit; 1. An integrated circuit (IC) comprising: at each source interval, a combiner combines values received from dedicated connections into parallel conductive paths, and further combines the result with an initial value for that combiner, or a value on a dedicated path from an adjacent upstream combiner, or a value received from a delay circuit, and posts the combined result to a register coupled to a dedicated path to an adjacent downstream combiner, or a delay circuit, or both; and when the last downstream combiner has produced a complete combination of values for the output of an aperture function at a particular location of an input R×C array, the combined value is passed to a finalization circuit, which processes the value and posts the result to an output port as one value in an output stream.
2. 2. The integrated circuit (IC) of claim 1, wherein the aperture function is for a convolutional neural node, and at each source interval, the combiner sums the products of the weights with the inputs, adds that sum of products to an initial bias, or a value on a dedicated path from an adjacent upstream combiner, or a value received from a delay circuit, and posts the sum to an output register.
3. 2. The integrated circuit (IC) of claim 1, wherein the aperture function produces truncated results for aperture positions where the M×N input patch overlaps the left and right edges of the R×C input array, and for certain source intervals where the source input position represents the first or last column of the R×C input array, the truncated patch results are delayed and accessed by the combiner and merged with the flow of the complete internal patch.
4. 2. The integrated circuit (IC) of claim 1, wherein the aperture function produces truncated results for those particular locations where the M×N input patch overlaps the top edge of the R×C input array, and for particular source intervals where the source input location represents the first row of the R×C input array, the truncated patch results are delayed and accessed by the combiner and integrated with the flow of the complete internal patch.
5. 2. The integrated circuit (IC) of claim 1, wherein the aperture function produces truncated results for those particular locations where the M×N input patch overlaps the bottom edge of the R×C input array, and for particular source intervals where the source input location represents the first row of the R×C input array, the truncated patch results are delayed and merged with the flow of the complete interior patch.
6. The integrated circuit (IC) of claim 1 , wherein certain outputs of the aperture function are excluded from the output stream in a fixed or variable stepping pattern.
7. 1. A method for implementing an M×N aperture function on an R×C source array to generate an R×C destination array, comprising the steps of: providing an ordered stream of independent input values from a source array to an input port of an integrated circuit (IC); multiplying each input value in turn in parallel by all weight values required by the aperture function by a large scale multiplier circuit on an IC coupled to the input port; generating, by a large scale multiplier, a stream of products on a set of parallel conductive product paths on the IC, each product path being dedicated to a single product by an input weight value; providing to each of an M×N array of combiner circuits on the IC, each associated with a sub-function of the aperture function, those products generated from weight values associated with the sub-functions by a dedicated connection from the product stream to each combiner circuit; providing control signals to the combiner, the plurality of delay circuits, and the finalization circuit by a control circuit implementing a counter and generating control signals; combining, by a combiner, at each source cycle, a value received in the stream of products from a dedicated connection with an initial value for that combiner, or with a value on a dedicated path to an adjacent upstream combiner, or with a value received from one of a plurality of delay circuits, and posting the result to a register coupled to a dedicated path to an adjacent downstream combiner, or to one of a plurality of delay circuits; when the last downstream combiner has generated a complete combination of values for the output of the aperture function at a particular location in the input R×C array, providing that complete combination to a finalization circuit; processing the complete combination through a finalization circuit and posting the result to an output port as a single value in an ordered output stream; Continue operation of the IC until all input elements have been received and the final output value has been generated in the output stream. A method comprising:
8. 8. The method of claim 7, wherein the aperture function is for a convolutional neural node, and at each source interval, the combiner sums the products of the weights with the inputs, adds the sum of the products to an initial bias, or a value on a dedicated path from an adjacent upstream combiner, or a value received from a delay circuit, and posts the sum to an output register.
9. 8. The method of claim 7, wherein the aperture function produces truncated results for aperture positions where the M×N input patch overlaps the left and right edges of the R×C input array, and for certain source intervals where the source input position represents the first or last column of the R×C input array, the truncated patch results are delayed and accessed by the combiner and merged with the flow of the complete interior patch.
10. 8. The method of claim 7, wherein the aperture function produces a truncated result for a particular location where the M×N input patch overlaps the top edge of the R×C input array, and for a particular source interval where the source input location represents the first row of the R×C input array, the truncated patch result is delayed and accessed by the combiner and integrated with the flow of the complete interior patch.
11. 8. The method of claim 7, wherein the aperture function produces truncated results for those particular locations where the M×N input patch overlaps the bottom edge of the R×C input array, and for particular source intervals where the source input location represents the first row of the R×C input array, the truncated patch results are delayed and merged with the flow of the complete interior patch.
12. The method of claim 7 , wherein certain outputs of the aperture function are excluded from the output stream in a fixed or variable stepping pattern.
Citation Information
Patent Citations
Image processor, convolutional integration circuit and method therefor
JP2001222712A