Multiplier-accumulator circuit system and method of operating the circuit system
Through Winograd-type data processing technology, the input data and filter weights are transformed from M×M matrices to N×N matrices and processed within the execution pipeline, which solves the problem of insufficient data throughput of the multiplier-accumulator circuit system and achieves more efficient data processing.
Patent Information
- Application Number
- CN202080010361.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-20
- Filing Date
- 2020-02-29
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2040-02-29
AI Technical Summary
Existing multiplier-accumulator circuit systems have a problem of insufficient data throughput when processing data, especially when transforming input data from an M×M matrix to an N×N matrix, which cannot effectively improve processing efficiency.
Winograd-type data processing technology is used to transform the input data and filter weights from M×M matrices to N×N matrices, and process them through a multiplier-accumulator circuit system within the execution pipeline. Combined with a ZY conversion logic circuit system to generate the final output, efficient data processing is achieved in a single execution pipeline.
Through Winograd-type data processing technology, the data throughput of the multiplier-accumulator circuit system is improved, faster data processing is achieved, and the dependence on multiple execution pipelines is reduced.
Smart Images

Figure CN113396387B_ABST
Abstract
Description
[0001] Related applications
[0002] This nonprovisional application claims priority to and the benefit of U.S. Provisional Application No. 62 / 823,161, filed on March 25, 2019, entitled “Multiplier-Accumulator Circuit System with Processing Pipeline and Method of Operating the Circuit System.” The entire contents of the '161 provisional application are hereby incorporated herein by reference.
[0003] Introduction
[0004] Many inventions are described and illustrated herein. The present invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or permutation of such aspects and / or embodiments. Importantly, each aspect and / or embodiment thereof may be used alone or in combination with one or more of the other aspects and / or embodiments thereof.
[0005] In one aspect, the present invention relates to an integrated circuit having a multiplier-accumulator circuit system (and a method of operating the circuit system), the multiplier-accumulator circuit system including one or more execution or processing pipelines, the one or more execution or processing pipelines including a circuit system for implementing Winograd-type processing to improve the data throughput of the multiplier-accumulator circuit system and processing. In one embodiment, the circuit system and technology transform input data from an M×M matrix to an N×N matrix (where N and M are positive integers and N is greater than M (e.g., M=3 and N=4)), where the input data can be stored in a memory (e.g., a layer of image pixels including a two-dimensional array). In one embodiment, the circuit system and technology also transform input weights or weight values from an M×M matrix to an N×N matrix or block, where the input weights or weight values can also be stored in a memory in M×M blocks (e.g., a layer of input weights or values including a two-dimensional array). Here, the filter weights or coefficients of each M×M matrix or block are associated with an M×M matrix of input data. After the above conversion, the multiplier-accumulator circuitry processes the NxN input data using the associated NxN filter weights or coefficients.
[0006] In one embodiment, the multiplier-accumulator circuitry processes N×N input data using associated N×N input weights to generate or accumulate output data in a Q×Q matrix. After further processing (e.g., addition and / or subtraction operations), the multiplier-accumulator circuitry generates an output value. That is, the aggregation of N×N element values by the multiplier-accumulator circuitry (which, in one embodiment, includes an N×N execution pipeline) provides or generates Q×Q output data / pixels. In this embodiment, after further transformation / conversion (via ZY conversion logic circuitry), circuitry external to the N×N execution pipeline generates the final Q×Q output. Here, as the N×N product elements / values are accumulated with other N×N product elements / values from other input layers, the individual elements / values are accumulated together into a final Q×Q output pixel until after the ZY conversion operation has been performed. In this embodiment, the ZY conversion logic circuit system outside the associated N×N execution pipeline receives data, transforms the data to generate and output (one or more) output values (P×P matrix, such as 1×1 values), which are associated with the results of the multiplication and accumulation processing of the M×M input data by the multiplier-accumulator circuit system.
[0007] As discussed in more detail below, in another embodiment, the ZY conversion logic circuitry and the operations it implements are incorporated into an associated execution pipeline. In this embodiment, the multiplier-accumulator circuitry can accumulate individual elements / values of N×N execution pipelines within the execution pipeline, so that data processing is implemented by a single execution pipeline rather than multiple execution pipelines (e.g., N×N execution pipelines (e.g., 16 execution pipelines)).
[0008] In particular, the present invention may include a plurality of separate multiplier-accumulator circuits and a plurality of registers (including a plurality of shadow registers) that facilitate pipelining of multiplication and accumulation operations (see, for example, U.S. Patent Application No. 16 / 545,345 and U.S. Provisional Patent Application No. 62 / 725,306, filed on August 31, 2018, and August 20, 2019, respectively, entitled "Multiplier-Accumulator Circuits, Multiply-Accumulate Logic Block Architecture, and IC Including Logic Block Arrays"). The present invention may be implemented in conjunction with the inventions and / or embodiments of the '306 and '345 applications, the entire contents of which are incorporated herein by reference. In particular, the multiplier-accumulator circuit systems described and / or illustrated in the '306 and '345 applications facilitate cascading and reconfiguring multiplication and accumulation operations, thereby enabling multiple multiplier-accumulator circuits to perform operations more rapidly.
[0009] As described above, in one embodiment, the circuitry and techniques of the present invention read an M×M block of input weights from a memory and then transform or convert such M×M block of input weights into an N×N block associated with an N×N block of input data. In this embodiment, the input data and input weights are read from the memory and transformed or converted by the multiplier-accumulator circuitry during operation of the multiplier-accumulator circuitry / pipeline (i.e., during operation or operation of the multiplier-accumulator circuitry / pipeline).
[0010] In another embodiment, the input weights are pre-transformed and stored in memory as N×N blocks. In this alternative embodiment, the transformed or converted filter weights are stored in memory in N×N block form and then read from memory in N×N block form by the multiplier-accumulator circuit system. During operation and execution of the multiplier-accumulator circuit system / pipeline, the multiplier-accumulator circuit system uses the pre-transformed / pre-converted weights on the associated input data (which is converted from M×M block input data to N×N block input data during operation by the circuit system and techniques of the present invention). Such input weight transformation / conversion can be performed by an off-chip computing system and then stored in memory. However, during operation, the multiplier-accumulator circuit system / pipeline (i.e., during operation) still uses the input weights of the N×N block transformed by the circuit system and techniques of the present invention and the input data of the associated N×N block to accumulate N×N product data / elements.
[0011] In particular, the integrated circuit may be, for example, a processor, a controller, a state machine, a gate array, a system on a chip (SOC), a programmable gate array (PGA), and / or an FPGA. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The present invention may be implemented in conjunction with the embodiments shown in the accompanying drawings. These drawings illustrate different aspects of the invention, and where appropriate, reference numerals, nomenclature, or names representing the same circuits, architectures, structures, components, materials, and / or elements in different figures are similarly labeled. It should be understood that various combinations of structures, components, materials, and / or elements other than those specifically shown are contemplated and fall within the scope of the present invention.
[0013] Furthermore, many inventions are described and illustrated herein. The present invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or permutation of such aspects and / or embodiments. Additionally, various aspects of the present invention and / or embodiments thereof may be employed alone or in combination with one or more other aspects and / or embodiments thereof. For simplicity, specific permutations and combinations are not discussed and / or illustrated separately herein. In particular, embodiments or implementations described herein as “examples” should not be construed as being preferred or advantageous, for example, over other embodiments or implementations; rather, they are intended to reflect or indicate that the embodiment(s) are “example” embodiments.
[0014] In particular, the configurations, block / data widths, data path widths, bandwidths, data lengths, values, procedures, pseudo-code, operations, and / or algorithms described herein and / or shown in the accompanying drawings, and the text associated therewith, are exemplary. Indeed, the present invention is not limited to any specific or example circuit, logic, block, functional, and / or physical diagrams, block / data widths, data path widths, bandwidths, values, procedures, pseudo-code, operations, and / or algorithms shown and / or described according to, for example, the exemplary circuit systems, logic, block, functional, and / or physical diagrams.
[0015] Figure 1A is a schematic block diagram of a logical overview of a first operating mode of multiplier-accumulator execution pipelines, wherein each multiplier-accumulator execution pipeline includes a multiplier-accumulator circuit system shown in block diagram form; in particular, the multiplier-accumulator circuit system includes one or more multiplier-accumulator circuits (although individual multiplier-accumulator circuits are not specifically shown herein).
[0016] Figure 1B is a schematic block diagram of a physical overview of an exemplary embodiment of a multiplier-accumulator execution pipeline according to certain aspects of the present invention; in particular, each multiplier-accumulator execution pipeline includes associated / individual multiplier-accumulator circuitry; furthermore, in one exemplary implementation of this embodiment, all of the 64×(3×3) input pixels at dij that determine the 64×(1×1) output pixels at yij are processed by a single execution pipeline; here, the 3×3 set / array of input pixels / data is image data associated with the output data; indeed, each of the plurality of execution pipelines processes all of the dij input pixels / data of a separate set (in this exemplary embodiment, 64×(3×3)) that determines the yij output pixels / data (in this exemplary embodiment, 64×(1×1)) associated with all of the dij input pixels / data of the set processed by the associated multiplier-accumulator execution pipeline in the plurality of execution pipelines;
[0017] Figure 2Ais a schematic block diagram of a logical overview of an exemplary embodiment of multiplier-accumulator circuitry in a plurality of multiplier-accumulator execution pipelines implementing Winograd data processing techniques according to certain aspects of the present invention; as described above, the multiplier-accumulator circuitry includes one or more multiplier-accumulator circuits (although individual multiplier-accumulator circuits are not specifically shown herein);
[0018] Figure 2B is a schematic block diagram of a physical overview of an exemplary embodiment of a plurality of multiplier-accumulator execution pipelines according to certain aspects of the present invention, each pipeline including multiplier-accumulator circuitry (shown in block diagram form), wherein the plurality of multiplier-accumulator execution pipelines are configured to implement a Winograd data processing technique; in particular, in this example, 64×(4×4) input pixels / data at dij that determine the associated 64×(2×2) output pixels at yij are processed by a plurality (here, 16) of multiplier-accumulator execution pipelines (compared to the operating mode described and illustrated in FIG. 1 , in which one multiplier-accumulator execution pipeline processes the 64×(3×3) input pixels at dij that determine the associated 64×(1×1) output pixels at yij);
[0019] Figure 2C According to a particular aspect of the present invention Figure 2B an exemplary timing diagram illustrating a physical overview of the exemplary embodiment;
[0020] Figures 2D-2F A conversion table illustrating certain operations for implementing Winograd data processing techniques according to certain aspects of the present invention includes converting filter weights or coefficients via FH conversion circuitry to facilitate Winograd processing (FH processing, Figure 2D ), converting data (such as image data) into Winograd format via the DE conversion circuit system (DE processing, Figure 2E ), and converting the processed image data into a non-Winograd format, such as a floating point format, via a zy conversion circuit system (ZY processing; Figure 2F ), which may facilitate additional processing of the data (e.g., image data);
[0021] Figure 3A According to a specific aspect of the present invention, (such as Figure 2A and 2B Schematic block diagram of physical details of an exemplary dij-eij conversion and extraction circuitry / operation embodiment of a multiplier-accumulator execution pipeline implementing Winograd processing techniques (logical and physical overviews of FIG and FIG, respectively);
[0022] Figure 3B According to a particular aspect of the present invention Figure 3A Example pseudo code for the example dij-eij conversion and extraction embodiment shown;
[0023] Figure 3C According to a particular aspect of the present invention Figure 3A a schematic diagram of an exemplary unit of an exemplary dij-eij conversion circuit system of a multiplier-accumulator execution pipeline;
[0024] Figure 3D According to a particular aspect of the present invention Figure 3A a schematic diagram of an exemplary unit of an exemplary dij-eij extraction circuit system of a multiplier-accumulator execution pipeline of FIG.
[0025] Figure 3E A rough diagram showing the conversion / extraction pipeline (horizontal axis) and the timing of the various units (vertical axis) according to certain aspects of the present invention;
[0026] Figure 4A According to a specific aspect of the present invention, (such as Figure 2A and 2B a schematic block diagram of the physical details of an exemplary fkl-hkl conversion and extraction circuitry / operation embodiment of a plurality of multiplier-accumulator execution pipelines implementing Winograd processing techniques (as shown in the logical and physical outlines of FIG and FIG);
[0027] Figure 4B According to a particular aspect of the present invention Figure 4A Example pseudo code for an example fij-hij conversion and extraction embodiment of a multiplier-accumulator execution pipeline;
[0028] Figure 4C According to a particular aspect of the present invention Figure 4A Schematic diagram of two exemplary units of an exemplary fkl-hkl conversion circuit system of an execution pipeline of FIG; in particular, the fkl-hkl conversion logic circuit system in this exemplary embodiment includes 16 left units and 16 right units, wherein the fkl-hkl conversion logic circuit system includes (i) data registers for fkl and hkl weight values; (ii) control logic for sorting, and (iii) for (according to Figure 2D Adder logic for conversion of the conversion table shown;
[0029] Figure 4D According to a particular aspect of the present invention Figure 4C Schematic block diagram of an exemplary embodiment of a multiplexer (Mux) circuit system and an adder circuit system of an exemplary fkl-hkl conversion circuit system;
[0030] Figure 4E A rough diagram showing the conversion / extraction pipeline (horizontal axis) and the timing of the various units (vertical axis) according to certain aspects of the present invention; in this exemplary embodiment, the horizontal axis shows 32 fkl-hkl conversion units in "X" positions {0, 1, ... 31};
[0031] Figure 5A According to a specific aspect of the present invention, (such as Figure 2A and 2B a schematic block diagram of a logical overview of an exemplary zij-yij insertion and conversion circuit system / operation embodiment of a multiple multiplier-accumulator execution pipeline implementing a Winograd processing technique (shown as a logical and physical overview of FIG and FIG);
[0032] Figure 5B According to a particular aspect of the present invention Figure 5A Example pseudo code for an example zij-yij insertion and conversion embodiment of an execution pipeline of
[0033] Figure 5C is a schematic diagram of an exemplary embodiment of a Zij insertion circuit system and a Zij-Yij conversion circuit system according to certain aspects of the present invention, which includes Figure 5A One unit of the zij insertion logic circuitry (left portion of the figure) and one unit of the zij-yij conversion logic circuitry (right portion of the figure) of the exemplary zij-yij insertion and conversion circuitry of FIG.
[0034] Figure 5D a rough diagram showing a conversion / extraction pipeline (horizontal axis) and the timing of the various units (vertical axis) according to certain aspects of the present invention; wherein the horizontal axis shows 16 zij insertion units in "X" positions {0, 1, ... 15}, 16 zij-yij conversion units in "X" positions {16, 17, ... 31}, and the lower axis shows the zij elements / values of the 4×4 Z blocks in registers ZREG_X and ZREG_Y at their fixed positions in the pipeline;
[0035] Figure 5Eis a schematic block diagram of a logical overview of another exemplary embodiment of multiplier-accumulator circuitry in block diagram form implementing a Winograd data processing technique according to certain aspects of the present invention, wherein accumulation of a first set of 64 input planes (also referred to as layers) is on the left, with the accumulated values stored in the Y region of L2 memory, and accumulation of a second set of 64 input values is on the right side of the diagram; in particular, in this exemplary embodiment, 4×4 Z block values are applied to (or converted by) zij-yij conversion logic circuitry before being written to the Y region of L2 memory, wherein the Y accumulated values from the first 64 input planes are read from L2 and loaded into the accumulation input port of the zij-yij conversion logic as the values are applied to and converted by the conversion logic;
[0036] Figure 6A is a schematic block diagram of a physical overview of another exemplary embodiment of a multiplier-accumulator circuit system for implementing an execution pipeline of a Winograd data processing technique according to certain aspects of the present invention, wherein the architecture combines DE conversion circuit system and ZY conversion circuit system in the multiplier-accumulator execution pipeline and performs their operations; in addition, pre-processed and pre-transformed input weights are read from a memory in N×N blocks by the multiplier-accumulator circuit system;
[0037] Figure 6B is a schematic diagram of four slices of a multiplier-accumulator execution pipeline (one of a plurality of pipelines) of an exemplary embodiment of a multiplier-accumulator circuit system for an execution pipeline implementing a Winograd data processing technique according to certain aspects of the present invention; details of the four slices of the pipeline stage are shown, wherein each of the four slices processes one of an input stream of 4×4 input data blocks D (received from the right side of the figure);
[0038] Figure 6C is an exemplary embodiment of a multiplier-accumulator circuit system for an execution pipeline implementing a Winograd data processing technique according to certain aspects of the present invention (e.g., Figure 6B ) a schematic diagram of a four-stage pipeline with four slices;
[0039] Figures 7A-7C A conversion table illustrating certain operations for implementing Winograd data processing techniques according to certain aspects of the present invention includes converting filter weights or coefficients via FH conversion circuitry to facilitate Winograd processing (FH processing, Figure 7A ), converting data (such as image data) into Winograd format via the DE conversion circuit system (DE processing, Figure 7B), and converting the processed image data into a non-Winograd format, such as a floating point format, via a zy conversion circuit system (ZY processing; Figure 7C ), which may facilitate additional processing of the data (e.g., image data);
[0040] Figure 8 and Figure 9 illustrating an exemplary downsampling mode of operation according to certain aspects of the present invention using any exemplary embodiment of a multiplier-accumulator circuit system for an execution pipeline implementing a Winograd data processing technique according to aspects of the present invention; and
[0041] Figure 10A and 10B The present invention is illustrated in schematic block diagram form as a mode selection circuit system according to certain aspects of the present invention, which, for example, in conjunction with inference operations, controls (i.e., enables and / or disables) or determines the operability of certain circuit systems (e.g., conversion circuit systems), data paths, and / or techniques for processing input data (e.g., image data) (e.g., a first operating mode and / or a second operating mode - respectively, refer to FIG. Figure 1A / 1B and Figure 2A / 2B); in this regard, the mode selection circuitry controls or determines the operation of the conversion circuitry, the multiplier-accumulator circuitry of one, more, or all of the multiplier-accumulator circuitry in the execution pipeline, and, in certain embodiments, controls the functionality / operability of the memory (e.g., reading and / or writing data in a 2D array format—see, e.g., Figure 10A ).
[0042] In particular, the pseudo-code, operations, configurations, block / data widths, datapath widths, bandwidths, data lengths, values, processes, and / or algorithms described and / or illustrated in the figures are exemplary. Indeed, the present invention is not limited to any particular pseudo-code, operations, block / data widths, datapath widths, bandwidths, values, processes, and / or algorithms illustrated and / or implemented according to, for example, the exemplary logical or physical schematic configurations and / or exemplary conversion logic circuitry.
[0043] Once again, many inventions are described and illustrated herein. The present invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or permutation of such aspects and / or embodiments. Each aspect and / or embodiment of the present invention may be used alone or in combination with one or more of the other aspects and / or embodiments of the present invention. For simplicity, many of these combinations and permutations are not discussed or illustrated separately herein. Specific embodiments
[0044] In a first aspect, the present invention relates to a multiplier-accumulator circuit system and a technique for operating such a circuit system, including a circuit system for implementing Winograd-type data processing to improve the data throughput of the multiplier-accumulator circuit system and the processing (and a method for performing Winograd-type data processing to improve the data throughput of the multiplier-accumulator circuit system and the processing). In one embodiment, the circuit system and the technique transform input data (such as image data) from an M×M matrix to an N×N matrix (where N and M are positive integers and N is greater than M (such as M=3 and N=4)), where the input data can be stored in a memory (e.g., a layer of image pixels including a two-dimensional array). In one embodiment, the circuit system and the technique also transform input filter weights, values, or coefficients from an M×M matrix to an N×N matrix or block, where the input filter weights, values, or coefficients can also be stored in a memory in M×M blocks (e.g., a layer of input filter weights or values including a two-dimensional array). Here, the filter weights or coefficients of each M×M matrix or block are associated with an M×M matrix of input data. After the above conversion, the multiplier-accumulator circuitry processes the NxN input data using the associated NxN filter weights or coefficients.
[0045] In one embodiment, the multiplier-accumulator circuitry processes N×N input data using associated N×N weights or coefficients to generate or accumulate output data of a Q×Q matrix. After further processing (e.g., addition and / or subtraction operations), the multiplier-accumulator circuitry generates an output value. That is, the aggregation of N×N element values by the multiplier-accumulator circuitry (which, in one embodiment, includes an N×N execution pipeline) provides or generates output data / pixels of a Q×Q matrix. In this embodiment, after further transformation / conversion (via ZY conversion logic circuitry) to convert the output data from Winograd format to a non-Winograd format (e.g., floating point format), circuitry external to the N×N execution pipeline generates a final Q×Q output, which facilitates or allows the value to be accumulated to, for example, an output value associated with the processing of the multiplier-accumulator circuitry of the M×M input data. Here, when the N×N product elements / values are accumulated with other N×N product elements / values from other input layers, after the ZY conversion operation is performed, the respective elements / values are accumulated together into the final Q×Q output pixel. In this embodiment, the ZY conversion logic circuitry outside the execution pipeline receives data, transforms the data to generate and output (one or more) output values (a P×P matrix, such as 1×1 values), which are associated with the results of the multiplication and accumulation processing of the M×M input data by the multiplier-accumulator circuitry.
[0046] In another embodiment, the ZY conversion logic circuitry and the operations it implements are incorporated into an execution pipeline. In this embodiment, the multiplier-accumulator circuitry can accumulate individual elements / values of an N×N execution pipeline within the execution pipeline so that processing can be accomplished by a single execution pipeline of the multiplier-accumulator execution pipeline rather than multiple execution pipelines (e.g., N×N execution pipelines (e.g., 16 execution pipelines)).
[0047] As described above, in one embodiment, the present invention may include multiple separate multiplier-accumulator circuits and multiple registers (including multiple shadow registers) that facilitate pipelining of multiplication and accumulation operations (see, for example, the aforementioned '306 and '345 applications). The present invention may be implemented in conjunction with the inventions and / or embodiments of the '306 and '345 applications to facilitate cascading multiplication and accumulation operations and reconfiguring the operations so that multiple multiplier-accumulator circuit systems can perform the operations more quickly.
[0048] In one embodiment, the circuitry and techniques of the present invention read an M×M block of filter weights or coefficients from a memory and then transform or convert such M×M blocks of filter weights / coefficients into an N×N block, where each N×N block of filter weights / coefficients is associated with at least one N×N block of input data. In this embodiment, the input data and weights are read from the memory and transformed or converted into Winograd format by the multiplier-accumulator circuitry during operation of the multiplier-accumulator circuitry / pipeline (i.e., during operation of the circuitry that performs the pipeline (in situ) or on the fly). In this manner, the filter weights or coefficients are first converted into Winograd format and then provided to the multiplier-accumulator circuitry for processing.
[0049] In another embodiment, the filter weights or coefficients are pre-transformed or converted into Winograd format and stored in memory as N×N blocks. In this way, the filter weights or coefficients are immediately suitable for processing using Winograd techniques. Therefore, in this alternative embodiment, the transformed input weights are stored in memory in N×N blocks and then read from memory in N×N blocks by the multiplier-accumulator circuitry. The multiplier-accumulator circuitry employs the pre-transformed weights on the associated input data (which is transformed from M×M blocks of input data to N×N blocks of input data by the circuitry and techniques of the present invention during operation or on-the-fly) during operation and execution of the multiplier-accumulator circuitry / pipeline.
[0050] In particular, the transformation of the filter weights or coefficients may be performed by an off-chip computing system and then stored in memory. During operation, the multiplier-accumulator circuitry / pipeline (i.e., while running) accumulates N×N product data / elements using the weights of the N×N blocks transformed by the circuitry and techniques of the present invention and the input data of the associated N×N blocks.
[0051] References Figure 1A and Figure 1B As shown in the logical and physical overview, in one embodiment, input data (e.g., image data / pixels) is stored in memory (e.g., organized in planes or layers), the input data comprising a two-dimensional array of input or image data / pixels (e.g., M×M, where M=3). Each two-dimensional array of input or image data / pixel (e.g., a 3×3 array or set of data) is associated / associated with or contributes to an output data value. In one embodiment, the image data / pixels are organized and / or stored in memory in "depth" planes or layers (e.g., K depth planes, where, in one embodiment, K=64; where each plane comprises a plurality of pixels (e.g., 9 pixels per plane, in one embodiment)), and the output data is stored in memory after processing, and in one embodiment, is organized in output "depth" planes or layers (e.g., L output depth planes / layers, where, in one embodiment, L=64). The memory may also store input weights or coefficients associated with the input data. In one embodiment, the input weights or coefficients are stored in memory in M×M blocks or arrays, where the K×L blocks or arrays cover combinations (e.g., all combinations) of input and output layers.
[0052] refer to Figure 1BIn a first operating mode of the multiplier-accumulator circuit system, a single multiplier-accumulator execution pipeline of the execution pipeline is employed to accumulate a 1×1 pixel output value in a single output layer by aggregating the sum of K×M×M multiplications of input data values and associated input weight values from K layers. In short, in this operating mode, 3×3 (M×M) multiplications and accumulations are performed by the multiplier-accumulator circuit system of the multiplier-accumulator execution pipeline to generate a Vijkl value (see FIG. 1 - "∑M" label). For example, all image data related / associated with or contributing to the single output data value is applied to or employed by the multiplier-accumulator execution pipeline, which generates the Vijkl value after processing. Thereafter, in one embodiment, the execution pipeline further performs accumulation of Vijkl values for each input plane (see index K) to produce Yijl values (see "∑K" label). The result of these accumulation operations performed by such processing is a single pixel value Yijl (1×1), which is written to the output plane (e.g., in parallel or simultaneously with other single pixels being written to other output planes along with other output depth values (e.g., "L" index value)). As described above, this processing performed by one multiplier-accumulator execution pipeline can continue for each pixel of the plane. In addition, each execution pipeline of the plurality of execution pipelines (see, e.g., Figure 1B , which shows one execution pipeline among multiple) processes a separate set of all dij input pixels / data (in this exemplary embodiment, 64×(3×3)), which all dij input pixels / data of the separate set determine the associated yij output pixels / data (in this exemplary embodiment, 64×(1×1)).
[0053] In particular, there are multiple planes that make up a layer (which may include non-visual image data and information such as identification of objects in the layer), and multiple layers that make up a frame.
[0054] refer to Figure 2A 、 2B and 2C. In another embodiment, in a second operating mode of the multiplier-accumulator circuitry, an N×N execution pipeline is employed to generate an output layer (which includes multiple planes and, in one embodiment, includes additional information such as recognition-related information) wherein a two-dimensional array of input or image data / pixels is transformed from an M×M array (e.g., M=3) to an N×N array (e.g., N=4). Here, the DE conversion circuitry implements logical operations for converting the M×M array of input data (e.g., image data / pixels) to generate an N×N array of input or image data / pixels (see Figures 3A-3D ).
[0055] Similarly, a two-dimensional array of input / filter weights is transformed or converted from an M×M array (e.g., M=3) to an N×N array (e.g., N=4). In one embodiment, FH conversion circuitry (e.g., in a pipelined architecture) is employed to convert the M×M array of filter weights or coefficients to generate an N×N array of filter weights or coefficients that are appropriately correlated / associated with the associated positions of the input values (see Figures 4A-4D ). In one embodiment, the FH conversion logic circuitry is arranged between a memory that initially receives and stores an M×M array of filter weights or coefficients and a multiplier-accumulator execution pipeline. In operation, the filter weights or coefficients are read from the memory and provided to the FH conversion circuitry, which converts the weights or coefficients from an M×M array (e.g., M=3) to an N×N array (e.g., N=4). The filter weights or coefficients are then provided to the multiplier-accumulator execution pipeline, where the image data / pixels are processed by the multiplier-accumulator circuitry of the execution pipeline.
[0056] In another embodiment, the memory stores an N×N array of input weights or weight values that are pre-computed (e.g., off-chip or by circuitry external to the multiplier-accumulator execution pipeline) and stored in the memory as an N×N array of filter weights or coefficients. In this embodiment, the FH conversion logic circuitry is not disposed between the memory and the multiplier-accumulator execution pipeline, and / or the FH conversion operation is performed before the filter weights or coefficients are stored in the memory. As in the previous embodiment, the filter weights are converted first and then the data is used in the multiplier-accumulator execution pipeline. In particular, however, storing the pre-computed N×N array of input weights or weight values in memory (rather than calculating such values during the multiplier-accumulator circuitry / pipeline operation (i.e., on the fly)) increases the memory storage required for such input weights or weight values, which in turn increases the capacity requirements of the memory used in this alternative embodiment (e.g., in this exemplary embodiment, the increase may be on the order of N×N / M×M, or approximately 16 / 9).
[0057] Continue to refer Figure 2A and Figure 2B In a second operating mode of the multiplier-accumulator circuitry, the plurality of execution pipelined multiplier-accumulator circuitry performs accumulation of a value Uijklm (from an Eijk*Hklm multiplication) from an input plane (index K) to a Zijlm value, as indicated by the ΣK notation, where an N×N (e.g., 4×4) multiplication is substituted or replaced. Figure 1A and 1BThe multiplier-accumulator circuitry shown is M×M (e.g., 3×3). Each two-dimensional array / set of data includes input or image data / pixels (e.g., all input or image data / pixels) that are related / associated with or contribute to the output data value. That is, in the second operating mode, the multiplier-accumulator circuitry of each execution pipeline in the plurality of pipelines performs a plurality (e.g., 16) of multiplications and, in one embodiment, implements or performs accumulation operations in the zij-yij conversion block, thereby writing the four output pixels at Yijl (2×2) to the output plane (in parallel with the other Yijl 2×2 pixels written to other output planes (other L index values)).
[0058] It is worth noting that Figure 2A and Figure 2B 2 are a logical and physical overview, respectively, of a second mode of operation of a multiplier-accumulator execution pipeline of a multiplier-accumulator circuit system according to certain aspects of the present invention. Figure 2C According to a particular aspect of the present invention Figure 2B An example timing summary of the physical summary shown. In addition, Figure 3A 、 4A and 5A are according to a particular aspect of the present invention. Figure 2B The physical details of the physical outline are shown. In addition, Figure 3C 、 3D , 4C, 4D and 5C show gate / RTL details of the conversion logic circuitry of an exemplary embodiment of certain aspects of the invention shown herein. Figure 3E 、 4E 5D and 5D illustrate specific sequencing details of conversion logic circuitry according to certain aspects of the present invention.
[0059] in particular, Figure 3A Shown Figure 2A and Figure 2BAdditional details of DE conversion logic according to certain aspects of the present invention are shown. In this exemplary embodiment, a memory (e.g., L2 SRAM memory) storing 4×4D blocks can be partitioned or divided into 16 physical blocks, so that 16 sets of data can be accessed in parallel for pipeline use by the 16 multiplier-accumulator circuitry of the multiplier-accumulator circuitry. Each set of input data (e.g., image data) comprises four 4×4D blocks, which can be read / accessed from each physical block of memory in 64-word bursts (e.g., in one embodiment, each access of the L2 SRAM memory (e.g., which may take 1 nanosecond)). The 4×4D blocks can be converted into 4×4E blocks by 16 dij-eij conversion circuitry that performs the conversion operation. The 4×4E blocks are separated into 16 streams, which are sorted by the respective eij values / elements of the 4×4E blocks. In one embodiment, this operation is performed by eij extraction logic circuitry (one eij extraction circuitry per stream (16 in this exemplary embodiment)). Each of the 16 eij streams may be directed to an e-shift-in block of one of the 16 multiplier-accumulator execution pipelines of the multiplier-accumulator circuitry.
[0060] It is worth noting that Figure 3C and Figure 3D Additional details of one unit of the dij-eij conversion logic circuitry and one unit of the eij extraction logic circuitry are shown, respectively, according to certain aspects of the present invention. In one example embodiment, the dij-eij conversion logic circuitry includes (i) 16 left units, (ii) data registers for the dij and eij data words, (iii) control logic for sequencing the operations for processing, and (iv) for converting (according to Figure 2E In this exemplary embodiment, the eij extraction logic circuitry also includes (i) 16 right cells, (ii) data registers for the dij and eij data words, (iii) control logic for sequencing operations, and (iv) control logic for converting (according to Figure 2E The dij-eij conversion circuitry also includes adder logic for the dij-eij conversion table. In addition, the dij-eij conversion circuitry also includes vertical "EX_IN" and "EX_OUT" ports that carry the extracted eij value / element to the associated or appropriate multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry. Note that some dij-eij conversion processing may be implemented or executed in the eij extraction unit.
[0061] Figure 3EA rough diagram of the conversion / extraction pipeline (horizontal axis) and the timing of each unit (vertical axis) according to certain aspects of the present invention is shown. In this exemplary embodiment, the horizontal axis shows 16 dij-eij conversion units in "X" positions {0, 1, ... 15} and 16 eij extraction units in "X" positions {16, 17, ... 31}. The lower axis shows the dij elements / values of the 4×4D data block in registers DREG_X and DREG_Y at their fixed positions in the pipeline. The 4×4E data block is passed from left to right, and each element / value is accumulated into the eij element. This accumulation is controlled by a pattern of "1" characters, each "1" specifying the time and location of the accumulation. In this exemplary embodiment, a total of 64 accumulations are required to convert a 4×4D block into a 4×4E block. For example, at T=9, the e23 element is subtracted from d11 using the DREG_X register at the X=10 unit position.
[0062] refer to Figure 2A 、 2B 4A, in one embodiment, FH conversion logic is arranged in or incorporated into the execution pipeline circuit system to convert the filter weights or coefficients into Winograd format. Specifically, Figure 4A According to certain aspects of the present invention, Figure 2A and Figure 2B Further details of the FH conversion logic. A memory (e.g., L2 SRAM memory) stores 3×3F blocks with filter weights or coefficients (e.g., finite impulse response (FIR) type). In this example embodiment, the memory can be partitioned or divided into 16 physical blocks so that 16 sets of data can be read or accessed in parallel by / for the 16 multiplier-accumulator execution pipelines of the multiplier-accumulator circuit system. Here, each set of data includes four 3×3F blocks, which require 36 accesses from each physical L2 block, each access requiring 1ns in this example. The 3×3F blocks are converted into 4×4H blocks (in Winograd format) by the conversion circuit system (in the embodiment shown, 16 fkl-hkl conversion logic circuits). These blocks can be written to a memory (e.g., L1 SRAM memory) shared by the 16 multiplier-accumulator execution pipelines. Thereafter, each of the 16 hkl elements / values of the 4×4H block is stored in a memory (e.g., L0 memory, such as SRAM) of one of the 16 multiplier-accumulator execution pipelines of the multiplier-accumulator circuit system and is available to the multiplier-accumulator circuit system of each of the execution pipelines for processing input data (e.g., image / pixel data).
[0063] In this exemplary embodiment, the tidying is performed by addressing sequences when reading hkl elements / values from L1 memory and writing hkl elements / values in memory (e.g., 16 L0 memories, which are SRAM in one embodiment). However, as an alternative, the tidying can be performed by addressing sequences with Figure 3A The eij extraction logic circuitry shown is performed using a hkl extraction logic circuitry similar to the eij extraction logic circuitry shown. In particular, the timing of transfers between memories (e.g., from L2 memory to L1 memory, and from L1 memory to L0 memory) is not as critical as the transfer of input and output data between memories (e.g., L2 memory) and the execution pipeline of the multiplier-accumulator circuitry. The weight value or data can be read from the memory once and transferred to the pipeline of the multiplier-accumulator circuitry, and then reused for each of thousands of blocks of 2×2 input pixels.
[0064] Figure 4C Detail of an exemplary embodiment of two units of a fkl-hkl conversion circuit system according to certain aspects of the present invention is shown. In particular, there is no Figure 3A and Figure 3D In one exemplary embodiment, the fkl-hkl conversion logic system includes 16 left cells and 16 right cells. In addition, the fkl-hkl conversion logic system includes (i) data registers for fkl and hkl weight values, (ii) control logic for sorting, and (iii) for (according to Figure 2D The adder logic is converted by the conversion table shown in FIG.
[0065] Notice, Figure 4C The embodiment in has a 10-bit precision hkl accumulation path and uses a saturating adder to handle overflow (see Figure 4D ). Alternative embodiment (combined with Figure 7A As described above, a 12-bit precision hkl accumulation path is used, so there is no need to include a saturating adder to handle overflow; that is, the 12-bit precision accumulation path can have sufficient numerical range to avoid overflow.
[0066] Figure 4EA rough diagram of the conversion / extraction pipeline (horizontal axis) and the timing of each unit (vertical axis) according to a particular aspect of the present invention is shown. In this exemplary embodiment, the horizontal axis shows the 32 fkl-hkl conversion units of the conversion circuit system at the "X" positions {0,1,...31}. The lower axis shows the fkl elements / values of the 3×3F blocks in registers DREG_X and DREG_Y at their fixed positions in the pipeline. The 4×4H blocks are passed from left to right, and the individual elements / values are accumulated to the hij elements / values. The accumulation is controlled by a graphic of "1 / 2 / 4" characters, where each "1 / 2 / 4" specifies the time and position of the accumulation. The values of the "1 / 2 / 4" characters specify a scaling factor of *1.0, *0.5, or *0.25, respectively. In this embodiment, a total of 64 accumulations are used to convert the 3×3F blocks into 4×4H blocks. For example, at time T=0, the h23 element at the X=1 location is subtracted by 0.5*f12 using the DREG_X register.
[0067] refer to Figure 2A 、 2B and 2C, in a second mode of operation, an N×N multiplier-accumulator execution pipeline of a multiplier-accumulator circuit system is employed to accumulate Q×Q pixel output data / values in a single output layer, wherein each execution pipeline aggregates the sum of K multiplications of input data values and associated input weight values for K input layers. In some embodiments, the aggregation of N×N element data / values for the Q×Q output data / pixel is implemented / performed outside the N×N multiplier-accumulator execution pipeline. Here, the N×N product data / element is accumulated with other N×N product data / elements from other input layers—however, in this embodiment, after performing ZY conversion logic operations on the accumulated N×N product data / elements, the individual elements / values are accumulated together into the final Q×Q output data / pixel (see Figures 5A-5D ).
[0068] In short, Figure 5A According to certain aspects of the present invention, Figure 2A and Figure 2B. In this exemplary embodiment, each of the 16 zij streams leads to a z-shift-out block of one of the 16 multiplier-accumulator execution pipelines of the multiplier-accumulator circuitry. A 4×4 Z block can be assembled from the 16 streams, which are sorted by the individual zij elements / values of the 4×4 Z block, which can be achieved by insertion logic circuitry (here, 16 zij insertion logic circuitry). The 4×4 Z block is converted to a 2×2 Y block by the zij-yij conversion logic circuitry. A memory (e.g., L2 SRAM memory) can store the 2×2 Y block (e.g., in a split or partitioned form) into 16 physical blocks, so that data for 16 sets can be written or stored in parallel for the 16 multiplier-accumulator execution pipelines. Here, each set of data may include four 2x2Y blocks, which will include 16 accesses from each physical block of memory (eg, L2 SRAM memory), with each access comprising, for example, 1 ns in this exemplary embodiment.
[0069] Note that in one embodiment, only 1 / 4 of the available L2 SRAM memory is used to write the Y block data; both the D block data and execution pipelines use a 64ns pipeline cycle time to process 16×64 4×4D input blocks with each 2×2 pixel step. In this exemplary embodiment, the lower Y access bandwidth of the L2 SRAM memory can facilitate reducing the number of physical blocks of Y memory from 16 to 4.
[0070] However, as an alternative, the additional bandwidth may be used in situations where more than 64 input planes are accumulated. For example, if there are 128 input planes (and 64 MAC elements / values per multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry), the first 64 input planes may be accumulated to a specific area of memory (such as the "Y" area of the L2 SRAM memory). Then, as the second 64 input planes are accumulated in the multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry, the Y value of the first plane is read from Y2 and passed to the accumulation port on the zij-yij conversion logic circuitry. The values of these two sets may be added together and rewritten or stored to the Y area of the L2 SRAM memory. This is done in Figure 5A The path outlined by the dashed line marked with a "V" is shown. (See also Figure 5E, which shows the accumulation of the first 64 input planes (also referred to herein as layers) on the left side of the diagram in each schematic block diagram. The accumulated values are stored in the Y region of the L2 memory. The second set of 64 input planes are accumulated in the right side of the diagram. Before being written to the Y region of the L2 memory, the 4×4 Z block values pass through the zij-yij conversion logic circuitry. As they pass through the conversion logic, the Y accumulated values from the first 64 input planes are read from the L2 and loaded into the accumulation input port of the zij-yij conversion logic.
[0071] It is worth noting that the reference Figure 5A When the input layer depth DD is greater than the pipeline depth, a read-modify-write (RMW) option can be used. For this option, the previously written Yij group (the four words marked "64a") is read and passed to the accumulator input of the zij-yij converter to be added to the four words of the "64b" operation. This can be shared with the write of the yij group, because only eight L2 memory cycles are required (four yij writes and four yij reads) outside each 16.
[0072] Figure 5C Details of one unit of the zij insertion logic circuit system (left portion of the figure) and one unit of the zij-yij conversion logic circuit system (right portion of the figure) according to certain aspects of the present invention are shown. The zij insertion logic circuit system includes (i) 16 left units, (ii) data registers for zij and yij data words, (iii) control logic for sequencing, and (iv) for (according to Figure 2F It also includes vertical "INSRT_IN" and "INSRT_OUT" ports that carry the inserted zij elements / values from the appropriate execution pipeline of the multiplier-accumulator circuitry. The zij insertion logic circuitry may also include an accumulation port (lower left side in the figure) - for example, in the case where there are more input planes than execution pipelines or pipeline stages. The zij-yij conversion logic circuitry includes (i) 16 left cells, (ii) data registers for the dij and eij data words, (iii) control logic for sorting, and (iv) for (according to Figure 2F Note that some of the zij-yij conversion processing may be implemented or performed in the zij insertion unit; it is worth noting that in certain embodiments (including embodiments in which some of the zij-yij conversion processing is implemented in the insertion unit), the zij-yij insertion unit may include some of the same circuitry as the zij-yij conversion unit.
[0073] Figure 5DA rough diagram of the conversion / extraction pipeline (horizontal axis) and the timing of each unit (vertical axis) according to certain aspects of the present invention is shown. Here, the horizontal axis shows 16 zij insertion units in the "X" positions {0, 1, ... 15} and 16 zij-yij conversion units in the "X" positions {16, 17, ... 31}. The lower axis shows the zij elements / values of the 4×4 Z blocks at fixed positions in the pipeline in registers ZREG_X and ZREG_Y. The 2×2Y blocks are passed from left to right and each element / value is accumulated to the yij element. This accumulation can be controlled by a pattern of "1" characters, each "1" specifying the time and position of the accumulation. A total of 36 accumulations are required to convert the 4×4 Z block into a 2×2Y block. For example, at time T=0, the y01 element is subtracted from z23 using the DREG_X register at the X=1 unit position.
[0074] It is worth noting that the reference Figure 1A 、 1B , 2A and 2B, the output data / pixel groups are respectively provided as 1×1 output element groups ( Figure 1A and Figure 1B ) and a 2×2 output element group ( Figure 2A and Figure 2B ) rather than being shown more generally as P×P and Q×Q arrays.
[0075] In short, reference Figure 2C , according to aspects of the present invention Figure 2B The exemplary timing of the circuitry in the exemplary physical schematic in FIGURE 1 illustrates the operation of the 16 parallel multiplier-accumulator execution pipelines of the multiplier-accumulator circuitry and the connection paths to the memory. Each pair of waveforms depicts the first and last of the 16 pipelines, which exhibit similar behavior to the middle 14 pipelines in the exemplary multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry. In this exemplary embodiment, each operation group processes 16×64 input data words corresponding to a 4×4 D block in each of the 64 layers, with each pipeline utilizing 64 clock cycles (e.g., each cycle may be 1 ns in this exemplary embodiment). The top waveform illustrates the movement of the D block from memory (e.g., L2 SRAM memory) through the DE conversion logic circuitry (via read and write operations) to the DE conversion operation. This transfer step has a pipeline delay of 16 ns; the conversion process can begin when a portion of the data is available (here, when ¼ of the data is available).
[0076] It is noteworthy that some stages have 16ns pipeline delays and 64ns pipeline cycle rates; in other words, in this exemplary embodiment, each stage can accept a new 16×64 word operation every 64ns interval, yet can overlap its processing by 48ns with the next stage. The DE conversion operation (implemented by the DE conversion circuitry) produces a 4×4E block. The extraction logic circuitry separates the 16 eij elements / values from each 4×4 block, passing each to one of the 16 execution pipelines. A 64ns 4×4E block requires 64ns to shift in—this stage (and the following two stages) have the same pipeline delay and pipeline cycle time.
[0077] Continue to refer Figure 2C , when the input data of the E block has been shifted into or applied to the multiplier-accumulator circuitry of the multiplier-accumulator execution pipeline, the multiplier-accumulator pipelines combine to perform 16×64×64 MAC operations (labeled “MAC operations”). Here, the 64 multipliers and 64 adders of the multiplier-accumulator circuitry in each of the 16 multiplier-accumulator pipelines each perform one operation per nanosecond over a 64ns interval. This is an accumulation on the “K” and “L” indices of the input and output planes. The 64ns 4×4 Z block requires 64ns to shift out, and this stage can overlap with the ZY insertion stage by up to 48ns. Similarly, the ZY conversion stage can overlap with the L2 write stage by up to 48ns. Each 2×2 pixel block consumes 64ns of pipeline cycle time—the next 2×2 block is shown in dark gray in the timing waveform. Therefore, processing all 128k pixels in this example would require 1ms (~1 million ns). In this exemplary embodiment, the entire 16×64 word operation has a pipeline delay of 18×16ns, or 288ns. In this exemplary example, the pipeline delay of 288ns is approximately 3472 times smaller than the total operation delay of 1ms, and therefore has a relatively small impact on the overall throughput of the system.
[0078] refer to Figure 2D , simply put, in Figure 2D There are nine fij elements / values that make up a 3x3 FIR (Finite Impulse Response) filter matrix "F". These elements / values are converted into a 4x4 "H" matrix having 16 elements / values hij. Figure 2D The upper left figure shows the details of this transformation. Each hij element is created by summing one to nine fij elements. Black text on a white background indicates "plus," and white text on a black background indicates "minus." Some elements / values are scaled by 1 / 2 or 1 / 4.
[0079] refer to Figure 2EIn one embodiment, each 2×2 input pixel / data block "D" is processed into a 2×2 output block. The 2×2 input data block is surrounded by a ring of 12 adjacent pixels, which will be used for the filter operation, but will themselves be processed in a different iterative loop step. Therefore, there are a total of 16 elements / values dij that make up the 4×4 input data block "D". These values / elements are converted into a 4×4 "E" matrix with 16 elements eij. Each eij element is generated by summing four dij elements. Black text on a white background means "addition" and white text on a black background means "subtraction".
[0080] refer to Figure 2F In this embodiment, a 4×4 “H” matrix and a 4×4 input data block “D” are multiplied together element-by-element (value-by-value) into a 4×4 output block “Z” having 16 zij elements. These elements are converted into a 2×2 “Y” matrix having 4 elements / values yij. The details of this conversion are shown in the lower left figure of the figure. Each yij element is generated by summing nine zij elements. Black text on a white background means “addition” and white text on a black background means “subtraction”. These yij elements / values, along with the yij elements / values generated from input blocks with pixels belonging to other input planes, are accumulated into a 2×2 output pixel block.
[0081] Note that when the zij-yij conversion occurs in the converter block between the execution pipeline and the memory (as in the first embodiment), the 4×4 output block “Z” generated in the multiplier-accumulator execution pipeline is not immediately accumulated into the 2×2 output pixels (as in the 3×3 filter of the first operating mode of the execution pipeline - see Figure 1A and Figure 1B and associated text). This means that each execution pipeline operates on only one of the 4×4 elements, while the 16 associated execution pipelines of the multiplier-accumulator circuitry operate in parallel or together to process the entire 4×4 block.
[0082] It is noted that the memory used to store data can be, for example, a block or array of dynamic and / or static random access memory cells, such as DRAM, SRAM, flash memory, and / or MRAM; in particular, all memory types and combinations thereof are intended to fall within the scope of the present invention. In one embodiment, the third and / or fourth memory stores input data, input weight values, and output data values in SRAM (e.g., a third memory, such as an L2 SRAM memory) and / or DRAM (e.g., a fourth memory, an L3 DRAM memory). Additionally, the third and / or fourth memory can store transformed input data (after the input data is transformed by the DE conversion logic operation) for an N×N array of input or image data / pixels. In one embodiment, both the "D" input data and the "Y" output data can be stored in the third (L2 SRAM) memory—each data participating in a different multiplier-accumulate (MAC) operation (e.g., 64 different MAC operations), so that the more limited L2 memory bandwidth is sufficient for the much higher bandwidth of the multiplier-accumulator execution pipeline. In contrast, the bandwidth of weight data required for the execution pipeline is much higher and needs to be stored in the first and / or second memory SRAM (e.g., L0 SRAM memory and L1 SRAM memory), which, in one embodiment, can retain: (i) "F" weight values for a first operating mode of the N×N multiplier-accumulator execution pipeline of the multiplier-accumulator circuit system, or (ii) "H" weight values for a second operating mode of the N×N multiplier-accumulator execution pipeline of the multiplier-accumulator circuit system.
[0083] As described above, in one embodiment, the DE conversion operation and / or the ZY conversion operation may be performed separately (and not on the fly)—although such an implementation would require additional read / write operations (e.g., 2x more read / write operations for L2 operations), which would also increase the capacity requirements of the memory (e.g., the third memory (L2 SRAM memory)).
[0084] In the case where the filter weights or coefficients are transformed on the fly (i.e., during pipeline operation of the multiplier-accumulator), the first and second memories may also store the transformed weight values or data. In one embodiment, the third and / or fourth memories may also be, for example, blocks or arrays of dynamic and / or static random access memory cells, such as DRAM, SRAM, flash memory, and / or MRAM; indeed, all memory types and combinations thereof are intended to fall within the scope of the present invention. In a preferred embodiment, the first and / or second memories are SRAM (e.g., L0 SRAM memory and L1 SRAM memory).
[0085] It is noted that in the exemplary embodiments described herein (text and figures), the multiplier-accumulator execution pipeline (which includes the multiplier-accumulator circuitry) is sometimes labeled "NMAX" or "NMAX pipeline" or "MAC pipeline."
[0086] refer to Figure 6A 、 6B 6C, in another embodiment, the architecture incorporates the DE conversion logic / circuitry and the ZY conversion logic / circuitry into the multiplier-accumulator execution pipeline, or performs their operations within the multiplier-accumulator execution pipeline. That is, in one embodiment of the architecture, input data stored in memory (e.g., in a layer comprising a two-dimensional M×M array of image data / pixels) is read from memory by the multiplier-accumulator execution pipeline and undergoes a transformation or conversion within the pipeline (e.g., into an N×N matrix). However, in this embodiment, the FH conversion logic or the operations performed thereby may be performed before the filter weights are applied or provided to the multiplier-accumulator execution pipeline. That is, in one embodiment, the FH conversion logic transforms or converts the M×M input weight block from an M×M matrix to an N×N matrix before the filter weights are applied or applied in the multiplier-accumulator execution pipeline. Thereafter, the circuitry of each multiplier-accumulator pipeline processes the N×N input data using the associated N×N filter weights.
[0087] As described above, in this embodiment, the ZY conversion logic is incorporated into the multiplier-accumulator execution pipeline. That is, the operation / processing of the ZY conversion circuit system is performed in the execution pipeline. The multiplier-accumulator circuit system can accumulate the individual elements / values of the N×N execution pipeline within the execution pipeline, so that the processing can be implemented via a single execution pipeline instead of N×N execution pipelines (for example, 16 execution pipelines). In this way, the individual elements / values in the multiplier-accumulator execution pipeline are accumulated together into the final Q×Q output data / pixel. That is, in this embodiment, the accumulation of the individual elements / values of N×N is implemented in the execution pipeline, so the single execution pipeline (relative to Figure 2A and Figure 2B The N×N (e.g., 16) execution pipelines shown accumulate N×N product data / elements after the ZY conversion operation.
[0088] refer to Figure 6A and Figure 6BIn this embodiment, the filter weights are converted or transformed prior to pipeline operation and stored as N×N blocks in memory. In this embodiment, the pre-processed and pre-transformed filter weights are read from memory in N×N blocks by the multiplier-accumulator circuitry. The multiplier-accumulator circuitry of each multiplier-accumulator execution pipeline uses transformed weights or coefficients with associated input data (which is transformed from M×M blocks of input data to N×N blocks of input data during operation by the circuitry and techniques of the DE conversion logic circuitry) during operation and execution of the multiplier-accumulator circuitry / pipeline. This weight conversion or transformation can be performed separately by circuitry different from the circuitry of the present invention (e.g., by an off-chip processor or computing system).
[0089] In case the input weight values are transformed on the fly (i.e. during execution of a pipeline operation), the weight values may be stored again in the first and / or second memory, which in a preferred embodiment are SRAMs (e.g. L0 SRAM memory and L1 SRAM memory).
[0090] It is worth noting that Figure 6A The physical outline of the multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry employing transformed or converted filter weights of the multiplier-accumulator circuitry with associated input data (which is transformed on the fly (i.e., during operation of the multiplier-accumulator circuitry) from M×M blocks of input data to N×N blocks of input data by the circuitry and techniques of the DE conversion logic circuitry) during operation and execution of the multiplier-accumulator circuitry / pipeline. Here, the throughput can be compared to Figure 2B The 16 pipelines shown are identical, achieved by employing nearly the same number of multipliers and accumulators implemented, organized, and / or configured in different arrangements.
[0091] also, Figure 6B Detail of four slices of a pipeline stage is shown, where each of the four slices processes one of the four input streams of 4×4 input data blocks D (received from the right in the figure). Here, the "H" input from the top receives the appropriate values of the 4×4H filter matrix for 4×4 multiplication. In addition, the processing includes DE conversion (via conversion circuitry) performed by the "add3" block, 4×4 element-by-element multiplication performed by the "mul" block, and ZY conversion performed by the "add4" block. The output data block Y is passed to the left for further accumulation. In particular, Figure 6BThe multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry shown shows the "add3 / mul / add4" blocks executed within a single pipeline cycle (for clarity). In one embodiment, these operations are separated or divided into two or three cycles (incorporating or implementing additional pipeline registers) and implemented by appropriate circuitry. This alternative approach can increase execution speed at the expense of slightly more complex sequencing.
[0092] Figure 6C shows the aggregation into a single block, Figure 6B This block is capable of accepting a 4x4D block and a 4x4H block and producing a 2x2Y block (including DE and ZY conversion via appropriate circuitry) in each pipeline cycle. Figure 6C The 64 blocks in the are aggregated into a single execution path, which provides the same Figure 2A The 16 pipelines (including DE and ZY conversion logic circuit systems) have the same or similar performance. Figure 1A and Figure 1B Each of the 16 pipelines includes a multiplier-accumulator circuit system having 16 multiplication / accumulation elements / circuits, so the total number of multiplication elements / circuits in the two embodiments is similar (e.g., both structures include 1024 multiplication elements / circuits).
[0093] It is important to note that the pseudocode, operations, configurations, block / data widths, datapath widths, bandwidths, data lengths, values, processes, and / or algorithms described and / or illustrated in the figures and text are exemplary only. Indeed, the present invention is not limited to the specific pseudocode, operations, block / data widths, datapath widths, bandwidths, values, processes, and / or algorithms illustrated and / or implemented according to the exemplary logical or physical schematic configurations, such as (one or more) execution pipelines and / or exemplary conversion circuitry.
[0094] refer to Figure 7A It should be noted that in cases where memory capacity is an issue, performing the conversion is advantageous because the filter (weight) elements / values are moved from L2 memory to L1 / L0 memory. The number of elements / values increases from nine to 16, which increases the capacity requirements of the L1 / L0 memory. This may be an appropriate implementation because the L2 memory occupies a larger chip area than the L1 / L0 memory. This running solution is also applied to the DE conversion and ZY conversion of input and output data (see Figure 7B and Figure 7C )—If data is kept in L2 memory in E and Z form in L2, it will require 4 times the L2 capacity.
[0095] Another problem may arise from the accuracy of the 4×4 “H” matrix generated from the 3×3 “F” matrix. The format of the fij elements / values is typically 8-bit signed integers. The conversion of fij elements / values to hij elements / values means that up to nine 8-bit fij integers (scaled by 1 / 4) must be added together to the hij elements. The hij number format must add two extra bits to reduce the chance of overflow (if overflow does occur, this can be detected in advance and the convolutional neural network (CNN) stage can handle it through the first operating mode). In addition, it may be necessary to accommodate two fractional bits (in the case of weights 1 / 2 and 1 / 4) to handle the 1 / 4 scaling operation during the fij-hij conversion.
[0096] This increases the bit format to three possible values. Four of the hij elements / values require a 12-bit signed integer format, eight of the hij elements / values require a 10-bit signed integer format, and four of the hij elements / values require an 8-bit signed integer format. This is an average of 10 bits per hij element. The L1 / L0 memories can be designed for the 10-bit average case with special shared logic so that the extra two bits required for the h11, h21, h12, h22 elements / values are stored in the memory cells of the h00, h03, h30, h33 elements (see Figure 7A ). This reduces the storage area used or required by the L0 / L1 memory.
[0097] Incremental precision may also be required in the data accumulation path, however, this is typically implemented using 16-bit and 32-bit signed integer precision for input and output data values. Therefore, existing formats can generally handle an additional two or four bits of precision range. If input or output overflow is a concern and the format cannot extend two or four bits, the conversion and accumulation hardware can be enhanced with additional saturation logic. When overflow occurs, some accuracy is lost, but the CNN results will be nearly identical.
[0098] Downsampling can be an important operation that may be required during CNN processing. This reduces the number of pixels in the input plane as they are transferred to the output plane. Generally, the number of output planes increases so that the number of pixels per level remains almost constant.
[0099] The present invention may employ downsampling processing / techniques. In brief, reference Figure 8, downsampling for the first mode of operation is shown on the left. Typically, 1 / 4 of the pixels are filtered with a 3×3 filter operation to produce output pixels. The 3×3 filter operation is performed using the remaining (unfiltered) input pixels, however the remaining (unfiltered) input pixels themselves are not filtered and written to the output plane. Note that for clarity purposes, the unwritten output locations are shown in white in the figure. In practice, the actual pixels are written to adjacent memory locations so that the output plane occupies a contiguous area of memory.
[0100] The downsampling on the right for the second mode of operation (i.e., implementing the Winograd processing technique) may not be able to handle downsampling cases efficiently because it operates on a 2×2 block of input pixels. However, it can process 1 / 4 of the 2×2 input block (black pixels) as shown on the right side of the figure. In other words, 1 / 4 of the pixels are filtered (black pixels) with a 4×4 filter operation to produce output pixels (black pixels). The input pixels are used to perform the 4×4 filter operation, but they themselves are not filtered and written to the output plane. Note that the unwritten output locations are shown in white in the figure - this helps to make Figure 8 Once again, actual pixels (black pixels) are written to adjacent memory locations so that the output planes occupy contiguous areas of memory.
[0101] This alternative downsampling method reduces the number of pixels by a factor of 4 and can be implemented in conjunction with the Winograd technique in the second mode of operation. Different phases for sampling the input pixels will require different training to adjust the weights so that the CNN stages have similar filtering functions. However, the cost of this additional training effort is offset by the improved performance of the downsampled CNN stages.
[0102] Note that the same approach can be applied to the CNN stage that performs upsampling, which increases the number of pixels per image plane. The ordering for this will look similar to the downsampling operation, but in reverse. The additional output pixels will be generated by interpolation of adjacent pixels.
[0103] refer to Figure 9 , for downsampling operations, the addressing and sequencing logic for executing the pipeline manages the L2 memory that reads 4×4D blocks of input and writes 2×2Y blocks of output. Here, stride = 1 (no downsampling) is shown on the left, and stride = 2 (downsampling) is shown on the right.
[0104] For stride = 1, one stride's input pixels (ΔDh × Dw) are read and converted into a series of 4×4D blocks. The blocks are converted into 4×4E blocks, which are passed to the NMAX execution pipeline. The resulting 4×4Z blocks are converted into 2×2Y blocks and written into one stride's output pixels (ΔYh × Yw).
[0105] For the modified stride=2, one stride of input pixels (ΔDh×Dw) is read and converted into a series of 4×4D blocks, however only half of the 2×2 pixel blocks are transferred; the control logic suppresses alternate 2×2 blocks.
[0106] The block is converted into a 4x4E block, which is passed to the NMAX execution pipeline. Again, only half of the 2x2 pixel block is transferred; the control logic suppresses alternate 2x2 blocks. The resulting 4x4Z block is converted into a 2x2Y block. Again, only half of the 2x2 pixel block is transferred; the control logic suppresses alternate 2x2 blocks. The 2x2 output block is written to a step of output pixels (ΔYh × Yw)—typically, the Yw width is scaled by 1 / 2 so that the output blocks are contiguous (no gaps).
[0107] Many inventions are described and illustrated herein. While specific embodiments, features, attributes, and advantages of the present invention have been described and illustrated, it should be understood that many other and different and / or similar embodiments, features, attributes, and advantages of the present invention will be apparent from the description and illustrations. Thus, the embodiments, features, attributes, and advantages of the present invention described and illustrated herein are not exhaustive, and it should be understood that such other, similar, and different embodiments, features, attributes, and advantages of the present invention are also within the scope of the present invention.
[0108] For example, although the illustrated embodiments and the text associated therewith describe and illustrate multiple memories (e.g., L3 memory, L2 memory, L1 memory, L0 memory), one or more of these memories may be omitted (e.g., L3 memory and / or L0 memory), and / or one or more of these memories may be combined / merged with one or more other memories (e.g., L3 memory may be combined into L2 memory, L2 memory may be combined with L1 memory, and / or L1 memory may be combined with L0 memory). Thus, the present invention is not limited to the exemplary embodiments set forth herein, including with respect to different memories.
[0109] In addition, reference Figure 10A and Figure 10B , one or more integrated circuits include circuitry for enabling and implementing one of a plurality of operating modes, the plurality of operating modes including, for example, a first operating mode (see, e.g. Figure 1A or Figure 1B ) and a second operating mode (see e.g. Figure 2A and Figure 2B For example, the mode control circuit system outputs a mode or modal control signal "MODE" to enable the first operation mode (eg, Figure 1A and Figure 1B) circuit systems and techniques that, in a first operating mode, employ a single execution pipeline and / or each execution pipeline in an execution pipeline to accumulate 1×1 pixel output values in a single output layer by aggregating the sum of K×M×M multiplications of input data values and associated input weight values from K layers. In one exemplary embodiment implementing the first operating mode, the 64×(3×3) input pixels at dij that determine the 64×(1×1) output pixels at yij are all processed by the single execution pipeline.
[0110] As described above, in the first operating mode, the multiplier-accumulator circuitry of the multiplier-accumulator execution pipeline performs M×M (e.g., 3×3) multiplications and accumulations, which produces yij values (see Figure 1B In one embodiment, when generating the yij value, all image data related to / associated with the single output data value or contributing to the set of the single output data value is applied to or employed by one of the multiplier-accumulator execution pipelines. As described above, processing may continue for each pixel of the plane. In addition, each execution pipeline of the plurality of execution pipelines (see, e.g., Figure 1B , which shows one execution pipeline among multiple) processing decisions associated with a separate set of yij output pixels / data (in this embodiment, 64×(1×1)) for all dij input pixels / data (in this exemplary embodiment, 64×(3×3)).
[0111] The mode control signal may output a mode or modal control signal "MODE" to enable the second operating mode (see, for example, Figure 2A and Figure 2B ) circuitry and techniques, including the conversion circuitry described in detail above. For example, the multiplier-accumulator circuitry in the plurality of execution pipelines performs accumulation of values Uijklm (from Eijk*Hklm multiplications) from an input plane (index K) to Zijlm values, as indicated by the ∑K notation—where N×N (e.g., 4×4) multiplications replace or substitute Figure 1A and Figure 1BThe multiplier-accumulator circuitry shown is M×M (e.g., 3×3). That is, in one example of the second operating mode, the 64×(4×4) input pixels / data at dij that determine the associated 64×(2×2) output pixels at yij are processed by 16 execution pipelines. The multiplier-accumulator circuitry of the execution pipelines performs multiple multiplications (e.g., 16), and in one embodiment, the accumulation operation is implemented or performed in the zij-yij conversion block, whereby the four output pixels at Yijl (2×2) are written to the output plane (in parallel with the other Yijl 2×2 pixels being written to other output planes (other L index values). Here, when the multiplier-accumulator circuitry in the multiple execution pipelines is enabled in this operating mode, one or more conversion circuitry can be incorporated into the data path (if necessary) to perform data processing operations using Winograd techniques (e.g., as described herein).
[0112] In one embodiment, the mode selection circuitry can be one-time programmable; in another embodiment, the mode selection circuitry is programmable more than once (i.e., multiple times). The mode selection circuitry can be programmed, for example, in situ (i.e., during operation of the integrated circuit), at the time of manufacture, and / or at or during power-up, at or during startup, at or during initialization, at or during initialization, at or during reinitialization, at or during configuration, at or during reconfiguration, etc. For example, the mode selection circuitry can receive a mode selection signal from internal or external circuitry (i.e., external to one or more integrated circuits (e.g., a host / processor)), the internal or external circuitry including one or more data storage elements (e.g., one or more memory cells, registers, flip-flops, latches, memory blocks / arrays), one or more input pins / conductors, a lookup table (LUT) (of any type), a processor or controller, and / or separate control logic. In response, the mode selection circuitry may employ (one or more) such signals to enable or disable the selected processing circuitry (as the case may be) and thereby implement (e.g., in situ and / or at or during power-up, startup or startup, initialization or initialization, re-initialization or re-initialization, configuration or configuration, reconfiguration or reconfiguration, etc.) one of the processing modes (e.g., the Winograd technique).
[0113] In fact, the present invention is neither limited to any single aspect or its embodiment, nor to any combination and / or permutation of these aspects and / or embodiments. In addition, each aspect of the present invention and / or its embodiment can be used alone or in combination with one or more of the other aspects of the present invention and / or its embodiment.
[0114] It is noteworthy that the various circuits, circuit systems, and techniques disclosed herein can be described using computer-aided design tools in terms of their behavior, register transfer, logic components, transistors, layout geometry, and / or other characteristics, and expressed (or represented) as data and / or instructions embodied in various computer-readable media. The formats of files and other objects in which such circuit, circuit system, layout, and routing expressions can be implemented include, but are not limited to, formats supporting behavioral languages such as C, Verilog, and HLDL, formats supporting register-level description languages such as RTL, and formats supporting geometric description languages such as GDSII, GDSIII, GDSIV, CIF, MEBES, and any other formats and / or languages now known or later developed. Computer-readable media in which such formatted data and / or instructions can be embodied include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media) and carrier waves that can be used to transmit such formatted data and / or instructions via wireless, optical, or wired signaling media, or any combination thereof. Examples of transmission of such formatted data and / or instructions via a carrier wave include, but are not limited to, transmission over the Internet and / or other computer networks (upload, download, email, etc.) via one or more data transmission protocols (such as HTTP, FTP, SMTP, etc.).
[0115] In practice, when received in a computer system via one or more computer-readable media, such data and / or instruction-based representations of the aforementioned circuits may be processed by a processing entity (e.g., one or more processors) in the computer system in conjunction with the execution of one or more other computer programs (including but not limited to netlist generation programs, placement and routing programs, etc.) to generate a representation or image of the physical manifestation of such circuits. Such representation or image may then be used in device fabrication, such as by enabling the generation of one or more masks for forming various components of the circuits during device fabrication.
[0116] In addition, various circuits, circuit systems and techniques disclosed herein can be represented via simulation using computer-aided design and / or test tools. Circuits, circuit systems, layouts and routings and / or simulations of the techniques implemented thereby can be implemented by computer systems, in which the features and operations of such circuits, circuit systems, layouts and the techniques implemented thereby are simulated, copied and / or predicted via the computer systems. The present invention also relates to such simulations of the circuits, circuit systems and / or the techniques implemented thereby of the invention, and is therefore intended to fall within the scope of the present invention. Computer-readable media corresponding to such simulations and / or test tools are also intended to fall within the scope of the present invention.
[0117] It is worth noting that references herein to "one embodiment" or "embodiment" (etc.) mean that the specific features, structures or characteristics described in conjunction with the embodiment may be included, adopted and / or combined in one, some or all embodiments of the present invention. The use or appearance of the phrases "in one embodiment" or "in another embodiment" (etc.) in the specification does not refer to the same embodiment, nor does it refer to separate or alternative embodiments that are necessarily mutually exclusive with one or more other embodiments, nor is it limited to a single exclusive embodiment. This also applies to the term "implementation". The present invention is not limited to any single aspect or embodiment thereof, nor is it limited to any combination and / or permutation of these aspects and / or embodiments. In addition, each aspect of the present invention and / or its embodiments may be used alone or in combination with one or more other aspects of the present invention and / or its embodiments. For simplicity, specific permutations and combinations are not discussed and / or shown separately herein.
[0118] Moreover, embodiments or implementations described herein as "exemplary" should not be construed as being ideal, preferred, or advantageous, for example, over other embodiments or implementations, but rather are intended to convey or represent that the embodiment(s) are exemplary embodiment(s).
[0119] Although the present invention has been described in terms of certain specific aspects, numerous other modifications and variations will be apparent to those skilled in the art. It should therefore be understood that the present invention may be practiced in other ways than those specifically described without departing from the scope and spirit of the present invention. The embodiments of the present invention are therefore to be considered in all respects as illustrative / exemplary and not restrictive.
[0120] The terms “comprises,” “comprising,” “includes,” “including,” “have,” and “having,” or any other variations thereof, are intended to cover a non-exclusive inclusion such that a process, method, circuit, article, or apparatus that comprises a list of components or elements includes not only those components or elements but may also include other components or elements not expressly listed or inherent to such process, method, article, or apparatus. Furthermore, the terms “connect,” “connected,” “connecting,” or “connection” as used herein should be broadly interpreted to include directly or indirectly (e.g., through one or more conductors and / or intermediate devices / elements (active or passive) and / or through inductive or capacitive coupling) unless otherwise intended (e.g., use of the terms “directly connect” or “directly connected”).
[0121] The terms "a" and "an" herein do not denote a limitation of quantity, but rather denote the presence of at least one of the referenced item. Additionally, the terms "first," "second," etc., herein do not denote any order, quantity, or importance, but are instead used to distinguish one element / circuit / feature from another.
[0122] Furthermore, the term "integrated circuit" specifically refers to any integrated circuit, including, for example, a general-purpose integrated circuit, a processor, a controller, a state machine, a gate array, a SoC, a PGA, and / or an FPGA. The term "integrated circuit" also refers to, for example, a processor, a controller, a state machine, and a SoC including an embedded FPGA.
[0123] Furthermore, the term "circuitry" refers in particular to a circuit (whether integrated or otherwise), a group of such circuits, one or more processors, one or more state machines, one or more processors implementing software, one or more gate arrays, programmable gate arrays and / or field programmable gate arrays, or a combination of one or more circuits (whether integrated or otherwise), one or more state machines, one or more processors, one or more processors implementing software, one or more gate arrays, programmable gate arrays and / or field programmable gate arrays. The term "data" refers in particular to (one or more) current or voltage signals (plural or singular) in analog or digital form, which may be a single bit(s) or multiple bits(s).
[0124] In the claims, the term "MAC circuit" refers to a multiplier-accumulator circuit of a multiplier-accumulator pipeline. For example, in U.S. Patent Application No. 16 / 545,345 Figure 1A The multiplier-accumulator circuit is described and illustrated in the exemplary embodiments of the 1C and the text associated therewith. However, it is noted that the term "MAC circuit" is not limited to circuits according to, for example, U.S. Patent Application No. 16 / 545,345. Figure 1A -1C exemplary embodiments show and / or describe the specific circuits, logic, blocks, functional and / or physical diagrams, block / data widths, data path widths, bandwidths and processes, as described above, and this U.S. patent application is incorporated herein by reference.
[0125] It is noteworthy that the limitations of the claims are not drafted in a means-plus-function format or a step-plus-function format.
Claims
1. An integrated circuit comprising: a plurality of multiplier-accumulator execution pipelines, each multiplier-accumulator execution pipeline including multiplier-accumulator execution circuitry to perform a plurality of multiplication and accumulation operations; mode selection circuitry electrically coupled to the plurality of multiplier-accumulator execution pipelines to configure the plurality of multiplier-accumulator execution pipelines to operate in a first mode or in a second mode, wherein: In a first mode, the multiplier-accumulator circuitry of one of the plurality of multiplier-accumulator execution pipelines is configured to: (i) receive image data of a first image data set, (ii) process the first image data set by performing a plurality of multiplication and accumulation operations using filter weights associated with the first image data set, and (iii) generate output data corresponding to the processed image data; as well as In a second mode, the mode selection circuitry enables the first conversion circuitry, a first plurality of the plurality of multiplier-accumulator execution pipelines, and the second conversion circuitry, and: the first conversion circuitry being configured to (i) receive image data, (ii) convert the image data into a Winograd format, and (iii) output a set of Winograd image data to the multiplier-accumulator execution circuitry of each multiplier-accumulator execution pipeline of the first plurality of multiplier-accumulator execution pipelines, the multiplier-accumulator circuitry of each of the plurality of multiplier-accumulator execution pipelines being coupled to an output of the first conversion circuitry and configured to (i) receive the Winograd image data set from the first conversion circuitry, (ii) process the Winograd image data set by performing a plurality of multiplication and accumulation operations using filter weights associated with the received Winograd image data set, and (iii) generate output data, wherein the output data of each multiplier-accumulator execution pipeline is in Winograd format; and The second conversion circuitry is coupled to an output of each of the first plurality of multiplier-accumulator execution pipelines to: (i) receive the output data from the multiplier-accumulator execution circuitry of each of the first plurality of multiplier-accumulator execution pipelines, and (ii) convert the output data from each of the first plurality of multiplier-accumulator execution pipelines to a non-Winograd format.
2. The integrated circuit according to claim 1 , further comprising: A memory is provided for storing image data and filter weights, wherein the memory stores the image data as a plurality of M×M arrays of image data.
3. The integrated circuit of claim 2, wherein: The first conversion circuitry is coupled to the memory to receive image data and, in operation, converts the image data from the plurality of M×M arrays of image data to a plurality of N×N arrays of image data, where N and M are integers and N is greater than M, and where the N×N arrays of image data are in Winograd format.
4. The integrated circuit of claim 1 , further comprising: Memory, and a third conversion circuit system, wherein the mode selection circuit system electrically couples the third conversion circuit system between the memory and the first plurality of multiplier-accumulator execution pipelines in the second mode, and wherein the third conversion circuit system is configured in the second mode to receive filter weights from the memory, convert the filter weights into a Winograd format, and output the filter weights in the Winograd format to the first plurality of multiplier-accumulator execution pipelines.
5. The integrated circuit of claim 4, wherein: The third conversion circuitry is operable to convert the filter weights from a plurality of M×M arrays of data to a plurality of N×N arrays of data, where N and M are integers and N is greater than M, wherein the N×N arrays of image data are in Winograd format.
6. The integrated circuit of claim 1 , wherein: The second conversion circuit system is operable to receive output data in a Winograd format from each of the first plurality of multiplier-accumulator execution pipelines and convert the output data from each of the first plurality of multiplier-accumulator execution pipelines to a non-Winograd format, wherein the non-Winograd format is a floating point format.
7. The integrated circuit of claim 1 , further comprising: A first memory stores each image data set and a second memory stores the filter weights, and wherein the second memory stores the filter weights associated with each image data set as a two-dimensional array of filter weights.
8. The integrated circuit of claim 7, further comprising: a third conversion circuit system coupled to the second memory and the multiplier-accumulator execution circuit systems of the first plurality of multiplier-accumulator execution pipelines, wherein, in the second mode, the third conversion circuit system is enabled to: (i) receive filter weights from the second memory, (ii) convert the filter weights into a Winograd format, and (iii) output the filter weights in the Winograd format to the multiplier-accumulator execution circuit systems of the first plurality of multiplier-accumulator execution pipelines.
9. The integrated circuit of claim 8, wherein: The third conversion circuitry is operable to receive the filter weights in a floating point format from the second memory.
10. The integrated circuit of claim 8, wherein: The third conversion circuitry includes extraction circuitry.
11. The integrated circuit of claim 8, wherein: The third conversion circuit system is operable to convert the filter weights from a plurality of M×M arrays of data into a plurality of N×N arrays of data, where N and M are integers and N is greater than M, and the filter weights in the plurality of N×N arrays of data are in Winograd format.
12. The integrated circuit of claim 1 , wherein: The second conversion circuit system is enabled in a second mode to convert output data from the multiplier-accumulator circuit system of each multiplier-accumulator execution pipeline of the first plurality of multiplier-accumulator execution pipelines from data in an N×N array to data in a P×P array, where N and P are integers and N is greater than P, and the output data in the P×P array of data is in a non-Winograd format.
13. The integrated circuit of claim 12, wherein: The non-Winograd format is a floating point format.
14. The integrated circuit of claim 1 , further comprising: Insertion circuitry is coupled between the second conversion circuitry and the memory, wherein the mode selection circuitry enables the insertion circuitry in the second mode.
15. The integrated circuit of claim 1 , further comprising: a first memory for storing the image data, and A second memory is used to store the filter weights, wherein: The first conversion circuit system is configured in a second mode to (i) receive image data in a floating-point format from a first memory, (ii) convert the image data into a Winograd format, and (iii) output a set of Winograd image data to the multiplier-accumulator execution circuit system of each multiplier-accumulator execution pipeline in the first plurality of multiplier-accumulator execution pipelines.
16. The integrated circuit of claim 1 , further comprising: A memory for storing image data and filter weights, wherein the memory comprises: A first memory, configured to store the image data as a plurality of M×M arrays of image data; and A second memory is used to store the filter weights as a plurality of M×M arrays of filter weights.
17. The integrated circuit of claim 16, wherein: a first conversion circuitry coupled to the first memory and enabled in a second mode to receive the image data and convert the image data from a plurality of M×M arrays of image data into a plurality of N×N arrays of image data, where N and M are integers and N is greater than M, wherein the N×N arrays of image data are in Winograd format; as well as The third conversion circuit system is coupled to the second memory and is enabled in the second mode to receive filter weights and convert the filter weights from a plurality of M×M arrays of filter weights to a plurality of N×N arrays of filter weights, wherein the N×N arrays of filter weights are in Winograd format.
18. The integrated circuit of claim 17, wherein: The first memory stores image data in a floating point format, and the second memory stores filter weights in a floating point format.
Citation Information
Patent Citations
Multiplier-accumulator circuit, logic tile architecture for multiply-accumulate, and IC including logic tile array
US10693469B2
Efficient sparse parallel winograd-based convolution scheme
US20170344876A1