MAC processing pipeline with conversion circuitry and method of operating the same
By using a conversion circuit system in the multiplier-accumulator circuit system to convert data and filter weights from floating point format to fixed point format, and then converting them back to floating point format after Winograd processing, the problem of insufficient data throughput is solved and processing efficiency is improved.
Patent Information
- Application Number
- CN202080057694.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-24
- Filing Date
- 2020-09-29
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-09-29
AI Technical Summary
The prior art is difficult to efficiently convert floating point data formats into fixed point data formats that are convenient for Winograd type processing, resulting in insufficient data throughput of the multiplier-accumulator circuit system.
The conversion circuit system is used to convert the input data and filter weights from the floating-point data format to the fixed-point data format, and then converted back to the floating-point data format after Winograd processing, so as to facilitate operation in the multiplier-accumulator circuit system.
Improves the data throughput of the multiplier-accumulator circuit system, and enhances processing efficiency and throughput.
Smart Images

Figure CN114270305B_ABST
Abstract
Description
[0001] Related applications
[0002] This nonprovisional application claims priority to and the benefit of U.S. Provisional Application No. 62 / 909,293, entitled “Multiplier-Accumulator Circuitry Processing Pipeline and Methods of Operating Same,” filed on October 2, 2019. The '293 provisional application is incorporated herein by reference in its entirety.
[0003] introduce
[0004] Many inventions are described and illustrated herein. These inventions are not limited to any single aspect or embodiment thereof, nor to any combination and / or permutation of these aspects and / or embodiments. Importantly, each aspect and / or embodiment of the present invention may be used alone or in combination with one or more of the other aspects and / or embodiments of the present invention.
[0005] In one aspect, the present invention relates to one or more integrated circuits having multiplier-accumulator circuitry (and methods of operating such circuitry) that include one or more execution or processing pipelines based on or using a floating-point data format, the one or more integrated circuits including circuitry for implementing Winograd-type processing to, for example, increase the data throughput of the multiplier-accumulator circuitry and processing. For example, in one embodiment, one or more integrated circuits having one or more multiplier-accumulator circuit (MAC) execution pipelines (e.g., for image filtering) may include conversion circuitry for converting or transforming data (e.g., input data and / or filter weight or coefficient data) from a floating-point data format to a data format that facilitates Winograd data processing. In this embodiment, input data and associated filter weights or coefficients may be stored in a memory, where the data is provided to conversion circuitry to convert or transform the data from a floating-point data format (sometimes identified herein and in the figures as “FP” or “FPxx,” where “xx” indicates / reflects an exemplary data width) to a data format that is convenient for implementation in or consistent with use of Winograd-type processing (e.g., a fixed-point format, such as a block-scaled fractional format (sometimes identified herein and in the figures as “BSF” or “BSFxx,” where “xx” indicates / reflects an exemplary data width).
[0006] In one embodiment, one or more integrated circuits include conversion circuitry or conversion circuitry that converts or processes input data from a floating-point data format to, for example, a fixed-point data format (sometimes referred to as D / E, DE, or D-to-E conversion logic or conversion circuitry). Thereafter, the input data in the fixed-point data format is input to circuitry that implements Winograd-type processing to, for example, increase the data throughput of the multiplier-accumulator circuitry and processing. The output of the Winograd processing circuitry is applied to or input to the conversion circuitry to convert or transform the input data after Winograd processing from the fixed-point data format to a floating-point data format before such data is input to the multiplier-accumulator circuitry via extraction logic circuitry. In this manner, the input data (e.g., image data) processed by the multiplier circuitry and accumulator circuitry of each of the one or more multiplier-accumulator pipelines is in floating-point format.
[0007] In another embodiment, one or more integrated circuits include a conversion circuit system (sometimes referred to as F / H, FH, or F to H conversion logic or circuit system) for converting or processing filter weight or coefficient values or data to adapt for use in or with Winograd type processing. For example, where the filter or weight values or data are in floating point format, the weight values or data can be converted or modified to a data format that is convenient for implementation in or consistent with Winograd type processing (e.g., a fixed point format, such as a block scaled fractional (BSF) format). Here, the filter weight data in a fixed point data format is input to the circuit system for implementing Winograd type processing. Thereafter, the output of the Winograd processing circuit system is applied to or input to the conversion circuit system to convert or transform the filter weight data after Winograd processing from a fixed point data format to a floating point data format before the filter weight data is input to the multiplier-accumulator circuit system via the extraction logic circuit system. The multiplier circuits of the multiplier-accumulator circuitry of the pipeline executed by the multiplier-accumulator (s) utilize such weight values or data in floating-point data format along with input data also in floating-point data format. In this manner, multiplication operations of the multiplier-accumulator operations are performed in floating-point data format conditions.
[0008] A multiplier-accumulator circuit system (and method of operating such circuit system) of an execution or processing pipeline (e.g., for image filtering) may include floating-point execution circuit system (e.g., multiplier circuitry and accumulator / adder circuitry) that implements floating-point addition and multiplication based on data having one or more floating-point data formats. The floating-point data format may be user- or system-defined and / or may be one-time programmable (e.g., at the time of manufacture) or more than one-time programmable (e.g., (i) at power-up or via power-up, startup, or the conduction / completion of an initialization sequence / processing sequence, and / or (ii) in-situ or during normal operation of the integrated circuit or multiplier-accumulator circuit system of the processing pipeline). In one embodiment, the circuit system of the multiplier-accumulator execution pipeline includes an adjustable precision data format (e.g., a floating-point data format). Additionally or alternatively, the circuit system of the execution pipeline may process data concurrently to increase the throughput of the pipeline. For example, in one implementation, the present invention may include multiple separate multiplier-accumulator circuits (sometimes referred to herein as "MACs") and multiple registers (including, in one embodiment, multiple shadow registers) that facilitate pipelining of multiplication and accumulation operations, where the circuitry performing the pipelining processes data concurrently to increase the throughput of the pipeline.
[0009] After the multiplier-accumulator execution or processing pipeline processes the data, the present invention, in one embodiment, employs conversion circuitry (sometimes referred to as Z / Y, ZY, or Z to Y conversion logic or circuitry) consistent with Winograd type processing to process the output data (e.g., output image data). In this regard, in one embodiment, after the multiplier-accumulator circuitry of the multiplier-accumulator execution pipeline (one or more) outputs the output data (e.g., image / pixel data), such data is applied to or input into data format conversion circuitry and processing, which converts or transforms the output data, which is in a floating-point data format, into data having a data format that is convenient for implementation in or consistent with Winograd type processing (e.g., a fixed-point data format such as BSF). Thereafter, the output data, which is in a fixed-point data format, is applied to or input into Winograd conversion circuitry of the ZY conversion circuitry, where the output data is processed consistent with Winograd type processing to convert the output data from Winograd format to a non-Winograd format. Notably, in one embodiment, the ZY conversion circuitry is directly incorporated or integrated into the multiplier-accumulator execution pipeline(s). Thus, the operation of the ZY conversion circuitry is performed within the MAC execution or processing pipeline.
[0010] Thereafter, the present invention may employ circuitry and techniques for converting or transforming output data (e.g., in a fixed-point data format) into a floating-point data format. That is, in one embodiment, after the output data (e.g., in a fixed-point data format such as BSF format) is processed by Winograd conversion circuitry, the present invention may employ conversion circuitry and techniques for converting or transforming the output data (e.g., image / pixel data) into a floating-point data format before outputting the output data to, for example, a memory (e.g., L2 memory such as SRAM). In one embodiment, the output data is written to and stored in the memory in a floating-point data format.
[0011] The present invention may employ or implement aspects of circuit systems and techniques (and methods of operating such circuit systems) having multiplier-accumulator execution or processing pipelines for implementing Winograd-type processing, for example, as described and / or illustrated in U.S. non-provisional patent application No. 16 / 796,111, filed on February 20, 2020, and entitled “Multiplier-Accumulator Circuitry having Processing Pipelines and Methods of Operating Same,” and / or U.S. provisional patent application No. 62 / 823,161, filed on March 25, 2019, and entitled “Multiplier-Accumulator Circuitry having Processing Pipeline and Methods of Operating and Using Same.” Here, the circuit systems and methods of the pipelines and processes implemented therein described and illustrated in the '111 and '161 applications can be supplemented or modified to further include conversion circuit systems and techniques to convert or transform floating-point data into a data format that facilitates a certain implementation of Winograd-type processing (e.g., a fixed-point format, such as a block-scaled data format), and vice versa (i.e., conversion circuit systems and techniques to convert or transform data having a data format that facilitates implementation of Winograd-type processing (e.g., a fixed-point format) into a floating-point data format that can then be performed or processed by or in conjunction with a multiplier-accumulator. It is noteworthy that the '111 and '161 applications are incorporated herein by reference in their entireties. In addition, while the '111 and '161 applications are incorporated herein by reference in their entireties, specific references to these applications will be made from time to time during the following discussion.
[0012] In addition, the present invention may adopt or implement aspects of circuit systems and techniques (and methods of operating such circuit systems) for floating-point multiplier-accumulator execution or processing pipelines having floating-point execution circuit systems (e.g., adder circuit systems) that implement one or more floating-point data formats, as described and / or illustrated in, for example, U.S. non-provisional patent application No. 16 / / 900,319, filed on June 12, 2020, entitled “Multiplier-Accumulator Circuitry and Pipeline using Floating Point Data, and Methods of Using Same,” and U.S. provisional patent application No. 62 / 865,113, filed on June 21, 2019, entitled “Processing Pipeline having Floating Point Circuitry and Methods of Operating and Using Same.” Here, the circuit systems and methods described and illustrated in the '319 and '113 applications, in which the pipelines and processing implemented therein, may be employed in the circuit systems and techniques described and / or illustrated herein in conjunction with one or more execution or processing pipelines based on or using floating-point data formats, including for implementing Winograd-type processing to increase data throughput of the multiplier-accumulator circuitry and processing. Notably, the '319 and '113 applications are incorporated herein by reference in their entireties.
[0013] In one embodiment, processing circuitry that performs pipelining (including multiple multiplier-accumulator circuits) can process data concurrently to increase the throughput of the pipeline. For example, in one implementation, the present invention may include multiple separate multiplier-accumulator circuits (sometimes referred to herein as "MACs" or "MAC circuits") and multiple registers (including, in one embodiment, multiple shadow registers) that facilitate pipelining of multiplication and accumulation operations, wherein the circuitry that performs pipelining processes data concurrently to increase the throughput of the pipeline - see, for example, the multiplier-accumulator circuitry, architecture, and integrated circuits described and / or illustrated in U.S. Patent Application No. 16 / 545,345, filed on August 20, 2019, and U.S. Provisional Patent Application No. 62 / 725,306, filed on August 31, 2018. The multiplier-accumulator circuitry described and / or illustrated in the '345 and '306 applications facilitates linking multiplication and accumulation operations, thereby allowing multiple multiplier-accumulator circuitry to perform such operations more rapidly (see, e.g., U.S. Patent Application No. 16 / 545,345). Figure 1 A- Figure 1C). The '345 and '306 applications are incorporated herein by reference in their entirety.
[0014] Additionally or alternatively, the present invention may be used and / or implemented in conjunction with circuitry and techniques for multiplier-accumulator execution or processing pipelines (and methods of operating such circuitry) having circuitry and / or architecture for processing data concurrently or in parallel to increase the throughput of the pipelines, for example, as described and / or illustrated in U.S. Patent Application No. 16 / 816,164 and U.S. Provisional Patent Application No. 62 / 831,413; the '164 and '413 applications are incorporated herein by reference in their entireties. Here, multiple processing or execution pipelines may process data concurrently to increase the throughput of the data processing and overall pipeline.
[0015] It is worth noting that the integrated circuit(s) may be, for example, a processor, a controller, a state machine, a gate array, a system on a chip ("SOC"), a programmable gate array ("PGA"), and / or a field programmable gate array ("FPGA"), and / or a processor, controller, state machine, and SOC including an embedded FPGA. FPGA refers to both discrete FPGAs and embedded FPGAs. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention may be implemented in conjunction with the embodiments illustrated in the accompanying drawings. These drawings illustrate different aspects of the invention, and where appropriate, reference numerals, nomenclature, or names illustrating similar circuits, architectures, structures, components, materials, and / or elements are similarly labeled in different figures. It should be understood that various combinations of structures, components, materials, and / or elements other than those specifically illustrated are contemplated and fall within the scope of the present invention.
[0017] In addition, many inventions are described and illustrated herein. The present invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or permutation of these aspects and / or embodiments. In addition, each aspect of the present invention and / or its embodiments may be used alone or in combination with one or more of the other aspects of the present invention and / or its embodiments. For simplicity, certain permutations and combinations are not discussed and / or illustrated separately herein. It is worth noting that embodiments or implementations described herein as "exemplary" should not be understood as being preferred or advantageous, for example, relative to other embodiments or implementations; rather, they are intended to reflect or indicate that the embodiment(s) are (one or more) "example" embodiments.
[0018] It is important to note that the configurations, block / data widths, datapath widths, bandwidths, data lengths, values, processes, pseudo-codes, operations, and / or algorithms described herein and / or illustrated in the accompanying drawings, and the text associated therewith, are exemplary. Indeed, the present invention is not limited to any specific or exemplary circuit, logic, block, functional, and / or physical diagrams illustrated and / or described with respect to, for example, exemplary circuit, logic, block, functional, and / or physical diagrams, the number of multiplier-accumulator circuits employed in an execution pipeline, the number of execution pipelines employed in a particular processing configuration, the organization / allocation of memory, block / data widths, datapath widths, bandwidths, values, processes, pseudo-codes, operations, and / or algorithms.
[0019] Furthermore, while the illustrative / exemplary embodiments include multiple memories (e.g., L3 memory, L2 memory, L1 memory, L0 memory) assigned, allocated, and / or used to store certain data and / or in certain organizations, one or more memories may be added, and / or one or more memories may be omitted and / or combined / merged—for example, the L3 memory or L2 memory and / or organization may be changed, supplemented, and / or modified. The present invention is not limited to the illustrative / exemplary embodiments of memory organization and / or allocation set forth in the application. Again, the present invention is not limited to the illustrative / exemplary embodiments set forth herein.
[0020] Figure 1 A schematic block diagram of a logical overview of an exemplary embodiment of a multiplier-accumulator circuit system (including a plurality of multiplier-accumulator circuits (not individually illustrated)) of a plurality of multiplier-accumulator processing or execution pipelines (“MAC pipelines” or “MAC processing pipelines” or “MAC execution pipelines”) is illustrated in block diagram form in conjunction with memories that store input data, filter weights or coefficients, and output data; in one embodiment, each MAC pipeline includes a plurality of multiplier-accumulator circuits connected in series (although individual multiplier-accumulator circuits are not specifically illustrated herein); notably, Figure 1 Also illustrated is exemplary pseudo-code for a schematic block diagram of a logical overview of an exemplary embodiment of a MAC processing pipeline;
[0021] Figure 2AA schematic block diagram of a logical overview of an exemplary embodiment of a multiplier-accumulator circuit system (including a plurality of multiplier-accumulator circuits (not individually illustrated)) implementing a Winograd data processing technique for multiple multiplier-accumulator processing or execution pipelines ("MAC pipelines" or "MAC processing pipelines" or "MAC execution pipelines") is illustrated in conjunction with a memory that stores input data, filter weights or coefficients, and outputs data according to certain aspects of the present invention; in one embodiment, each MAC pipeline includes a plurality of multiplier-accumulator circuits connected in series (although individual multiplier-accumulator circuits are not specifically illustrated herein); notably, Figure 2A Also illustrated is exemplary pseudo-code for a schematic block diagram of a logical overview of an exemplary embodiment of a MAC processing pipeline implementing the Winograd data processing technique according to certain aspects of the present invention;
[0022] Figure 2B A schematic block diagram illustrates a physical overview of an exemplary embodiment of multiple MAC processing pipelines according to certain aspects of the present invention, wherein multiple multiplier-accumulator execution pipelines are configured to implement the Winograd technique for data processing; notably, each MAC pipeline includes multiplier-accumulator circuitry (illustrated in block diagram form), wherein each multiplier-accumulator circuitry includes multiple multiplier-accumulator circuits (not individually illustrated); in this illustrative exemplary embodiment, multiple (here, sixteen) multiplier-accumulator execution pipelines process 64x (4x4) input pixels / data at dij to determine the associated 64x (2x2) output pixels at yij;
[0023] Figure 2C According to some aspects of the present invention Figure 2B An exemplary timing overview diagram of a physical overview of an exemplary embodiment illustrated in FIG;
[0024] Figure 3A Schematic diagram of an implementation of Winograd processing techniques according to certain aspects of the present invention (e.g., in Figure 2A and Figure 2B1 , a schematic block diagram of an exemplary input data (e.g., image data) conversion circuitry and extraction circuitry for a multiplier-accumulator circuitry execution or processing pipeline (as illustrated in the logical and physical overview in FIG), including certain physical and logical overviews of the operation of an exemplary D to E (dij to eij) conversion and extraction circuitry and a MAC execution pipeline implementing a Winograd processing technique; in brief, an L3 / L2 memory is read to obtain a 4x4 pixel block of input data (Dij values); the 4x4 pixel block of input data is converted via the conversion circuitry in the dij to eij pipelines (in the illustrative embodiment, sixteen dij to eij pipelines) to a 4x4 block of Eij values, wherein each pipeline of the conversion and extraction circuitry is associated or linked to a MAC processing pipeline; thereafter, the 4x4 Eij block is extracted and shifted to a 16x64 In the MAC processing pipeline, the Eij value is multiplied by the filter coefficient (Hij value) stored in the L1 memory and accumulated to the Zij value; in this embodiment, each Eij value is used multiple times (typically 64 times or more) in the 16x64 MAC pipeline;
[0025] Figure 3B According to some aspects of the present invention Figure 3A Example pseudo code for the example D to E conversion circuitry and E extraction circuitry (dij to eij conversion and eij extraction embodiments) illustrated in FIG, wherein the D to E conversion and extraction are specifically identified therein;
[0026] Figure 3C FIGURE 1 illustrates a data flow diagram of exemplary input data (e.g., image data) conversion and extraction circuitry (D to E conversion circuitry and E extraction circuitry) implementing a plurality of pipelines (sixteen pipelines in the illustrative embodiment) of Winograd processing techniques according to certain aspects of the present invention, wherein each pipeline of the conversion and extraction circuitry is associated or related to a MAC processing pipeline (see FIGURE 2). Figure 3A ); In short, data (D IN) are input into the sixteen D / E conversion pipelines of the circuit system to convert these data from a floating point data format to a fixed point data format (e.g., a block-scaled data format); here, the input data / values (Dij data / values) all in floating point data format are processed and converted to a fixed point data format (e.g., an integer format (INT16) or a block-scaled format (BSF16)); the conversion circuit system includes a "P" unit or circuit (hereinafter, the units or circuits are collectively referred to as "units" in this document) to receive the Dij data / values in floating point data format and determine the maximum exponent (EMAX) of such input data; the "R" unit performs a time delay function / operation for exponent search; the "Q" unit performs floating point data format to fixed point data format conversion (specifically, FP16 to BSF16 format conversion of the Dij data / values using the EMAX value); thereafter, the fixed point data 16. The data in the fixed-point data format is input to the Winograd conversion circuitry (via the "X" unit); here, the "X00-X31" units perform D / E conversion of the Winograd processing technique using the fixed-point data format (e.g., BSF16) representation of the input data / value (Dij data / value); thereafter, the data / value is input to an additional conversion circuitry to generate data / value in a floating-point data format using the data / value in the fixed-point data format output by the Winograd conversion circuitry; here, the "S" unit converts the input data from the fixed-point data format to the floating-point data format using the EMAX value (in this illustrative embodiment, a BSP16 to FP16 format conversion of the Eij data / value using the EMAX value); it is noted that in the accompanying drawings, a particular "unit" is sometimes labeled or identified as "unit x," where x is P, R, Q, S, or X (e.g., Figure 4A Unit P in );
[0027] Figure 3D According to some aspects of the present invention Figure 3A and Figure 3C An exemplary overview and exemplary timing diagram of a data flow diagram of an exemplary D-to-E conversion circuit system and E extraction circuit system embodiment of multiple pipelines illustrated in FIG;
[0028] Figure 3E and Figure 3F Each illustrates an embodiment according to certain aspects of the present invention. Figure 3A and Figure 3C An overview of an exemplary D-to-E conversion circuit system and E extraction circuit system embodiment with multiple pipelines and an exemplary timing diagram of a data flow diagram are shown in FIG. Figure 3D the selected portion identified in ;
[0029] Figure 4A FIGURES illustrate some aspects of the present invention. Figure 3Ca schematic block diagram of an exemplary D-to-E conversion circuitry and an exemplary P-unit / circuit of an E extraction circuitry (hereinafter, the units / circuits (i.e., units or circuits) are collectively referred to herein as "units" or "units"), wherein in the exemplary embodiment, the P-units (16 P-units in this illustrative embodiment) include circuitry for identifying and / or determining a maximum exponent (EMAX) of input data / values (Dij data / values) in a floating point data format (FP16 format in this illustrative embodiment); in this embodiment, the EMAX register is loaded with the exponent of the first Dij value, and the exponent of each of the remaining Dij data / values is compared to EMAXP and, if greater, EMAXP is replaced;
[0030] Figure 4B FIGURES illustrate some aspects of the present invention. Figure 3C a schematic block diagram of an exemplary D-to-E conversion circuitry and an exemplary R-unit of the E extraction circuitry, wherein in the exemplary embodiment, the R-unit includes circuitry that temporally delays input data / values (Dij data / values) by a predetermined amount of time to identify and / or determine a maximum exponent (EMAX) of the input data / values (Dij data / values) before the input data / values (Dij data / values) are input to or provided to the floating-point to fixed-point conversion circuitry (in this illustrative embodiment, via the Q-unit);
[0031] Figure 4C FIGURES illustrate some aspects of the present invention. Figure 3C a schematic block diagram of an exemplary D-to-E conversion circuitry and an exemplary Q unit of the E extraction circuitry, wherein in the exemplary embodiment, the Q units (sixteen units in this illustrative embodiment) convert input data / values (Dij data / values) from a floating point data format to a fixed point data format using an EMAX value; in this illustrative embodiment, the Q units convert the Dij data / values from an FP16 data format to a BSF16 data format;
[0032] Figure 4D and Figure 4E According to some aspects of the present invention, Figure 3C a schematic block diagram of exemplary X units (X00-X31 units) of exemplary D-to-E conversion circuitry and E extraction circuitry (dij-to-eij conversion and extraction circuitry) of FIG. 1 , wherein the D-to-E conversion circuitry employs Winograd processing techniques to generate input data / values (Eij) using fixed-point data format representations of the input data / values (Dij);
[0033] Figure 4F FIGURES illustrate some aspects of the present invention. Figure 3CSchematic diagram of an exemplary D-to-E conversion circuit system and an exemplary S unit of the E extraction circuit system, wherein in the exemplary embodiment, the S units (sixteen units in total in this illustrative embodiment) use the EMAX value to perform data format conversion from a fixed-point data format to a floating-point data format after Winograd processing (in this illustrative embodiment, a BSF16 to FP16 format conversion of input data / values using the EMAX value) to generate input data / values (Eij) in a floating-point data format for pipeline processing by the MAC (see, for example, Figure 2B and Figure 3A );
[0034] Figure 5A Illustrated is some conversion circuitry (e.g., Q unit of D to E conversion circuitry - see FIG) for converting input values / data (Dij values / data) from a floating point data format (FP16 format in the illustrative embodiment) to a fixed point data format (BSF16 format in the illustrative embodiment) in accordance with an embodiment of the present invention. Figure 3C ) processing or operation; here, the BSF16 format allows each group of sixteen Dij values to adopt a common exponent (EMAX), thereby reducing and / or eliminating problems with alignment and normalization operations (which may be problems associated with adding values using the FP16 format); each Dij value can adopt additional bits that can be used for the fractional portion of the data - thereby increasing the precision of the BSF16 value; it is worth noting that both formats use the same or substantially the same amount of memory capacity and memory bandwidth; however, data in a fixed-point data format such as a block-scaled format (e.g., BSF16) can be added via two's complement integers; it is worth noting that the floating-point data format to fixed-point data format conversion circuitry can be used in conjunction with the filter weight (F to H) conversion circuitry and the output data (Z to Y) conversion circuitry Figure 5A The process or operation of the conversion circuitry described and illustrated in for converting values / data from a floating point data format (FP16 format in the illustrative embodiment) to a fixed point data format (BSF16 format in the illustrative embodiment) - although the bit width may be different; Figure 5A The processes or operations illustrated in are exemplary and not limiting. Other techniques than those described herein may be used to convert values / data from a floating-point data format to a fixed-point data format (here, BSF).
[0035] Figure 5BIllustrated is a diagram of certain conversion circuitry (e.g., S unit of D to E conversion circuitry—see FIG. 1 ) for converting input values / data (Eij values / data) in a fixed-point data format (BSF16 format in the illustrative embodiment) to input values / data in a floating-point data format (FP16 format in the illustrative embodiment) in accordance with an embodiment of the present invention. Figure 3C ) processing or operation; in this way, the input value / data (Eij value / data) is in a predetermined data format for the MAC to perform pipeline processing (e.g., Figure 2B and Figure 3A As described above, the two formats use the same or substantially the same amount of memory capacity and memory bandwidth; it is worth noting that the fixed-point data format to floating-point data format conversion circuitry in the filter weight (F to H) conversion circuitry and the output data (Z to Y) conversion circuitry can be used in conjunction with the fixed-point data format to floating-point data format conversion circuitry in the filter weight (F to H) conversion circuitry and the output data (Z to Y) conversion circuitry. Figure 5B The process or operation of the conversion circuitry described and illustrated in for converting values / data from a fixed-point data format (in the illustrative embodiment, a BSF16 format) to a floating-point data format (in the illustrative embodiment, an FP16 format) - although the bit width may be different; Figure 5B The processes or operations illustrated in are exemplary and not limiting. Other techniques than those described herein may be used to convert values / data from a fixed-point data format (here, BSF) to a floating-point data format.
[0036] Figure 6A Schematic diagram of an implementation of Winograd processing techniques according to certain aspects of the present invention (e.g., in Figure 2A and Figure 2B Schematic block diagram of a physical and logical overview of an exemplary filter weight data / value (F to H) conversion and extraction circuitry and operational embodiment of multiple multiplier-accumulator execution pipelines (as illustrated in the logical and physical overview in FIG), wherein each pipeline of the conversion and extraction circuitry is associated or associated with a MAC processing pipeline (see FIG. Figure 3A );
[0037] Figure 6B According to some aspects of the present invention Figure 6A Example pseudo code for an exemplary F to H (fij to hij) conversion and extraction embodiment of a multiplier-accumulator execution pipeline;
[0038] Figure 6C FIGURE 1 illustrates a data flow diagram of an exemplary F-to-H conversion circuit system and a plurality of filter weight conversion pipelines (sixteen in the illustrative embodiment) implementing Winograd processing techniques according to certain aspects of the present invention; in short, filter weight data / values (F IN) is input into the F to H conversion circuitry pipeline (sixteen pipelines in the illustrative embodiment) to convert the filter weight data from a floating point data format to a fixed point data format (e.g., a block scaled data format); here, the filter weight or coefficient data / values (Fij data / values), all in a floating point data format, are processed and converted into a fixed point data format (e.g., a block scaled format (BSF16)); the conversion circuitry includes a "K" unit / circuit to receive the Fij data / values in a floating point data format and determine or identify a maximum exponent (EMAX) of the filter weight or coefficient data / values; an "M" unit performs a time delay function / operation for an exponent search; an "L" unit performs a floating point data format to a fixed point data format conversion (specifically, an FP16 to BSF16 format conversion of the Fij data / values using an EMAXK value); thereafter, the data in the fixed point data format is input into the W Winograd processing / conversion circuitry (via "X" units); where the "X00-X31" units perform an F to H conversion of the Winograd processing technique using a fixed-point data format (e.g., BSF16) of the filter weight data / values (Fij filter weight data / values); thereafter, the filter weight data / values are input into additional conversion circuitry to convert the filter weights or coefficients (output by the Winograd processing circuitry) from the fixed-point data format to a floating-point data format; where the "N" unit uses the EMAXL values of the filter weights / coefficients to convert the data from the fixed-point data format to a floating-point data format (in this illustrative embodiment, a BSF16 to FP16 format conversion of the Hij data / values using the EMAXL values) and outputs the filter weights / coefficients to memory local to or associated with a particular MAC processing pipeline (see, e.g., Figure 6A ), memories such as L1 memory and L0 memory (e.g., SRAM); units / circuits are collectively referred to herein as "units" or "cells," and in the accompanying drawings, a specific "unit" is sometimes labeled or identified as "unit x," where x is K, L, M, N, or X (e.g., Figure 7A Unit K in ); it is worth noting that Figures 7A-7D An exemplary circuit block diagram is shown. Figure 6C 1 and 2. Units of a logic block diagram of an exemplary F-to-H conversion circuit system are illustrated in FIG.
[0039] Figure 6D FIGURES illustrate some aspects of the present invention. Figure 6ASchematic diagram of two exemplary units of an exemplary F to H (fij to hij) conversion circuit system of an execution pipeline of FIG; It is noteworthy that the fij to hij conversion circuit system in this exemplary embodiment includes sixteen left-side units and sixteen right-side units, wherein the fij to hij conversion circuit system includes (i) data registers for fij and hij weight values, (ii) control logic for sorting, and (iii) adder logic for conversion; In addition, Figure 6D further illustrating a schematic block diagram of an exemplary embodiment of multiplexer (mux) circuitry and adder circuitry of an exemplary fij to hij conversion circuitry according to certain aspects of the present invention;
[0040] Figure 7A FIGURES illustrate some aspects of the present invention. Figure 6C a schematic block diagram of an exemplary K unit of an exemplary F to H conversion circuitry of FIG. 1 , wherein in the exemplary embodiment, the K units (sixteen in this illustrative embodiment) include circuitry for identifying and / or determining a maximum exponent (EMAX) of filter weight data / values (Fij data / values) in a floating point data format (FP16 format in this illustrative embodiment); in this embodiment, an EMAXK register is loaded with the exponent of a first Fij value, and the exponent of each of the remaining Fij values is compared to EMAXK and, if greater, replaced by EMAXK;
[0041] Figure 7B FIGURES illustrate some aspects of the present invention. Figure 6C a schematic block diagram of an exemplary M unit of an exemplary F to H conversion circuitry of FIG. 1 , wherein in the exemplary embodiment, the M unit includes circuitry for temporally delaying the filtered weight data / values (Fij data / values) by a predetermined amount of time to identify and / or determine a maximum or maximal exponent (EMAX) of the filtered weight data / values (Fij data / values) before the filtered weight data / values (Fij data / values) are input to or provided to the floating point data format to fixed point data format conversion circuitry (in this illustrative embodiment, data format conversion is implemented via the L unit);
[0042] Figure 7C FIGURES illustrate some aspects of the present invention. Figure 6CFIG. 1 is a schematic block diagram of an exemplary L unit of an exemplary F-to-H conversion circuitry of FIG. 1 , wherein in the exemplary embodiment, the L units (sixteen units in this illustrative embodiment) use EMAXK values to convert filter weight data / values (Fij data / values) from a floating point data format to a fixed point data format (FP16 data format to BSF16 data format in this illustrative embodiment); thereafter, in accordance with certain aspects of the present invention, the filter weight data / values in the fixed point data format are applied to Winograd processing circuitry (e.g., Figure 6D ) to generate Figure 6C The filter weight data / value (Hij data / value) after Winograd processing;
[0043] Figure 7D FIGURES illustrate some aspects of the present invention. Figure 6C a schematic diagram of an exemplary N-unit of an exemplary F-to-H conversion circuit system, wherein in the exemplary embodiment, the N units (sixteen units total in this illustrative embodiment) perform data format conversion of filter weight data / values from a fixed-point data format to a floating-point data format (in this illustrative embodiment, BSF16 to FP16 format conversion of the filter weight data / values using the EMAXL value) after Winograd processing to generate filter weight data / values in a floating-point data format (Hij), which is thereafter applied to a MAC processing pipeline (e.g., stored in an SRAM memory and used by a multiplier circuit of the MAC processing pipeline);
[0044] Figure 8A Illustrated is certain conversion circuitry (e.g., N-unit of F / H conversion circuitry—see FIG. 1 ) for converting filter weight values / data (Hij values / data) from a fixed-point data format (BSF12 format in the illustrative embodiment) to a floating-point data format (FP16 format in the illustrative embodiment) in accordance with an embodiment of the present invention. Figure 6C ) processing or operation; In this way, the filter weight value / data (Hij value / data) is in a predetermined data format for MAC to perform pipeline processing (e.g., Figure 2B and Figure 6A It is worth noting that the fixed-point data format to floating-point data format conversion circuit system in the input data / value (D to E) conversion circuit system and the output data (Z to Y) conversion circuit system can be used in conjunction with the fixed-point data format to floating-point data format conversion circuit system. Figure 8A The process or operation of the conversion circuitry described and illustrated in for converting values / data from a fixed-point data format (in the illustrative embodiment, a BSF16 format) to a floating-point data format (in the illustrative embodiment, an FP16 format) - although the bit width may be different;
[0045] Figure 8B The embodiment of the present invention is shown in FIG. Figures 6A-6D and Figures 7A-7D The circuit system and data flow of the filter weight value / data are converted from a first fixed-point data format (e.g., INT8 format) to a second fixed-point data format (e.g., BSF12 format) by the conversion circuit system or operation; in short, the Fij value can be stored in a memory (e.g., L3 DRAM memory and / or L2 DRAM memory) in a first fixed-point data format that requires less memory capacity and memory bandwidth than other fixed-point data formats (e.g., INT8 requires less memory capacity and memory bandwidth than BSF12). SRAM memory); after the filter weight data / values are written to and stored in a memory "local" to the MAC processing pipeline (e.g., L1 or L0 memory - SRAM memory local to the associated pipeline), the format size of the filter weight data / values (Fij values) may be less important / significant; with this in mind, by allocating 72 bits to nine BSF7 values (block-scaled fractional format), 9 of which can be used for a shared exponent EMAX, the 3x3 blocks of 8-bit filter weight (Fij) values are converted and modified to provide a larger or better dynamic range without increasing the memory capacity footprint or memory bandwidth requirements; the BSF7 values are input to or provided to the conversion circuit system to convert the filter weights in the BSF7 data format to the BSF12 data format; thereafter, according to an embodiment of the present invention, the filter weights can be converted to a BSF12 data format via the conversion circuit system (e.g., N unit of the F / H conversion circuit system - see Figure C) as shown Figure 8A The filter weight values / data are processed as indicated in the flowchart to convert the filter weight values / data (Hij values / data) from a fixed-point data format (e.g., BSF12 format) to a floating-point data format (e.g., FP16 format) for use by the multiplier circuit of the MAC processing pipeline; in this manner, the filter weight values / data (Hij values / data) are in a predetermined data format for the MAC to perform pipeline processing (e.g., Figure 1 、 Figure 2A and Figure 2B );
[0046] Figure 9A Schematic diagram of an implementation of Winograd processing techniques according to certain aspects of the present invention (e.g., in Figure 2A and Figure 2B a schematic block diagram of a physical and logical overview of an exemplary processed output data (Z to Y or zij to yij) insertion and conversion circuitry and an operational embodiment of a plurality of multiplier-accumulator execution pipelines (as illustrated in the logical and physical overview in FIG), wherein each pipeline of the conversion and extraction circuitry is associated or related to a MAC processing pipeline;
[0047] Figure 9B According to some aspects of the present invention Figure 9A Example pseudo code for an exemplary Z to Y or zij to yij insertion and conversion embodiment of the execution pipeline of
[0048] Figure 9C FIGURE 1 illustrates a data flow diagram of exemplary Z-to-Y conversion circuitry and Z-insertion circuitry for multiple pipelines (sixteen in the illustrative embodiment) implementing the Winograd processing technique according to certain aspects of the present invention; in short, data from the MAC processing pipeline (Z IN ) are input to the circuit system of the 16×16 Z to Y conversion pipeline to convert these data from a floating point data format (e.g., FP32) to a fixed point data format (e.g., a block-scaled data format); here, by converting the Zij data / value (“Z IN ”) are input into the 16x16 array of “T” units, and the data / values (Zij data / values) all in floating point data format are processed and converted into fixed point data format (e.g., block scaled format (BSF32)), wherein, thereafter, the inserted Zij data / values are input into or provided to sixteen Z to Y conversion pipelines including “U” units to receive the data / values in floating point data format and determine the maximum exponent (EMAXU) of such input data; the “R” units perform time delay functions / operations for exponent search; the “V” units perform floating point data format to fixed point data format conversion (specifically, FP32 to BSF32 format conversion of the Zij data / values using EMAXU values); thereafter, the Zij output data in fixed point data format are input into the Winograd conversion circuitry (in one embodiment, the “X” units (i.e., X00-X31)); here, the “X00-X31” units 3. The unit performs a Z-to-Y conversion using a fixed-point data format (e.g., BSF32) representation of data / values (Zij data / values) via a Winograd processing technique to generate Yij output data / values; thereafter, the output data / values (Yij) are input into an additional conversion circuitry to generate output data / values in a floating-point data format (e.g., FP32) using the (Yij) data / values in the fixed-point data format from the Winograd conversion circuitry; here, the "W" unit converts the data from the fixed-point data format to the floating-point data format (in this illustrative embodiment, a BSF32 to FP32 format conversion of the Yij data / values using the EMAXV value) using the EMAXV value; it is noted that units / circuits are collectively referred to herein as "units" or "cells," and that in the accompanying drawings, particular "units" are sometimes labeled or identified as "unit x," where x is R, T, U, V, W, or X (e.g., Figure 10AUnit T in );
[0049] Figure 9D A schematic diagram illustrating exemplary embodiments of Z (Zij) insertion circuitry and Z to Y (zij to yij) conversion circuitry according to certain aspects of the present invention includes Figure 9A One unit of the zij insertion logic circuitry (left portion of the figure) and one unit of the zij to yij conversion circuitry (right portion of the figure) of the exemplary Z to Y (zij to yij) insertion / conversion circuitry of FIG.
[0050] Figure 10A FIGURES illustrate some aspects of the present invention. Figure 9A and Figure 9C Schematic block diagram of an exemplary T-cell of an exemplary Z-to-Y conversion circuitry and a Z-insertion circuitry, wherein in this exemplary embodiment, the T-cell (in this embodiment a 16×16 array) receives data (Z IN ) for insertion into an exemplary Z-to-Y conversion circuit system;
[0051] Figure 10B FIGURES illustrate some aspects of the present invention. Figure 9A and Figure 9C a schematic block diagram of an exemplary U unit of exemplary Z-to-Y conversion circuitry and Z insertion circuitry, wherein in the exemplary embodiment, the U units (sixteen in this illustrative embodiment) include circuitry for receiving the output of the T units and identifying and / or determining a maximum exponent (EMAX) of the output data / values (Zij data / values) in a floating point data format (FP32 format in this illustrative embodiment); where, in one embodiment, the EMAX register is loaded with the exponent (EMAXU) of the first Zij value, and the exponent of each of the remaining Zij values is compared to EMAXU and, if greater, replaced with EMAXU;
[0052] Figure 10C FIGURES illustrate some aspects of the present invention. Figure 9A and Figure 9C a schematic block diagram of an exemplary R-unit of exemplary Z-to-Y conversion circuitry and Z-insertion circuitry, wherein in the exemplary embodiment, the M-unit includes circuitry for temporally delaying data / values (Zij data / values) by a predetermined amount of time to identify and / or determine a maximum or maximum exponent (EMAXU) of the output data / values (Zij data / values) of the MAC processing pipeline before the output data / values (Zij data / values) are input to or provided to the floating point data format to fixed point data format conversion circuitry (in this illustrative embodiment, via the V-unit);
[0053] Figure 10D FIGURES illustrate some aspects of the present invention. Figure 9A and Figure 9C FIG. 1 is a schematic block diagram of an exemplary Z-to-Y conversion circuitry and an exemplary V-unit of a Z-insertion circuitry, wherein in the exemplary embodiment, the V-units (sixteen in this illustrative embodiment) convert data / values (Zij data / values) from a floating point data format to a fixed point data format (FP32 data format to BSF32 data format in this illustrative embodiment) using EMAXU values; thereafter, in accordance with certain aspects of the present invention, the data in the fixed point data format is applied to Winograd conversion circuitry (e.g., Figure 6D to generate Yij data / values; and
[0054] Figure 10E FIGURES illustrate some aspects of the present invention. Figure 9A and Figure 9C Schematic diagram of an exemplary Z-to-Y conversion circuit system and an exemplary W unit of the Z insertion circuit system, wherein in the exemplary embodiment, after Winograd processing, the W units (a total of sixteen units in this illustrative embodiment) use EMAXV values to perform data format conversion of data / values from fixed point to floating point (in this illustrative embodiment, BSF32 to FP32 format conversion of processed output data / values using EMAXV values) to generate output data (Yij) in floating point data format.
[0055] Once again, many inventions are described and illustrated herein. The invention is not limited to the illustrative exemplary embodiments, including with respect to: (i) the illustrated specific floating-point format(s), specific fixed-point format(s), block / data width or length, data path width, bandwidth, values, processes, and / or algorithms, or (ii) exemplary logical or physical overview configurations, exemplary circuitry configurations, and / or exemplary Verilog code.
[0056] Furthermore, the present invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or permutation of such aspects and / or embodiments. Each aspect and / or embodiment of the present invention may be used alone or in combination with one or more of the other aspects and / or embodiments thereof. For the sake of brevity, many such combinations and permutations are not separately discussed or illustrated herein. DETAILED DESCRIPTION
[0057] In a first aspect, the present invention relates to one or more integrated circuits having a plurality of multiplier-accumulator circuits connected or configured in one or more data processing pipelines (e.g., linear pipelines) that process input data (e.g., image data) using Winograd type processing techniques. The one or more integrated circuits also include (i) a conversion circuit system for converting the data format of the input data and / or filter coefficients or weights from a floating-point data format to a fixed-point data format and (ii) a circuit system for implementing Winograd type processing to, for example, increase the data throughput of the multiplier-accumulator circuit system and the processing. The one or more integrated circuits may also include a conversion circuit system for converting the data format of the input data and / or filter coefficient / weight data after processing by the Winograd processing circuit system from a fixed-point data format to a floating-point data format. The output of the Winograd processing circuit system is applied to or input into the conversion circuit system to convert or transform the input data and / or filter coefficient / weight data after Winograd processing from a fixed-point data format to a floating-point data format before such data is input to the multiplier-accumulator circuit system. In this manner, data processing performed by the multiplier circuits and the accumulation circuits of each of the one or more multiplier-accumulator pipelines is in a floating point format.
[0058] Notably, the input data and associated filter weights or coefficients may be stored in a memory in a floating point data format, where such data is provided to a conversion circuit system to convert or transform the data from the floating point data format into a fixed point data format (e.g., a block-scaled fractional format) that is convenient for implementation in or consistent with use of Winograd type processing.
[0059] In one embodiment, one or more integrated circuits include conversion circuitry or conversion circuitry for converting or processing input data from a floating-point data format to a fixed-point data format (sometimes referred to as D / E, DE, or D to E conversion logic or conversion circuitry). Thereafter, the input data in the fixed-point data format is input to circuitry for implementing Winograd-type processing to, for example, increase the data throughput of the multiplier-accumulator circuitry and processing. The output of the Winograd processing circuitry is applied to or input to the conversion circuitry to convert or transform the input data from the fixed-point data format to a floating-point data format before the input data is input to the multiplier-accumulator circuitry via the extraction logic circuitry. The input data that can be processed by the multiplier-accumulator circuitry of one or more multiplier-accumulator pipelines is in a floating-point format.
[0060] In another embodiment, one or more integrated circuits include a conversion circuit system (sometimes referred to as F / H, FH, or F to H conversion logic or circuit system) for converting filter weight or coefficient values / data into a fixed-point format for processing via Winograd processing circuit system. Thereafter, the output of the Winograd processing circuit system is input into the conversion circuit system to convert the filter weight data from a fixed-point data format to a floating-point data format before the data is input into the multiplier-accumulator circuit system via the extraction logic circuit system. The multiplier circuit of the multiplier-accumulator circuit system of the pipeline executed by the (one or more) multiplier-accumulators uses such weight values or data in a floating-point data format together with the input data (which is also in a floating-point data format). In this way, the multiplication operation of the multiplier accumulation operation is performed in a floating-point data format condition.
[0061] A multiplier-accumulator circuit system (and a method of operating such circuit system) of an execution or processing pipeline (e.g., for image filtering) may include floating-point execution circuit system (e.g., floating-point multiplier circuitry and accumulator / adder circuitry) that implements floating-point multiplication and floating-point addition based on inputs having one or more floating-point data formats. The floating-point data format may be user- or system-defined and / or may be one-time programmable (e.g., at the time of manufacture) or more than one-time programmable (e.g., (i) at power-up or via power-up, startup, or the conduction / completion of an initialization sequence / processing sequence, and / or (ii) in-situ or during normal operation of the multiplier-accumulator circuit system of the integrated circuit or processing pipeline). In one embodiment, the circuit system of the multiplier-accumulator execution pipeline includes an adjustable-precision data format (e.g., a floating-point data format). Additionally or alternatively, the circuit system of the execution pipeline may process data concurrently to increase the throughput of the pipeline. For example, in one implementation, the present invention may include multiple separate multiplier-accumulator circuits (sometimes referred to herein as "MACs") and multiple registers (including, in one embodiment, multiple shadow registers) that facilitate pipelining of multiplication and accumulation operations, where the circuitry performing the pipelining processes data concurrently to increase the throughput of the pipeline.
[0062] The present invention may also include conversion circuitry incorporated into the MAC pipeline or coupled to the output of the MAC pipeline to convert output data in a floating-point data format to a fixed-point data format, wherein the output data is input to the Winograd processing circuitry. After Winograd processing, the output data may be converted to a floating-point data format and, in one embodiment, stored in a memory. For example, the present invention employs conversion circuitry consistent with Winograd-type processing (sometimes referred to as Z / Y, ZY, or Z-to-Y conversion logic or circuitry) to generate processed output data (e.g., output image data). However, before the multiplier-accumulator circuitry of the multiplier-accumulator execution pipeline(s) outputs the processed data, such data is applied to or input to data format conversion circuitry and processing that converts or transforms the processed data in a floating-point data format into processed data having a data format (e.g., a fixed-point data format, such as BSF) that is convenient for implementation in or consistent with the use of Winograd-type processing. Here, the processed data in fixed-point data format is processed by ZY conversion circuitry consistent with Winograd-type processing. Notably, in one embodiment, the ZY conversion circuitry is incorporated into or integrated into the multiplier-accumulator execution pipeline(s). That is, the ZY conversion operation of the processed data is performed in the MAC execution / processing pipeline.
[0063] Thereafter, the present invention may employ circuitry and techniques for converting or transforming output data (in a fixed-point data format) after Winograd processing from the fixed-point data format to a floating-point data format. That is, in one embodiment, after the processed output data has a data format that is convenient for implementation in or consistent with the use of Winograd-type processing (e.g., a fixed-point data format such as a block-scaled fractional format), the present invention may employ conversion circuitry and techniques for converting or transforming the data (e.g., pixel data) to a floating-point data format before output to a memory (e.g., an L2 memory such as an SRAM). In one embodiment, the output data (e.g., image data) is written to and stored in the memory in a floating-point data format.
[0064] refer to Figure 1 and Figure 2A-2CAs shown in the logical overview illustrated in FIG, in one exemplary embodiment, input data (e.g., image data / pixels in floating point format) is stored in memory in layers consisting of a two-dimensional array (e.g., MxM, where M=3) of input or image data / pixels. The input data / values (Dij) can be organized and / or stored in memory in a "depth" plane or layer (e.g., K depth layers, where in one embodiment, K=64), and the output data / values (Yij) are organized and / or stored in memory in an output "depth" plane or layer (e.g., L output depth planes, where in one embodiment, L=64) after processing. The memory storing the input data / values and the output data / values (which can be the same physical memory or different physical memories) also stores filter weight data / values (e.g., in floating point format) associated with the input data. In one embodiment, the filter weight or coefficient data / values (Fkl) are organized and / or stored in memory in M×M blocks or arrays, where the KxL blocks or arrays cover combinations (e.g., all combinations) of input and output layers.
[0065] One or more (or all) of the data can be stored in a floating point data format, input into the circuit system of the processing pipeline, and / or output from the circuit system of the pipeline. That is, in one embodiment, the input data (e.g., image data / pixel) is stored in a memory and / or input into the circuit system of the processing in a floating point data format. In addition or in lieu thereof, the filter weights or coefficient data / values used in the multiplier accumulation process are stored in a memory and / or input into the circuit system of the processing in a floating point data format. In addition, the processed output data as set forth herein can be stored in a memory and / or output in a floating point data format. All combinations of data stored in a memory, input into the circuit system of the processing pipeline, and / or output from the pipeline in a floating point data format are intended to fall within the scope of the present invention. However, in the discussion below, input data (e.g., image data / pixels), filter weight data / values, and processed output data are stored in floating-point data format, input into or accessed through the multiplier-accumulator circuits of the processing pipeline in floating-point data format, and output from the processing pipeline in floating-point data format.
[0066] It is noteworthy that the pipelined multiplier-accumulator circuitry includes a plurality of separate multiplier-accumulator circuits (sometimes referred to herein and / or in the figures as "MACs" or "MAC circuits") that pipeline the multiplication and accumulation operations. In the illustrative embodiments described herein (text and figures), the pipeline of multiplier-accumulator circuits is sometimes referred to or labeled as a "MAC execution pipeline," a "MAC processing pipeline," or a "MAC pipeline."
[0067] In a first operating mode of the multiplier-accumulator circuitry, a single execution pipeline is employed to accumulate 1x1 pixel output values in a single output layer by aggregating the sum of KxMxM multiplications of input data values and associated filter weight values from K layers of input data. In short, in the first operating mode, 3x3 (MxM) multiplications and accumulations are performed by the NMAX execution pipeline of the multiplier-accumulator circuitry to obtain Vijkl values (see Figure 1 - "ΣM" notation). Here, the MAC execution pipeline further performs accumulations of the input planes (see, index K) to obtain Yijl values (see "ΣK" notation). The result of these accumulation operations is a single pixel value Yijl (1x1), which is written to the output plane (in parallel with other single pixels written to other output planes with other output depth values (e.g., "L" index value)).
[0068] In a second mode of operation of the multiplier-accumulator circuitry, an NxN execution pipeline is employed wherein the two-dimensional array of input or image data / pixels is transformed from an MxM array (e.g., M=3) to an NxN array (e.g., N=4). (See Figure 2A-2C ). In this exemplary embodiment, the D to E conversion logic operates (via the D to E conversion circuitry) to generate, among other things, an NxN array of input or image data / pixels. (See Figures 3A-3D ). Similarly, the two-dimensional array of filter weights or weight values is transformed or converted from an MxM array (e.g., M=3) to an NxN array (e.g., N=4) via the F-to-H conversion circuitry. The NxN array of filter coefficients or filter weights correctly correlates with the associated input data / values.
[0069] In one embodiment, the F-to-H conversion circuitry is disposed between a memory and the multiplier-accumulator execution pipeline. In another embodiment, the memory stores an N×N array of filter weights or weight values, which are pre-calculated (e.g., off-chip) and stored in the memory as an N×N array of filter weights or filter weight values. In this manner, the F-to-H conversion circuitry is not disposed between the memory and the multiplier-accumulator execution pipeline and / or performs / implements an F-to-H conversion operation on such data immediately prior to employing the filter weights in the multiplier-accumulator execution pipeline. However, storing an N×N array of pre-calculated filter weight values / data in memory (rather than having the multiplier-accumulator circuitry / pipeline calculate such values / data during operation (i.e., on the fly)) may increase the memory required to store such filter weights or weight values, which in turn may increase the capacity requirements of the memory employed in this alternative embodiment (e.g., an increase that may be on the order of N×N / M×M, or approximately 16 / 9 in this exemplary embodiment).
[0070] Continue to refer Figure 2A and Figure 2B In a second mode of operation of the multiplier-accumulator circuitry, the MAC processing pipeline implements or performs accumulation of values Uijklm (from Eijk*Hklm multiplications) from the input plane (index K) into Zijlm values, as shown by the ΣK notation - where the NxN (e.g., 4x4) multiplications are substituted or replaced Figure 1 In the second mode of operation, each MAC processes or executes a pipeline implementing or performing multiplications and accumulations (sixteen in the illustrative embodiment), after which the processed data (Zij) is output to the zij to yij conversion circuitry, where after further processing, four data / pixels are output Yijl (2x2) and written to the output (memory) plane (e.g., in parallel with the other Yijl 2x2 pixels being written to other output planes (other L index values)).
[0071] In short, reference Figure 2C , according to various aspects of the present invention Figure 2B The exemplary timing diagram of the circuit system in the exemplary physical overview in FIG illustrates the operation of 16 parallel multiplier-accumulator execution pipelines of the multiplier-accumulator circuit system and the connection paths to the memory. Here, each pair of waveforms illustrates the first and last of the 16 pipelines, with similar behavior to the middle 14 pipelines of the exemplary multiplier-accumulator execution pipeline of the multiplier-accumulator circuit system. Each operation group processes a 16x64 input data word corresponding to a 4x4 D block in each of the 64 layers - where each pipeline uses 64 clock cycles (e.g., each cycle can be 1ns in this exemplary embodiment). The top waveform illustrates the D block moving from the memory (e.g., L2 SRAM memory) to the D to E conversion operation through the D to E conversion circuit system. This transfer step has a pipeline latency of 16ns; the conversion step can begin when 1 / 4 of the data is available.
[0072] It is worth noting that some stages have 16ns pipeline latency and 64ns pipeline cycle rate; in other words, each stage can accept a new 16x64 word operation every 64ns interval, but may overlap the processing of the next stage by 48ns. The D-to-E conversion operation produces a 4x4 E block. The extraction logic circuitry separates the 16 Eij data / values into each 4x4 block, passing each to one of the 16 execution pipelines. The 64ns of the 4x4 E block require 64ns to shift in - this stage (and the two stages below) have pipeline latency and pipeline cycle time that are the same.
[0073] Continue to refer Figure 2C, when the E block has been shifted into the execution pipeline of the multiplier-accumulator circuitry, the pipeline performs 16x64x64 MAC operations (labeled "MAC Operations"); in each of the 16 pipelines, the 64 multipliers and 64 adders perform one operation per nanosecond with a 64ns interval. This is an accumulation on the "K" and "L" indices on the input and output planes. The 64ns of the 4×4 Z block requires a 64ns shift out - this stage can overlap the Z to Y insertion stage by 48ns. Similarly, the Z to Y conversion stage can overlap the L2 write stage by 48ns. Each 2x2 pixel block consumes 64ns of pipeline cycle time - the next 2x2 block is shown in dark gray in the timing waveform. Therefore, in this example, processing all 128k pixels will require 1ms (~1 million ns). In this exemplary embodiment, the entire 16x64 word operation has a pipeline latency of 18x16ns, or 288ns. In this exemplary illustration, the pipeline latency of 288 ns is approximately 3,472 times smaller than the total operation latency of 1 ms, and therefore has a relatively small impact on the overall throughput of the system.
[0074] refer to Figure 2A 、 Figure 2B and Figure 3A In one embodiment, input data conversion circuitry (sometimes referred to as D / E, DE, or D to E conversion logic or conversion circuitry) is employed to (i) convert input data / values from a floating-point data format to a fixed-point data format, (ii) thereafter, the input data in the fixed-point data format is input to circuitry implementing Winograd type processing, and (iii) convert or transform the input data output by the Winograd processing circuitry from the fixed-point data format to a floating-point data format before inputting the input data (e.g., pixel / image data) to the multiplier-accumulator circuitry via extraction logic circuitry. Here, the input data / values processed by the multiplier-accumulator circuitry of one or more multiplier-accumulator pipelines are in floating-point format.
[0075] refer to Figure 3AIn this embodiment, a 4×4 pixel / data block of input data / values (Dij) is read from L3 / L2 memory and input to the D-to-E conversion circuitry (sixteen dij to eij pipelines in the illustrated embodiment) to generate a 4×4 pixel / data block of input data / values (Eij). The D-to-E conversion circuitry extracts the 4×4 Eij block of data via the extraction circuitry and shifts the input data / values into the MAC processing pipeline (a 16×64 array of MAC pipelines in the illustrated exemplary embodiment). As discussed in more detail below, the Eij data / values are multiplied by Hij values (input from L1 / L0 memory) via the multiplier circuitry of the MAC processing pipeline and accumulated as Zij data / values via the accumulator circuitry of the MAC processing pipeline. Notably, in operation, each Eij value / data is used multiple times (typically 64 times or more) within the 16×64 array of multiplier-accumulator execution pipelines of the MAC pipeline array.
[0076] Figure 3C is a schematic block diagram of a logic overview of a D-to-E conversion logic / circuitry system according to certain aspects of the present invention, wherein input data (e.g., image or pixel data) is first converted from a floating point format to a fixed point format (in the illustrative example, a fractional format scaled by block) via a Q unit, the fixed point format being a data format that is convenient for implementation in or consistent with the use of Winograd type processing (e.g., a fixed point format, such as a fractional format scaled by block). The input data is then processed via a Winograd conversion circuitry (X unit) to implement the Winograd type processing. Thereafter, according to the present invention, the data is converted from the fixed point data format to a floating point format for implementation in the multiplier-accumulate processing of the MAC processing pipeline. Notably, Figure 3D According to some aspects of the present invention Figure 3C An exemplary timing diagram overview of the overview illustrated in .
[0077] refer to Figure 3C and Figure 3D-3FIn one exemplary embodiment, the D to E conversion logic receives input data from a memory (e.g., an L2 SRAM memory) that stores 4x4 D blocks of data in a floating point data format. In one embodiment, the memory is partitioned or divided into 16 physical blocks so that sixteen sets of image data (i.e., 4x4 D blocks of data) can be accessed in parallel for use by the sixteen multiplier-accumulator execution pipelines of the multiplier-accumulator circuitry. Each set of data consists of four 4x4 D blocks of data that can be read or accessed in 64 words from each physical block of memory. The D to E conversion logic includes conversion circuitry that converts or transforms data (in this illustrative embodiment, read from the memory) from a floating point data format to a data format that is convenient for implementation in or consistent with use with Winograd type processing (e.g., a fixed point data format, such as a block-scaled fractional format). (See Figure 3C 、 Figure 4A 、 Figure 4B and Figure 4C ). The data is then processed by the Winograd processing circuit system of the D to E (dij to eij) conversion circuit system. (See, Figure 4D and Figure 4E ).
[0078] refer to Figure 3C 、 Figures 4A-4C , data (D IN ) are input into the 16 D / E conversion pipelines of the circuit system to convert this data from a floating point data format to a fixed point data format (e.g., a block-scaled data format). In this exemplary embodiment, the input data / values (Dij data / values), all in floating point data format, are processed and converted to a fixed point data format (e.g., an integer format (INT16) or a block-scaled format (BSF16)). The P units / circuits (hereinafter collectively referred to as units) of the input data conversion circuit system receive the Dij data / values in floating point data format and determine the maximum exponent (EMAX) of such input data. The 16x16 array of R units implements or performs a time delay function / operation to accommodate the exponent search (i.e., maximum exponent - EMAX). The Q units use the EMAX value to convert the input data / values (Dij) from a floating point data format (FP16) to a fixed point data format (BSF16).
[0079] Thereafter, the input data (Dij) in fixed-point data format is input into the Winograd transformation circuitry (via the "X" unit). Figure 3C 、 Figure 4D and Figure 4E, the X00-X31 units of the input data conversion circuitry perform or implement Winograd processing using a fixed-point data format (e.g., BSF16) representation of the input data / value (Dij data / value). The Winograd processing circuitry of the input data conversion circuitry (D to E conversion circuitry) in this exemplary embodiment includes (i) sixteen left-side units (i.e., X00 to X15 units), (ii) data registers for dij and eij data words, (iii) control logic for sequencing operations for processing, and (iv) adder logic for conversion. In this exemplary embodiment, the eij extraction logic circuitry further includes (i) sixteen right-side units (i.e., X16 to X31 units), (ii) data registers for dij and eij data words, (iii) control logic for sequencing operations, and (iv) adder logic for conversion.
[0080] refer to Figure 3C and Figure 4F After Winograd processing, but before the data is input to the multiplier-accumulator circuit of the MAC processing pipeline, the input data (Eij) is converted from a fixed-point data format to a floating-point data format via the S-unit. Here, the S-unit converts the input data from a fixed-point data format to a floating-point data format using the EMAX value (in this illustrative embodiment, a BSF16 to FP16 format conversion of the Eij data / value using the EMAX value). Furthermore, the extraction logic circuitry outputs the input data (Eij) in floating-point data format to the multiplier-accumulator circuit of the MAC execution pipeline. In this manner, the data processed by the multiplier and accumulator circuits of the MAC pipeline is in floating-point format.
[0081] Thus, in this exemplary embodiment, the conversion circuitry converts the 4x4D block of data into a fixed point format ( Figure 3C and Figures 4A-4C ), which is then processed through the Winograd circuit system ( Figure 3C 、 Figure 4D and Figure 4E ) is converted to a 4×4E block of data. Then, in this illustrative embodiment, the 4×4E block of data is converted from a fixed-point data format to a floating-point data format, wherein the 4×4E block is separated into sixteen streams classified by the respective Eij data / data of the 4×4E block of data. (See Figure 3C 、 Figure 4E and Figure 4F ). Separately, the eij extraction logic circuit system (sixteen eij extraction logic circuits in this illustrative embodiment - see Figure 4EEach of the sixteen eij streams may involve an “e-shift DIN” circuit block in one of the sixteen MAC execution pipelines of the multiplier-accumulator circuitry.
[0082] It is worth noting that, with reference to the present application Figure 3C The D to E (dij to eij) conversion circuitry also includes a vertical “E” that carries the extracted data input (eij) in floating point data format to the associated or appropriate MAC execution pipeline of the multiplier-accumulator circuitry. OUT Output( Figure 4F Here, the circuit system of unit S (see Figure 4F ) converts the data to floating point format. In an alternative embodiment, the Figure 4E Some dij to eij conversion processing is implemented or performed in the eij extraction unit.
[0083] As mentioned above, Figures 4A-4F The circuit block diagram shows the Figure 3C Details of the units / circuits of the logic block diagram illustrated in FIG for converting input data (which may be stored in a memory (e.g., a layer consisting of a two-dimensional array of image data / pixels)) from a floating point data format to a fixed point data format (e.g., a block scaled fractional (BSF) format). Figure 4F The circuit block diagram shows the Figure 3C Detail of the elements of the logic block diagram for converting Winograd transformed data in a fixed-point data format (e.g., Block Scaled Fractional (BSF) format) to a floating-point data format is shown in FIG. In this manner, input data input to the multiplier-accumulator execution pipeline for processing by the multiplier-accumulator circuitry is in a floating-point data format, and as such, multiplication and accumulation operations are performed by the multiplier-accumulator circuitry under the floating-point data condition.
[0084] Figure 5A Illustrated is an exemplary process or steps for converting or changing input data / values (Dij) from a floating point data format to a block-scaled fractional data format. In an exemplary embodiment, the data width is 16 - thus, FP16 to BSF16 conversion. Figure 3C 、 Figures 4A-4F and Figure 5AIn one embodiment, the input data is converted from a floating point data format to a fixed point data format (BSF in this illustrative embodiment), and in this way, each set of input data / values (16 Dij values) can use a common exponent (EMAX), thereby eliminating and / or reducing the alignment and normalization operations required to add values using the FP16 format. In addition, each input data / value (Dij) can then have an additional bit available for the fraction field of the data / value - thereby potentially increasing the precision of the BSF16 value.
[0085] Continue to refer Figure 5A , the processing or operation of the conversion circuitry for converting data / values from a floating point data format (FP16 format in the illustrative embodiment) to a fixed point data format (BSF16 format in the illustrative embodiment) includes determining a maximum exponent of the data / value (e.g., by comparing the exponent of each data / value (e.g., on a rolling basis). Additionally, the fraction field is right-shifted for each data / value having a smaller exponent. The fraction field of each data / value is rounded to conform to the BSF precision of the fraction field (which may be predetermined). Finally, the processing may include a two's complement operation (inverting and incrementing bits) where the data / value is negative.
[0086] As described above, the filter weight (F to H) conversion circuit system and the output data (Z to Y) conversion circuit system can be respectively adopted. Figure 5A The processing or operation described and illustrated in the foregoing is used to convert the filter weight data / values and output data / values from a floating point data format to a fixed point data format - although the bit width may be different. In addition, Figure 5A The processes or operations illustrated in are exemplary and not limiting. Other techniques than those described herein may be used to convert values / data from a floating-point data format to a fixed-point data format (here, BSF).
[0087] refer to Figure 3C 、 Figure 4F and Figure 5B In one embodiment, after the pipeline of the D / E conversion circuitry generates sixteen Eij data / values, the Winograd-processed input data / values (Eij) are converted to floating-point format before being input to the multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry. Here, the multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry performs a multiplication operation of a multiplier-accumulate operation.
[0088] Simply put, Figure 5BThe diagram illustrates the process or steps used to change or convert input data / values (Eij) from a block-scaled fractional (BSF) format to a floating-point (FP) format (in the illustrative embodiment, a BSF16 format to an FP16 format). Again, although both formats can utilize the same amount of memory capacity and memory bandwidth, in this embodiment, the D / E conversion pipeline is modified with circuitry to convert the data from fixed-point to a floating-point format (e.g., a BSF16 to FP16 change / conversion). In this manner, the input data (e.g., image / pixel data) processed by the circuitry of the multiplier-accumulator pipeline is in a floating-point format.
[0089] Continue to refer Figure 5B The processing or operation of the conversion circuitry for converting data / values from a fixed-point data format (BSF16 format in the illustrative embodiment) to a floating-point data format (FP16 format in the illustrative embodiment) includes determining whether the data / value is negative and, if so, performing a two's complement operation on the fraction field (inverting and incrementing the bits). Additionally, the processing may include performing a priority encoding operation on the fraction field of the BSF data / value and a left shift operation on the fraction field of negative values / data. The fraction field of each data / value may be rounded relative to the shifted fraction field. Finally, the processing generates an exponent field using the maximum exponent of the data / value—e.g., from the left shift amount.
[0090] As described above, the filter weight (F to H) conversion circuit system and the output data (Z to Y) conversion circuit system can be respectively adopted in the conversion circuit system. Figure 5B The processing or operation described and illustrated in the foregoing is used to convert the filter weight data / values and output data / values from a fixed-point data format to a floating-point data format - although the bit width may be different. In addition, Figure 5B The processes or operations illustrated in are exemplary and not limiting. Other techniques than those described herein may be used to convert values / data from a fixed-point data format (here, BSF) to a floating-point data format.
[0091] refer to Figure 6A and Figure 6CIn another embodiment, the present invention includes a filter weight conversion circuit system (sometimes referred to as F / H, FH, or F to H conversion logic or circuit system). In one embodiment, the filter weight conversion circuit system (i) converts the filter weight values / data from a floating point data format to a fixed point data format, (ii) thereafter, the filter weight values / data in the fixed point data format are input into the circuit system that implements Winograd type processing, and (iii) converts or transforms the filter weight values / data as the output of the Winograd processing circuit system from the fixed point data format to the floating point data format. After converting the filter weight values / data from the fixed point data format to the floating point data format, the filter weight values / data in the floating point data format are input into the multiplier-accumulator circuit system. The multiplier circuit of the multiplier-accumulator circuit system that performs the pipeline by (one or more) multipliers-accumulators uses such weight values / data together with the input data (which is also in the floating point data format). In this way, the multiplication operation of the multiplier accumulation operation is performed under the floating point data format condition.
[0092] refer to Figure 6A In one embodiment, a memory (e.g., L2 SRAM memory) stores 3×3F blocks of data for filter weight data / values (e.g., FIR filter weights). In this illustrative exemplary embodiment, the memory is partitioned or divided into sixteen physical blocks so that sixteen groups of data can be read or accessed in parallel by or for use by the sixteen multiplier-accumulator execution pipelines of the multiplier-accumulator circuitry. Each group of data consists of four 3×3 F blocks of data, requiring 36 accesses from each physical L2 block in this illustrative exemplary embodiment, each access taking 1 ns. The 3×3 F blocks of data are converted into 4×4 H blocks of data by F-to-H (fkl-to-hkl) conversion circuitry (sixteen pipelines in this illustrative embodiment). The 4x4 H-block of filter weight data can be written to a memory (e.g., L1 SRAM memory) shared by the associated multiplier-accumulator execution pipelines (sixteen pipelines in this illustrative embodiment) of the multiplier-accumulator circuitry. In one embodiment, each of the sixteen filter weight data / values (hkl) of the 4x4 H-block of data is written and stored in a memory (e.g., L0 SRAM memory) local to and associated with one of the sixteen multiplier-accumulator execution pipelines for use by its multiplier circuitry.
[0093] In one embodiment, this sorting and organization is achieved and performed by a memory addressing sequence when reading the filter weight data / values (hkl) in L1 memory and writing the filter weight data / values (hkl) to a local and associated memory (e.g., L0 memory associated with a particular multiplier-accumulator execution pipeline). However, alternatively, the sorting and organization may be performed via extraction logic circuitry (in the filter weight data / values (hkl), similar to Figure 3A and Figure 4E The exemplary extraction circuitry and operations of the D to E (dij to eij) conversion circuitry illustrated in FIG. 2 are implemented and categorized and organized. The weight values or data may be read from memory once and transferred to the MAC processing pipeline of the multiplier-accumulator circuitry and then reused to process each of the blocks (e.g., thousands) of 2x2 input / image data or pixels.
[0094] Figure 6C is a logical overview of an F to H conversion circuitry according to certain aspects of the present invention, wherein filter weight data / values are first converted from a floating point format to a fixed point format (in the illustrative example, a block-scaled fractional format or an integer format), which is a data format that is convenient for implementation in or consistent with the use of Winograd type processing. The filter weight (F to H) conversion circuitry may employ Figure 5A The processing or operation described in to convert the filter weight data from a floating point data format to a fixed point data format.
[0095] Continue to refer Figure 6C , the filter weight data / values are then processed through the Winograd conversion circuitry to implement Winograd type processing. Thereafter, the data is converted from fixed point data format to floating point data format for implementation in the multiplier-accumulator processing of the MAC processing pipeline. As discussed herein, the filter weight (F to H) conversion circuitry may employ Figure 5B and / or Figure 8A The processing or operation described in to convert the filter weight data from a fixed-point data format to a floating-point data format.
[0096] Continue to refer Figure 6C , filter weight data / value (F IN) is input into the F to H conversion circuit system of the pipeline (in the illustrative embodiment, sixteen F to H conversion pipelines) to convert the filter weight data from a floating point data format to a fixed point data format (e.g., a block-scaled data format). The filter weight or coefficient data / values (Fij data / values), all in floating point data format, are processed and converted to a fixed point data format (e.g., a block-scaled format (BSF16)) (via the K unit / circuit, the M unit / circuit, and the L unit / circuit). The K unit / circuit receives the Fij data / values in floating point data format and determines or identifies the maximum exponent (EMAX) of the filter weight or coefficient data / values. The M unit / circuit performs a time delay function / operation to be related to the time used to perform or implement the exponent search mentioned above. The L unit performs a floating point data format to a fixed point data format conversion (in this illustrative embodiment, an FP16 to BSF16 format conversion of the Fij data / values using the EMAXK value).
[0097] Thereafter, the data in fixed-point data format is input into the Winograd processing / conversion circuit system. Here, the "X00-X31" units use a fixed-point data format (e.g., BSF16) representation of the filter weight data / value (Fij filter weight data / value) to perform the F to H conversion of the Winograd processing technique. After the Winograd conversion circuit system processes the filter weight data / value, the filter weight data / value (Hij) is input into the conversion circuit system to convert the filter weights or coefficients from a fixed-point data format (e.g., BSF16 or INT16) to a floating-point data format (FP16). In this regard, the N unit converts the filter weight data / value (Hij) from a fixed-point data format to a floating-point data format using the maximum exponent (EMAXL value) of the filter weight / coefficient (in this illustrative embodiment, the BSF16 to FP16 format conversion of the Hij data / value using the EMAXL value). In one embodiment, the filter weight data / value (Hij) in floating-point data format is output to a memory local to or associated with one or more specific MAC processing pipelines (see, for example, Figure 6A ), for example, L1 memory and L0 memory (for example, SRAM).
[0098] Figure 6D The diagram illustrates details of the units of an exemplary F / H conversion pipeline with respect to the weight data or values of the Winograd conversion circuitry. In particular, Figure 6D The diagram shows details of two units of the fkl to hkl conversion circuit system according to certain aspects of the present invention. It is worth noting that in this illustrative embodiment, no circuits such as Figure 4EThe F to H (fkl to hkl) conversion circuitry in this exemplary embodiment includes sixteen left-side units and sixteen right-side units. In addition, the F to H (fkl to hkl) conversion circuitry includes (i) data registers for fkl and hkl weight values, (ii) control logic for sorting, and (iii) adder logic for conversion.
[0099] Continue to refer Figure 6D , after the weight data or value is converted from the floating point data format to the fixed point data format, the filter weight data / value (Fij) is shifted into the X00-X31 units for processing by the Winograd conversion circuit system and is held or temporarily stored in the FREG_X and FREG_Y registers. The sixteen filter weight data / values (Hij) are shifted through these 32 positions and the appropriate Fij data / values are accumulated into the Winograd representation of the filter weight data / value (Hij). In this illustrative embodiment, the sixteen Hij data / values are then shifted to the conversion circuit system (N unit) where the data format of the filter weight data / value (Hij) is converted from the fixed point data format to the floating point data format using the maximum exponent (EMAXL value) of the filter weight / coefficient (in this illustrative embodiment, the BSF16 to FP16 format conversion of the Hij data / value using the EMAXL value). (See Figure 6C ).
[0100] Thereafter, the F to H conversion circuitry shifts or outputs the filter weight data / values (Hij) in floating point data format to an L1 SRAM memory which, in one embodiment, is shared by multiple multiplier-accumulator execution pipelines of the multiplier-accumulator circuitry. Figure 6A The filter weight data / values (Hij) in floating point data format may then be output and written to a memory local to or associated with one or more specific MAC processing pipelines, such as an L0 memory (e.g., SRAM). In one embodiment, each MAC processing pipeline is associated with an L0 memory, where the local memory is dedicated to the circuitry / operation of the associated MAC processing pipeline.
[0101] It is worth noting that Figure 6DThe embodiment illustrated in FIG includes an accumulation path for the filter weight data / value (hkl) after processing by the Winograd processing circuitry with 10 bits of precision, and utilizes a saturating adder to handle overflow. An alternative embodiment may utilize an accumulation path for the filter weight data / value (hkl) with 12 bits of precision, and not utilize a saturating adder to handle overflow; here, the 12-bit accumulation path has sufficient numerical range to avoid overflow.
[0102] refer to Figure 6C , the F to H conversion circuitry for filter weight data / values is similar to the conversion circuitry described in conjunction with the D to E conversion circuitry (and conversion pipeline). In short, the K unit receives filter weight data / values (Fij) in a floating point data format (FP16), evaluates the data / values to determine or identify the maximum exponent (EMAXK) of the filter weight or coefficient. The array of M units (16×16 in the illustrative embodiment) implements or performs a time delay function / operation to accommodate an exponential search or is associated with an exponential search (i.e., maximum exponent - EMAXK). The L unit converts the filter weight data / value (Fij) from a floating point data format (FP16) to a fixed point data format (BSF16) using the EMAXK value received from the K unit. The Winograd conversion circuitry of the F to H conversion circuitry is implemented by the X00-X31 units using the filter weight data / values (Fij) in a fixed point data format (BSF16). The N-unit receives the output of the Winograd transformation circuitry (i.e., the filter weight data / value (Hij)) and converts the data / value from a fixed-point data format to a floating-point data format (e.g., BSF16 to FP16 format conversion of the Hij data / value) using the EMAXL value. In this manner, the filter weight data or value provided to or input to the multiplier-accumulator circuitry in the multiplier-accumulator execution pipeline is in floating-point format, and thus the multiplication operations performed by the multiplier-accumulator circuitry are in floating-point condition.
[0103] It is noteworthy that in the context of filter weight data / values, the data format conversion from floating point format to fixed point format can be similar to the data format conversion described above in conjunction with the D to E conversion circuitry and pipeline in the context of data format conversion of input data (e.g., image data / pixels). However, in one embodiment, after the F to H conversion pipeline generates sixteen Hij values, the following can be used: Figure 8A The fixed-point format to floating-point format (eg, BSF12 format to FP16 format) is implemented as illustrated in the exemplary process illustrated in FIG. Here, Figure 8AThe diagram shows additional processing for changing the Hij values from BSF12 format to FP16 format. The F to H conversion circuitry includes additional circuitry for performing this BSF12 to FP16 data format conversion (e.g., N unit - see Figure 6C and Figure 7D ).
[0104] Continue to refer Figure 8A The processing or operation of the conversion circuitry for converting data / values from a fixed-point data format (BSF16 format in the illustrative embodiment) to a floating-point data format (FP16 format in the illustrative embodiment) includes determining whether the data / value is negative and, if so, performing a two's complement operation on the fraction field (inverting and incrementing the bits). Additionally, the processing may include performing a priority encoding operation on the fraction field of the BSF data / value and a left shift operation on the fraction field of negative values / data. The fraction field of each data / value may be rounded relative to the shifted fraction field. Finally, the processing generates an exponent field using the maximum exponent of the data / value—e.g., from the left shift amount.
[0105] In another embodiment, the F to H conversion circuitry and operations may be modified relative to certain portions discussed above. Here, because the memory capacity and memory bandwidth used by the filter weight data or values (Fij values) are significant, the Fij data / values are typically stored in L3 DRAM memory and L2 SRAM memory in INT8 format. Once the data is read from the L1 SRAM memory to the L0 SRAM memory that may be in the multiplier-accumulator execution pipeline (e.g., MAC pipeline array) of the multiplier-accumulator circuitry, the data format size of the Fij data / value is less important. With this in mind, the floating point data format to fixed point data format conversion circuitry and processing may be implemented Figure 8B Here, Figure 8B Illustrated is a 3x3 block of filter weight data that modifies 8-bit filter weight data / values (Fij) to increase the dynamic range of the filter weight data / values without increasing memory footprint or memory bandwidth requirements. 72 bits are allocated to nine BSF7 values (block-scaled fractional format), of which 9 bits are available for a shared exponent, EMAX. The BSF7 values are passed to the F-to-H conversion circuitry, where the filter weight data / values (Fij) are initially converted to BSF12 before being output to the F-to-H conversion circuitry and operation.
[0106] After the data is processed by the circuitry of the MAC execution or processing pipeline, the present invention employs ZY conversion circuitry consistent with Winograd-type processing to generate image or pixel data. In one embodiment, output data conversion circuitry is incorporated into the MAC processing pipeline or coupled to the output of the MAC processing pipeline to convert output data in a floating-point data format to a fixed-point data format. Regardless, in one embodiment, after the data is processed by the multiplier-accumulator execution or processing pipeline, the output data conversion circuitry converts the output data (e.g., image / pixel data) in a floating-point data format to output data in a fixed-point data format such as BSF. Thereafter, the output data in the fixed-point data format is applied to or input into Winograd conversion circuitry of the ZY conversion circuitry, where the output data is processed consistent with Winograd-type processing to convert the output data from Winograd format to a non-Winograd format.
[0107] In one embodiment, after being processed by the Winograd transformation circuitry, the output data is written to and stored in a memory (e.g., L2 memory such as SRAM). In another embodiment, after being processed by the Winograd transformation circuitry, the output data is applied to or input into a transformation circuitry to convert the output data from a fixed-point data format to a floating-point data format. That is, in one embodiment, after being processed by the Winograd transformation circuitry into a non-Winograd format having a fixed-point data format (e.g., BSF), the output data is applied to or input into a transformation circuitry to convert or transform the data (e.g., image / pixel data) to a floating-point data format before being output to, for example, a memory. Here, in one embodiment, the output data (e.g., image / pixel data) is written to and stored in the memory in a floating-point data format.
[0108] Notably, in one embodiment, the Z-to-Y conversion circuitry is incorporated or integrated into the multiplier-accumulator execution pipeline(s). That is, the Z-to-Y conversion of processed data is performed in the MAC execution / processing pipeline, where, among other things, the output data from the MAC processing pipeline has already been processed by the Winograd transformation circuitry.
[0109] refer to Figure 2A-2CIn one mode of operation, an NxN multiplier-accumulator execution pipeline employing multiplier-accumulator circuitry accumulates QxQ pixel output data / values to a single output layer, where each execution pipeline aggregates the sum of K multiplications of K input layers' input data / values and associated filter weights. In one embodiment, the aggregation of the NxN component data / values of the QxQ output data / pixel is implemented / performed externally to the NxN multiplier-accumulator execution pipeline. The NxN product data / values are accumulated with other NxN product data / values from other input layers - however, in this embodiment, the individual data / values are accumulated together into the final QxQ output data / pixel after the Z to Y conversion logic operations are performed on the accumulated NxN product data / values. (See Figure 9A ).
[0110] refer to Figure 2B and Figure 9A In one exemplary embodiment, Z-to-Y conversion circuitry is connected to the output of each MAC processing pipeline, or incorporated or integrated therein, with each of the sixteen zij data / value streams being directed from the Z shift-out block of one of the sixteen MAC pipelines of the multiplier-accumulator circuitry. A 4×4 Z block of data / values can be assembled from 16 streams, each of which is sorted by the individual zij data / values of the 4×4 Z block of data / values, as implemented by zij insertion logic circuitry (16 in this illustrative embodiment). The 4×4 Z block of data / values is converted to a 2x2 Y block by Winograd conversion circuitry in the Z-to-Y (zij to yij) conversion circuitry. A memory (e.g., L2 SRAM memory) can store the 2x2 Y block of data / values in a partitioned or divided fashion across the sixteen physical blocks, allowing sixteen sets of data to be written or stored in parallel across the sixteen MAC execution pipelines. Here, each set of data may consist of four 2x2 Y blocks, which will include 16 accesses from each physical block of memory (eg, L2 SRAM memory), where each access includes, for example, 1 ns in this exemplary embodiment.
[0111] In one embodiment, the Z-to-Y conversion circuitry includes circuitry for converting the output of the Winograd conversion circuitry from data in a fixed-point data format to a floating-point data format. In this manner, the output data in the floating-point data format can be used by circuitry external to the circuitry of the present invention. For example, the output data in the floating-point data format can be stored in a memory (e.g., internal and / or external memory) for use in additional processing (e.g., image generation).
[0112] Note that only 1 / 4 of the available L2 SRAM memory is used to write the Y block data / values; the D block data and execution pipelines each use a 64ns pipeline cycle time to process 16x64 4x4 D input blocks for each 2x2 pixel step. In this exemplary embodiment, the lower Y access bandwidth of the L2 SRAM memory can facilitate reducing the number of physical blocks of Y memory from 16 to 4.
[0113] Alternatively, however, the additional bandwidth may be used where more than 64 input planes are accumulated. For example, if there are 128 input planes (and in this exemplary implementation, 64 MAC data / values per multiplier-accumulator execution pipeline of the multiplier-accumulator circuitry), then the first 64 input planes may be accumulated into a particular region of memory (e.g., the "Y" region of the L2 SRAM memory). Then, as the second 64 input planes are accumulated in the pipeline of the multiplier-accumulator circuitry, the Y values for the first plane are read from Y2 and passed to the accumulation port on the Z to Y (zij to yij) conversion circuitry. The two sets of values may be added together and rewritten or stored in the Y region of the L2 SRAM memory. This is done in the region labeled "V" by Figure 9A The paths marked by dashed lines are shown in the diagram. The accumulated values are stored in the Y region of L2 memory. The second set of 64 input planes are accumulated on the right side of the diagram. The 4x4 Z block values pass through the zij to yij conversion circuitry before being written to the Y region of L2 memory. As they pass through the conversion logic, the Y accumulated values from the first 64 input planes are read from L2 and loaded into the accumulation input port of the Z to Y (zij to yij) conversion circuitry.
[0114] Figure 9C is a logical overview of ZY conversion circuitry according to certain aspects of the present invention, wherein data output by circuitry executing a multiplier-accumulator pipeline is first converted from a floating-point data format to a fixed-point data format (in an illustrative example, a block-scaled fractional format), which is a data format that is convenient for implementation in or consistent with the use of Winograd type processing (e.g., a fixed-point data format such as a block-scaled fractional format). The output data (Z to Y) conversion circuitry may employ Figure 5A, to convert the output data / value from a floating point data format to a fixed point data format. The output data (Zij) is then processed by Winograd conversion circuitry to convert the data from a Winograd environment to a non-Winograd environment, where the data is in a fixed point data format. In one embodiment, according to aspects of the present invention, the data (from the fixed point data format) is converted back to a floating point data format. Notably, the output data (Z to Y) conversion circuitry may employ Figure 5B to convert output data / values from a fixed-point data format to a floating-point data format.
[0115] refer to Figure 9C The ZY conversion circuitry includes circuitry that initially converts the data format of the output data (Zij) from a floating-point data format to a fixed-point data format. In one embodiment, this data format conversion is performed by the T unit, the U unit, the R unit, and the V unit (using the EMAXU value). Winograd processing is performed by the X00-X31 units that output the processed output data (Yij) in the fixed-point data format. After processing by the Winograd conversion circuitry, the output data (Yij data / value) is converted from the fixed-point data format to a floating-point data format via the W unit using the EMAXV value.
[0116] Figure 9D The cells of the zij insertion logic circuitry (left portion of the figure) and one cell of the zij to yij conversion circuitry (right portion of the figure) of an embodiment of a ZY conversion circuitry are illustrated in circuit block diagram form. The zij insertion logic circuitry includes (i) 16 cells on the left side, (ii) data registers for the zij and yij data words, (iii) control logic for sorting, and adder logic for conversion. It also includes vertical "INSRT_IN" and "INSRT_OUT" ports that carry inserted zij data / values from the appropriate execution pipeline of the multiplier-accumulator circuitry. The zij insertion logic circuitry may also include an accumulation port (lower left portion of the figure) - for example, if there are more input planes than execution pipeline stages. The zij to yij conversion circuitry includes (i) 16 cells on the left side, (ii) data registers for the dij and eij data words, (iii) control logic for sorting, and (iv) adder logic for conversion. Note that some of the zij to yij conversion processing may be implemented or performed in the zij insertion units; notably, in some embodiments they include some of the same circuitry as the zij to yij conversion units.
[0117] Continue to refer Figure 9D, after the output data is converted from floating point to fixed point data format, the sixteen Zij data / values are inserted into the X00-X15 locations. From there, they are shifted from left to right and loaded (at two different times) into the DREG_X and DREG_Y registers. The four Yij data / values are shifted through these 32 locations and the appropriate Zij value is accumulated into the Yij data / value. The four Yij data / values are then shifted through the ADD block (adder circuitry - see Figure 9A ) and the result is written to L2 memory. For the case where there are more than 64 input image planes, the FPADD block (floating point adder circuitry) adds the previously accumulated results.
[0118] Figure 9C The output data conversion circuitry is illustrated, wherein the Zij input values and Yij output values are in floating point format (here, FP32 format). In this illustrative embodiment, the output data conversion circuitry includes Winograd conversion circuitry having sixteen conversion pipelines, wherein the Zij input values are in INT32 data format. Figure 9C , data from the MAC processing pipeline (Z IN ) is input into the circuitry of the 16x16 Z to Y conversion pipeline to convert this data from a floating point data format (e.g., FP32) to a fixed point data format (e.g., BSF32 or INT32). IN ”) are input into the 16×16 array of the T unit, and the output data / values (Zij data / values) all in floating-point data format are processed and converted into fixed-point data format, wherein the inserted Zij data / values are then input into or provided to sixteen pipelines including the U unit to receive the data / values in floating-point data format and determine the maximum exponent (EMAXU) of these data. The R unit performs or implements a time delay function / operation to accommodate the time used for the maximum exponent search. Thereafter, the V unit performs floating-point data format to fixed-point data format conversion (in one embodiment, FP32 to BSF32 format conversion of the Zij data / values using the EMAXU value).
[0119] Continue to refer Figure 9C, the output data / value (Zij) in a fixed-point data format is input to the Winograd conversion circuitry (in one embodiment, the "X" unit (i.e., X00-X31)); here, the "X00-X31" unit uses the fixed-point data format (e.g., BSF32) representation of the data / value (Zij data / value) to perform a Z to Y conversion via Winograd processing techniques to generate Yij output data / value. The data / value (Yij) output by the Winograd conversion circuitry is input to an additional conversion circuitry to generate output data / value in a floating-point data format (e.g., FP32) using the data / value (Yij) in a fixed-point data format from the Winograd conversion circuitry. In this regard, the W unit converts the output data (Yij) from a fixed-point data format to a floating-point data format using the EMAXV value (in this illustrative embodiment, a BSF32 to FP32 format conversion).
[0120] It is worth noting that, as indicated above, units / circuits are collectively referred to herein as "units" or "cells," and in the accompanying drawings, a particular "unit" is sometimes labeled or identified as "unit x," where x is R, T, U, V, W, or X (e.g., Figure 10A Unit T in ).
[0121] refer to Figure 9C and Figures 10A-10E The array of T units (16×16 in the illustrative embodiment) receives the inserted data / value (Zij) in FP32 format from the MAC processing pipeline. The U unit analyzes the data / value (Zij) to identify or determine the maximum exponent (EMAXU). The array of R units (16×16 in the illustrative embodiment) performs a delay function / operation to accommodate the search for the maximum exponent by the U unit. The V unit converts the data format of the output data / value (Zij) from floating point (e.g., FP32) to fixed point (e.g., BSF32) using the maximum exponent identified by the U unit (EMAXU value) and output. The output data / value (Zij) in fixed-point data format is used by the X00-X31 units to implement a Winograd transformation of the Z to Y transformation circuitry. Finally, the W unit receives the output (Yij) of the Winograd transformation circuitry and converts the data / value from fixed-point data format (BSF32) to floating-point data format (FP32) using the EMAXV value. In this manner, the output data (Yij) is in a floating point data format, and therefore the output of the multiplication operation performed by the multiplier-accumulator circuitry is in a floating point condition. Figures 10A-10EThe exemplary T circuit / cell, U circuit / cell, R circuit / cell, V circuit / cell and W circuit / cell of the exemplary logic block diagram of the Z to Y conversion circuit system illustrated in FIG9 are respectively illustrated in the form of circuit block diagrams. Figure 9D The X circuit / cell illustrated in
[15] implements Winograd processing.
[0122] It is worth noting that the reference Figure 1 、 Figure 2A and Figure 2B , in the exemplary embodiment, the output data / pixel group is respectively illustrated as a 1×1 output pixel group ( Figure 1 ) and 2×2 output pixel groups ( Figure 2A and Figure 2B ) rather than the more common PxP and QxQ arrays.
[0123] It is worth noting that the memory used to store data can be, for example, a block or array of dynamic and / or static random access memory cells such as DRAM, SRAM, flash memory, and / or MRAM; it is worth noting that all memory types and combinations thereof are intended to fall within the scope of the present invention. In one embodiment, the third and / or fourth memory stores input data, filter weight values, and output data values in SRAM (e.g., third memory, e.g., L2 SRAM memory) and / or DRAM (e.g., fourth memory, L3 DRAM memory). In addition, the third and / or fourth memory can store transformed input data (after the input data undergoes transformation via the D to E conversion logic operation) of the NxN array of input or image data / pixels. In one embodiment, both the "D" input data and the "Y" output data can be stored in the third (L2 SRAM) memory - each piece of data participates in a different multiplier-accumulate (MAC) operation (e.g., 64 different MAC operations), so that the more limited L2 memory bandwidth is sufficient to meet the much higher bandwidth of the MAC pipeline. In contrast, the weight data bandwidth required by the MAC pipeline is much higher and such data needs to be stored in the first / or second memory SRAM (e.g., L0 SRAM memory and L1 SRAM memory). In one embodiment, it can save: (i) the "F" weight value for the first operating mode of the N×N multiplier-accumulator execution pipeline of the multiplier-accumulator circuit system, or (ii) the "H" weight value for the second operating mode of the N×N multiplier-accumulator execution pipeline of the multiplier-accumulator circuit system.
[0124] It is worth noting that in one embodiment, the DE and ZY conversion operations can be performed separately (rather than on the fly) - although such an implementation may require additional read / write operations (e.g., more read / write operations for L2 operations, more than 2x), which may also increase memory capacity requirements (e.g., the third memory (L2 SRAM memory)).
[0125] In the case of transforming the filter weight values on the fly, the first and second memories may also store the transformed filter weight values or data. In one embodiment, the third and / or fourth memories may also be blocks or arrays of dynamic and / or static random access memory cells such as DRAM, SRAM, flash memory, and / or MRAM; in fact, all memory types and their combinations are intended to fall within the scope of the present invention. In a preferred embodiment, the first and / or second memories are SRAM (e.g., L0 SRAM memory and L1 SRAM memory).
[0126] It is worth noting that in the illustrative embodiments described herein (text and figures), the multiplier-accumulator circuitry is sometimes labeled "NMAX" or "NMAX pipeline" or "MAC pipeline."
[0127] Many inventions are described and illustrated herein. While certain embodiments, features, attributes, and advantages of the present invention have been described and illustrated, it should be understood that many other and different and / or similar embodiments, features, attributes, and advantages of the present invention are apparent from the description and illustrations. Therefore, the embodiments, features, attributes, and advantages of the invention described and illustrated herein are not exhaustive, and it should be understood that such other, similar, and different embodiments, features, attributes, and advantages of the present invention are within the scope of the present invention.
[0128] For example, a multiplier-accumulator circuit system (and a method of operating such circuit system) of an execution or processing pipeline (e.g., for image filtering) may include floating-point execution circuit system (e.g., floating-point multiplier circuitry and accumulator / adder circuitry) that implements floating-point multiplication and floating-point addition based on inputs having one or more floating-point data formats. The floating-point data format may be user- or system-defined and / or may be one-time programmable (e.g., at the time of manufacture) or more than one-time programmable (e.g., (i) at power-up or via power-up, startup, or the conduction / completion of an initialization sequence / processing sequence, and / or (ii) in-situ or during normal operation of the multiplier-accumulator circuit system of the integrated circuit or processing pipeline). In one embodiment, the circuit system of the multiplier-accumulator execution pipeline includes an adjustable precision data format (e.g., a floating-point data format). Additionally or alternatively, the circuit system of the execution pipeline may process data concurrently to increase the throughput of the pipeline. For example, in one implementation, the present invention may include multiple separate multiplier-accumulator circuits (sometimes referred to herein as "MACs") and multiple registers (including, in one embodiment, multiple shadow registers) that facilitate pipelining of multiplication and accumulation operations, where the circuitry performing the pipelining processes data concurrently to increase the throughput of the pipeline.
[0129] Furthermore, although the memory cells in certain embodiments are illustrated as static memory cells or storage elements, the present invention may employ dynamic or static memory cells or storage elements. In fact, as described above, such memory cells may be latches, flip-flops, or any other static / dynamic memory cells or memory cell circuits or storage elements now known or later developed. Furthermore, although the illustrative / exemplary embodiments include multiple memories (e.g., L3 memory, L2 memory, L1 memory, L0 memory) assigned, allocated, and / or used to store certain data and / or in certain organizations, one or more memories may be added, and / or one or more memories—e.g., L3 memory or L2 memory—may be omitted and / or combined / merged, and / or the organization may be changed, supplemented, and / or modified. The present invention is not limited to the illustrative / exemplary embodiments of memory organization and / or allocation set forth in the application. Once again, the present invention is not limited to the illustrative / exemplary embodiments set forth herein.
[0130] It is noted that in describing and illustrating certain aspects of the present invention, certain drawings have been simplified for clarity in order to describe, focus, highlight and / or illustrate certain aspects of the circuitry and techniques of the present invention.
[0131] The circuits and operational implementations of the input data conversion circuitry, filter weight conversion circuitry, and output data conversion circuitry described and / or illustrated herein are example embodiments. Different circuits and / or operational implementations of these circuitry may be employed, which are intended to fall within the scope of the present invention. Indeed, the present invention is not limited to the illustrative / exemplary embodiments of the input data conversion circuitry, filter weight conversion circuitry, and / or output data conversion circuitry set forth herein.
[0132] In addition, although the input data conversion circuit system, filter weight conversion circuit system and output data conversion circuit system in the illustrative exemplary embodiments describe the bit width of the floating-point data format and the fixed-point data format of the input data and filter weights, such (one or more) bit widths are exemplary. Here, although several exemplary embodiments and features of the present invention are described and / or illustrated in the context of conversion circuit systems and / or processing pipelines (including multiplier circuit systems) with specified bit widths and precisions, the embodiments and inventions are applicable to other contexts and other precisions (e.g., FPxx, where: xx is an integer and is greater than or equal to 8, 10, 12, 16, 24, etc.). For the sake of brevity, these other contexts and precisions will not be illustrated separately, but will be very clear to those skilled in the art based on, for example, this application. Thus, the present invention is not limited to (i) the specific fixed-point data format(s) illustrated (e.g., integer formats (INTxx) and block-scaled fractional formats (e.g., BSFxx)), block / data widths, datapath widths, bandwidths, values, processes, and / or algorithms, nor is it limited to (ii) the exemplary logical or physical overview configurations of specific circuit systems and / or overall pipelines, and / or exemplary module / circuitry configurations, overall pipelines, and / or exemplary Verilog code.
[0133] In addition, the present invention may include additional conversion techniques to convert data from a fixed-point data format to a floating-point data format before storage in a memory. For example, in the case where filter weight data is stored in an integer format (INTxx) or a block scaled-fractional format ("BSFxx") in a memory, the conversion circuit system may include a circuit system that converts the fixed-point data to floating-point data. In fact, in one embodiment, in the case where input data and / or filter weight data are provided in a fixed-point data format suitable for Winograd processing performed by a Winograd conversion circuit system of the input data conversion circuit system and / or the filter weight conversion circuit system, such circuit system may not convert the input data and / or filter weight data from a floating-point data format to a fixed-point data format. Thus, in this embodiment, the input data conversion circuit system and / or the filter weight conversion circuit system may not include a circuit system that converts data from a floating-point data format to a fixed-point data format.
[0134] As described above, the present invention is not limited to (i) the specific floating-point format(s), specific fixed-point format(s), operations (e.g., addition, subtraction, etc.), block / data width or length, datapath width, bandwidth, values, processes, and / or algorithms illustrated, nor is it limited to (ii) the exemplary logical or physical overview configurations, exemplary module / circuitry configurations, and / or exemplary Verilog code.
[0135] It is noteworthy that the various circuits, circuit systems, and techniques disclosed herein can be described using computer-aided design tools and expressed (or represented) in terms of their behavior, register transfers, logic components, transistors, layout geometry, and / or other characteristics as data and / or instructions embodied in various computer-readable media. The formats of files and other objects in which such circuit, circuit system, layout, and wiring representations can be implemented include, but are not limited to, formats supporting behavioral languages (such as C, Verilog, and HLDL), formats supporting register-level description languages (such as RTL), and formats supporting geometric description languages (such as GDSII, GDSIII, GDSIV, CIF, MEBES), as well as any other formats and / or languages now known or later developed. The computer-readable media in which such formatted data and / or instructions can be embodied include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media) and carrier waves that can be used to transmit such formatted data and / or instructions via wireless, optical, or wired signaling media, or any combination thereof. Examples of transmission of such formatted data and / or instructions via a carrier wave include, but are not limited to, transmission (upload, download, email, etc.) over the Internet and / or other computer networks via one or more data transmission protocols (e.g., HTTP, FTP, SMTP, etc.).
[0136] In practice, when received within a computer system via one or more computer-readable media, the data and / or instruction-based representations of the circuits described above can be processed within the computer system by a processing entity (e.g., one or more processors) in conjunction with the execution of one or more other computer programs, including but not limited to netlist generation programs, placement and routing programs, etc., to generate representations or images of the physical manifestations of these circuits. This representation or image can then be used in device fabrication, for example, by enabling the generation of one or more masks for forming various components of the circuits in the device fabrication process.
[0137] In addition, the various circuits, circuit systems, and techniques disclosed herein can be represented by simulations using computer-aided design and / or test tools. The simulations of circuits, circuit systems, layout and routing, and / or the techniques implemented thereby, can be implemented by computer systems, wherein the characteristics and operations of these circuits, circuit systems, layouts, and the techniques implemented thereby are simulated, copied, and / or predicted via the computer systems. The present invention also relates to such simulations of circuits of the present invention, circuit systems, and / or the techniques implemented thereby, and therefore are intended to also fall within the scope of the present invention. Computer-readable media corresponding to such simulations and / or test tools are also intended to fall within the scope of the present invention.
[0138] It is worth noting that reference herein to "one embodiment" or "an embodiment" (or similar) means that a particular feature, structure or characteristic described in conjunction with the embodiment may be included, adopted and / or incorporated in one, some or all embodiments of the present invention. The use or appearance of the phrase "in one embodiment" or "in another embodiment" (or similar) in the specification does not refer to the same embodiment, nor does it refer to separate or alternative embodiments that are necessarily mutually exclusive with one or more other embodiments, nor is it limited to a single exclusive embodiment. The same applies to the term "implementation". The present invention is not limited to any single aspect or embodiment thereof, nor is it limited to any combination and / or permutation of these aspects and / or embodiments. In addition, each aspect of the present invention and / or its embodiments may be used alone or in combination with one or more of the other aspects of the present invention and / or its embodiments. For the sake of brevity, certain permutations and combinations are not discussed and / or illustrated separately herein.
[0139] Moreover, embodiments or implementations described herein as "exemplary" should not be construed as ideal, preferred, or advantageous, for example, over other embodiments or implementations; rather, they are intended to convey or indicate that the embodiment or embodiments are example embodiment(s).
[0140] Although the present invention has been described in certain specific aspects, many additional modifications and variations will be apparent to those skilled in the art. It should therefore be understood that the present invention may be practiced in other ways than those specifically described without departing from the scope and spirit of the present invention. The embodiments of the present invention are therefore to be considered in all respects as illustrative / exemplary and not restrictive.
[0141] The terms "comprise," "including," "comprising," "include," "having," and "having" or any other variations thereof are intended to cover a non-exclusive inclusion, so that a process, method, circuit, article, or apparatus that includes a list of parts or elements may include not only those parts or elements but also other parts or elements not expressly listed or inherent to such process, method, article, or apparatus. Furthermore, the terms "connect," "connected," "connected to," or "connector" as used herein should be interpreted broadly to include directly or indirectly (e.g., via one or more conductors and / or intermediate devices / elements (active or passive) and / or via inductive or capacitive coupling) unless otherwise intended (e.g., use of the terms "directly connected" or "directly connected").
[0142] The terms "a" and "an" herein do not denote a limitation of quantity, but rather denote the presence of at least one of the referenced item. Additionally, the terms "first," "second," etc., herein do not denote any order, quantity, or importance, but are instead used to distinguish one element / circuit / feature from another.
[0143] Furthermore, the term "integrated circuit" means, among other things, any integrated circuit, including, for example, a general-purpose or non-application specific integrated circuit, a processor, a controller, a state machine, a gate array, a SoC, a PGA, and / or an FPGA. The term "integrated circuit" also means any integrated circuit (e.g., a processor, a controller, a state machine, and a SoC) - including an embedded FPGA.
[0144] Furthermore, the term "circuitry" means, among other things, a circuit (whether integrated or otherwise), a group of such circuits, one or more processors, one or more state machines, one or more processors implementing software, one or more gate arrays, programmable gate arrays, and / or field programmable gate arrays, or a combination of one or more circuits (whether integrated or otherwise), one or more state machines, one or more processors, one or more processors implementing software, one or more gate arrays, programmable gate arrays, and / or field programmable gate arrays. The term "data" means, among other things, (one or more) current or voltage signals (plural or singular), whether in analog or digital form, which may be a single bit (or similar) or multiple bits (or similar).
[0145] In the claims, the term "MAC circuit" means a multiplier-accumulator circuit having a multiplier circuit coupled to an accumulator circuit. For example, in U.S. Patent Application No. 16 / 545,345 Figure 1 A- Figure 1The multiplier-accumulator circuit is described and illustrated in the exemplary embodiments of C and the text associated therewith. However, it is noted that the term "MAC circuit" is not limited to circuits according to, for example, U.S. Patent Application No. 16 / 545,345. Figure 1 A- Figure 1 The specific circuit, logical, block, functional and / or physical diagrams, block / data widths, data path widths, bandwidths, and processes illustrated and / or described in the exemplary embodiments of C, as noted above, are incorporated herein by reference.
[0146] In the context of this application, the term "in situ" means during normal operation of the integrated circuit - as well as after power-up, startup, or completion of its initialization sequence / process. The term "data processing operation" means any operation that processes data (e.g., image and / or audio data) including, for example, digital signal processing, filtering, and / or other forms of data manipulation and / or transformation, whether now known or later developed.
[0147] It is noteworthy that the limitations of the claims are not drafted in a means-plus-function format or a step-plus-function format.
Claims
1. An integrated circuit comprising: a multiplier-accumulator execution pipeline for receiving (i) image data and (ii) filter weights, wherein the multiplier-accumulator execution pipeline includes a plurality of multiplier-accumulator circuits to process the image data via a plurality of multiplication and accumulation operations using associated filter weights; a first conversion circuitry coupled to an input of the multiplier-accumulator execution pipeline, wherein the first conversion circuitry comprises: input for receiving a plurality of sets of image data, wherein each set of image data comprises a plurality of image data, wherein the image data of each set of image data in the plurality of sets of image data comprises a floating point data format, block-scaled fractional conversion circuitry coupled to an input of the first conversion circuitry to convert image data of each of the plurality of sets of image data to a block-scaled fractional BSF data format, Winograd conversion circuitry is coupled to the block-scaled fractional conversion circuitry and is configured to (i) receive image data for each of the plurality of sets of image data in a block-scaled fractional BSF data format and (ii) convert each of the plurality of sets of image data into a corresponding set of Winograd image data, wherein: the image data of each of the plurality of sets of image data converted via the Winograd conversion circuitry to the corresponding set of Winograd image data comprises a block-scaled fractional BSF data format, floating point format conversion circuitry coupled to the Winograd conversion circuitry to (i) receive image data for each of the plurality of Winograd sets of image data and (ii) convert the image data for each of the plurality of Winograd sets of image data to a floating point data format, and output, for outputting image data in the plurality of sets of Winograd image data to the multiplier-accumulator execution pipeline, wherein the image data in each set of Winograd image data includes a floating-point data format; and wherein, in operation, the multiplier-accumulator circuit of the multiplier-accumulator execution pipeline is configured to: (i) perform the plurality of multiplication and accumulation operations using (a) image data in a floating-point data format from the plurality of sets of Winograd image data output from the first conversion circuit system and (b) filter weights, and (ii) generate output data based on the plurality of multiplication and accumulation operations; in: The block-scaled fractional conversion circuitry is configured to convert each image data having a floating point data format for each set of image data in the plurality of sets of image data into a block-scaled fractional BSF data format including an exponent field and a fraction field, wherein a value of the exponent field for each image data of a particular set of image data is the same for each image data associated with the particular set of image data; and The block-scaled fractional conversion circuitry further includes determination circuitry for determining a maximum exponent of image data for each of the plurality of sets of image data.
2. The integrated circuit of claim 1 , wherein: The floating point format conversion circuitry (i) receives data of a maximum exponent of image data for each set of image data and (ii) converts the image data of each set of Winograd image data into a floating point data format using the data of the maximum exponent of image data associated with the set of image data.
3. The integrated circuit of claim 2, wherein: The first conversion circuit system further includes the following circuit system, which, for image data of each group of image data in the multiple groups of image data, right shifts the fractional field of the image data having a smaller exponent relative to the maximum exponent of the image data associated with the group of image data, rounds the fractional field of the image data to block-scaled fractional BSF precision, and performs a two's complement operation on the fractional field of the image data if the image data is a negative value.
4. The integrated circuit of claim 1 , wherein: The block-scaled fractional conversion circuitry converts the data format of the image data of each set of image data into a block-scaled fractional (BSF) data format using a maximum exponent of the image data of the associated set of image data, wherein: The Winograd conversion circuitry converts the image data of each set of image data into a corresponding set of Winograd image data using the image data having the block-scaled fractional BSF data format.
5. The integrated circuit of claim 4, wherein: The floating point format conversion circuitry (i) receives data of a maximum exponent of image data for each set of image data and (ii) converts the image data of each set of Winograd image data into a floating point data format using the data of the maximum exponent of image data associated with the set of image data.
6. The integrated circuit of claim 5, wherein: The floating-point format conversion circuit system includes the following circuit system, which performs a binary complement operation on a fraction field of each set of image data in each set of the multiple sets of Winograd image data if the image data is a negative value.
7. The integrated circuit of claim 1 , wherein: The floating point format conversion circuitry of the first conversion circuitry is configured to convert image data of each set of Winograd image data of the plurality of sets of Winograd image data to a floating point data format using a maximum exponent of the image data associated with the set of Winograd image data.
8. The integrated circuit of claim 7, wherein: The floating-point format conversion circuit system of the first conversion circuit system includes the following circuit system, which, for each set of Winograd image data in the multiple sets of Winograd image data, performs a binary complement operation on a fractional field of the image data if the image data is a negative value, performs a priority encoding operation on the fractional field of the image data, shifts the fractional field of the image data left if the image data is a negative value, and rounds the fractional field to the position where the image data is left-shifted.
9. The integrated circuit of claim 1 , wherein: Each multiplier-accumulator circuit of the multiplier-accumulator execution pipeline includes a floating-point multiplier and a floating-point adder to perform the plurality of multiplication and accumulation operations.
10. An integrated circuit comprising: a multiplier-accumulator execution pipeline for receiving (i) image data and (ii) filter weights, wherein the multiplier-accumulator execution pipeline includes a plurality of multiplier-accumulator circuits to process the image data via a plurality of multiplication and accumulation operations using associated filter weights; a first conversion circuitry coupled to an input of the multiplier-accumulator execution pipeline, wherein the first conversion circuitry comprises: an input for receiving image data of a plurality of sets of image data, wherein each set of image data comprises a plurality of image data, wherein the image data of each set of image data in the plurality of sets of image data comprises a floating point data format, block-scaled fractional conversion circuitry coupled to an input of the first conversion circuitry to convert image data of each of the plurality of sets of image data to a block-scaled fractional BSF data format, Winograd conversion circuitry is coupled to the block-scaled fractional conversion circuitry and is configured to (i) receive image data for each of the plurality of sets of image data in a block-scaled fractional BSF data format and (ii) convert each of the plurality of sets of image data into a corresponding set of Winograd image data, wherein: the image data of each of the plurality of sets of image data converted via the Winograd conversion circuitry to a corresponding set of Winograd image data comprises a block-scaled fractional BSF data format, floating-point format conversion circuitry coupled to the Winograd conversion circuitry of the first conversion circuitry to (i) receive image data for each of a plurality of sets of Winograd image data and (ii) convert the image data for each of the plurality of sets of Winograd image data into a floating-point data format, and output, for outputting image data of the plurality of sets of Winograd image data to the multiplier-accumulator execution pipeline, wherein the image data of each set of Winograd image data includes a floating-point data format; Second conversion circuitry coupled to an input of the multiplier-accumulator execution pipeline, wherein the second conversion circuitry comprises: An input for receiving a plurality of filter weights, wherein each filter weight comprises a plurality of filter weights, Winograd conversion circuitry for converting each set of filter weights in the plurality of sets of filter weights into Winograd filter weights for a corresponding set in the plurality of sets of Winograd filter weights, floating-point format conversion circuitry coupled to the Winograd conversion circuitry of the second conversion circuitry to (i) receive the filter weights of each set of Winograd filter weights and (ii) convert the filter weights of each set of Winograd filter weights in the plurality of sets of Winograd filter weights into a floating-point data format, and output, for outputting the filter weights of each set of Winograd filter weights in the plurality of sets of Winograd filter weights to the multiplier-accumulator execution pipeline, wherein the filter weights of each set of Winograd image data include a floating-point data format; and wherein, in operation, the multiplier-accumulator circuit of the multiplier-accumulator execution pipeline is configured to: (i) perform the plurality of multiplication and accumulation operations using (a) image data in a floating-point data format from the plurality of sets of Winograd image data output from the first conversion circuit system and (b) filter weights in a floating-point data format from each set of the plurality of sets of Winograd filter weights, and (ii) generate output data of the plurality of sets of output data based on the plurality of multiplication and accumulation operations; in: the block-scaled fractional conversion circuitry of the first conversion circuitry being configured to convert each image data having a floating point data format for each set of image data in the plurality of sets of image data into a block-scaled fractional BSF data format including an exponent field and a fraction field, wherein a value of the exponent field for each image data of a particular set of image data is the same for each image data associated with the particular set of image data; and The block-scaled fractional conversion circuitry further includes determination circuitry for determining a maximum exponent of image data for each of the plurality of sets of image data.
11. The integrated circuit of claim 10, wherein: The floating-point format conversion circuitry of the first conversion circuitry (i) receives data of a maximum exponent of image data for each set of image data and (ii) converts the image data of each set of Winograd image data into a floating-point data format using the data of the maximum exponent of image data associated with the set of image data in the plurality of sets of image data.
12. The integrated circuit of claim 11 , wherein: The first conversion circuitry further includes circuitry that, for image data of each set of image data in the plurality of sets of image data, right shifts a fractional field of image data having a smaller exponent relative to a maximum exponent of the image data associated with the set of image data, rounds the fractional field of the image data to block-scaled fractional BSF precision, and performs a two's complement operation on the fractional field of the image data if the image data is a negative value.
13. The integrated circuit of claim 10, wherein: The Winograd conversion circuitry of the first conversion circuitry converts the image data of each set of image data into a corresponding set of Winograd image data using the image data having the block-scaled fractional BSF data format.
14. The integrated circuit of claim 10, wherein: The second conversion circuitry further includes circuitry for determining a maximum index of filter weights for each of the plurality of sets of filter weights, and The floating-point format conversion circuit system of the second conversion circuit system (i) receives data of the maximum exponent of the filter weights of each group of filter weights in the multiple groups of filter weights and (ii) converts the filter weights of each group of Winograd image data into a floating-point data format using the data of the maximum exponent of the filter weights associated with the group of filter weights in the multiple groups of filter weights.
15. The integrated circuit of claim 14, wherein: The second conversion circuit system further includes the following circuit system for the filter weights of each group of filter weights in the multiple groups of filter weights, which right shifts the fractional field of the filter weight having a smaller exponent relative to the maximum exponent of the filter weights associated with the group of filter weights, rounds the fractional field to the block-scaled fractional BSF precision of the filter weight, and performs a two's complement operation on the fractional field of the filter weight if the image data is a negative value.
16. The integrated circuit of claim 10, wherein: The second conversion circuitry further includes fixed-point format conversion circuitry coupled between an input of the second conversion circuitry and an input of the Winograd transform circuitry of the second conversion circuitry to convert the data format of the filter weights of the set of filter weights to a fixed-point data format using a maximum exponent of the image data associated with each set of filter weights, wherein: The fixed-point format conversion circuitry includes circuitry for determining a maximum exponent of filter weights for each of the plurality of sets of filter weights; and The Winograd transformation circuitry of the second transformation circuitry converts the filter weights of each group of filter weights into Winograd filter weights of a corresponding group using the filter weights having a fixed-point data format.
17. The integrated circuit of claim 10, wherein: the floating point format conversion circuitry of the first conversion circuitry (i) receiving data of a maximum exponent of image data for each set of image data and (ii) converting the image data of each set of image data in the plurality of sets of Winograd image data into a floating point data format using the data of the maximum exponent of image data associated with the set of Winograd image data, The second conversion circuitry further includes circuitry for determining a maximum index of filter weights for each of the plurality of sets of filter weights, and The floating-point format conversion circuit system of the second conversion circuit system (i) receives data of the maximum exponent of the filter weights of each group of filter weights in the multiple groups of filter weights and (ii) converts the filter weights of the group of filter weights into a floating-point data format using the data of the maximum exponent of the filter weights associated with each group of Winograd filter weights in the multiple groups of Winograd filter weights.
18. The integrated circuit of claim 17, wherein: Each multiplier-accumulator circuit in the multiplier-accumulator execution pipeline includes a floating-point multiplier and a floating-point adder to perform the plurality of multiplication and accumulation operations.
19. The integrated circuit of claim 17, wherein: The floating-point format conversion circuitry of the first conversion circuitry further includes circuitry that, for each set of Winograd image data, performs a two's complement operation on a fraction field of the image data if the image data is a negative value, performs a priority encoding operation on the fraction field of the image data, left-shifts the fraction field of the image data if the image data is a negative value, and rounds the fraction field of the image data to a position to which the image data is left-shifted, and The floating-point format conversion circuit system of the second conversion circuit system further includes the following circuit system, which performs a binary complement operation on the fraction field of the filter weight of each set of Winograd filter weights in the multiple sets of Winograd filter weights if the filter weight is a negative value, shifts the fraction field of the filter weight to the left if the image data is a negative value, and rounds the fraction field of the filter weight to the position where the filter weight is left-shifted.
20. An integrated circuit comprising: a multiplier-accumulator execution pipeline for receiving (i) image data and (ii) filter weights, wherein the multiplier-accumulator execution pipeline includes a plurality of multiplier-accumulator circuits to process the image data via a plurality of multiplication and accumulation operations using associated filter weights; a first conversion circuitry coupled to an input of the multiplier-accumulator execution pipeline, wherein the first conversion circuitry comprises: an input for receiving image data of a plurality of sets of image data, wherein each set of image data comprises a plurality of image data, wherein the image data of each set of image data in the plurality of sets of image data comprises a floating point data format, block-scaled fractional conversion circuitry coupled to an input of the first conversion circuitry to convert image data of each of the plurality of sets of image data to a block-scaled fractional BSF data format, Winograd conversion circuitry is coupled to the block-scaled fractional conversion circuitry and is configured to (i) receive image data for each of the plurality of sets of image data in a block-scaled fractional BSF data format and (ii) convert each of the plurality of sets of image data into a corresponding set of Winograd image data, wherein: the image data of each of the plurality of sets of image data converted via the Winograd conversion circuitry to a corresponding set of Winograd image data comprises a block-scaled fractional BSF data format, floating-point format conversion circuitry coupled to Winograd conversion circuitry of the first conversion circuitry to (i) receive image data for each of a plurality of sets of Winograd image data and (ii) convert the image data for each of the plurality of sets of Winograd image data into a floating-point data format, and output, for outputting image data of the plurality of sets of Winograd image data to the multiplier-accumulator execution pipeline, wherein the image data of each set of Winograd image data includes a floating-point data format; Second conversion circuitry coupled to an input of the multiplier-accumulator execution pipeline, wherein the second conversion circuitry comprises: An input for receiving a plurality of filter weights, wherein each filter weight comprises a plurality of filter weights, Winograd conversion circuitry for converting each set of filter weights in the plurality of sets of filter weights into Winograd filter weights for a corresponding set in the plurality of sets of Winograd filter weights, floating-point format conversion circuitry coupled to the Winograd conversion circuitry of the second conversion circuitry to (i) receive the filter weights of each set of Winograd filter weights and (ii) convert the filter weights of each set of Winograd filter weights in the plurality of sets of Winograd filter weights into a floating-point data format, and output, for outputting the filter weights of each set of Winograd filter weights in the plurality of sets of Winograd filter weights to the multiplier-accumulator execution pipeline, wherein the filter weights of each set of Winograd image data include a floating-point data format; wherein, in operation, the multiplier-accumulator circuit of the multiplier-accumulator execution pipeline is configured to: (i) perform the plurality of multiplication and accumulation operations using (a) image data in floating point data format from the plurality of sets of Winograd image data output from the first conversion circuit system and (b) filter weights in floating point data format from each set of the plurality of sets of Winograd filter weights, and (ii) generate output data of the plurality of sets of output data based on the plurality of multiplication and accumulation operations; and A third conversion circuit system is coupled to an output of the multiplier-accumulator execution pipeline, wherein the third conversion circuit system comprises: an input for receiving output data of the plurality of sets of output data from a multiplier-accumulator circuit of the multiplier-accumulator execution pipeline, Winograd conversion circuitry for converting each of the plurality of sets of output data into a corresponding set of non-Winograd output data, floating point format conversion circuitry coupled to the Winograd conversion circuitry of the third conversion circuitry to convert the output data of each set of non-Winograd output data to a floating point data format, and output, for outputting the output data having a floating point data format from the floating point format conversion circuit system of the third conversion circuit system; in: the block-scaled fractional conversion circuitry of the first conversion circuitry being configured to convert each image data having a floating point data format for each set of image data in the plurality of sets of image data into a block-scaled fractional BSF data format including an exponent field and a fraction field, wherein a value of the exponent field for each image data of a particular set of image data is the same for each image data associated with the particular set of image data; and The block-scaled fractional conversion circuitry further includes determination circuitry for determining a maximum exponent of image data for each of the plurality of sets of image data.
21. The integrated circuit of claim 20, further comprising: A memory is coupled to the floating point format conversion circuitry of the third conversion circuitry to store output data in the floating point data format from among the plurality of sets of output data.
22. The integrated circuit of claim 20, wherein: Each multiplier-accumulator circuit in the multiplier-accumulator execution pipeline includes a floating-point multiplier and a floating-point adder to perform the plurality of multiplication and accumulation operations.
23. The integrated circuit of claim 20, wherein: the floating point format conversion circuitry of the first conversion circuitry (i) receiving data of a maximum exponent of image data for each set of image data and (ii) converting the image data of each set of image data in the plurality of sets of Winograd image data into a floating point data format using the data of the maximum exponent of image data associated with the set of Winograd image data, The second conversion circuitry further includes circuitry for determining a maximum index of filter weights for each of the plurality of sets of filter weights, The floating-point format conversion circuit system of the second conversion circuit system (i) receives data of the maximum exponent of the filter weights of each group of filter weights in the multiple groups of filter weights and (ii) converts the filter weights of the group of filter weights into a floating-point data format using the data of the maximum exponent of the filter weights associated with each group of Winograd filter weights in the multiple groups of Winograd filter weights.
24. The integrated circuit of claim 23, wherein: Each multiplier-accumulator circuit in the multiplier-accumulator execution pipeline includes a floating-point multiplier and a floating-point adder to perform the plurality of multiplication and accumulation operations.
Citation Information
Patent Citations
Multiplier-accumulator circuit, logic tile architecture for multiply-accumulate, and IC including logic tile array
US10693469B2
Multiplier-accumulator processing pipelines and processing component, and methods of operating same
US11314504B2
Multiplier-Accumulator Circuitry having Processing Pipelines and Methods of Operating Same
US20200310818A1
Filtering unit for floating-point texture data
US20080211827A1
System and method for an optimized winograd convolution accelerator
US20190042923A1