Mac processing pipeline, circuitry to control and configure the same, and methods of operating the same
By controlling and configuring the circuit system to manage the connection and operation of the multiplier-accumulator circuit, and employing pipelined processing and FPGA programmable configuration, the problem of insufficient throughput of the multiplier-accumulator circuit system is solved, achieving more efficient data processing, especially in image data processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FLEX LOGIX TECHNOLOGIES INC
- Filing Date
- 2021-03-30
- Publication Date
- 2026-08-04
AI Technical Summary
In the prior art, multiplier-accumulator circuit systems suffer from insufficient throughput when processing data, especially when processing image data, making it difficult to achieve efficient parallel or simultaneous processing.
A control and configuration circuit system is used to manage the connection and operation of multiple multiplier-accumulator circuits. Multiplication and accumulation operations are performed in a pipelined manner, including the use of local memory for skip reading and writing of filter weights to reduce latency and overhead time. At the same time, FPGA is used for programmable configuration to achieve more efficient data processing.
This improves the processing throughput of the multiplier-accumulator circuit system, enabling faster and more efficient data processing, especially in image data processing, where it increases data processing efficiency and throughput.
Smart Images

Figure CN115280281B_ABST
Abstract
Description
[0001] Related applications
[0002] This non-provisional application claims priority and interest in U.S. Provisional Application No. 63 / 012,111, filed April 18, 2020, entitled "MAC Processing Pipelines, Circuit to Control and Configure Same, and Methods of Operating Same". Provisional Application '111 is hereby incorporated herein by reference in its entirety.
[0003] introduce
[0004] This document describes and illustrates numerous inventions. The invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or substitution of such aspects and / or embodiments. Importantly, each aspect and / or embodiment of the invention can be used alone or in combination with one or more other aspects and / or embodiments of the invention. All combinations and substitutions thereof are intended to fall within the scope of this invention.
[0005] In one aspect, the present invention relates to a circuit system (and a method of operating and configuring such a circuit system) for configuring and controlling a multiplier-accumulator circuit system comprising multiple multiplier-accumulator execution or processing pipelines. The circuit system of this aspect of the invention configures (e.g., is programmable once or more than once) and controls the multiplier-accumulator circuit system to implement one or more execution or processing pipelines to process data, such as processing data in parallel or simultaneously. In one embodiment, the circuit system configures and controls rows / groups of multiple separate multiplier-accumulator circuits (“multiplier-accumulator circuits”, sometimes referred to herein as “MAC” or “MAC circuits”, or in plural form “MAC” or “MAC circuits”) or interconnected (serialized) multiplier-accumulator circuits to pipeline multiplication and accumulation operations. Multiple multiplier-accumulator circuits may also include multiple registers (including multiple shaded registers), wherein the circuit system further controls these registers to implement or facilitate the pipelined multiplication and accumulation operations performed by the multiplier-accumulator circuits, thereby increasing the throughput of the multiplier-accumulator execution or processing pipeline related to the processing of relevant data (e.g., image data).
[0006] It is worth noting that, in one embodiment of the present invention, the invention employs the techniques described in U.S. Patent Application No. 16 / 545,345 and U.S. Provisional Patent Application No. 62 / 725,306. Figure 1AOne or more multiplier-accumulator circuits are described and illustrated in the exemplary embodiment of FIG1C and in its associated text. Here, the multiplier-accumulator circuits described and / or illustrated in '345 and '306 applications facilitate the connection of multiplication and accumulation operations and the reconfiguration of their circuit systems and the operations performed therefrom; in this way, multiple multiplier-accumulator circuits can be configured and / or reconfigured to process data (e.g., image data) in a manner that thereby allows for faster and / or more efficient processing and operations. '345 and '306 applications are incorporated herein by reference in their entirety.
[0007] In one embodiment, the circuitry for controlling and configuring the multiplier-accumulator circuitry (sometimes referred to as the control / configuration circuitry) includes circuitry for controlling, configuring, and / or programming (e.g., one-time or more-than-one-time programmable) one or more execution or processing paths of the multiplier-accumulator circuitry, including pipelines (sometimes referred to herein as “MAC processing pipelines” or “MAC execution pipelines”). For example, in one embodiment, the control / configuration circuitry may configure or connect a selected number of multiplier-accumulator circuits or rows / groups of multiplier-accumulator circuits to, among other things, implement a predetermined multiplier-accumulator execution or processing pipeline or its architecture. Here, the control / configuration circuitry can configure or determine the pipeline architecture or configuration implemented via interconnected (serialized) rows / groups of multiplier-accumulator circuits or interconnected multiplier-accumulator circuits for performing multiplication and accumulation operations and / or via connections of multiplier-accumulator circuits (or rows / groups of interconnected multiplier-accumulator circuits) for performing multiplication and accumulation operations. Thus, in one embodiment, the control / configuration circuitry configures or implements the architecture for executing or processing the pipeline by controlling or providing connections between (one or more) rows of multiplier-accumulator circuits and / or interconnected multiplier-accumulator circuits. Although multiple multiplier-accumulator circuits are described / illustrated as “rows” of multiplier-accumulator circuits, multiple can be described as “columns” of multiplier-accumulator circuits, wherein the layout of the multiple multiplier-accumulator circuits is vertical; both are intended to fall within the scope of the invention.
[0008] It is worth noting that the present invention can be performed or processed using one or more multiplier-accumulators and processing techniques described and illustrated in applications '345 and '306. As stated above, applications '345 and '306 are incorporated herein in their entirety.
[0009] In addition, in one embodiment, the control / configuration circuitry can facilitate or control the writing of filter weights or values to multiple multiplier-accumulator circuits to perform multiplication. For example, the control / configuration circuitry can connect multiple multiplier-accumulator circuits (or rows / groups of interconnected (serialized) multiplier-accumulator circuits) to memory to control or facilitate the provisioning / writing of filter weights or values to multiple multiplier-accumulator circuits to perform multiplication. In one embodiment, each multiplier-accumulator circuit includes a "local" memory (e.g., a register or a set of registers) to "locally" store filter weights or values for use in relation to the multiplication operation. Each "local" memory can be dedicated to the associated multiplier-accumulator circuit.
[0010] In one embodiment, each multiplier-accumulator circuit includes multiple "local" memory / register banks to store at least two sets of filter weights or values, including a first set of filter weights and a second set of filter weights. The first set of filter weights is used for a first multiplication operation (e.g., in relation to implementing the current multiplication operation), and the second set of filter weights is used for a multiplication operation immediately following the current multiplication operation (i.e., immediately after the completion of the "current" multiplication operation using the first set of filter weights). In this embodiment, in operation, the multiplier-accumulator circuit can read the first filter weight from the first memory during the first set of multiplication operations and read the second filter weight from the second memory during the second set of multiplication operations (associated with the processing of the second set of data) after the completion of the first set of multiplication operations (associated with the processing of the second set of data). In this way, the multiplier-accumulator circuit can ping-pong read operations on a single set of multiplication operations between multiple "local" memory / register banks. In other words, the multiplier-accumulator circuit can read / access the first set of filter weights (i.e., stored in the first "local" memory / register set) for the current multiplication operation related to processing the first set of data, and upon completion, immediately read / access the next / second set of filter weights (i.e., stored in the second "local" memory / register set) related to processing the second set of data. Here, because the second set of filter weights can be written and stored in the second memory / register set before the multiplication operation using the first set of filter weights stored in the first memory / register set is completed, there is no delay or overhead time arising from access to and availability of the next set of filter weights for the multiplier-accumulator circuit during data processing performed by the multiplier-accumulator circuit in the pipeline. A memory output selection circuit (e.g., a multiplexer) can responsively control which memory / register set (and when) among the multiple memory / register sets is connected to the multiplier circuit system of the multiplier-accumulator circuit. In this way, the memory output selection circuit responsively controls the jump read operation and the connection of the memory / register bank to the multiplier circuit system of the multiplier-accumulator circuit.
[0011] Furthermore, the control / configuration circuitry can alternately write a new or updated set of filter weights to multiple "local" memory / register sets. For example, the control / configuration circuitry can skip write operations between two "local" memory / register sets associated with each multiplier-accumulator circuit in the multiplier-accumulator execution or processing pipeline, based on a set of multiplication operations. That is, after completing a set of multiplication operations using a first set of filter weights stored in the first memory / register set, a new set of filter weights can be written to the first memory / register set for use in processing after a set of multiplication operations using a second set of filter weights (i.e., a set of filter weights stored in the second "local" memory / register set). Here, the control / configuration circuitry provides or writes a new set of filter weights to the first "local" memory / register set, while the multiplier-accumulator circuit uses the second set of filter weights for multiplication operations. The new or third set of filter weights (stored in the first "local" memory / register set associated with each multiplier-accumulator circuit) is then immediately available to the multiplier-accumulator circuit after data processing is completed using the second set of filter weights or values stored in the second "local" memory / register set. In this way, no delay or overhead time is introduced into the multiplier-accumulator execution or processing pipeline during data processing due to writing, updating, or providing "new" filter weights or values to be used by the circuit to the "local" memory / register set.
[0012] As described above, in one embodiment, the control / configuration circuitry controls or configures one or more connections between rows of multiplier-accumulator circuits and / or interconnected multiplier-accumulator circuits to implement an architecture for performing or processing pipelines. For example, the control / configuration circuitry may determine and / or configure (i) which multiplier-accumulator circuits are employed and / or interconnected to perform multiplication and accumulation operations and / or (ii) the number of interconnected multiplier-accumulator circuits (and / or rows / groups of interconnected (serialized) multiplier-accumulator circuits) used to perform predetermined multiplication and accumulation operations, and / or (iii) the pipeline architecture or configuration implemented via (a) connections of multiplier-accumulator circuits used to perform multiplication and accumulation operations and / or (b) connections between rows of interconnected multiplier-accumulator circuits. In one embodiment, the control / configuration circuitry system includes multiple control / configuration circuits associated with and dedicated to interfacing with multiple interconnected (e.g., cascaded) multiplier-accumulator circuits and / or one or more rows interconnected (e.g., cascaded) multiplier-accumulator circuits. For example, the configuration can be implemented in-situ (i.e., during operation of the integrated circuit) to, for example, meet or exceed time-based system requirements or constraints. Furthermore, the configuration implemented by the control / configuration circuitry system can be one-time programmable (e.g., at manufacturing time, via, for example, a programmable fuse array) or multiple-time programmable (including, for example, at startup / power-on, initialization, and / or in-situ).
[0013] It is worth noting that the MAC processing or execution pipeline can be organized or set on one or more integrated circuits. In one embodiment, the integrated circuit is a discrete field-programmable gate array (FPGA) or an embedded FPGA (hereinafter collectively referred to as "FPGA" unless otherwise stated). In short, an FPGA is an integrated circuit that is configured and / or reconfigured by a user, operator, customer, and / or designer before and / or after manufacturing (hereinafter collectively referred to as "configuration," etc., e.g., "configurable" and "configurable") unless otherwise stated. An FPGA may include programmable logic components (commonly referred to as "logic cells," "configurable logic blocks" (CLBs), "logic array blocks" (LABs), or "logic blocks"—hereinafter collectively referred to as "logic blocks").
[0014] In one embodiment of the invention, one or more (or all) logic blocks of the FPGA include multiple multiplier-accumulator circuits to implement multiplication and accumulation operations, for example, in a pipelined manner. Control / configuration circuitry may be included in or may include switch interconnect networks within the logic blocks. Switch interconnect networks may be configured as hierarchical and / or mesh interconnect networks. Logic blocks may include data storage elements, input pins, and / or lookup tables (LUTs) associated with the switch interconnect networks, which, when programmed, determine the configuration and / or operation of the switches / multiplexers and, among other things, communication between circuitry systems (e.g., logic components) of logic blocks (including MAC circuitry and / or MAC processing pipelines) and / or between circuitry systems of multiple logic blocks (e.g., between MAC circuitry and / or MAC processing pipelines of multiple logic blocks).
[0015] In one embodiment, a switch interconnect network can provide connections to / from a logic circuit system of associated or different logic blocks, or from (separate) multiplier-accumulator circuitry that processes or executes a pipelined multiplier-accumulator circuit. In this way, MAC circuitry and / or MAC processing pipelines of multiple logic blocks can be used simultaneously, for example, to process related data (e.g., related image data). In practice, such connections can be configurable and / or reconfigurable – for example, in-situ (i.e., during normal operation of the integrated circuit) and / or during or after power-on, startup, initialization, re-initialization, configuration, reconfiguration, etc. In one embodiment, the switch interconnect network can employ one or more embodiments, features, and / or aspects of the interconnect networks described and / or illustrated in '345 and '306 applications. Furthermore, the switch interconnect network can interface to and / or include one or more embodiments, features, and / or aspects of the interface connectors described and / or illustrated in '345 and '306 applications (e.g., see '345 application's...). Figures 7A-7C It is worth noting that certain details of the NLINK circuits described and illustrated herein may relate to the circuit systems described and / or illustrated in '345 and '306 applications, referred to as NLINX (e.g., NLINX conductors, NLINX interfaces, NLINX interface connectors, etc.). As mentioned above, '345 and '306 applications are incorporated herein by reference in their entirety.
[0016] It is worth noting that (one or more) integrated circuits can be, for example, processors, controllers, state machines, gate arrays, system-on-a-chip (SoC), programmable gate arrays (PGAs), and / or FPGAs and / or processors, controllers, state machines, and SoCs, including embedded FPGAs. Field-programmable gate arrays or FPGAs refer to both discrete FPGAs and embedded FPGAs.
[0017] In one embodiment, the invention may also be employed or implemented in concurrent and / or parallel processing techniques in multiplier-accumulator execution or processing pipelines (and methods of operating such circuit systems), which increase pipeline throughput, as described and / or illustrated in U.S. Patent Application No. 16 / 816,164 and U.S. Provisional Patent Application No. 62 / 831,413, both of which are incorporated herein by reference in their entirety.
[0018] In short, refer to Figure 1A In one embodiment, the multiplier-accumulator circuitry system in the execution pipeline is configured in a linear or cascaded pipeline architecture. In this configuration, Dijk data is fixed in place during execution, while Yijl data rotates during execution (between the MACs in the pipeline). m x m (e.g., 64 x 64) Fkl filter weights are distributed on L0 memory (in this illustrative embodiment, 64 L0 SRAMs – one L0 SRAM in each of the m (e.g., m = 64) MAC processing circuits in the pipeline). In each execution cycle, m (e.g., m = 64) Fkl values are read and passed to the MAC elements or circuitry. MEM Once the Dijk shift chain in memory (here, L2 memory - such as SRAM) is loaded, the Dijk data value is stored or held in a processing element for m (e.g., m = 64) execution cycles.
[0019] Furthermore, during processing, the Yijlk MAC value is shifted from the Yijk shift chain (see Y). MEM Once loaded from memory, the memory is rotated through all m (e.g., m=64) MAC processors over m (e.g., m=64) execution cycles and will be unloaded using the same shift chain.
[0020] Furthermore, in this exemplary embodiment, “m” (e.g., 64 in the illustrative embodiment) MAC processing circuits in the execution pipeline operate simultaneously, causing the multiplier-accumulator processing circuit to perform mxm (e.g., 64x64) multiplication-accumulation operations every m (e.g., 64) cycle intervals (here, a cycle may be denoted as 1 ns). Subsequently, during the same m (e.g., 64) cycle intervals, the next set of input pixels / data (e.g., 64) is shifted in and the previous output pixels / data is shifted out. Notably, every m (e.g., 64) cycle intervals process the Dd / Yd (depth) columns of input and output pixels / data at a specific (i, j) position (index of the width Dw / Yw and height Dh / Yh dimensions). The m-cycle execution interval is repeated for each Dw*Dh depth column in this stage. In this exemplary embodiment, the filter weights or weight data are loaded from, for example, external memory or the processor into memory (e.g., L1 / L0 SRAM memory) before the stage processing begins (see, for example, applications '345 and '306). In this particular embodiment, Dw = 512, Dh = 256, and Dd = 128 for the input stage, and Yw = 512, Yh = 256, and Yd = 64 for the output stage. Note that in one embodiment, only 64 of the 128 Dd inputs are processed in each 64x64 MAC execution operation.
[0021] Continue to refer to Figure 1A The method, implemented with the configuration shown, can be adapted to any image / data plane dimension (Dw / Yw and Dh / Yh) by simply adjusting the number of iterations of the basic 64x64 MAC accumulation operation. The loop indices “I” and “j” are adjusted via a control and sequencing logic circuitry system to accommodate the image / data plane dimension. Furthermore, the method can be adjusted and / or extended to handle Yd column depths larger than the number of MAC processing elements in the execution pipeline (e.g., 64 in an illustrative example). In one embodiment, this can be achieved by dividing the depth column of the output pixels into blocks (e.g., 64) and repeating the process for each of these blocks. Figure 1A This is achieved by accumulating MAC addresses.
[0022] In fact, it can be further expanded. Figure 1A The method shown is for handling column depths Dd larger than the number of MAC processing elements / circuits in the execution pipeline (64 in one illustrative example). In one embodiment, this can be achieved by initially accumulating a first block of m data (e.g., m = 64) of the input pixel Dijk to each output pixel Yijl. Subsequently, the partially accumulated value Yijl is read (from memory Y). memReturning to the execution pipeline, the initial value is used as the next block of m input data / pixels Dijk to continue accumulating to each output pixel Yijl. The memory storing or holding the continuously accumulated values (e.g., L2 memory) can be organized, partitioned, and / or resized to accommodate any additional read / write bandwidth to support processing operations.
[0023] refer to Figure 1B Integrated circuits may include multiple multi-bit MAC execution pipelines, which are organized into one or more clusters of processing components. Here, components may include "resources," such as bus interfaces (e.g., PHY and / or GPIO), to facilitate communication with circuitry external to the components and with memory (e.g., SRAM and DRAM), used for storage and by the circuitry of the components. For example, see reference... Figure 1B In one embodiment, the component (labeled "X1") includes four clusters, each cluster comprising multiple multi-bit MAC execution pipelines (16 64-MAC execution pipelines in this illustrative embodiment). It is worth noting that, for reference, the lower right figure is illustrated... Figure 1A A MAC execution pipeline (in this illustrative embodiment, this includes 64 MAC processing circuits).
[0024] Continue to refer to Figure 1B The memory hierarchy in this exemplary embodiment includes L0 memory (e.g., SRAM) that stores filter weights or coefficients to be used by the multiplier-accumulator circuitry in relation to the multiplication operations thus performed. In one embodiment, each MAC execution pipeline includes L0 memory to store filter weights or coefficients associated with data processed by the circuitry of the MAC execution pipeline. L1 memory (a larger SRAM resource) is associated with each cluster of the MAC execution pipeline. These two memories may store, retain, and / or hold the filter weight values Fijklm used in the accumulation operations.
[0025] It is worth noting that, Figure 1BImplementations may employ L2 memory (e.g., SRAM memory larger than L1 or L0 memory). An on-chip network (NOC) couples the L2 memory to the PHY (physical interface) to provide connectivity to external memory (e.g., L3 memory—such as one or more external DRAM components). The NOC is also coupled to a PCIe PHY, which in turn is coupled to an external host. The NOC is also coupled to a GPIO input / output PHY, allowing simultaneous operation of multiple X1 components. Control / configuration circuitry (sometimes referred to as “NLINK” or “NLINK circuitry”) is connected via programmable or configurable interconnect paths to the multiplier-accumulator circuitry system (which includes multiple (here, 64) multiplier-accumulator circuits or MAC processors) to configure the entire execution pipeline, among other things, by providing or “bootstrapping” data between one or more MAC pipelines. Furthermore, the control / configuration circuitry can configure the interconnect between the multiplier-accumulator circuitry and one or more memories—one or more memories that can be shared by one or more (or all) clusters of the MAC execution pipeline—including external memories (e.g., L3 memories, such as external DRAM). These memories can store, for example, input image pixels Dijk, output image pixels Yijl (i.e., image data processed by the circuitry of the MAC pipeline(s), and filter weight values Fijklm associated with such data processing).
[0026] It is worth noting that although illustrative or exemplary embodiments describe and / or illustrate the assignment, allocation, and / or use of multiple different memories (e.g., L3 memory, L2 memory, L1 memory, L0 memory) for storing certain data and / or in certain organizations, one or more other memories may be added, and / or one or more memories may be omitted and / or combined / merged—for example, the L3 memory or L2 memory and / or organization may be changed. All combinations are intended to fall within the scope of this invention.
[0027] Furthermore, in the illustrative embodiments described herein (text and figures), the multiplier-accumulator circuit system and / or multiplier-accumulator pipeline are sometimes referred to as “NMAX”, “NMAX pipeline”, “MAC processing pipeline”, “MAC execution pipeline”, or “MAC pipeline”.
[0028] Continue to refer to Figure 1BThe integrated circuit (one or more) comprises multiple clusters (e.g., two, four, or eight), each cluster comprising multiple multiplier-accumulator circuits (“MAC”) execution pipelines (e.g., 16). Each MAC execution pipeline may include multiple separate multiplier-accumulator circuits (e.g., 64) to perform multiplication and accumulation operations. In one embodiment, the multiple clusters are interconnected to form processing circuitry / components (such components are typically identified as “X1” or “X1 component” in the figures). The processing circuitry / components may include memory (e.g., SRAM, MRAM, and / or flash memory), a switching interconnect network, and the switching interconnect network interconnects the circuitry of the component (e.g., the multiplier-accumulator circuitry of the X1 component and / or one or more MAC execution pipelines) and / or the circuitry of the component with the circuitry of one or more other X1 components or other circuitry external to the associated X1 component. Here, the multiplier-accumulator circuitry of one or more MAC execution pipelines of the multiple clusters of the X1 component may be configured to process related data (e.g., image data) simultaneously. In other words, multiple separate multiplier-accumulator circuits in multiple MAC execution pipelines can process related data simultaneously to, for example, increase the data throughput of the X1 component.
[0029] It is worth noting that the X1 component may also include an interface circuitry (e.g., a PHY and / or GPIO circuitry) to interface with, for example, external memory (e.g., DRAM, MRAM, SRAM, and / or flash memory).
[0030] In one embodiment, the MAC execution pipeline can be of any size or length (e.g., 16, 32, 64, 96, or 128 multiplier-accumulator circuits). In practice, the size or length of the pipeline can be configurable or programmable (e.g., one-off or multi-off – such as in-situ (i.e., during operation of the integrated circuit) and / or during or after power-on, startup, initialization, re-initialization, configuration, reconfiguration, etc.). Here, the size or length of the MAC execution pipeline (i.e., the number of MACs connected in the pipeline, such as a linear or serially interconnected pipeline) can be increased or decreased – one-off (during manufacturing or testing) or multi-off, e.g., in-situ and / or during or after power-on, startup, initialization, re-initialization, configuration, reconfiguration, etc.
[0031] In another embodiment, one or more integrated circuits include multiple components or X1 components (e.g., 2, 4, ...), wherein each component includes multiple clusters with multiple MAC execution pipelines. For example, in one embodiment, an integrated circuit includes multiple components or X1 components, wherein each X1 component includes multiple clusters with multiple MAC execution pipelines (e.g., 4 clusters). As described above, each cluster includes multiple execution or processing pipelines (e.g., 16, 32, or 64), which can be configured or programmed to simultaneously process, operate, and / or function to simultaneously process related data (e.g., image data). In this way, the related data is processed simultaneously by each execution pipeline of the multiple clusters to, for example, reduce the processing time of the related data and / or increase the data throughput of the X1 component.
[0032] As discussed in applications '164' and '413, multiple execution or processing pipelines of one or more clusters of multiple X1 components can be interconnected to process data (e.g., image data), both of which are incorporated herein by reference in their entirety. In one embodiment, such execution or processing pipelines can be interconnected in a ring configuration or architecture to process related data simultaneously. Here, multiple MAC execution pipelines of one or more (or all) clusters of multiple X1 components (which may be integrated / manufactured on a single die or multiple dies) can be interconnected in a ring configuration or architecture (wherein, bus interconnect components) to process related data simultaneously. For example, multiple MAC execution pipelines of one or more (or all) clusters of each X1 component are configured to process one or more stages of an image frame, such that the circuitry of each X1 component processes one or more stages of each of the multiple image frames. In another embodiment, multiple MAC execution pipelines of one or more (or all) clusters of each X1 component are configured to process one or more portions of each stage of each image frame, such that the circuitry of each X1 component is configured to process a portion of each stage of each of the multiple image frames. In yet another embodiment, multiple MAC execution pipelines of one or more (or all) clusters of each X1 component are configured to process all stages of at least one overall image frame, such that the circuitry of each X1 component is configured to process all stages of at least one image frame. Here, each X1 component is configured to process all stages of one or more image frames, such that the circuitry of each X1 component processes different image frames.
[0033] As mentioned above, numerous inventions have been described and illustrated herein. The invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or substitution of such aspects and / or embodiments. Furthermore, each aspect and / or embodiment of the invention may be used alone or in combination with one or more other aspects and / or embodiments of the invention. For the sake of brevity, certain substitutions and combinations have not been discussed and / or illustrated in detail herein. Attached Figure Description
[0034] This invention can be implemented in conjunction with the embodiments illustrated in its accompanying drawings. These drawings illustrate different aspects of the invention, and where appropriate, reference numerals, names, or designations illustrating similar circuits, architectures, structures, components, materials, and / or elements are similarly labeled in different figures. It should be understood that various combinations of structures, components, materials, and / or elements beyond those specifically shown are conceived, and these various combinations also fall within the scope of this invention.
[0035] Furthermore, numerous inventions are described and illustrated herein. The invention is neither limited to any single aspect or embodiment thereof, nor to any combination and / or substitution of such aspects and / or embodiments. Additionally, each aspect and / or embodiment of the invention may be used alone or in combination with one or more other aspects and / or embodiments of the invention. For the sake of brevity, certain substitutions and combinations are not discussed and / or illustrated separately herein. It is important to note that the embodiments or implementations described herein as “exemplary” should not be construed as being preferred or advantageous, for example, relative to other embodiments or implementations; rather, they are intended to reflect or indicate that one or more embodiments are “example” embodiments.
[0036] It is worth noting that the configurations, block / data / signal widths, data / signal path widths, bandwidths, data lengths, values, processing, pseudocodes, operations, and / or algorithms illustrated in the description and / or figures and associated text herein are exemplary. In fact, the invention is not limited to any particular or exemplary circuit, logic, block, function, and / or physical diagram illustrated and / or described, such as exemplary circuits, logic, blocks, functions, and / or physical diagrams; the number of multiplier-accumulator circuits employed in the execution pipeline; the number of execution pipelines employed in a particular processing configuration; the organization / allocation of memory; block / data widths; data path widths; bandwidths; values; processing; pseudocodes; operations, and / or algorithms. Furthermore, although illustrative / exemplary embodiments include assigning, allocating, and / or using memory for storing certain data (e.g., filter weights) and / or multiple memories (e.g., L3 memory, L2 memory, L1 memory, L0 memory) in certain organizations. In practice, the organization of the memory can be altered, in which one or more memories can be added, and / or one or more memories can be omitted and / or combined / merged with other memories—such as (i) L3 or L2 memories and / or (ii) L1 or L0 memories. Likewise, the invention is not limited to the illustrative / exemplary embodiments set forth herein.
[0037] Figure 1AThis is a schematic block diagram of the logic overview of an exemplary multiplier-accumulator execution pipeline connected in a linear pipeline configuration according to one or more aspects of the present invention, wherein the multiplier-accumulator execution pipeline includes a multiplier-accumulator circuit system (“MAC”), which is illustrated in block diagram form; it is worth noting that the multiplier-accumulator circuit system includes one or more multiplier-accumulator circuits (although individual multiplier-accumulator circuits are not specifically illustrated herein); multiple MACs are illustrated in block diagram form; exemplary multiplier-accumulator circuits are illustrated in illustration A in schematic block diagram form; it is worth noting that in this exemplary embodiment, “m” (e.g., 64 in one illustrative embodiment) multiplier-accumulator circuits are linearly pipelined to operate simultaneously, thereby processing the circuitry at every m (e.g.) For example, m x m (e.g., 64 x 64) multiplication-accumulation operations are performed in 64 period intervals (e.g., the period can be denoted as 1 ns); it is worth noting that every m (e.g., 64) period intervals process the Dd / Yd (depth) columns of input and output pixels / data at a specific (i, j) position (the indices of the width Dw / Yw and height Dh / Yh dimensions in this exemplary embodiment -- Dw = 512, Dh = 256 and Dd = 128, and Yw = 512, Yh = 256 and Yd = 64), wherein the interval is repeated for r (e.g., 64) period intervals for each Dw*Dh depth column in this stage; furthermore, in one embodiment, the filter weights or weight data are loaded into memory (e.g., L1 / L0) before the multiplier-accumulator circuitry begins processing. In an L1 SRAM memory (see, for example, '345 and '306 applications); in one embodiment, the L1 SRAM memory can provide data to a plurality of L0 SRAM memories, wherein each linear pipeline (like...) Figure 2D As illustrated in the block diagram, it is associated with a dedicated LO SRAM memory—a dedicated LO SRAM memory is one of a plurality of LO SRAM memories associated with an L1 SRAM memory (in one embodiment, dedicated to one of a plurality of clusters);
[0038] Figure 1BThis is a high-level block diagram layout of an integrated circuit or part of an integrated circuit (sometimes referred to as an X1 component) comprising multiple multi-bit MAC execution pipelines having multiple multiplier-accumulator circuits, each of the multiple multiplier-accumulator circuits implementing multiplication and accumulation operations; the multi-bit MAC execution pipelines and / or the multiple multiplier-accumulator circuits can be configured to implement one or more processing architectures or technologies (alone or in combination with one or more X1 components); in this illustrative embodiment, the multi-bit MAC execution pipelines are organized into clusters (in this illustrative embodiment, four clusters, each cluster comprising multiple multi-bit MAC execution pipelines (in this illustrative embodiment, each cluster comprises 16 64-MAC execution pipelines (each MAC may also be individually referred to as a MAC circuit or MAC processor hereinafter)); in one embodiment, the multiple multiplier-accumulator circuits are configurable or programmable (one-time or multiple times, e.g., at startup and / or in-situ) to implement one or more pipelined processing architectures or technologies (see, e.g., the lower right corner). Figure 1B An extended view of a portion of the high-level block diagram is a single MAC execution pipeline (in the illustrative embodiment, including, for example, 64 multiplier-accumulator circuits or a MAC processor), which is related to a schematic block diagram of a logical overview of an exemplary multiplier-accumulator circuit system arranged in a linear execution pipeline configuration - see [link to block diagram]. Figure 1AThe processing components in this illustrative embodiment include memory (e.g., L2 memory, L1 memory, and L0 memory (e.g., SRAM)), a bus interface (e.g., PHY and / or GPIO), and multiple switches / multiplexers. The bus interface facilitates communication with external circuitry and with memory (e.g., SRAM and DRAM) used by the circuitry of the component. The multiple switches / multiplexers are electrically interconnected to form a switch interconnect network “Network on Chip” (“NOC”) to facilitate a cluster of multiplier-accumulator circuitry for interconnecting the MAC execution pipeline. In one embodiment, the NOC includes open... The interconnect network (e.g., a hybrid-mode interconnect network (i.e., a hierarchical switch matrix interconnect network and interconnect networks such as meshes, tori, etc. (collectively referred to below as "mesh networks" or "mesh interconnect networks")), associated data storage elements, input pins, and / or lookup tables (LUTs) determine the operation of the switch / multiplexer when programmed; in one embodiment, one or more (or all) clusters include one or more computing elements (e.g., multiple multiplier-accumulator circuit systems labeled "NMAX rows" - see, for example, '345 and '306 applications); notably, in one embodiment, each MAC An execution pipeline (in one embodiment, consisting of multiple cascaded multiplier-accumulator circuits) is connected to an associated L0 memory (e.g., SRAM memory) dedicated to that processing pipeline; the associated L0 memory stores filter weights used by the multiplier circuitry system of each multiplier-accumulator circuit of that particular MAC processing pipeline when performing multiplication operations, wherein each MAC processing pipeline of a given cluster is connected to an associated L0 memory (in one embodiment, the L0 memory is dedicated to the multiplier-accumulator circuitry of that MAC processing pipeline); multiple (e.g., 16) MAC clusters The AC execution pipeline (and particularly the L0 memory of each MAC execution pipeline in the cluster) is coupled to an associated L1 memory (e.g., SRAM memory); here, the L1 memory is connected to and shared by each MAC execution pipeline in the cluster to receive filter weights to be stored in the L0 memory associated with each MAC execution pipeline in the cluster; in one embodiment, the L1 memory is associated with and dedicated to multiple pipelines of the MAC cluster; that is, in one embodiment, the L1 memory is connected to and provides data to multiple L0 memories, wherein each MAC pipeline (e.g., ...) is connected to the L0 memory of each MAC execution pipeline. Figure 2DThe NOC (as illustrated in the block diagram) is associated with a dedicated L0 memory—one of a plurality of L0 SRAM memories associated with the L1 memory (in one embodiment, the L1 memory is associated with and dedicated to one of a plurality of clusters of X1 components); notably, each 64-MAC execution pipeline's move-in and move-out paths are coupled to L2 memory (e.g., SRAM memory), where the L2 memory is also coupled to both L1 and L0 memories; the NOC couples the L2 memory to a PHY (physical interface), which can be connected to L3 memory (e.g., external DRAM); the NOC is also coupled to a PCIe or PHY, which in turn can provide interconnection or communication with circuitry outside the X1 processing components (e.g., external processors, such as host processors); in one embodiment, the NOC can also connect to multiple X1 components (e.g., via GPIO input / output PHYs), which allows multiple X1 components to process associated data (e.g., image data), as discussed herein, according to one or more aspects of the invention;
[0039] Figure 2A and Figure 2B The illustration shows a schematic block diagram of an exemplary memory architecture for storing filter weights according to one or more aspects of the present invention, wherein the multiplier circuitry of each multiplier-accumulator circuit is connected to a memory to receive filter weights for multiplication; in one embodiment, each memory is associated with a multiplier circuitry of a multiplier-accumulator circuit, wherein each multiplier-accumulator circuit includes a dedicated memory (e.g., SRAM) to locally store filter weights (typically referred to as L0 memory); wherein the exemplary memory architecture includes multiple memories that selectively output to the multiplier circuitry of the associated multiplier-accumulator circuit, and a memory output selection circuitry (e.g., a multiplexer) can be connected to an appropriate memory (storing filter weights associated with data currently being processed by the multiplier-accumulator circuit); in one exemplary embodiment, the memory includes multiple separately addressable memories, wherein read / write operations of each memory are independently controllable relative to read / write operations of the other memories of the multiple memories (see [link to documentation]). Figure 2B Therefore, data (i.e., filter weights) can be read from the first memory (i.e., read operation), while data (i.e., filter weights) can be written to the second memory simultaneously (i.e., write operation); at a given time (e.g., for a set of filter weights associated with a set of data under processing), only the output of one memory is provided to the multiplier circuit system of the associated multiplier-accumulator circuit - see, for example, see Figure 2B The memory output selection circuitry system (e.g., multiplexer) in the memory;
[0040] Figure 2C The illustration depicts a schematic block diagram of an exemplary memory architecture for storing filter weights or coefficients used by a multiplier circuit system of associated multiplier-accumulator circuits, according to one or more aspects of the present invention; in this embodiment, the memory (e.g., SRAM) comprises multiple memory groups (two in this illustrative embodiment) dedicated to and associated with one of the multiple multiplier-accumulator circuits to locally store filter weights (again, generally referred to as L0 memory); the memory groups are separately addressable memories, wherein read / write operations of each memory group are independently controllable relative to the other memory groups, such that, for example, memory group 1 can be read (where filter weight data is applied to the multiplier circuit system of the associated MAC), while memory group 0 is written (i.e., storing the filter weights of another set of filter weights); in operation, at a given time (e.g., for a set of filter weights associated with a set of data being processed), data from only one group is output to the multiplier circuit system of the associated multiplier-accumulator circuit via a multiplexer, the multiplexer... The multiplexer is controlled by a memory bank selection control signal; notably, in one embodiment, the write data bus may be connected to another memory (e.g., L1 memory – which may be SRAM), which provides multiple sets of filter weights for storage in the appropriate / selected memory bank; in operation, in one embodiment, the multiplier-accumulator circuitry may skip read operations between two memory banks on a set of multiplication operations to reduce or minimize time delays or overheads arising from the availability of “new” filter weights or values to the multiplier-accumulator circuitry during data processing performed by the multiplier-accumulator circuitry; similarly, the control / configuration circuitry system may skip write operations between two “local” memory banks associated with the multiplier-accumulator circuitry on a set of multiplication operations, such that filter weights are read from one memory bank (e.g., memory bank 1) while another memory bank (e.g., memory bank 0) is receiving and storing the next set of filter weights to be used by the multiplier-accumulator circuitry during data processing of the next set of data; as Figure 1B As reflected in the document, in one embodiment, the L1 memory may be associated with a group of L0 memories for multiple MACs (e.g., MACs of a cluster associated with the L1 memory).
[0041] Figure 2D Combined with an exemplary multiplier-accumulator with multiple MAC processors, it performs or processes a pipeline (see [link]). Figure 1A and Figure 1B A schematic block diagram illustrating a logical overview of a linear pipeline configuration according to one or more aspects of the invention is shown. Figure 2CA schematic block diagram of an exemplary memory architecture is provided, wherein each MAC processor includes a multiplier-accumulator circuit having a multiplier circuit (“MUL”) for performing / implementing multiplication operations and an accumulator circuit (“ADD”) for performing / implementing accumulation operations; furthermore, each MAC processor includes two memory banks (e.g., SRAM memory, such as L0 memory) dedicated to the multiplier-accumulator circuit to store filter weights used by the associated multiplier circuitry system of the multiplier-accumulator circuit; in this illustrative embodiment, the MAC execution or processing pipeline includes n multiplier-accumulator processing circuits (e.g., n = 64), and in one embodiment, each circuit includes two dedicated memory banks to store at least two different sets of filter weights (each set of filter weights is associated with and used to process a set of data), wherein each memory bank can be alternately read to process a given set of associated data and alternately written to after processing a given set of associated data; it is worth noting that the linear flow of this exemplary MAC execution or processing pipeline In the logical overview of the pipeline configuration, n processing (MAC) circuits (e.g., n = 64) are connected and operate simultaneously in the execution pipeline, whereby the multiplier-accumulator processing circuit performs 64x64 multiplication-accumulation operations every 64-cycle interval (here, the cycle can be, for example, nominal 1 ns); thereafter, during the same 64-cycle interval, the next 64 input pixels / data are shifted in, and the previous output pixels / data are shifted out; processing is performed every 64-cycle interval at a specific (i, j) position (width Dw / Yw and height Dh / Yh). The input and output pixel / data Dd / Yd (depth) columns (indices of the dimension index); for each Dw*Dh depth column of this stage, the execution interval is repeated for 64 cycles; before the start of stage processing, the filter weights or weight data are loaded into memory (e.g., L1 and / or L0 (i.e., L1 / L0) SRAM memory) from, for example, external memory or processor (see, for example, '345 and '306 applications); it is worth noting that the standalone MAC may sometimes be referred to herein as a MAC processor, MAC circuitry or MAC processing circuitry;
[0042] Figure 2EThe illustration shows a schematic block diagram of an exemplary multiplier-accumulator execution or processing pipeline according to one or more aspects of the present invention. This exemplary multiplier-accumulator execution or processing pipeline includes a plurality of cascaded MACs, wherein the output of each accumulator circuit (“ADD”) of the MACs is coupled to the input of the next accumulator circuit (“ADD”) of the MACs in the linear processing pipeline. In this manner, the accumulated value (“Y”) generated by the MACs (see MAC_r[p]) is rotated, transferred, or moved (e.g., in each of the cascaded MACs of the pipeline) in an execution sequence (i.e., a set of associated execution cycles). Before, during, or at the end of each execution cycle, each accumulated value generated by the MAC (see MAC_r[p] - "Rotation Current Y") is output to the next MAC in the linear pipeline before, during, or at the end of each execution cycle and used in the accumulation operation of the accumulator circuit ("ADD") of the next MAC; it is worth noting that each MAC according to one or more aspects of the invention includes a multiplier circuit ("MUL") for performing / implementing multiplication operations and an accumulator circuit ("ADD") for performing / implementing accumulation operations; in this exemplary embodiment, the MAC processor may include or be derived from a memory bank (see Figure 2A ) or multiple memory groups (see Figure 2B For example, two SRAM memory groups - see Figure 2C The memory bank can be dedicated to the MAC processing circuitry to store the filter weights used by the multiplier circuitry of the associated MAC; as mentioned above, a standalone MAC can sometimes be referred to as a "MAC processor" or "MAC processing circuitry"; it is worth noting that linear processing pipelines and MACs incorporated into such pipelines can be configured to rotate, transfer, or move (before, during, or at the end of an execution cycle) input data values (rather than accumulating values - maintained, stored, or held in a specific MAC during each execution cycle of the execution sequence), as in U.S. Provisional Patent Application No. 63 / 156,263, filed March 3, 2021, "MAC Processing Pipelines, Circuitry to Configure Same, and Methods of Operating..." The '263 application is incorporated herein by reference in its entirety as described and illustrated in the Same; In this context, in operation, after initial data input values are input or loaded into the MAC of a linear MAC processing pipeline, the input data values are periodically rotated, transferred, or moved from one MAC of the linear pipeline to the next MAC of the pipeline and used in the multiplication operation of the multiplier circuit of the next MAC of the pipeline, as described and / or illustrated in the '263 application;
[0043] Figure 3A The illustration shows a schematic block diagram of exemplary circuitry connected to a multiplier-accumulator circuitry system according to various aspects of the invention to configure and control the multiplier-accumulator circuitry system (including multiple MACs interconnected in a series / linear pipeline (64 MACs, multiplier-accumulator circuits, or MAC processors in this illustrative embodiment)); in one embodiment, this circuitry (sometimes referred to as "control / configuration circuitry," "NLINK," or "NLINK circuitry") is associated with and dedicated to interfacing with multiple interconnected (e.g., series) multiplier-accumulator circuits and / or one or more rows of interconnected (e.g., series) multiplier-accumulator circuits; in one embodiment, an integrated circuit having one or more control / configuration circuits performs pipelined interface with multiple multi-bit MACs, each pipeline including multiple interconnected (e.g., series) multiplier-accumulator circuits (see, e.g., Figure 1BEach control / configuration circuit, according to certain aspects of the invention, controls one of a plurality of multi-bit MAC execution pipelines; in this illustrative embodiment (implementing an exemplary processing pipeline embodiment), (i) the L0 memory control, address, and data signal paths in the NLINK circuitry provide control, address, and data signals generated by the logic circuitry system (not shown) to the L0 memory in each MAC of the pipeline during one or more execution sequences to manage or control memory operations; (ii) the input data signal paths in the NLINK circuitry and MAC execution pipeline of this embodiment represent the signal paths of input data to be processed via the MAC during one or more execution sequences; (iii) the accumulated data paths in the NLINK circuitry and execution pipeline of this embodiment represent the total number of ongoing / accumulated MAC accumulations generated by the MAC during one or more execution sequences; and (iv) the output data paths in the NLINK circuitry and MAC execution pipeline of this embodiment represent the number of outputs generated by one or more execution sequences. According to (i.e., input data processed via the multiplier-accumulator circuitry or MAC processor executing the pipeline); "row interconnects" connect two MACs connected in series to form or provide a linear MAC pipeline, wherein the signal lines are: "OD" for output data path, "AD" for accumulated data path, "ID" for input data path, "MC" for L0 memory control path, "MA" for L0 memory address path, and "MD" for L0 memory address path; the circles with numbers in each data path indicate the number of bits and / or conductors or lines (e.g., data or control) of a particular path or port - see illustration; to avoid doubt, the block / data widths, data path and / or port widths set forth herein are merely exemplary and are not limiting in any way; it is worth noting that a "port" is a physical point of entry / exit from control / configuration or NLINK circuitry; all physical forms of entry or exit from control / configuration or NLINK circuitry are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit);
[0044] Figure 3B The diagram shows... Figure 3A The diagram above is a schematic block diagram illustrating an exemplary embodiment of a MAC execution or processing pipeline.
[0045] Figures 3C-3F The illustrations depict embodiments according to certain aspects of the present invention. Figure 3A The schematic block diagram of an exemplary control / configuration circuit and multiplier-accumulator circuit system (e.g., multiple MACs connected in series to a linear MAC processing pipeline) is shown by Figure 3A The selected / specific part indicated in the text; Figure 3C and Figure 3DThe illustrations in the diagram provide general examples of (i) the input data flow in a MAC pipeline and an associated NLINK circuit ("shift of data input DI_x within a single pipeline" refers to the path of the input data), (ii) the multiplication and accumulation operation flow in a MAC pipeline and an associated NLINK circuit ("shift of MAC_x within a single pipeline" refers to the path of the accumulated data in the multiplication-accumulation operation), and (iii) the output data flow of an accumulation operation with a previous result in a MAC pipeline associated with the associated NLINK circuit ("shift of MAC_Sx within a single pipeline" refers to the path of the output data corresponding to the accumulation operation with the previous result). As mentioned above, a “row interconnect” connects two rows of MAC circuits connected in series (to form a linear MAC pipeline consisting of MAC circuits in two rows connected in series, wherein in this illustrative embodiment, n = 64); the signal lines of the row interconnect include: “OD” for output data path, “AD” for accumulated data path, “ID” for input data path, “MC” for L0 memory control path, “MA” for L0 memory address path, and “MD” for L0 memory address path; the circles with numbers in each data path indicate the number of bits and / or conductors or lines (e.g., data or control) of a particular path or port – see [link to relevant documentation] Figure 3C and Figure 3D The “illustrative” in the text; to avoid ambiguity, the block / data width, data path and / or port width illustrated are merely exemplary and do not constitute any limitation; it is worth noting that a “port” is a physical point of entry / exit from a control / configuration or NLINK circuit; all physical forms of entry / exit from a control / configuration or NLINK circuit are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit); again, a “port” is a physical point of entry / exit from a control / configuration or NLINK circuit; all physical forms of entry / exit from a control / configuration or NLINK circuit are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit);
[0046] Figure 4Exemplary embodiments of multiple control / configuration circuits or NLINK circuits in conjunction with a MAC pipeline according to certain aspects of the present invention are illustrated, wherein each NLINK circuit is connected to (and in one embodiment, dedicated to) an associated MAC pipeline (including multiple MACs connected in series); in one embodiment, the NLINK circuit may be configured to connect to one or more other NLINK circuits (where each NLINK circuit is associated with multiple MACs configured in one or more MAC pipelines), for example, to control, configure, and connect a linear MAC execution or processing pipeline configured to perform, for example, multiplication-accumulation operations; it is worth noting that this is an illustrative embodiment. For example, according to certain aspects of the invention, each control / configuration (NLINK) circuit-MAC pipeline pair forms a separate operable processing pipeline configured to perform, for example, multiplication-accumulation processing; in an exemplary embodiment, each control / configuration circuit (labeled "NLINK") includes MAC_I / MAC_O ports to form a loop path (e.g., a ring path) for intermediate accumulated values to traverse the linear MAC execution pipeline (including multiple connected (e.g., cascaded) multiplication-accumulation circuits (labeled "MAC pipelines" and illustrated in the block diagram, depicting an exemplary multiplier-accumulator circuit system (including, for example, 64 MACs) - see [link]). Figure 1A and Figure 1B Multiple MACs; it is worth noting that the DI_I and MAC_SI ports will read from memory (e.g., external memory, such as L2 memory) and input data from there into the MAC pipeline, and the MAC_SO port can be used to read data processed via the MAC pipeline and subsequently write it to external memory (e.g., external memory, such as L2 memory); it is worth noting that, as mentioned above, "port" is a physical point of entry / exit from the control / configuration or NLINK circuit; all physical forms of entry or exit from the control / configuration or NLINK circuit are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit);
[0047] Figures 5A-5C The illustration depicts different signal paths in an exemplary interconnect architecture of multiple control / configuration or NLINK circuits connected in series according to certain aspects of the invention, wherein each NLINK circuit is connected to (and in one embodiment, dedicated to) an associated multiplier-accumulator circuit pipeline, which, when connected to the control / configuration or NLINK circuit, is part of a composite / larger linear MAC pipeline formed by MACs connected in series with each illustrative pipeline architecture connected in series, wherein... Figure 5AAn exemplary interconnect architecture of multiple control / configuration or NLINK circuits connected in series according to certain aspects of the present invention is illustrated to form a single shift data path for input data (DI) operands traversing multiple processing circuitry systems in a cluster (or a portion thereof); here, the NLINK circuits are configured to be connected to one or more other NLINK circuits via the DI_O and DI_I ports of the control / configuration or NLINK circuits to form a single shift chain of the processing pipeline, wherein each control / configuration or NLINK circuit is connected (to and, in one embodiment, dedicated to) multiple associated multiplier-accumulator circuits; Figure 5B An exemplary interconnect architecture of multiple control / configuration or NLINK circuits connected in series according to certain aspects of the invention is illustrated to form a cyclic shift data path (e.g., a ring path) for intermediate accumulated values to traverse multiple processing circuitry systems in a cluster (or a portion thereof); here, the NLINK circuits are configured to connect to one or more other NLINK circuits via the MAC_I and MAC_O ports of the control / configuration or NLINK circuits to form a cyclic shift chain of a processing pipeline, wherein each control / configuration or NLINK circuit is connected to (and, in one embodiment, dedicated to) multiple associated multiplier-accumulator circuits; and Figure 5C The illustration depicts an exemplary interconnect architecture of multiple control / configuration or NLINK circuits connected in series according to certain aspects of the invention, forming a shift data path for a final accumulated value to traverse multiple processing circuitry systems in a cluster (or a portion thereof); here, the NLINK circuits are configured to connect to one or more other NLINK circuits via the MAC_SI and MAC_SO ports of the control / configuration or NLINK circuits to form a shift chain in the processing pipeline, wherein each control / configuration or NLINK circuit is connected to (and, in one embodiment, dedicated to) multiple associated multiplier-accumulator circuits; it is noteworthy that, for clarity, [the following is omitted] Figures 5A-5C The numerous connections, signals, and signal paths / lines between control / configuration or NLINK circuits;
[0048] Figures 6A-6C A more detailed schematic block diagram of an exemplary control / configuration circuit or NLINK circuit according to certain aspects of the present invention is illustrated, which is configured to route input data DI_I and DI_O among a plurality of NLINK circuits, such as Figure 5AAs shown, a single shift chain forms the input data DI operand, traversing multiple (here, 64) multiplier-accumulator circuits or a MAC processor's processing pipeline (organized in the illustrated embodiment as multiple rows of associated interconnected (e.g., cascaded) multiplier-accumulator circuits), each of which is associated with control / configuration circuitry or NLINK circuitry; in this illustrative embodiment, this could be a "bottom," "end," or "edge" NLINK circuit (see integral, complete, or composite pipelines including multiple NLINK circuits and associated pipelines). Figure 5A The control / configuration of NLINK A or one of the NLINK circuits (see NLINK A) Figure 6A It is configured to receive input data from, for example, a memory (via the DI_I port) and route such data to its associated processing pipeline, to the DI_O port in the NLINK circuit, to connect to (and output data to) the DI_I signal and port in the adjacent NLINK circuit (e.g., see monolithic, complete, or composite pipelines). Figure 5A The NLINK B (including multiple NLINK circuits and associated pipelines) and / or the DI_I signal in the NLINK circuit to connect to (and input data from) the DI_O signal and port in the adjacent NLINK circuit (here, "below" - as shown in the image) Figure 5A (as illustrated in the diagram); each adjacent NLINK circuit is associated with and dedicated to different multiple multiplier-accumulator circuits or MAC processors (e.g., multiplier-accumulator circuits organized into two rows of interconnected (e.g., cascaded) multiplier-accumulator circuits and arranged adjacent to them; in this illustrative embodiment, located in (monolithic, interlocked, combined, or complete pipelines - including multiple NLINK circuits and associated pipelines - see [reference]). Figure 5A One or more (or all) NLINK circuits within an end or edge NLINK circuit are configured to connect to the DI_O signals and ports of an adjacent NLINK circuit and route signal paths to the DI_O ports of that NLINK circuit via their associated execution pipeline, to connect to (and output data to) the DI_I signals and ports of the adjacent NLINK circuit – according to Figure 6B It is worth noting that, Figure 6BThe illustration in the NLINK circuit provides a general example of the input data flow or path through the NLINK circuit and its associated MAC pipeline to another / different NLINK circuit and its associated MAC pipeline (“upward” refers to the path of input data to, for example, an adjacent NLINK circuit). In this embodiment, the other / different NLINK circuit and its associated MAC pipeline are adjacent to the NLIBK circuit and its associated MAC pipeline; in this illustrative embodiment, it may be a “top” or second “end” or “edge” NLINK circuit (see integral, joined, combined, or complete pipeline). Figure 5A The NLINK X in the context of NLINK X—which includes multiple NLINK circuits and associated pipelines—see [link to NLINK X]. Figure 5A One of the control / configuration circuits or NLINK circuits (see Figure 6C It is configured to connect from another NLINK circuit (see...) Figure 5A and Figure 6B It receives data and routes this data to its associated processing pipeline, to the DI_O port in the NLINK circuit, to connect and output data to, for example, memory and ports in adjacent NLINK circuits (here, set up above NLINK - see...). Figure 5A It is worth noting that the second end or edge NLINK circuit is relative to Figure 6A The first end or edge of the NLINK circuit; as described above, a "port" (e.g., DI_I port, DI_O port) is a physical point for entering and exiting the control / configuration or NLINK circuit; in an exemplary embodiment, a circle containing a number is located in the data path (see...). Figure 6A The illustrations in the diagram indicate the number of bits and / or conductors or lines (e.g., data or control) for a particular path or port; for the avoidance of doubt, the block / data widths, data path widths, and / or port widths illustrated therein are merely exemplary and are not limiting in any way; it is worth noting that, for clarity, many other connections, signal paths, and signals (e.g., memory address, data and control paths, accumulated data paths, and output data paths) between NLINK circuits (and between NLINK circuits and their associated execution pipelines) have been omitted; however, Figures 3A-3F The diagram illustrates some such connections, signal paths, and signals; however, it is worth noting that in one embodiment, certain connections, signal paths, and signals are not modified (relative to...). Figures 3A-3F To implement the pipeline architecture of this exemplary embodiment (e.g., memory address, data and control paths, accumulated data paths, and output data paths);
[0049] Figures 7A-7CA more detailed schematic block diagram of an exemplary control / configuration circuit or NLINK circuit according to certain aspects of the present invention is illustrated. The exemplary control / configuration circuit or NLINK circuit is configured to route MAC_I and MAC_O among multiple NLINK circuits, such as... Figure 5B As shown, a cyclic shift path (e.g., a loop) is formed for intermediate accumulated values, traversing a processing pipeline of multiple (here, 64) multiplier-accumulator circuits or MAC processors (here, multiplier-accumulator circuits organized into two rows of interconnected (e.g., cascaded) multiplier-accumulator circuits) associated with each control / configuration circuit or NLINK circuit. Each control / configuration circuit or NLINK circuit is incorporated into an integral, cohesive, combined, or complete pipeline formed via the connection of the MAC circuits of each MAC processor associated with the interconnected NLINK circuits (see [link to pipeline]). Figure 5B In this illustrative embodiment, the accumulated data path illustrates a circuit configuration of an NLINK circuit to provide connections between multiple NLINK circuits (via an associated MAC processing pipeline associated with such NLINK circuits) to implement a cyclic shift path (e.g., a loop) for intermediate accumulated values; specifically, in this illustrative embodiment, it may be a "bottom" or first end or edge NLINK circuit (e.g., an NLINK A-monolithic, complete, or composite pipeline-including multiple NLINK circuits and associated pipelines-see) Figure 5B ) control / configuration circuit or NLINK circuit (see Figure 7A The NLINK circuit is configured to receive accumulated data from NLINK "X" via the MAC_I port and route this data to the associated processing pipeline, to the MAC_O port in NLINK circuit A, to connect to (and output data to) the MAC_I port of the adjacent NLINK circuit (e.g., NLINK B, set "above" NLINK - see above). Figure 5B and Figure 7B Each adjacent NLINK circuit is associated with a different plurality of multiplier-accumulator circuits or MAC processors (e.g., multiplier-accumulator circuits organized into two rows of interconnected (e.g., in series) multiplier-accumulator circuits and positioned adjacent to them (here, above); in this illustrative embodiment, located in (monolithic, interlocking, combined, or complete pipelines - including multiple NLINK circuits and associated pipelines - see [link]); Figure 5BOne or more NLINK circuits within the end or edge NLINK circuit are configured to connect to the MAC_I and MAC_O ports of adjacent NLINK circuits and route the accumulated data signal path to the MAC_O port of that NLINK circuit via the associated execution pipeline, to connect to (and output data to) the MAC_O signal and port of the adjacent NLINK circuit (see [link]). Figure 7B In this illustrative embodiment, it can be the control / configuration circuitry of the top, end, or edge NLINK circuitry or one of the NLINK circuits (NLINKX - see [link]) of a complete pipeline (including multiple NLINK circuits and associated pipelines). Figure 5B and Figure 7B It is configured to receive data from the NLINK circuit below (see below). Figure 7B This data is then routed to its associated processing pipeline, to the MAC_O port in the NLINK circuit, to connect and output the data to an adjacent NLINK circuit (here, the settings are configured below - see NLINK). Figure 5B To complete the cyclic shift path; therefore, multiple NLINKS and associated MAC pipelines are interconnected to form a “loop” relating to the accumulated data path via the MAC_O port and port, as well as the accumulated data path in the NLINK circuitry (see [link]). Figure 5B It is worth noting that each NLINK circuit is associated with multiple different multiplier-accumulator circuits or MAC processors (e.g., multiplier-accumulator circuits organized into two rows of interconnected (e.g., cascaded) multiplier-accumulator circuits), and these multiple different multiplier-accumulator circuits or MAC processors can be incorporated into a monolithic, coupled, combined, or complete pipeline via the configuration of the NLINK circuits; it is worth noting that... Figure 7BThe illustrations in the NLINK circuit provided a general illustration of the inflow and outflow of multiplication and accumulation data from the previous NLINK (and its associated MAC pipeline) to the next NLINK (and its associated MAC pipeline) as well as the single MAC pipeline (the "upward" path of the accumulated data for multiplication-accumulation operations) and associated LINKs; in one exemplary embodiment, circles with numbers in the data path indicate the number of bits and / or conductors or lines (e.g., data or control) of a particular path or port; for the avoidance of doubt, the block / data widths, data path and / or port widths illustrated are only... This is exemplary and not limited in any way; as previously stated, a “port” (e.g., MAC_I port, MAC_O port) is a physical point of entry and exit from the control / configuration or NLINK circuitry; it is worth noting that, for clarity, many other connections, signal paths, and signals (e.g., memory address, data and control paths, input data paths, and output data paths) between control / configuration or NLINK circuitry (and between NLINK circuitry and their associated execution pipelines) have been omitted; however, certain connections, signal paths, and signals (in one embodiment, without modification to implement the architecture of that embodiment) are present in... Figures 3A-3F The diagram shows (e.g., memory addresses, data and control paths, input data paths and / or output data paths);
[0050] Figures 8A-8C A more detailed schematic block diagram of an exemplary control / configuration circuit or NLINK circuit according to certain aspects of the present invention is illustrated. The exemplary control / configuration circuit or NLINK circuit is configured to route the MAC_SI and MAC_SO signals among multiple NLINK circuits, such as... Figure 5C As shown, a shift path (e.g., a loop) is formed for the final accumulated value, traversing multiple (here, 64) multiplier-accumulator circuits or a MAC data processor's processing pipeline (organized in the illustrated embodiment as multiple associated two-row interconnected (e.g., cascaded) multiplier-accumulator circuits), wherein each associated two-row interconnected (e.g., cascaded) multiplier-accumulator circuit is associated with and dedicated to control / configuration circuitry or NLINK circuitry; in this illustrative embodiment, this could be bottom, end, or edge NLINK circuitry (see overall, complete, or composite pipelines). Figure 5C The NLINK A (which includes multiple NLINK circuits and associated pipelines) is either a control / configuration circuit system or one of the NLINK circuits (see [link]). Figure 8AIt is configured to receive data from memory (via the MAC_SI port) and route such data to the associated processing pipeline, to the MAC_SO port in the NLINK circuit, to connect to (and output data to) the MAC_SO signal and port in the adjacent NLINK circuit (see, for example, monolithic, complete, or composite pipelines). Figure 5C Connect the MAC_SO port in the NLINK B circuit and / or the NLINK circuit to the MAC_SO port in the adjacent NLINK circuit (here, the settings are in the NLINK section below - see below) to connect to (and input data from) the MAC_SO port in the adjacent NLINK circuit. Figure 5C Each adjacent NLINK circuit is associated with a different plurality of multiplier-accumulator circuits or MAC processors (e.g., multiplier-accumulator circuits organized into two rows of interconnected (e.g., in series) multiplier-accumulator circuits and arranged adjacently (here, above); in this illustrative embodiment, located in (monolithic, interlocking, combined, or complete pipelines - including multiple NLINK circuits and associated pipelines - see [link]); Figure 5C One or more (or all) NLINK circuits within an end or edge NLINK circuit are configured to connect to the MAC_SI and MAC_SO ports of an adjacent NLINK circuit and route signal paths to the MAC_SO port of that NLINK circuit via an associated execution pipeline to connect to (and output data to) the MAC_SO signal and port of the adjacent NLINK circuit (see [link to NLINK circuit]). Figure 8B It is worth noting that, Figure 8B The illustration in the NLINK circuit provides a general illustration of the data flow from the previous NLINK (and its associated MAC pipeline) to the subsequent NLINK (and its associated MAC pipeline) through the NLINK circuit and its associated pipeline (“up” refers to the output data path of the accumulation operation with the previous result in the MAC pipeline and its associated NLINK circuit). In this embodiment, the subsequent NLINK is adjacent to it; in this illustrative embodiment, it may be a “top” or second-end or edge NLINK circuit (see overall, complete or composite pipeline). Figure 5C The NLINK X in the context of NLINK X (which includes multiple NLINK circuits and associated pipelines) is either a control / configuration circuit or one of the NLINK circuits (see [link]). Figure 8C It is configured to receive data from another NLINK circuit (see...) Figure 8B This data is routed to its associated processing pipeline, to the MAC_SO port in the NLINK circuit, to connect and output the data to, for example, memory and ports in adjacent NLINK circuits (here, the settings are configured above - see NLINK). Figure 5CIt is worth noting that, for clarity, many other connections, signal paths, and signals (e.g., memory address, data and control paths, input data paths, and accumulated data paths) between control / configuration or NLINK circuits (and between NLINK circuits and their associated execution pipelines) have been omitted; however, some of these connections, signal paths, and signals (in one embodiment, they are not modified to implement the architecture of that embodiment - e.g., memory address, data and control paths, input data paths, and accumulated data paths) in Figures 3A-3F The middle diagram shows circles containing numbers located in the data path (see...). Figure 8A The illustrations in the figures indicate the number of bits and / or conductors or lines (e.g., data or control) of a particular path or port in an exemplary embodiment; for the avoidance of doubt, the block / data widths, data path widths, and / or port widths illustrated therein are merely exemplary and are not limiting in any way; as stated above, a “port” (e.g., MAC_SI port, MAC_SO port) is a physical point for entering and exiting control / configuration or NLINK circuitry; all physical forms of entering or exiting control / configuration or NLINK circuitry are intended to fall within the scope of the invention (e.g., conductors or metal wiring in an integrated circuit); and
[0051] Figure 9 The illustration depicts an exemplary configurable processing circuitry system according to embodiments of certain aspects of the present invention to implement additional data processing operations, including, for example, preprocessing of data operands and postprocessing of accumulated results. Notably, the configurable processing circuitry can be organized into four circuit blocks (a0, a1, a2, a3), wherein each circuit block of the configurable processing circuitry system can be configured to perform one or more operations. Furthermore, the configurable processing circuitry system includes additional programmable / configurable circuitry to establish, configure, or "bootstrap" data paths to implement one or more preprocessing operations (e.g., preprocessing data operands) and / or one or more postprocessing operations (e.g., further / follow-up processing of accumulated results from a MAC processing pipeline, which can be a coupled, combined, or complete MAC processing pipeline, wherein multiple smaller MAC processing pipelines are coupled or combined via the configuration of NLINK circuitry associated with smaller MAC processing pipelines – see [link to documentation]). Figures 5A-5CThe configurable processing circuitry system can be one-time programmable (e.g., at manufacturing time, via, for example, a programmable fuse array) or multiple-time programmable (including, for example, at startup / power-on, initialization, and / or in-situ (i.e., during operation of the integrated circuit)); in one embodiment, the configuration is programmed via a multiplexer prior to the operation or implementation of the execution sequence of the processing pipeline to establish data paths into or bypass one or more selected processing circuits; notably, the configurable processing circuitry system and its connections are superimposed (for illustrative purposes) on Figure 3A , Figure 3C , Figures 6A-6C Figure 7 and Figures 8A-8C Detailed schematic block diagrams of exemplary configurations of the control / configuration circuitry or NLINK circuitry shown (see the left side of the "NLINK (top)" section for each exemplary configuration of the control / configuration circuitry or NLINK circuitry); in one embodiment, the configurable processing circuitry system may be, for example, via an interconnection network (see...). Figure 1B Accessible by any (or all) NLINK circuitry, wherein the interconnect network is configurable to connect a configurable processing circuitry system to one or more NLINK circuitry and the associated MAC processing pipeline; in one exemplary embodiment, circles with numbers located in the signal path indicate the number of bits and / or conductors or lines (e.g., data) of a particular path or port; for the avoidance of doubt, the block / data width, data path, or port width illustrated herein are merely exemplary and are not intended to limit in any way.
[0052] As described above, the pseudocode, operations, configurations, block / data widths, data path widths, bandwidths, data lengths, values, procedures, and / or algorithms depicted and / or illustrated in the accompanying drawings are exemplary, and the present invention is not limited to any particular or exemplary circuit, logic, block, function, and / or physical diagram illustrated and / or described, for example, the number of multiplier-accumulator circuits employed in an execution pipeline, the number of execution pipelines employed in a particular processing configuration / architecture, memory organization / allocation, block / data widths, data path widths, bandwidths, values, procedures, pseudocode, operations, and algorithms. Furthermore, although illustrative / exemplary embodiments include assigning, allocating, and / or using memory for storing certain data (e.g., filter weights) and / or multiple memories (e.g., L3 memory, L2 memory, L1 memory, L0 memory) in certain organizations. In practice, the organization of memory can be changed, in which one or more memories can be added, and / or one or more memories can be omitted and / or combined / merged with other memories - for example, (i) L3 memory or L2 memory and / or (ii) L1 memory or L0 memory.
[0053] To reiterate, this document describes and illustrates numerous inventions. The invention is neither limited to any single aspect or embodiment thereof, nor to any combination and / or substitution of such aspects and / or embodiments. Each aspect and / or embodiment of the invention may be used alone or in combination with one or more other aspects and / or embodiments of the invention. For the sake of brevity, many such combinations and substitutions are not discussed or illustrated separately herein. Detailed Implementation
[0054] In a first aspect, the present invention relates to circuit systems (and methods of operating and configuring such circuit systems) for configuring and controlling multiplier-accumulator circuit systems, which include multiple multiplier-accumulator execution or processing pipelines. The circuit systems of this aspect of the invention configure (e.g., programmable once or more than once) and control the multiplier-accumulator circuit systems to implement one or more execution or processing pipelines to process data, such as processing data in parallel or simultaneously. In one embodiment, the control / configuration circuitry controls the loading of filter weights into memory used by the multiplier-accumulator circuitry to implement multiplication operations. In this respect, the control / configuration circuitry can facilitate or control the writing of filter weights or values to multiple multiplier-accumulator circuits and the storage of such filter weights in memory accessible to the circuitry for multiplication operations. For example, the control / configuration circuitry can connect multiple multiplier-accumulator circuits (or rows / groups of interconnected (serialized) multiplier-accumulator circuits) to memory to facilitate the storage of filter weights in “local” memory accessible by the multiplier-accumulator circuitry. Reference Figure 2A In one embodiment, each multiplier-accumulator circuit includes a "local" memory (e.g., SRAM or register) to "locally" store filter weights or values used in relation to the multiplication operations of the associated MAC's multiplier circuitry. In this respect, the memory connects and outputs the filter weights / values to the associated MAC's multiplier circuitry. It is noteworthy that, in one embodiment, each "local" memory may be dedicated to the associated MAC.
[0055] refer to Figure 2B In one embodiment, each multiplier-accumulator circuit includes multiple "local" memory / register sets (e.g., memory a-memory x) to store multiple sets of different filter weights (each set of filter weights may be associated with a different set of data for MAC processing). For example, refer to... Figure 2B and Figure 2CMultiple “local” memory / register sets may include a first memory / register set and a second memory / register set to store two different sets of filter weights, including a first set of filter weights for the “current” multiplication operation and a second set of filter weights for the multiplication operation immediately following the current multiplication operation (i.e., immediately following the completion of the “current” multiplication operation using the first set of filter weights).
[0056] In operation, the multiplier-accumulator circuit can read a first set of filter weights from a first memory during a first set of multiplication operations, and upon completion of the first set of multiplication operations (associated with processing the first set of data), read a second set of filter weights from a second memory associated with a second set of multiplication operations (associated with processing the second set of data using the second set of filter weights). In this way, the multiplier-accumulator circuit can skip read operations between multiple "local" memory / register sets on a single set of multiplication operations. That is, the multiplier-accumulator circuit can read / access the first set of filter weights (i.e., stored in the first "local" memory / register set) for use in the current multiplication operation associated with processing the first set of input data, and upon completion, immediately read / access the next set / second set of filter weights stored in the second "local" memory / register set associated with processing the second set of data. Here, because the second set of filter weights can be written to and stored in the second memory / register set during or before the completion of the multiplication operation using the first set of filter weights stored in the first memory / register set, there is no delay or overhead time arising from the reading, access, and availability of the next set or second set of filter weights from the multiplier-accumulator circuit (stored in the second "local" memory / register set) during data processing performed by the multiplier-accumulator circuit or the multiplier-accumulator circuit processing pipeline (e.g., the second set of input data).
[0057] It is worth noting that the memory output selection circuit (e.g., one or more multiplexers) can responsively control which memory / register bank (and when) is connected to the multiplier circuitry of the multiplier-accumulator circuit. In this way, the memory output selection circuit responsively controls the jump read operation and the connection of the memory / register bank to the multiplier circuitry of the multiplier-accumulator circuit.
[0058] Furthermore, the control / configuration circuitry can alternately write a new or updated set of filter weights to multiple "local" memory / register sets. For example, see [reference to previous section]. Figure 2B and Figure 2CThe control / configuration circuitry can skip write operations between two "local" memory / register sets associated with each multiplier-accumulator circuit in the multiplier-accumulator execution or processing pipeline, based on a set of multiplication operations. That is, during the reading of the first set of filter weights or while completing a set of multiplication operations using the first set of filter weights stored in the first memory / register set, a new set of filter weights can be written (e.g., immediately) to the first memory / register set for use in processing after a set of multiplication operations using the second set of filter weights (i.e., a set of filter weights stored in the second "local" memory / register set). Here, the control / configuration circuitry provides or writes the new / next set of filter weights to the first "local" memory / register set, while the multiplier-accumulator circuitry performs multiplication operations using the second set of filter weights (stored in the second "local" memory / register set). This new, next, or third set of filter weights (stored in the first "local" memory / register set associated with each multiplier-accumulator circuit) can overwrite the first set of filter weights. Furthermore, during or immediately after data processing (via the associated multiplier-accumulator circuitry) is completed using the second set of filter weights or values stored in the second "local" memory / register set, a new, next, or third set of filter weights can then be made available to the multiplier-accumulator circuitry. In this way, there is no delay or overhead introduced into the multiplier-accumulator execution or processing pipeline during data processing due to writing, updating, or providing "new" filter weights or values to the "local" memory / register set used by the circuitry.
[0059] refer to Figure 2DIn one embodiment, multiple multiplier-accumulator circuits (e.g., n = 64) are configured (via control / configuration circuitry) in a linear multiplier-accumulator execution or processing pipeline. In this embodiment, the MAC is associated with and connected to multiple "local" memory / register sets (associated with and dedicated to a particular MAC) to store multiple different sets of filter weights for use in relation to multiplication operations associated with processing a given set of input data via the MAC's multiplier circuitry system. Here, each MAC processor includes two memory / register sets (e.g., L0, such as SRAM). In this embodiment, the two memory / register sets are independent memory sets such that in each execution cycle, one set from each MAC can be read (using a shared read address bus), placing the read data on the associated RD[p] signal line input to the multiplexer ("mux"), while the other memory sets can be written to (filter weights to be used in the next execution cycle). The read data is moved / written into the F register (D_r[p]) for use in the execution cycle. The F register (D_r[p]) writes new filter weights (Fkl values) for each execution cycle.
[0060] As described above, during an execution cycle, other memory / register banks (i.e., banks not read from during the execution cycle) can be used via write operations to store filter weights (using the WA address bus, which in one embodiment is shared / common among memory / register banks). Here, during the current processing operation, write data (i.e., filter weight values) can be written to a memory bank not accessed by the multiplier-accumulator circuitry. In one embodiment, filter weight data (e.g., the next set of filter weights to be used in processing) can be read from a larger memory (e.g., L1 SRAM external to the MAC processor) and subsequently stored in a memory / register bank (L0 SRAM) without interfering with the current / ongoing set of execution cycles of the current processing operation.
[0061] It is worth noting that, reference Figure 2D and Figure 2E Each MAC or MAC processor in a linear pipeline can be implemented with or with a single memory / register set (e.g., see [link to implementation]). Figure 2A Two memory / register set embodiments (see example) Figure 2C ) or two or more memory / register sets (see, for example, examples) Figure 2B )interface.
[0062] Regarding the execution cycle, please refer to... Figure 2D and Figure 2EEach multiplier-accumulator circuit (also referred to as a "processing element" or "MAC processor") includes a shift chain (D_SI[p]) for the data input (DIJk data). In one embodiment, the next Dijk data is shifted in, while the current Dijk data is used for the current set of execution cycles. The current Dijk data is stored in the D_i[p] register without changing during the current set of execution cycles.
[0063] In addition, each multiplier-accumulator circuit includes a shift chain (MAC_SO[p]) for preloading Yijl sums. During the execution cycle of the current group, the next group of Yijl sums is shifted in, while the current group's Yijl sum is calculated / generated. In this embodiment, each multiplier-accumulator circuit also uses a shift chain (MAC_SO[p]) for unloading or outputting Yijl sums. During the execution cycle of the current group, the previous Yijl sum is shifted out, while the current Yijl sum is generated. It is worth noting that the simultaneous use of the Yijl shift chain (MAC_SO[p]) for both preloading and unloading will be discussed in more detail below.
[0064] In each execution cycle, the filter weight value (Fkl value) in the D_r[p] register is multiplied by the Dijk value in the D_i[p] register via the multiplier circuitry, and the result is output to the MULT_r[p] register. In the next pipeline cycle, this product (i.e., the D*F value) is added to the Yijl accumulated value in the MAC_r[p-1] register (in the previous multiplier-accumulator circuitry), and the result is stored in the MAC_r[p] register. This execution process is repeated for the execution cycle of the current group. It is worth noting that the Yijl accumulated value is shifted (rotated) during the execution cycle of the current group.
[0065] In one aspect of the invention, control / configuration circuitry configures (e.g., programmable once or more than once) and controls a multiplier-accumulator circuitry system to implement one or more execution or processing pipelines to process data, such as processing data in parallel or simultaneously. In one embodiment, the circuitry system configures and controls rows / groups of multiple separate multiplier-accumulator circuits or interconnected (serialized) multiplier-accumulator circuits to perform data processing via pipelined multiplication and accumulation operations, thereby increasing, for example, the throughput of the multiplier-accumulator execution or processing pipeline associated with processing data (e.g., image data).
[0066] In one embodiment, the circuitry for controlling and configuring the multiplier-accumulator circuitry system (sometimes referred to as the control / configuration circuitry system) includes circuitry for controlling, configuring, and / or programming (e.g., one-time or more-than-one-time programmable) one or more execution or processing paths of the multiplier-accumulator circuitry (including MAC processing pipelines). For example, in one embodiment, the control / configuration circuitry system may configure or connect a selected number of multiplier-accumulator circuits or rows / groups of multiplier-accumulator circuits to, among other things, implement a predetermined multiplier-accumulator execution or processing pipeline or its architecture. Here, the control / configuration circuitry system may configure or determine the pipeline architecture or configuration implemented via interconnected (serialized) rows / groups of multiplier-accumulator circuits or interconnected multiplier-accumulator circuits used for multiplication and accumulation operations and / or the connections of the multiplier-accumulator circuits (or rows / groups of interconnected multiplier-accumulator circuits) used for multiplication and accumulation operations. Therefore, in one embodiment, a control / configuration circuitry system (which may include multiple control / configuration circuits) configures or implements an architecture for executing or processing pipelines by controlling or providing connections between rows of multiplier-accumulator circuits and / or interconnected multiplier-accumulator circuits.
[0067] For example, control / configuration circuitry or NLINK circuitry is connected to a multiplier-accumulator circuitry system (including multiple (64 illustrated here) multiplier-accumulator circuits or MAC processors) to configure the overall execution pipeline, among other things, by providing, transferring, or “bootstrapping” data between one or more MAC pipelines via programmable or configurable interconnect paths. Furthermore, the control / configuration circuitry can configure the interconnects between the multiplier-accumulator circuitry system and one or more memories—including external memories (e.g., L3 memories, such as external DRAM)—that can be shared by one or more (or all) clusters of the MAC execution pipeline. These memories can store, for example, input image pixels Dijk, output image pixels Yijl (i.e., image data processed by the circuitry system of the one or more MAC pipelines, and filtering weights Fijklm associated with such data processing). (See also...) Figure 1A , Figure 1B , Figure 2D and Figure 2E ).
[0068] It is worth noting that the configuration can be implemented, for example, in situ (i.e., during the operation of the integrated circuit) to meet or exceed time-based system requirements or constraints, for example. Furthermore, the configuration implemented by the control / configuration circuitry system can be programmable once (e.g., at manufacturing time, via, for example, a programmable fuse array) or programmable multiple times (including, for example, at startup / power-on, initialization, and / or in situ).
[0069] refer to Figure 3A In one exemplary configuration, the control / configuration or NLINK circuitry (in one embodiment, dedicated to the associated execution pipeline) is connected to the execution pipeline via multiple ports, including (i) DI_I, MAC_SI, DI_O, and MAC_SO ports that connect the execution pipeline to external memory (e.g., L2 memory, such as SRAM), and (ii) MAC_I and MAC_O ports that connect multiple multiplier-accumulator circuits (or two rows of multiplier-accumulator circuits) in a ring configuration or architecture. Therefore, in this exemplary embodiment, the control / configuration or NLINK circuitry is configured to provide input data to the execution pipeline (e.g., from memory) and receive output / processed data from the execution pipeline (e.g., output to memory) – wherein the pipeline includes those MACs associated with and dedicated to the NLINK circuitry. Furthermore, the control / configuration or NLINK circuitry is not configured to interact with (one or more) other or adjacent control / configuration or NLINK circuits and / or other or adjacent MAC processing pipelines (e.g., see [link to NLINK circuitry]). Figure 1B Interface, communication, or interaction.
[0070] It is worth noting that, as mentioned above, a “port” is a physical point of entry and / or exit from a control / configuration or NLINK circuit; all physical forms of entry or exit from a control / configuration or NLINK circuit are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit).
[0071] refer to Figure 3A , Figure 3B , Figure 3E and Figure 3F The execution pipeline in this embodiment includes multiple multiplier-accumulator circuits (each labeled "MAC"), which are connected in series to form a linear execution pipeline of multiple rows of MACs interconnected via row interconnects. In operation, the execution pipeline receives input data at DI_I (see "Input Data Port" in NLINK (below), processes the data via a multiplication and accumulation operation, and outputs the processed data at MAC_SO (see "Output Data Port" in NLINK (above)). As described above, the control / configuration or NLINK circuitry configures the two rows of multiplier-accumulator circuits in a ring configuration or architecture via processing operations connecting the MAC_I and MAC_O ports and thus connecting the multiple multiplier-accumulator circuits of the execution pipeline.
[0072] refer to Figures 3A-3FThe signals traversed along the L0 memory data, address, and control paths in the NLINK and execution pipeline represent control and address information for managing and controlling the execution sequence of processing. These control and address signals are generated by a control logic circuitry system (not shown – and in one embodiment, it is external to the NLINK and execution pipeline). Furthermore, this control logic also manages or controls the write data (i.e., filter weights) and write sequences associated with each L0 memory (e.g., SRAM), which is associated with and dedicated to the multiplier-accumulator circuitry in the execution pipeline. The write data (filter weights) is read from a memory outside the execution pipeline (L1 memory – e.g., SRAM) (see, for example, see...). Figure 1A and Figure 1B It is worth noting that before the execution process using the filter weights associated with the sequence is initiated, the sequence is performed / completed to write data into one of the groups of L0 memories associated with each multiplier-accumulator circuit.
[0073] Continue to refer to Figures 3A-3F The signals traversed along the input data path in the NLINK and execution pipeline represent input data (e.g., image data) applied to or received by the MAC pipeline and processed / processed in the execution sequence. Input data may be stored in L2 memory (e.g., SRAM) and provided (i.e., read from memory) to the NLINK via a control logic circuitry system (not shown). In one embodiment, this input data is provided to the pipeline in groups or sets (e.g., a set or group of 64 elements, where each element is 17 bits in size / length). A group or set of input data is serially shifted in through the DI_I port and loaded into the pipeline in parallel with the D registers in each multiplier-accumulator circuit of the execution pipeline. The end of the serial path is the DI_O output port in the NLINK circuitry (see “Input Data Port” in NLINK (above)). In this illustrative pipeline architecture / configuration, the input data path in NLINK (above) may not be used.
[0074] The accumulated data path in the NLINK and execution pipeline represents the ongoing MAC accumulation total generated by the multiplier-accumulator circuitry during the execution sequence. In each cycle, each multiplier-accumulator circuitry multiplies the filter weight value (from L0 memory) by the (static) value in the D register, adds the total to its accumulation register Y, and passes Y to the right for the next cycle. The accumulated value of Y (64 here) is rotated counterclockwise, and at the end of each cycle interval (e.g., a 64-cycle interval), the accumulated value is loaded into the output shift register MAC_S. Note that the accumulated data ports (MAC_O port and MAC_I port) in the NLINK circuitry, and the data path between them, are configured and enabled to allow the Y value to be rotated through each MAC in the execution pipeline (see also...). Figure 2D ).
[0075] Continue to refer to Figures 3A-3F The signals on the output data path in the NLINK circuitry and execution pipeline represent the output data generated by the execution sequence (the MAC pipeline and an associated NLINK circuitry with an accumulation operation having a previous result; note that "shifting MAC_Sx within a single pipeline" refers to the output data path corresponding to the accumulation operation with a previous result)). The output data is loaded in parallel from the accumulation register Y (64 here) into the MAC_SO register (64 here). In one embodiment, this output data can be written to memory (e.g., L2 memory (SRAM)) outside the NLINK circuitry and execution pipeline via a control logic circuitry system (not shown). In one embodiment, the output data is returned as a set of 64 elements, each 35 bits in length / size. The set of output data is serially shifted out through the output data port (MAC_SO port) and stored in memory. The serial bus (MAC_SO and MAC_SI) can also be used to preload the initial accumulated value / total before starting the first execution cycle of each execution sequence. These initial accumulated totals are provided (e.g., serially shifted in) to the NLINK circuitry and the execution pipeline via the MAC_SI port, and loaded in parallel into the Y register of each multiplier-accumulator circuit in the execution pipeline (see also...). Figure 2D In one exemplary embodiment, the initial accumulated data size / length is 35 bits.
[0076] In one embodiment, multiple control / configuration or NLINK circuits can be interconnected to, for example, configure a data processing circuit including multiple MAC pipelines—each pipeline associated with one of the interconnected control / configuration or NLINK circuits. In one embodiment, multiple (or all) MAC pipelines, such as a cluster of MAC pipelines, are employed in the processing operation of the circuit via the control / configuration or NLINK circuits associated with these MAC pipelines. For example, refer to... Figure 4 In one embodiment, each of the multiple NLINK circuits (e.g., associated with multiple (or all) MAC pipelines of a cluster) is configured to provide data (e.g., image data) to the associated MAC pipeline (including multiple interconnected multiply-accumulate circuits) via DI_I and MAC_SI ports. In this embodiment, the MAC_I / MAC_O ports of each NLINK circuit are connected to provide a ring topology or architecture of interconnected multiply-accumulate circuits for the associated MAC pipelines. Additionally, the DI_I and MAC_SI ports are configured to provide data (data read from memory (e.g., L2)) to each MAC pipeline for processing and to output MAC_SO and write it back to memory (L2). In this embodiment, partially processed data (from a given MAC pipeline) is exchanged between memory (e.g., L2) and each of the multiple MAC pipelines via the NLINK circuits to perform processing operations.
[0077] It is worth noting that, as mentioned above, the input data does not need to be written back to memory; therefore, the input data port (DI_O port) can remain unconnected at the NLINK output (see [link]). Figure 3C ).
[0078] In one embodiment, multiple control / configuration or NLINK circuits and associated MAC pipelines can be configured and interconnected into a single shift chain, wherein multiple MAC pipelines are employed in processing operations (e.g., a cluster) via the control / configuration or NLINK circuits associated with such MAC pipelines and their configurations. Here, the data processing pipeline includes multiple MAC pipelines (each associated with one of the interconnected NLINK circuits), which are interconnected into a single shift chain via programmable or configurable connections between multiple control / configuration or NLINK circuits and via connections between multiple MAC pipelines provided / implemented by NLINK circuits.
[0079] For example, refer to Figure 5AIn one embodiment, a single shift chain is provided via the DI / I and DI / O ports of interconnected adjacent NLINK circuits (and by extension, each MAC pipeline associated with each NLINK circuit of the interconnected NLINK circuits). Here, the DI_O port of an NLINK circuit is connected to the DI_I port of an adjacent NLINK circuit (e.g., in the illustrated embodiment, the NLINK circuit positioned "above," where data flows from bottom to top (NLINK A to NLINK B, etc., to NLINK X)). It is worth noting that the data flow can be from top to bottom (i.e., from NLINK X, etc., to NLINK B, to NLINK A) – or any other direction or path – all of which are intended to fall within the scope of this invention.
[0080] Continue to refer to Figure 5A In the embodiment illustrated here, the DI_I port of the bottommost NLINK circuit is configured to receive input data (e.g., read from external memory, such as L2 memory – SRAM). The DI_O port of this bottommost NLINK circuit is input to the DI_I port of its adjacent NLINK circuit. Therefore, in this embodiment, the shift chain for processing operations uses the DI / I and DI / O ports of adjacent NLINK circuits to form a shift path for the data input (DI) operand, which traverses each multiplier-accumulator circuit of each MAC pipeline in a plurality of MAC pipelines.
[0081] In one embodiment, multiple MAC pipelines for a given cluster are incorporated into a shift chain. (See, for example, Figure 1B In another embodiment, all MAC pipelines of a given cluster are incorporated into a shift chain. It is worth noting that the DI_O port of the topmost NLINK circuit may or may not output to memory. In one embodiment, the DI_O port of the last NLINK circuit in the interconnected circuitry is unconnected (e.g., see...). Figure 5A (NLINK X in the middle).
[0082] refer to Figures 6A-6C ,exist Figure 5A In one embodiment of the shift chain shown, the NLINK circuitry is configured to provide appropriate DI_I-DI_O routing between them to load pipelines for combination (which consist of multiple pipelines; in the illustrative embodiment, each of the multiple pipelines includes two rows of multiply-accumulate circuitry—see example...) Figure 3A , Figure 3B , Figure 3E and Figure 3F The data here is related to... Figures 6A-6CThe diagram illustrates the connection of DI_I-DI_O routes between NLINK circuits, along with the associated input data signals and signal paths (see [link]). Figure 5A (DI_I-DI_O routing in the context of NLINK circuitry). These connections configure the DI_I port of one of the NLINK circuits (“first NLINK circuit” – e.g., NLINK A) to receive input data (e.g., from memory (e.g., L2 memory)) and provide that input data to the associated execution pipeline for processing, and then to the DI_O port. (See also: [link to NLINK circuitry]) Figure 6A The DI_O port of the first NLINK circuit is configured to output initial data to an adjacent NLINK circuit (“second NLINK circuit” – e.g., NLINK B) via an input data port (see upper NLINKS) and by extending the execution pipeline associated with the adjacent NLINK circuit. (See also...) Figure 6B Here, the DI_I port of the second NLINK circuit is connected to the DI_O port of the first NLINK circuit to provide some of the processed data to the execution pipeline associated with the second NLINK circuit. That is, the DI_O signal in the NLINK circuit is connected to the DI_I signal in the adjacent NLINK circuit (here, above the first NLINK circuit - see...). Figure 5A The DI_O port of the second NLINK circuit is configured to output data that has undergone further partial processing to another adjacent NLINK circuit (“the third NLINK circuit”) and the execution pipeline associated with the third NLINK circuit. And so on – for example, NLINK X. As described herein, a “port” (e.g., DI_I port, DI_O port) is a physical point of entry and exit from a control / configuration or NLINK circuit; all physical forms of entry or exit from a control / configuration or NLINK circuit are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit).
[0083] refer to Figure 6C It can be a top, end, or edge NLINK circuit (whole, complete, or combined pipeline). Figure 5A In NLINK X (which includes multiple NLINK circuits and associated pipelines), the control / configuration circuit or one of the NLINK circuits is configured to control / configure from another NLINK circuit (like...). Figure 6B It receives data (as in the example) and routes this data to its associated processing pipeline, to the DI_O port in the NLINK circuit, to connect and output the data to, for example, memory (e.g., L2 memory—such as SRAM). It is worth noting that... Figures 6A-6COther connections and configurations of the NLINK circuit (e.g., between the NLINK circuit and its associated execution pipeline) like Figure 3A As shown in those examples, but for clarity... Figures 6A-6C Not described in the text (e.g., memory addresses, data and control paths, accumulated data paths, and / or output data paths - see, for example, Figures 3A-3F ).
[0084] In short, during operation, the execution pipeline associated with each NLINK circuit interconnected with the NLINK circuits passes through the input data path before data processing. Figure 6A The output data path loads data via data read from the memory where the shifted data is stored, performing an accumulation operation with previous results in a MAC pipeline and an associated NLINK circuit ("shifting MAC_Sx within a single pipeline" refers to the output data path corresponding to the accumulation operation with previous results - see example...). Figure 3A Each pipeline of the interconnected NLINK circuitry is loaded into the overall / larger / composite pipeline (which is a combination of all pipelines associated with interconnected NLINK circuitry). (See also:) Figure 5A and Figures 6A-6C In one embodiment, multiple MAC pipelines of a given cluster are incorporated into a cyclic shift path or ring architecture via associated NLINK circuits interconnected in the architecture. In another embodiment, all MAC pipelines of a given cluster are incorporated into a cyclic shift path or ring architecture consisting of all MAC pipelines connected together via associated NLINK circuits, such as... Figure 5A As shown in the diagram.
[0085] It is worth noting that the size or length of a composite or combined MAC pipeline can be configured via NLINK circuitry to incorporate associated execution pipelines into the composite pipeline (which is a combination of all pipelines associated with interconnected NLINK circuitry). The size, length, or number of MACs in a MAC pipeline is configurable (larger / increased or smaller / decreased) via NLINK circuitry associated with (and, in one embodiment, dedicated to) multiple MAC pipelines including the composite or combined MAC pipeline. For example, a larger composite pipeline includes more interconnected NLINK circuitry (each associated with one or more MAC pipelines), which interconnects a larger number of MAC pipelines (and a larger number of MACs) into the composite or combined MAC pipeline. In contrast, a smaller composite pipeline includes fewer interconnected NLINK circuitry, which interconnects a smaller number of MAC pipelines (and a smaller number of MACs) into the composite or combined MAC pipeline. As mentioned above, the configuration implemented by the control / configuration circuitry system can be one-time programmable (e.g., at manufacturing time, via, for example, a programmable fuse array) or multiple-time programmable (including, for example, at startup / power-on, initialization, and / or in-situ).
[0086] In another embodiment, the MAC circuitry system is interconnected via the MAC_I and MAC_O ports of adjacent NLINK circuits, and a single shift path is configured in a cyclic shift path (loop) for intermediate accumulated values. (See reference...) Figure 5B The MAC_O port of an NLINK circuit (e.g., NLINK A) is connected to the MAC_I port of an adjacent NLINK circuit (e.g., in the illustrated embodiment, NLINK B – an NLINK circuit positioned "above" NLINK A, where the data flow is from bottom to top). In this illustrated embodiment, the MAC_I port of the bottommost NLINK circuit is configured to receive intermediate accumulated data from the topmost LINK circuit via its MAC_O port to complete a ring architecture (e.g., in the illustrated embodiment, NLINK X – an NLINK circuit positioned "top," end, or edge" of the NLINK, where the data flow completes the ring from top to bottom). The MAC_O port of the bottommost NLINK circuit is input to the MAC_I port of its adjacent NLINK circuit. Therefore, the cyclic shift path or ring architecture of the processing operation in this embodiment uses the MAC_I and MAC_O ports of adjacent NLINK circuits to form a path traversing each multiplier-accumulator circuit of each MAC pipeline across multiple MAC pipelines to transfer or provide intermediate accumulated values through / between multiple MAC pipelines.
[0087] In one embodiment, multiple MAC pipelines of a given cluster are incorporated into a cyclic shift path or ring architecture via associated NLINK circuits interconnected in the architecture, such as Figure 5B As reflected in [the document]. In another embodiment, all MAC pipelines of a given cluster are incorporated into a cyclic shift path or ring architecture via an associated NLINK circuit configuration.
[0088] refer to Figures 7A-7C In the Figure 5B In one embodiment of the single shift path in the cyclic shift path or ring architecture configuration shown, the NLINK circuitry is configured to provide appropriate MAC_I and MAC_O routing. First, the signal path labeled "Accumulated Data Path" is illustrated as... Figure 3A The configuration shown is a ring topology or architecture for multiplication-accumulation circuits, where the MAC_I / MAC_O ports of each NLINK circuit are connected to provide interconnection of associated MAC pipelines. The accumulated data path transfers the accumulated data from intermediate accumulations to adjacent NLINK circuits (“second NLINK circuits” – see example...). Figure 5B The NLINK (B) is extended and transferred to the execution pipeline associated with the adjacent NLINK circuit, where the intermediate accumulated input is fed into its MAC pipeline. The NLINK is configured to output the MAC_O signal in the topmost NLINK circuit via the accumulated data path through the NLINK circuit (instead of traversing the pipeline), to connect to the MAC_I signal in the bottommost NLINK circuit (see example NLINK B). Figure 5B In other words, the signals on the accumulated data path not applied to the pipeline illustrate the connection of the MAC_I and MAC_O routes between adjacent NLINK circuits (see NLINK X). Figure 5B (MAC_I and MAC_O routes in the code).
[0089] It is worth noting that, Figures 7A-7C Other connections and configurations of the NLINK circuit (e.g., between the NLINK circuit and its associated execution pipeline) like Figure 3A As shown in those examples, but for clarity... Figures 7A-7C Not described in the text (e.g., memory addresses, data and control paths, input data paths, and accumulated data paths - see example) Figures 3A-3F Furthermore, although a bottom-up data flow has been described ( Figures 5A-5C and Figures 7A-7CHowever, data flow can be top-down (i.e., NLINK X to NLINK B to NLINK A) – or any other direction or path – all of which are intended to fall within the scope of this invention. Furthermore, as described herein, a “port” (e.g., MAC_I and MAC_O ports) is a physical point of entry and exit from control / configuration or NLINK circuitry; all physical forms of entry and exit from control / configuration or NLINK circuitry are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit).
[0090] In short, in operation, the execution pipeline associated with each NLINK circuit interconnected with the NLINK circuits proceeds via the input data path (see, for example, ...) before data processing. Figures 6A-6C Data read from, for example, the memory into which input data is shifted, and output data (which may represent accumulated data of previous results) ("Shift MAC_Sx within a single pipeline" refers to the output data path corresponding to the accumulation operation with previous results – see example) Figure 3A This loads data into each pipeline of the interconnected NLINK circuitry that forms the overall / larger pipeline (which is a combination of all pipelines associated with the connected NLINK circuitry). (See also...) Figure 5B and Figures 7A-7C In one embodiment, all MAC pipelines of a given cluster are incorporated into a closed-loop or ring architecture consisting of all MAC pipelines connected together via associated NLINK circuitry.
[0091] It is worth noting that the size or length of the overall, complete, or combined pipeline can be configured via the NLINK circuitry to incorporate associated execution pipelines into an overall, larger, combined, or composite pipeline (which is a combination of all pipelines associated with the connected NLINK circuitry). As mentioned above, the configuration implemented by the control / configuration circuitry system can be one-time programmable (e.g., during manufacturing / testing, via, for example, a programmable fuse array) or multiple-time programmable (including, for example, during startup / power-on, initialization, and / or in-situ).
[0092] In another embodiment, a single shift path is configured in a cyclic shift path (loop) for the final accumulated value via the MAC_SI and MAC_SO ports that interconnect adjacent NLINK circuits. (See reference) Figure 5CThe MAC_SO port of the NLINK circuit is connected to the MAC_SI port of the adjacent NLINK circuit (e.g., in the illustrated embodiment, the NLINK circuit positioned "above" (e.g., NLINK B relative to NLINK A), where the data flow is from bottom to top). In this illustrated embodiment, the MAC_SI port of the bottommost NLINK circuit is configured to receive input data from memory (e.g., L2 memory, such as SRAM). Furthermore, the topmost LINK circuit is configured to output data to memory via its MAC_SO port. The MAC_SO port of the bottommost NLINK circuit (NLINK A) is input to the MAC_SI port of its adjacent NLINK circuit (NLINK B). Therefore, the single shift path implemented via a cyclic shift path for processing the final accumulated value of the operation in this embodiment utilizes the MAC_SI and MAC_SO ports of the adjacent NLINK circuits to form a path traversing each multiplier-accumulator circuit of each MAC pipeline across multiple MAC pipelines to transfer or provide intermediate accumulated values through / between multiple MAC pipelines. In one embodiment, all MAC pipelines of a given cluster are incorporated into a circular shift path or a ring architecture. In another embodiment, multiple (but not all) MAC pipelines of a given cluster are incorporated into a circular shift path or a ring architecture.
[0093] refer to Figures 8A-8C ,exist Figure 5C In one embodiment of a single shift path configured in a cyclic shift path or ring architecture shown, the NLINK circuitry is configured to provide appropriate MAC_SI and MAC_SO routing. First, certain configurations of the NLINK circuitry (e.g., ports identified as input data ports and accumulated data ports) are like... Figure 3A As shown, but for clarity, in Figures 8A-8C Not described in the text. The output data path diagram in the NLINK circuit illustrates the configuration of a cyclic shift path (ring) architecture used to implement the final accumulated value of each of the processing sequences(s) in the processing sequence(s).
[0094] The NLINK circuitry is configured to provide appropriate MAC_SI-MAC_SO routing to load pipelines for combination (consisting of multiple pipelines; where, in the illustrative embodiment, each of the multiple pipelines includes two rows of multiply-accumulate circuitry—see, for example...). Figure 3A , Figure 3B , Figure 3E and Figure 3F (Data from [reference]). Figure 8A In this embodiment, the output data path is configured on the MAC_SI port of one of the NLINK circuits (“first NLINK circuit” - for example, Figure 5CThe NLINK A in the first NLINK circuit receives data (e.g., from memory) and provides that data to the associated execution pipeline for processing, and also provides it to the MAC_SO port. The MAC_SO port of the first NLINK circuit is configured to... Figure 8A The output data port on the NLINK circuit shown in the diagram outputs data to an adjacent NLINK circuit ("second NLINK circuit" - for example, Figure 5C (NLINK A in the example). Subsequently, data is provided to the execution pipeline associated with the adjacent NLINK circuit via the MAC_SO / MAC_SI port of the adjacent NLINK circuit (see NLINK below in Figure 8). Here, the MAC_SO port of the second NLINK circuit (e.g., ...) Figure 5C The NLINK B in the second NLINK circuit is connected to the MAC_SI port of the pipeline. The MAC_SO port of this second NLINK circuit is configured to connect to and output the processed data to another adjacent NLINK circuit (“the third NLINK circuit”) and the execution pipeline associated with the third NLINK circuit. And so on. (For example, Figure 5C (NLINK X in NLINK). This configuration of the NLINK circuit connects MAC_SO and MAC_SI to facilitate pipelined operation.
[0095] As described herein, a “port” (e.g., MAC_SI port, MAC_SO port) is a physical point of entry and exit from a control / configuration or NLINK circuit; all physical forms of entry and exit from a control / configuration or NLINK circuit are intended to fall within the scope of this invention (e.g., conductors or metal wiring in an integrated circuit).
[0096] refer to Figure 8C This can be a monolithic, complete, or combined pipeline—including multiple NLINK circuits and associated pipelines—see [link to documentation]. Figure 5C The control / configuration circuit or one of the NLINK circuits at the top, end, or edge of the NLINK X is configured to receive data from another NLINK circuit (see [link]). Figure 8B This path is routed to the associated processing pipeline, to the MAC_SO port in the NLINK circuit, to connect to and output data to, for example, memory (e.g., L2 memory such as SRAM). It is worth noting that... Figures 8A-8C Certain other signal / data paths and configurations of the NLINK circuit (e.g., between the NLINK circuit and its associated execution pipeline) like Figure 3A As shown (e.g., input data path), but not for clarity. Figures 8A-8CThe details are as follows (e.g., memory addresses, data and control paths, input data paths, and accumulated data paths - see, for example, Figures 3A-3F ).
[0097] In short, during operation, the execution pipeline associated with each NLINK circuit interconnected with the NLINK circuits runs on the output data path before data processing (see [link to NLINK pipeline]). Figure 8A The data, which is the final accumulated data value, is loaded into each pipeline of the interconnected NLINK circuitry, forming a whole / larger pipeline (which is a combination of all pipelines associated with connected NLINK circuitry), via data read from the memory where the data has been moved in. (See also...) Figure 5C and Figures 8A-8C As described above, in one embodiment, all MAC pipelines of a given cluster are incorporated via associated NLINK circuitry into a cyclic shift path or ring architecture consisting of all MAC pipelines connected together, such as... Figure 5C As shown in the image.
[0098] It is worth noting that the size or length of an overall, complete, or combined pipeline can be configured via NLINK circuitry to incorporate associated execution pipelines into an overall, larger, combined, or composite pipeline (which is a combination of all pipelines associated with the connected NLINK circuitry) – or to configure NLINK circuitry to incorporate associated execution pipelines into a smaller, combined, or composite pipeline (which is a combination of fewer pipelines associated with the connected NLINK circuitry). As mentioned above, the configuration implemented by the control / configuration circuitry system can be programmable at one time (e.g., at manufacturing time, via, for example, a programmable fuse array) or programmable multiple times (including, for example, at startup / power-on, initialization, and / or in-situ).
[0099] In yet another embodiment, an exemplary pipeline and interconnect architecture of the control / configuration or NLINK circuitry and its associated MAC pipeline includes multiple such MAC execution pipelines connected in series via a series connection of multiple control / configuration or NLINK circuits, wherein the pipeline architecture includes two or three / all of the following: (i) providing a single shift chain or path via interconnecting the DI / I and DI / O ports of adjacent NLINK circuits and each MAC pipeline associated with each NLINK circuit of the interconnected NLINK circuits (see [link to documentation]). Figure 5A and Figures 6A-6C (ii) To interconnect the MAC circuit system via the MAC_I and MAC_O ports of adjacent NLINK circuits to configure a single shift path for intermediate accumulated values in a cyclic shift path (loop) (see Figure 5B and Figures 7A-7C(iii) and configuring a single shift path for the final accumulated value via the MAC_SI and MAC_SO ports of interconnecting adjacent NLINK circuits in a cyclic shift path (loop) (see Figure 5C and Figures 8A-8C In this embodiment, each cascaded control / configuration circuit is associated with (and in one embodiment, dedicated to and / or directly connected to) one of a plurality of MAC pipelines comprising a composite MAC execution pipeline, wherein, when the control / configuration or NLINK circuitry is connected, each multiplier-accumulator circuit pipeline is part of a composite / larger linear MAC pipeline formed by cascaded MACs associated with each cascaded MAC pipeline. (See also...) Figures 5A-5C Furthermore, certain signal / data paths and configurations of NLINK circuits (e.g., between NLINK circuits and their associated execution pipelines) can be like... Figures 3A-3F Implemented as described in the document (e.g., memory addresses, data, and control paths).
[0100] It is worth noting that in this embodiment (i.e., Figures 5A-5C In this context, the size or length of composite, complete, or combined pipelines can be configured via programming or configuration of the NLINK circuitry to incorporate associated execution pipelines into a larger composite pipeline (which is a combination of all pipelines associated with connected NLINK circuitry) or into a smaller composite pipeline (which includes fewer interconnected NLINK circuitry and fewer associated MAC pipelines). As mentioned above, the configuration implemented by the control / configuration circuitry system can be programmable at one time (e.g., at manufacturing time, via, for example, a programmable fuse array) or programmable multiple times (including, for example, at startup / power-on, initialization, and / or in-situ).
[0101] refer to Figure 9 According to embodiments of certain aspects of the present invention, the integrated circuit may include a configurable processing circuitry system to implement additional data processing operations, including, for example, preprocessing of data operands and postprocessing of accumulated results. Here, the configurable processing circuitry is organized into four circuit blocks (a0, a1, a2, a3), wherein each circuit block of the configurable processing circuitry system can be configured to perform one or more operations. Furthermore, the configurable processing circuitry system includes additional programmable / configurable circuitry systems to establish, configure, or “bootstrap” data paths to implement one or more preprocessing operations (e.g., preprocessing data operands) and / or one or more postprocessing operations (e.g., further / follow-up processing of accumulated results from a MAC processing pipeline – which may be a composite, complete, or combined MAC processing pipeline – see, for example, Figures 5A-5C ).
[0102] The configurable processing circuitry system includes circuitry that performs one or more floating-point and fixed-point operations (whether for preprocessing input data and / or post-processing of accumulated results). These operations include, for example, adding / subtracting register values, multiplying register values, converting values from integer data format to floating-point data format, converting values from floating-point data format to integer data format, adjusting the format precision of values (e.g., increasing or decreasing the precision of values), and performing one or more transformations using one or more unary functions (e.g., one or more of the following: inversion, square root, inverse square root, hyperbolic tangent, sigmoid, etc.). Continue to refer to Figure 9 The multiplexer in the configurable processing circuitry system configures and establishes data paths and / or directs data through these paths to one or more selected processing circuits to perform one or more selected operations and / or bypass one or more operations. In one embodiment, the path and / or directing are configured via control of the multiplexer in the configurable processing circuitry system prior to initiating the execution sequence (i.e., performing pipelined processing via MAC).
[0103] For example, in one embodiment, a configurable processing circuitry system may include an activation circuitry system to perform one or more operations or processes, including, for example, linear and / or nonlinear activation operations and / or threshold functions, as described and / or illustrated in U.S. Patent Application No. 63 / 144,553, filed February 2, 2021, entitled “MAC Processing Pipeline having Activation Circuitry, and Methods of Operating Same.” Here, the activation circuitry system may be connected to the output of a MAC processing pipeline to further process data initially processed by the MAC processing pipeline (e.g., filtered image data). The activation circuitry system may include one or more circuits to process this data via one or more operations, including, for example, linear and / or nonlinear activation operations and / or threshold functions. The one or more circuits, individually or in combination, may perform specific operations, including, for example, specific linear or nonlinear activation operations or threshold functions. The '553 application is incorporated herein by reference in its entirety.
[0104] It is worth noting that the configurable processing circuitry system can be one-time programmable (e.g., at manufacturing time, via, for example, a programmable fuse array) or multiple-time programmable (including, for example, at startup / power-on, initialization, and / or in-situ (i.e., during operation of the integrated circuit)). In one embodiment, the configuration is programmed via a multiplexer prior to the operation or implementation of the execution sequence of the processing pipeline to establish data paths to or bypass one or more selected processing circuits.
[0105] Configurable processing circuitry and its connections Figure 3A , Figure 3C , Figures 6A-6C , Figures 7A-7C and Figures 8A-8C A detailed schematic block diagram of an exemplary configuration of the control / configuration circuitry or NLINK circuitry is superimposed (for illustrative purposes) (see the left side of the “NLINK (Above)” section for each exemplary configuration of the control / configuration circuitry or NLINK circuitry). In one embodiment, for example, the configurable processing circuitry system is accessible to any (or all) NLINK circuitry, for example, via an interconnection network (see [link]). Figure 1B The interconnect network is configurable to connect a configurable processing circuitry system to one or more NLINK circuits and their associated MAC processing pipelines.
[0106] Configurable processing circuitry systems can be one-time programmable (e.g., at manufacturing time, via, for example, a programmable fuse array) or multiple-time programmable (including, for example, at startup / power-on, initialization, and / or in-situ (i.e., during operation of the integrated circuit)). In one embodiment, the configuration is programmed via a multiplexer prior to operation or implementation of the execution sequence of the processing pipeline to establish data paths to or bypass one or more selected processing circuits.
[0107] This document describes and illustrates numerous inventions. While certain embodiments, features, attributes, and advantages of the inventions have been described and illustrated, it should be understood that many other embodiments of the invention, as well as different and / or similar embodiments, features, attributes, and advantages, are apparent from the description and illustrations. Therefore, the embodiments, features, attributes, and advantages of the inventions described and illustrated herein are not exhaustive, and it should be understood that these other similar and different embodiments, features, attributes, and advantages of the invention are within the scope of this invention.
[0108] For example, in one embodiment, the linear processing pipeline and the MACs incorporated in this pipeline can be configured to rotate, transfer, or move (before, during, or at the end of an execution cycle) input data values (instead of accumulating values—which are maintained, stored, or held in a particular MAC during each execution cycle of the execution sequence), as described and illustrated in U.S. Provisional Patent Application No. 63 / 156,263, filed March 3, 2021, “MAC Processing Pipelines, Circuit to Configure Same, and Methods of Operating Same”; the entire '263 application is incorporated herein by reference. In short, in operation, after initial data input values are input or loaded into the MACs of the linear MAC processing pipeline, the input data values are rotated, transferred, or moved cycle-by-cycle from one MAC of the linear pipeline to the next MAC of the pipeline and used in multiplication operations of the multiplier circuitry of the next MAC in the processing pipeline, as described and / or illustrated in the '263 application.
[0109] For example, the degree or length of interconnections (i.e., the number of multiplier-accumulator circuits interconnected to perform multiplication and accumulation operations) can be adjusted (i.e., increased or decreased) via the configuration of the NLINK circuitry, for example, in-situ (i.e., during operation of the integrated circuit), to meet system requirements or constraints (e.g., time-based requirements for system performance). In practice, in one embodiment, rows of multiplier-accumulator circuits can be connected or disconnected via control of the circuitry (e.g., multiplexers) in the NLINK circuitry associated with the rows of multiplier-accumulator circuits to adjust the degree or length of interconnections (i.e., the number of multiplier-accumulator circuits interconnected to perform multiplication and accumulation operations in, for example, an execution or processing pipeline). (See, for example, applications '345 and '306, respectively...) Figures 7A-7C and Figures 6A-6C (and the text associated with it).
[0110] The MAC processing pipeline or architecture of the present invention, and the circuitry for configuring and controlling such a pipeline / architecture, can employ or implement simultaneous and / or parallel processing techniques, architectures, pipelines, and configurations, as described and / or illustrated in U.S. Patent Application No. 16 / 816,164, filed March 11, 2020, entitled “Multiplier-Accumulator Processing Pipelines and Processing Component, and Methods of Operating Same”, and U.S. Provisional Patent Application No. 62 / 831,413, filed April 9, 2019, entitled “Multiplier-Accumulator Circuitry and System having Processing Pipeline and Methods of Operating and Using Same”. Here, the control / configuration circuitry (including multiple control / configuration circuits) can be programmed to configure the pipeline to implement the simultaneous and / or parallel processing techniques described and / or illustrated in '164 and '413 applications, to, for example, increase the throughput of data processing; such applications are incorporated herein by way of inclusion in their entirety.
[0111] In one embodiment, the MAC processing pipeline or architecture of the present invention, and the circuitry for configuring and controlling such a pipeline / architecture, can employ Winograd processing techniques to process image data. Here, the conversion circuitry can convert the data format of the filter weights from a Gaussian floating-point data format to a block-scaled fractional format with appropriate characteristics, facilitating the implementation of Winograd processing techniques associated with the multiplier-accumulator circuitry executing the pipeline. Preprocessing and / or post-processing can be performed... Figure 9The configurable processing circuitry system described and illustrated herein is implemented. It is noteworthy that, among other things, detailed descriptions and / or illustrations of the circuitry system, structure, architecture, function, and operation of the multiplier-accumulator execution pipeline implementing the Winograd processing technology are in: (1) U.S. Patent Application No. 16 / 796,111, filed February 20, 2020, entitled “Multiplier-Accumulator Circuit having Processing Pipelines and Methods of Operating Same,” and / or (2) U.S. Provisional Patent Application No. 62 / 909,293, filed October 2, 2019, entitled “Multiplier-Accumulator Circuit Processing Pipeline and Methods of Operating Same.” These patent applications are incorporated herein by reference.
[0112] Furthermore, although the invention has been described and illustrated in the context of a multiplier-accumulator circuit system, the circuit system and operation of the invention may replace, or otherwise substitute / implement, the multiplication circuit system and the conversion circuit system to facilitate the connection of logarithmic addition and accumulation operations in accordance with the invention. For example, the invention may be employed in conjunction with U.S. Patent Application No. 17 / 092,175 (filed November 6, 2020) and U.S. Provisional Patent Application No. 62 / 943,336, filed December 4, 2019, entitled “Logarithmic Addition-Accumulator Circuitry, Processing Pipeline including Same and Method of Operating Same,” which are incorporated herein by reference in their entirety. In this regard, a pipeline implementing a logarithmic adder-accumulator circuit system (and a method of operating such a circuit system) may be employed in the processing pipeline of the invention and in the circuit system for configuring and controlling such a pipeline, wherein data (e.g., image data) is processed based on a logarithmic format, for example, in relation to inference operations. Therefore, although the invention has been described and illustrated in the context of a multiplier-accumulator circuit system, the circuit system and operation of the invention can replace the multiplication circuit system, or additionally replace / implement the logarithmic addition circuit system and the conversion circuit system to facilitate the connection of logarithmic addition and accumulation operations consistent with the invention.
[0113] Furthermore, the present invention can employ circuitry, functions, and operations to enhance the dynamic range of filter weights or coefficients, as described and / or illustrated in non-provisional patent application No. 17 / 074,670, filed October 20, 2020, entitled "MAC Processing Pipeline using Filter Weights having Enhanced Dynamic Range, and Methods of Operating Same," and / or U.S. provisional patent application No. 62 / 930,601, filed November 5, 2019, entitled "Processing Pipeline Circuitry using Filter Coefficients having Enhanced Dynamic Range and Methods of Operating Same." In other words, the present invention can use circuitry and techniques to enhance the dynamic range of the filter weights or coefficients of the '601 provisional application. Such circuitry and techniques can be... Figure 9 The configurable processing circuitry system described and illustrated herein is implemented. '601 Provisional Application is incorporated herein by reference in its entirety.
[0114] Furthermore, the present invention can employ various data formats for input data and filter weights. For example, the present invention can employ circuit systems, functions, and operations that implement (or modify) the data format of the input data and / or filter weights, as described and / or illustrated in (1a) U.S. Non-Provisional Patent Application No. 16 / 900,319 and (1b) U.S. Provisional Patent Application No. 62 / 865,113 and / or (2a) U.S. Non-Provisional Patent Application No. 17 / 140,169 and (2b) U.S. Provisional Patent Application No. 62 / 961,627. Such preprocessing and / or postprocessing can be performed... Figure 9 The configurable processing circuitry system described and illustrated herein is implemented. It is worth noting that these four (4) applications are incorporated herein by reference in their entirety.
[0115] It is worth noting that a series of multiple multiplier-accumulator circuit systems can be configured, selected, modified, and / or adjusted, for example, in situ (i.e., during the operation of the integrated circuit), to perform or provide specific operations and / or meet or exceed system requirements or constraints (e.g., time-based requirements or constraints).
[0116] Furthermore, while many embodiments described and illustrated herein connect or configure adjacent NLINKS circuits to form or provide a larger execution pipeline (as opposed to a pipeline associated with a single NLINKS circuit), embodiments may connect non-adjacent NLINKS circuits (and by extending non-adjacent rows of multiplier-accumulator circuits) to facilitate pipelined processing and provide a connectivity architecture; for example, the routing circuitry system of the NLINKS interface connector (e.g., one or more multiplexers) may be configured to connect the output of the last multiplier-accumulator circuit of a row of multiplier-accumulator circuits to the input of the first multiplier-accumulator circuit of one or more different rows (adjacent and / or non-adjacent) of multiplier-accumulator circuits. To avoid any doubt, all embodiments described and illustrated herein can be implemented via non-adjacent NLINKS circuits (and by extending non-adjacent rows of multiplier-accumulator circuits) – however, for the sake of brevity, these embodiments will not be illustrated or restated separately in the context of non-adjacent NLINKS circuits and non-adjacent rows of multiplier-accumulator circuits.
[0117] Importantly, the present invention is not limited to any single aspect or embodiment thereof, nor to any combination and / or substitution of such aspects and / or embodiments. Furthermore, each aspect and / or embodiment of the invention may be used alone or in combination with one or more other aspects and / or embodiments of the invention.
[0118] Furthermore, although memory cells in some embodiments are illustrated as static memory cells or storage elements, the present invention may employ dynamic or static memory cells or storage elements. In fact, as described above, such memory cells may be latches, flip-flops, or any other static / dynamic memory cells or memory cell circuits or storage elements now known or later developed.
[0119] It is worth noting that the various circuits, circuit systems, and techniques disclosed herein can be described using computer-aided design tools and expressed (or represented) as data and / or instructions embodied in various computer-readable media in terms of their behavior, register transfers, logic components, transistors, layout geometry, and / or other characteristics. Formats of documents and other objects in which such circuits, circuit systems, layouts, and wiring expressions can be implemented include, but are not limited to, formats supporting behavioral languages (such as C, Verilog, and HLDL), formats supporting register-level description languages (such as RTL), and formats supporting geometric description languages (such as GDSII, GDSIII, GDSIV, CIF, MEBES), as well as any other formats and / or languages now known or developed hereafter. Computer-readable media in which such formatted data and / or instructions can be embodied include, but are not limited to, various forms of non-volatile storage media (e.g., optical, magnetic, or semiconductor storage media) and carrier waves that can be used to transfer such formatted data and / or instructions via wireless, optical, or wired signaling media or any combination thereof. Examples of transferring such formatted data and / or instructions via carrier waves include, but are not limited to, transferring (uploading, downloading, emailing, etc.) over the Internet and / or other computer networks via one or more data transfer protocols (e.g., HTTP, FTP, SMTP, etc.).
[0120] In practice, when received within a computer system via one or more computer-readable media, the representation of the circuits described above based on such data and / or instructions can be processed within the computer system by a processing entity (e.g., one or more processors) in conjunction with the execution of one or more other computer programs, including but not limited to netlist generators, placement and routing programs, etc., thereby generating a representation or image of the physical appearance of these circuits. This representation or image can then be used in device fabrication, for example by enabling the generation of one or more masks for forming various components of the circuit during device fabrication processes.
[0121] Furthermore, the various circuits, circuit systems, and techniques disclosed herein can be represented by simulation using computer-aided design and / or testing tools. Simulation of circuits, circuit systems, layouts and routing, and / or the techniques implemented theretherein can be performed by a computer system, wherein the characteristics and operation of these circuits, circuit systems, layouts, and the techniques implemented theretherein are simulated, replicated, and / or predicted by the computer system. This invention also relates to such simulation of the circuits, circuit systems, and / or the techniques implemented theretherein, and is therefore intended to fall within the scope of this invention. Computer-readable media corresponding to such simulation and / or testing tools are also intended to fall within the scope of this invention.
[0122] It is important to note that references to "an embodiment" or "an embodiment" (or similar) herein refer to specific features, structures, or characteristics that may be included, employed, and / or incorporated in relation to the embodiment in one, some, or all embodiments of the invention. The phrases "in one embodiment" or "in another embodiment" (or similar) used or appearing in the specification do not refer to the same embodiment, nor do they imply separate or alternative embodiments that are necessarily mutually exclusive with one or more other embodiments, nor are they limited to a single exclusive embodiment. The same applies to the term "implementation." The invention is neither limited to any single aspect or embodiment thereof, nor to any combination and / or substitution of such aspects and / or embodiments. Furthermore, each aspect and / or embodiment of the invention may be used alone or in combination with one or more other aspects and / or embodiments of the invention. For the sake of brevity, certain substitutions and combinations are not discussed and / or illustrated separately herein.
[0123] Furthermore, the embodiments or implementations described herein as “exemplary” should not be construed as being ideal, preferred, or advantageous, for example, relative to other embodiments or implementations; rather, they are intended to convey or indicate that the embodiments or embodiments are (one or more) exemplary embodiments.
[0124] While the invention has been described in certain specific aspects, many additional modifications and variations will be apparent to those skilled in the art. Therefore, it should be understood that the invention may be practiced in ways different from those specifically described without departing from the scope and spirit of the invention. Consequently, the embodiments of the invention should be considered in all respects as illustrative / exemplary rather than restrictive.
[0125] The terms “comprising,” “including,” “containing,” “having,” and “having,” or any other variations thereof, are intended to cover a non-exclusive inclusion, so that a process, method, circuit, article, or apparatus that includes a list of parts or elements may include not only those parts or elements but also other parts or elements not expressly listed or inherent to those processes, methods, articles, or apparatus. Furthermore, the terms “connected,” “connected to,” “connected to,” or “connector” used herein should be interpreted broadly to include direct or indirect coupling (e.g., via one or more conductors and / or intermediate devices / elements (active or passive) and / or via inductive or capacitive coupling) unless otherwise intended (e.g., the terms “directly connected” or “directly connected” are used).
[0126] The terms “a” and “an” used in this document do not indicate a limitation of quantity, but rather indicate the presence of at least one of the cited items. Furthermore, the terms “first,” “second,” etc., do not indicate any order, quantity, or importance in this document, but are used to distinguish one element / circuit / feature from another.
[0127] Furthermore, the term "integrated circuit" means, among other things, any integrated circuit, including, for example, general-purpose or non-application-specific integrated circuits, processors, controllers, state machines, gate arrays, SoCs, PGAs, and / or FPGAs. The term "integrated circuit" also means, for example, processors, controllers, state machines, and SoCs—including embedded FPGAs.
[0128] Furthermore, the term "circuit system" means, among other things, a circuit (whether integrated or otherwise), a group of such circuits, one or more processors, one or more state machines, one or more processors implementing software, one or more gate arrays, programmable gate arrays and / or field-programmable gate arrays, or a combination of one or more circuits (whether integrated or otherwise), one or more state machines, one or more processors, one or more processors implementing software, one or more gate arrays, programmable gate arrays and / or field-programmable gate arrays. The term "data" means, among other things, one or more current or voltage signals (complex or singular), whether in analog or digital form, which can be a single bit (or similar) or multiple bits (or similar).
[0129] It is worth noting that the term "MAC circuit" refers to the multiplier-accumulator circuit in a multiplier-accumulator pipeline system. For example, in U.S. Patent Application No. 16 / 545,345... Figure 1A - An exemplary embodiment of Figure 1C and the associated text describe and illustrate a multiplier-accumulator circuit. In the claims, the term "MAC circuit" means, for example, a circuit like the one described in U.S. Patent Application No. 16 / 545,345. Figure 1A - The multiplier-accumulator circuit, etc., described and illustrated in the exemplary embodiment of Figure 1C and its associated text. However, it is worth noting that the term "MAC circuit" is not limited to that described, for example, in U.S. Patent Application No. 16 / 545,345. Figure 1A - The exemplary embodiments illustrated and / or described in Figure 1C include specific circuits, logic, blocks, functions and / or physical diagrams, block / data widths, data path widths, bandwidths and processing.
[0130] In the claims, "row" means row, column, and / or row and column. For example, in the claims, "row of MAC circuit" means (i) a row of MAC circuit, (ii) a column of MAC circuit, and / or (iii) a row and a column of MAC circuit—all of which are intended to fall within the meaning of a row of MAC circuit as relevant to the scope of the claims. In the claims, "column" means column, row, and / or column and row. For example, in the claims, "column of control / configuration circuit" means (i) a column of control / configuration circuit, (ii) a row of control / configuration circuit, and / or (iii) a column and a row of control / configuration circuit—all of which are intended to fall within the meaning of a column of control / configuration circuit as relevant to the scope of the claims.
Claims
1. An integrated circuit, comprising: Multiple multiplier-accumulator circuits are organized into multiple groups, wherein each group of multiplier-accumulator circuits includes multiple multiplier-accumulator circuits connected in series to perform multiple multiplication and accumulation operations, wherein each multiplier-accumulator circuit in each group includes: The multiplier multiplies the data by its weights and generates a product. An accumulator, a multiplier coupled to an associated multiplier-accumulator circuit, adds the input data and the product of the associated multiplier to generate a summation. Multiple control / configuration circuits, wherein each control / configuration circuit is directly connected to and associated with a set of multiplier-accumulator circuits, wherein each control / configuration circuit includes: Multiple data paths, each of which includes: The first terminal is directly connected to the input terminal of the first multiplier-accumulator circuit in a series-connected group of multiplier-accumulator circuits in the associated group. The second end is directly connected to the output of the last multiplier-accumulator circuit in a series-connected multiplier-accumulator circuit of an associated group, wherein the first end of the data path is coupled to the second end of the data path through a series-connected multiplier-accumulator circuit of a series-connected multiplier-accumulator circuit of an associated group. The third terminal can be configured to connect to the end of a corresponding data path of a different control / configuration circuit among multiple control / configuration circuits, and Fourth end; Each control / configuration circuit has multiple data paths, including: The first data path receives input data at the start of the execution sequence that will be processed in each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits.
2. The integrated circuit according to claim 1, wherein: Multiple control / configuration circuits include a first control / configuration circuit, a second control / configuration circuit, and a third control / configuration circuit. The third end of the first data path of the first control / configuration circuit is configured to be connected to the third end of the first data path of the second control / configuration circuit, and The fourth terminal of the first data path of the first control / configuration circuit is coupled to the first memory that stores the input data.
3. The integrated circuit according to claim 2, wherein: The fourth terminal of the first data path of the second control / configuration circuit is coupled to the third terminal of the first data path of the third control / configuration circuit.
4. The integrated circuit according to claim 2, wherein: Each control / configuration circuit's multiple data paths further include a second data path to output data from each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits at the end of the execution sequence. The third end of the second data path of the first control / configuration circuit is configured to be connected to the third end of the second data path of the second control / configuration circuit. The fourth terminal of the second data path of the second control / configuration circuit is coupled to the third terminal of the second data path of the third control / configuration circuit, and The integrated circuit further includes a configurable processing circuitry coupled to a fourth end of a second data path of a third control / configuration circuitry to: (i) receive output data from each of a plurality of serially connected multiplier-accumulator circuits in each group of a plurality of multiplier-accumulator circuits and (ii) process the output data, wherein the processed output data will subsequently be stored in a second memory.
5. The integrated circuit according to claim 4, wherein: The configurable processing circuitry system includes multiple configurable data paths and one or more circuits connected therein, wherein the one or more circuits process output data via adding / subtracting register values, multiplying register values, converting from integer data format to floating-point data format, converting from floating-point data format to integer data format, format precision adjustment, and / or one or more unary functions, including: inversion, square root, inverse square root, hyperbolic tangent, and / or sigmoid.
6. The integrated circuit according to claim 1, wherein: The multiple data paths of each control / configuration circuit further include a third data path to rotate a portion of the accumulated value to each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits during the execution sequence.
7. An integrated circuit, comprising: Multiple multiplier-accumulator circuits are organized into multiple groups, wherein each group of multiplier-accumulator circuits includes multiple multiplier-accumulator circuits connected in series to perform multiple multiplication and accumulation operations, wherein each multiplier-accumulator circuit in each group includes: The multiplier multiplies the data by its weights and generates a product. An accumulator is a multiplier coupled to an associated multiplier-accumulator circuit to add the input data and the product data of the associated multiplier to generate a summation data. Multiple control / configuration circuits, each of which is directly connected to and associated with a set of multiplier-accumulator circuits, wherein each control / configuration circuit includes: Multiple data paths, each of which includes: The first terminal is directly connected to the input terminal of the first multiplier-accumulator circuit in a series-connected group of multiplier-accumulator circuits in the associated group. The second end is directly connected to the output of the last multiplier-accumulator circuit in a series-connected multiplier-accumulator circuit of an associated group, wherein the first end of the data path is coupled to the second end of the data path through a series-connected multiplier-accumulator circuit of the associated group. The third terminal can be configured to connect to the end of a corresponding data path of a different control / configuration circuit in a plurality of control / configuration circuits, and Fourth end; Each control / configuration circuit has multiple data paths, including: The first data path is used to input accumulated values to each group of multiple multiplier-accumulator circuits and output accumulated values from each group of multiple multiplier-accumulator circuits during the execution sequence.
8. The integrated circuit according to claim 7, wherein: The first data path inputs a portion of the accumulated value to each group of multiple multiplier-accumulator circuits and outputs a portion of the accumulated value from each group of multiple multiplier-accumulator circuits during the execution sequence.
9. The integrated circuit according to claim 7, wherein: Multiple control / configuration circuits include a first control / configuration circuit, a second control / configuration circuit, and a third control / configuration circuit. The third end of the first data path of the first control / configuration circuit is configured to be connected to the third end of the first data path of the second control / configuration circuit. The fourth terminal of the first data path of the second control / configuration circuit is coupled to the third terminal of the first data path of the third control / configuration circuit, and The fourth terminal of the first data path of the third control / configuration circuit is connected to the fourth terminal of the first data path of the first control / configuration circuit.
10. The integrated circuit according to claim 9, wherein: Multiple control / configuration circuits are arranged in a row, and The first control / configuration circuit and the third control / configuration circuit are located at the first and second edges of the plurality of control / configuration circuits, respectively.
11. The integrated circuit according to claim 9, wherein: Each control / configuration circuit's multiple data paths further include a second data path to output data from each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits at the end of the execution sequence. The third end of the second data path of the first control / configuration circuit is configured to be connected to the third end of the second data path of the second control / configuration circuit. The fourth terminal of the second data path of the first control / configuration circuit is coupled to the second memory storing the output data, and The fourth terminal of the second data path of the second control / configuration circuit is coupled to the third terminal of the second data path of the third control / configuration circuit.
12. The integrated circuit according to claim 11, wherein: The fourth terminal of the second data path of the third control / configuration circuit is coupled to the second memory that stores the output data.
13. The integrated circuit according to claim 9, wherein: Each control / configuration circuit's multiple data paths include a second data path to input data from each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits at the start of the execution sequence. The third end of the second data path of the first control / configuration circuit is configured to be connected to the third end of the second data path of the second control / configuration circuit. The fourth terminal of the second data path of the first control / configuration circuit is coupled to the first memory storing the input data, and The fourth terminal of the second data path of the second control / configuration circuit is coupled to the third terminal of the second data path of the third control / configuration circuit.
14. The integrated circuit according to claim 13, wherein: Each control / configuration circuit's multiple data paths further include a third data path to output data from each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits at the end of the execution sequence. The third end of the third data path of the first control / configuration circuit is configured to be connected to the third end of the third data path of the second control / configuration circuit. The fourth terminal of the third data path of the first control / configuration circuit is coupled to the second memory storing the output data, and The fourth terminal of the third data path of the second control / configuration circuit is coupled to the third terminal of the second data path of the third control / configuration circuit.
15. The integrated circuit of claim 14, further comprising: A configurable processing circuitry system, coupled to a fourth end of a third data path of a third control / configuration circuit, is configured to: (i) receive output data from each of a plurality of serially connected multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits, and (ii) process the output data, wherein the processed output data will subsequently be stored in a second memory.
16. The integrated circuit according to claim 15, wherein: The configurable processing circuitry system includes multiple configurable data paths and one or more circuits connected therein, wherein the one or more circuits process output data via adding / subtracting register values, multiplying register values, converting from integer data format to floating-point data format, converting from floating-point data format to integer data format, format precision adjustment, and / or one or more unary functions, including: inversion, square root, inverse square root, hyperbolic tangent, and / or sigmoid.
17. An integrated circuit, comprising: Multiple multiplier-accumulator circuits are organized into multiple groups, wherein each group of multiplier-accumulator circuits includes multiple multiplier-accumulator circuits connected in series to perform multiple multiplication and accumulation operations, wherein each multiplier-accumulator circuit in each group includes: The multiplier multiplies the data by its weights and generates a product. An accumulator, a multiplier coupled to an associated multiplier-accumulator circuit, adds the input data and the product of the associated multiplier to generate a summation. Multiple control / configuration circuits, wherein each control / configuration circuit is directly connected to and associated with a set of multiplier-accumulator circuits, wherein each control / configuration circuit includes: Multiple data paths, each of which includes: The first terminal is directly connected to the input terminal of the first multiplier-accumulator circuit in a series-connected group of multiplier-accumulator circuits in the associated group. The second end is directly connected to the output of the last multiplier-accumulator circuit in a series-connected multiplier-accumulator circuit of an associated group, wherein the first end of the data path is coupled to the second end of the data path through a series-connected multiplier-accumulator circuit of the associated group. The third terminal can be configured to connect to the end of a corresponding data path of a different control / configuration circuit in a plurality of control / configuration circuits, and Fourth end; Each control / configuration circuit has multiple data paths, including: The first data path outputs data from each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits.
18. The integrated circuit according to claim 17, wherein: Multiple control / configuration circuits include a first control / configuration circuit, a second control / configuration circuit, and a third control / configuration circuit. The third end of the first data path of the first control / configuration circuit is configured to be connected to the third end of the first data path of the second control / configuration circuit. The fourth terminal of the first data path of the first control / configuration circuit is coupled to a memory storing the output data, and The fourth terminal of the first data path of the second control / configuration circuit is coupled to the third terminal of the second data path of the third control / configuration circuit.
19. The integrated circuit according to claim 18, wherein: The fourth terminal of the first data path of the third control / configuration circuit is coupled to a memory that stores the output data at the end of the execution sequence.
20. The integrated circuit of claim 18, further comprising: A configurable processing circuitry system is coupled between a fourth terminal of the first data path of the third control / configuration circuitry and a memory, wherein the configurable processing circuitry system is programmable to: (i) receive output data from each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits and (ii) postprocess the output data and store the postprocessed output data in the memory.
21. The integrated circuit according to claim 20, wherein: The first data path receives data from memory at the start of the execution sequence to input into each of the multiple cascaded multiplier-accumulator circuits in each group of multiple multiplier-accumulator circuits, and The fourth terminal of the first data path of the first control / configuration circuit is coupled to the memory to receive data from the memory at the start of the execution sequence.
22. The integrated circuit according to claim 18, wherein: The fourth terminal of the first data path of the third control / configuration circuit is coupled to the memory, and the integrated circuit further includes: A configurable processing circuitry system is coupled between a fourth end of a first data path of a third control / configuration circuit and a memory, wherein the configurable processing circuitry system includes at least one configurable data path and one or more circuits connected therein.
23. The integrated circuit according to claim 22, wherein: At least one configurable data path includes at least one multiplexer, and One or more circuits include one or more circuits that process output data by adding / subtracting register values, multiplying register values, converting from integer data format to floating-point data format, converting from floating-point data format to integer data format, format precision adjustment, and / or one or more unary functions, including: inversion, square root, inverse square root, hyperbolic tangent, and / or sigmoid.