Processor, method, and system with configurable spatial accelerators
Through the sequencer data flow operator architecture of configurable space accelerator, the throughput and energy consumption problems in high-performance computing are solved, efficient cyclic control signal generation and memory operation are achieved, and the performance and energy efficiency of the processor are improved.
Patent Information
- Application Number
- CN201811131626.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-30
- Filing Date
- 2018-09-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2038-09-27
AI Technical Summary
Existing processors are difficult to achieve high throughput and low energy consumption at the same time in high-performance computing. The program execution performance and energy efficiency of the traditional von Neumann architecture are insufficient, especially in the generation of cyclic control signals.
Using a sequencer data flow operator architecture with a configurable space accelerator (CSA), the cyclic control signals are directly executed through the data flow diagram, reducing memory prefetching and data speculation, and connecting processing elements with lightweight backpressure networks to achieve efficient cyclic iteration and memory operations.
Significantly improves performance and energy efficiency for high-performance computing applications, cyclic control signal generation speed is increased by 2 to 3 times, energy consumption is reduced by at least 50%, and compatibility with integer processing elements is maintained.
Smart Images

Figure CN109597646B_ABST
Abstract
Description
[0001] Statement Regarding Federally Sponsored Research and Development
[0002] This invention was made with government support under contract number H98230B-13-D-0124-0132 awarded by the Department of Defense. The government has certain rights in this invention. Technical Field
[0003] This disclosure generally relates to electronic devices, and more particularly, embodiments of this disclosure relate to sequencer data flow operators. Background Art
[0004] A processor or collection of processors executes instructions from an instruction set (e.g., Instruction Set Architecture (ISA)). An instruction set is part of a computer architecture related to programming and generally includes native data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I / O). It should be noted that the term "instruction" herein may refer to a macro-instruction (e.g., an instruction provided to the processor for execution) or a micro-instruction (e.g., an instruction resulting from the decoder of the processor decoding a macro-instruction). Brief Description of the Drawings
[0005] The present disclosure is illustrated by way of example and not limitation in the figures, in which like reference numerals indicate like elements and which include:
[0006] Figure 1 An accelerator tile according to an embodiment of the present disclosure is shown.
[0007] Figure 2 A hardware processor coupled to memory according to an embodiment of the present disclosure is shown.
[0008] Figure 3A A program source according to an embodiment of the present disclosure is shown.
[0009] Figure 3B A data flow diagram of a program source according to an embodiment of the present disclosure is shown. Figure 3A of the program source is shown.
[0010] Figure 3C An accelerator having a plurality of processing elements configured to run a data flow diagram according to an embodiment of the present disclosure is shown. Figure 3B of the data flow diagram is shown.
[0011] Figure 4 An example execution of a data flow diagram according to an embodiment of the present disclosure is shown.
[0012] Figure 5A A program source according to an embodiment of the present disclosure is shown.
[0013] Figure 5BShows a program source according to an embodiment of the present disclosure.
[0014] Figure 6 Shows an accelerator primitive including an array of processing elements according to an embodiment of the present disclosure.
[0015] Figure 7A Shows a configurable data path network according to an embodiment of the present disclosure.
[0016] Figure 7B Shows a configurable flow control path network according to an embodiment of the present disclosure.
[0017] Figure 8 Shows a hardware processor primitive including an accelerator according to an embodiment of the present disclosure.
[0018] Figure 9 Shows a processing element according to an embodiment of the present disclosure.
[0019] Figure 10 Shows a request address file (RAF) circuit according to an embodiment of the present disclosure.
[0020] Figure 11 Shows a plurality of request address file (RAF) circuits coupled between a plurality of accelerator primitives and a plurality of cache groups according to an embodiment of the present disclosure.
[0021] Figure 12 Shows a floating-point multiplier divided into three regions (a result region, three potential carry regions, and a gated region) according to an embodiment of the present disclosure.
[0022] Figure 13 Shows an in-flight configuration of an accelerator having a plurality of processing elements according to an embodiment of the present disclosure.
[0023] Figure 14 Shows a snapshot of an in-flight pipelined fetch according to an embodiment of the present disclosure.
[0024] Figure 15 Shows a compilation tool chain for an accelerator according to an embodiment of the present disclosure.
[0025] Figure 16 Shows a compiler for an accelerator according to an embodiment of the present disclosure.
[0026] Figure 17A Shows a sequential assembly code according to an embodiment of the present disclosure.
[0027] Figure 17B Shows according to an embodiment of the present disclosure, Figure 17A the data flow assembly code of the sequential assembly code.
[0028] Figure 17C A data flow diagram of data flow assembly code according to an embodiment of the present disclosure. Figure 17B
[0029] Figure 18A A C source code according to an embodiment of the present disclosure.
[0030] Figure 18B A data flow assembly code of a C source code according to an embodiment of the present disclosure. Figure 18A
[0031] Figure 18C A data flow diagram of data flow assembly code according to an embodiment of the present disclosure. Figure 18B
[0032] Figure 19A A C source code according to an embodiment of the present disclosure.
[0033] Figure 19B A data flow assembly code of a C source code according to an embodiment of the present disclosure. Figure 19A
[0034] Figure 19C A data flow diagram of data flow assembly code according to an embodiment of the present disclosure. Figure 19B
[0035] Figure 20A A C source code according to an embodiment of the present disclosure.
[0036] Figure 20B A data flow assembly code of a C source code according to an embodiment of the present disclosure. Figure 20A
[0037] Figure 20C A data flow diagram of data flow assembly code according to an embodiment of the present disclosure. Figure 20B
[0038] Figure 21 An implementation of integer arithmetic / logic data flow operators on a processing element according to an embodiment of the present disclosure.
[0039] Figure 22 An implementation of sequencer data flow operators on a processing element according to an embodiment of the present disclosure.
[0040] Figure 23 An example operation format of an implementation of integer arithmetic / logic data flow operators on a processing element according to an embodiment of the present disclosure.
[0041] Figure 24 An example operation format of an implementation of sequencer data flow operators on a processing element according to an embodiment of the present disclosure.
[0042] Figure 25 Shows an example arithmetic format of the sequencer data flow operator implementation on a processing element according to an embodiment of the present disclosure.
[0043] Figure 26 Shows an example arithmetic format of the sequencer data flow operator implementation on a processing element according to an embodiment of the present disclosure.
[0044] Figure 27 Shows circuit 2700 of the sequencer data flow operator implementation on a single processing element according to an embodiment of the present disclosure.
[0045] Figure 28 Shows a circuit that supports the one-way mode of the sequencer data flow operator implementation on a single processing element according to an embodiment of the present disclosure.
[0046] Figure 29 Shows a circuit that supports the simplified mode of the sequencer data flow operator implementation on a single processing element according to an embodiment of the present disclosure.
[0047] Figure 30 Shows a circuit that switches to the sequencer mode of the sequencer data flow operator implementation on a single processing element according to an embodiment of the present disclosure.
[0048] Figure 31 Shows a circuit that switches between the activation mode and the deactivation mode of the selective double-ended queue of the sequencer data flow operator implementation on a single processing element according to an embodiment of the present disclosure.
[0049] Figure 32 Shows a matrix multiplication code example according to an embodiment of the present disclosure.
[0050] Figure 33A - 33B Shows, according to an embodiment of the present disclosure, generating Figure 32 The first sequencer data flow operator implementation on multiple processing elements of A[i][k] and B[k][j] for matrix multiplication.
[0051] Figure 34 Shows, according to an embodiment of the present disclosure, generating Figure 32 The second optimized sequencer data flow operator implementation on multiple processing elements of A[i][k] and B[k][j] for matrix multiplication.
[0052] Figure 35 Shows the sequencer data flow operator implementation on multiple processing elements that transforms the sparse memory access mode into the dense memory access mode according to an embodiment of the present disclosure.
[0053] Figure 36 Shows a flowchart according to an embodiment of the present disclosure.
[0054] Figure 37 Shows a flowchart according to an embodiment of the present disclosure.
[0055] Figure 38 Shows the energy per operation graph versus throughput according to an embodiment of the present disclosure.
[0056] Figure 39 Shows an accelerator primitive including a processing element array and a local configuration controller according to an embodiment of the present disclosure.
[0057] Figure 40A - 40C Shows a local configuration controller for configuring a data path network according to an embodiment of the present disclosure.
[0058] Figure 41 Shows a configuration controller according to an embodiment of the present disclosure.
[0059] Figure 42 Shows an accelerator primitive including a processing element array, a configuration cache, and a local configuration controller according to an embodiment of the present disclosure.
[0060] Figure 43 Shows an accelerator primitive including a processing element array and a configuration and exception handling controller having a reconfiguration circuit according to an embodiment of the present disclosure.
[0061] Figure 44 Shows a reconfiguration circuit according to an embodiment of the present disclosure.
[0062] Figure 45 Shows an accelerator primitive including a processing element array and a configuration and exception handling controller having a reconfiguration circuit according to an embodiment of the present disclosure.
[0063] Figure 46 Shows an accelerator primitive including a processing element array and a mezzanine exception aggregator coupled to a primitive-level exception aggregator according to an embodiment of the present disclosure.
[0064] Figure 47 Shows a processing element having an exception generator according to an embodiment of the present disclosure.
[0065] Figure 48 Shows an accelerator primitive including a processing element array and a local extraction controller according to an embodiment of the present disclosure.
[0066] Figure 49A - 49C Shows a local extraction controller for configuring a data path network according to an embodiment of the present disclosure.
[0067] Figure 50 Shows an extraction controller according to an embodiment of the present disclosure.
[0068] Figure 51 Shows a flowchart according to an embodiment of the present disclosure.
[0069] Figure 52 Shows a flowchart according to an embodiment of the present disclosure.
[0070] Figure 53A Is a block diagram of a system according to an embodiment of the present disclosure that employs a memory ordering circuit between an insert memory subsystem and acceleration hardware.
[0071] Figure 53B Is according to an embodiment of the present disclosure and instead employs multiple memory ordering circuits Figure 53A Of the system.
[0072] Figure 54 Is a block diagram showing the general functions of memory operations into / out of acceleration hardware according to an embodiment of the present disclosure.
[0073] Figure 55 Is a block diagram showing the spatial correlation process of storage operations according to an embodiment of the present disclosure.
[0074] Figure 56 Is a detailed block diagram of the memory ordering circuit of FIG. 53 according to an embodiment of the present disclosure.
[0075] Figure 57 Is a flowchart of the microarchitecture of the memory ordering circuit of FIG. 53 according to an embodiment of the present disclosure.
[0076] Figure 58 Is a block diagram of an executable determiner circuit according to an embodiment of the present disclosure.
[0077] Figure 59 Is a block diagram of a polarity encoder according to an embodiment of the present disclosure.
[0078] Figure 60 Is a block diagram of a demonstration load operation that is both logical and binary according to an embodiment of the present disclosure.
[0079] Figure 61A Is a flowchart showing the logical execution of example code according to an embodiment of the present disclosure.
[0080] Figure 61B Is a flowchart showing the memory-level parallelism of the expanded version of the example code according to an embodiment of the present disclosure. Figure 61A Of the flowchart.
[0081] Figure 62A Is a block diagram of exemplary memory arguments for load operations and store operations according to an embodiment of the present disclosure.
[0082] Figure 62Bis a block diagram showing the load and store operations (e.g., Figure 57 those operations) of the microarchitecture of a memory ordering circuit through Figure 62A in accordance with an embodiment of the present disclosure.
[0083] Figure 63A , Figure 63B , Figure 63C , Figure 63D , Figure 63E , Figure 63F , Figure 63G and Figure 63H is a block diagram showing the functional flow of the load and store operations of a sample program in a queue of the microarchitecture through Figure 63B in accordance with an embodiment of the present disclosure.
[0084] Figure 64 is a flowchart of a method for ordering memory operations between accelerated hardware and an out-of-order memory subsystem in accordance with an embodiment of the present disclosure.
[0085] Figure 65A is a block diagram showing a general vector-friendly instruction format and its class A instruction template in accordance with an embodiment of the present disclosure.
[0086] Figure 65B is a block diagram showing a general vector-friendly instruction format and its class B instruction template in accordance with an embodiment of the present disclosure.
[0087] Figure 66A is a block diagram showing the fields of the general vector-friendly instruction format in Figure 65A and Figure 65B in accordance with an embodiment of the present disclosure.
[0088] Figure 66B is a block diagram showing the fields of a specific vector-friendly instruction format that makes up the full opcode field in Figure 66A in accordance with an embodiment of the present disclosure.
[0089] Figure 66C is a block diagram showing the fields of a specific vector-friendly instruction format that makes up the register index field in Figure 66A in accordance with an embodiment of the present disclosure.
[0090] Figure 66D is a block diagram showing the fields of a specific vector-friendly instruction format that makes up the extended operation field 6550 in Figure 66A in accordance with an embodiment of the present disclosure.
[0091] Figure 67 is a block diagram of a register architecture in accordance with an embodiment of the present disclosure.
[0092] Figure 68Ais a block diagram showing an exemplary in - order pipeline and an exemplary register renaming, out - of - order issue / execution pipeline in accordance with an embodiment of the present disclosure.
[0093] Figure 68B is a block diagram showing an in - order architectural core and an exemplary register renaming, out - of - order issue / execution architectural core included in a processor in accordance with an embodiment of the present disclosure.
[0094] Figure 69A is a block diagram of a single processor core along with its connections to the on - die interconnect network and its connection to a local subset of its level 2 (L2) cache in accordance with an embodiment of the present disclosure.
[0095] Figure 69B is in accordance with an embodiment of the present disclosure Figure 69A is an expanded view of a portion of the processor core in
[0096] Figure 70 is a block diagram of a processor that may have more than one core, may have an integrated memory controller, and may have integrated graphics in accordance with an embodiment of the present disclosure.
[0097] Figure 71 is a block diagram of a system in accordance with one embodiment of the present disclosure.
[0098] Figure 72 is a block diagram of a more specific exemplary system in accordance with an embodiment of the present disclosure.
[0099] Figure 73 Shown is a block diagram of a second more specific exemplary system in accordance with an embodiment of the present disclosure.
[0100] Figure 74 Shown is a block diagram of a system - on - a - chip (SoC) in accordance with an embodiment of the present disclosure.
[0101] Figure 75 is a block diagram in accordance with an embodiment of the present disclosure, contrasting with a software instruction converter for converting binary instructions in a source instruction set into binary instructions in a target instruction set. Detailed Description
[0102] In the following description, numerous specific details are set forth. However, it is understood that embodiments of the present disclosure may be practiced without these specific details. In other instances, well - known circuits, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0103] As used herein, the terms "one embodiment", "an embodiment", "example embodiment", etc. mean that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment necessarily includes that particular feature, structure, or characteristic. Moreover, such terms do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly described.
[0104] A processor (e.g., having one or more cores) may execute instructions (e.g., threads of instructions) to operate on data, such as performing arithmetic, logical, or other functions. For example, software may request an operation, and a hardware processor (e.g., one or more of its cores) may perform the operation in response to the request. A non-limiting example of an operation is an operation that inputs a plurality of vector elements and outputs a vector having a mixture of the plurality of elements. In certain embodiments, multiple operations are implemented by the execution of a single instruction.
[0105] For example, exascale performance as defined by the Department of Energy may require system-level performance of more than 10^18 floating-point operations per second (exaFLOP) or more within a given (e.g., 20MW) power budget. Certain embodiments herein are directed to a spatial array of processing elements (e.g., a configurable spatial accelerator (CSA)) that targets high-performance computing (HPC) of, for example, a processor. Certain embodiments of the spatial array of processing elements (e.g., CSA) herein are directed to the direct execution of (one or more) data flow graphs to produce a computationally intensive but energy-efficient spatial microarchitecture that far exceeds conventional roadmap architectures.
[0106] Certain embodiments of a spatial architecture (e.g., the spatial array disclosed herein) are an energy-efficient and high-performance way to accelerate user applications. In certain embodiments, a spatial array (e.g., a plurality of processing elements coupled together through, for example, a circuit-switching (e.g., interconnect) network) accelerates an application, such as by running a region of a single-stream program (e.g., faster than the core of a processor). Certain embodiments of the spatial architecture herein facilitate the mapping of sequential programs to the spatial array.
[0107] The key architectural interface for embodiments of an accelerator (such as a CSA) is the data flow operator, which is, for example, a direct representation as a node in a data flow graph. From an operational perspective, the data flow operator behaves in a streaming or data-driven manner. The data flow operator can run immediately when its incoming operands become available. CSA data flow execution can (e.g., only) depend on the high-location domain state, such as giving rise to a highly-scalable architecture with a distributed asynchronous execution model. The data flow operator can include arithmetic data flow operators, such as one or more of floating-point addition and multiplication, integer addition, subtraction, and multiplication, various forms of comparison, logical operators, and shifts. However, embodiments of the CSA can also include an enrichment of control operators, which assist in the management of data flow tokens in the program graph. Examples of these include: a "pick" operator, such as one that multiplexes two or more logical input channels into a single output channel; and a "switch" operator, such as one that operates as a channel demultiplexer (e.g., outputs a single channel from two or more logical input channels).
[0108] These operators can enable the compiler to implement control paradigms, such as conditional expressions and loops. Certain embodiments of the CSA can include a limited set of data flow operators (e.g., for a smaller number of operations) to yield a dense and energy-efficient PE microarchitecture. Certain embodiments can include data flow operators for complex arithmetic operations (which are common in HPC code). An example of a data flow operator is an ordinal data flow operator, such as to enable efficient control of a for loop (e.g., a loop construct). One embodiment of an ordinal data flow operator that implements a loop introduces a feedback path between the conditional and post-condition update portions of the loop. For example, for loop terms are often related, such that an exit condition term (e.g., "M < i < N" or "i < N") is often followed by a decrement or increment term (e.g., "i++", etc., where i is the loop counter variable). In some embodiments, this can form a bottleneck in the performance of the ordinal data flow operator implementation, which is addressed by introducing a composite ordinal operation (e.g., which is capable of performing the condition and update of a for loop pattern in a single operation (e.g., a single cycle)). In one embodiment, a for loop includes one or more (e.g., all) of the following parts: initialization, condition, and afterthought. In one embodiment, the initialization declares any variables required (e.g., and assigns a value(s) to them). For example, if multiple variables are being used in the initialization part, the variables can be of the same type. In one embodiment, the condition checks a certain condition and exits the loop when "false". In one embodiment, the afterthought is only executed once at the end of each loop iteration and then repeated.
[0109] The CSA data flow operator architecture is highly compliant with deployment-specific extensions. For example, in some embodiments, more complex mathematical data flow operators (such as trigonometric functions) may be included to accelerate certain math-intensive HPC workloads. Similarly, neural network tuning extensions may include data flow operators for vectorized low-precision arithmetic.
[0110] Certain embodiments herein provide a sequencer data flow operator architecture and a sequencer microarchitecture such that, for example, the generation of control signals for (e.g., the most common) for-loop constructs can achieve peak performance of one loop iteration per cycle (e.g., the cycles of an accelerator including the sequencer). Certain embodiments herein can greatly improve the performance of many high-performance computing (HPC) applications. Certain embodiments of the sequencer data flow operator separate the generation of such loop control signals from the actual data flow tokens of the loop constructs themselves such that, for many HPC applications, memory prefetching and / or data speculation (and associated energy waste) can be completely eliminated. Certain embodiments of the sequencer data flow operator can be formed by modifying one or more integer processing elements (PEs) and / or by adopting (e.g., smaller) configuration changes and microarchitecture extensions, where the illustrated sequencer PEs can still operate as (e.g., basic) integer PEs. Full binary compatibility with (e.g., basic) integer PEs can also be achieved to minimize software engineering costs. Certain embodiments herein may include a sequencer data flow operator (e.g., a circuit) that manipulates data (e.g., data tokens) such as 64-bit wide, 32-bit wide, etc. in a coarse-grained manner (e.g., as opposed to control tokens) and targets the highest achievable clock frequency (e.g., 1 - 1.5 GHz), while still using an energy-efficient circuit network topology / design.
[0111] Certain embodiments herein include a sequencer data flow operator (e.g., a circuit) that minimizes overhead in terms of energy, area, throughput, and latency. Certain embodiments herein include a sequencer data flow operator (e.g., a circuit) that minimizes the hardware resources utilized while achieving the highest possible performance.
[0112] Also included below is a description of the architectural philosophy and certain features of embodiments of a spatial array of processing elements (e.g., CSA). As with any transformative architecture, programmability can be a risk. To mitigate this issue, embodiments of the CSA architecture are co-designed with a compilation toolchain (which is also discussed below).
[0113] Introduction
[0114] Exascale computing goals can require huge system-level floating-point performance (e.g., ExaFLOP) within an oversubscribed power budget (e.g., 20 MW). However, improving both the performance and energy efficiency of program execution using traditional von Neumann architectures has become a challenge: out-of-order scheduling, simultaneous multithreading, complex register files, and other constructs provide performance but at a high energy cost. Some embodiments of this disclosure meet both performance and energy requirements. Exascale computing power-performance goals can require high throughput and low energy consumption per operation. Some embodiments of this disclosure provide this by providing a large number of low-complexity, energy-efficient processing (e.g., computing) elements that largely eliminate the control overhead of prior processor designs. Guided by this observation, some embodiments of this disclosure include a spatial array of processing elements (e.g., a Configurable Spatial Accelerator (CSA)), e.g., an array of Processing Elements (PEs) connected by a set of lightweight backpressure (e.g., communication) networks. An example of a CSA primitive is shown in Figure 1 Some embodiments of processing (e.g., computing) elements are dataflow operators, e.g., multiples of dataflow operators that process input data (e.g., no processing occurs otherwise) only when (i) the input data has arrived at the dataflow operator and (ii) there is space available for storing the output data. Some embodiments (e.g., of an accelerator or CSA) do not utilize triggered instructions.
[0115] A coarse-grained spatial architecture (e.g., an embodiment of the Configurable Spatial Accelerator (CSA) shown in Figure 1 ) is a composition of lightweight Processing Elements (PEs) connected by an interconnection network. For example, a program that can be viewed as a control dataflow graph can be mapped onto the architecture by configuring the PEs and the network. In general, a PE can be configured as a dataflow operator such that, once all input operands have arrived at the PE, an operation occurs and the result is forwarded downstream (e.g., to one or more destination PEs) in a pipelined fashion. A dataflow operator (e.g., a primitive operation) can be a load or a store, e.g., as shown in the Request Address File (RAF) in Figure 10 . A dataflow operator can be selected based on per-operator consumption of incoming data.
[0116] Some embodiments of this disclosure extend the capabilities of a spatial array (e.g., a CSA) to perform parallel accesses to memories, e.g., in a memory subsystem, via one or more hazard detection circuits.
[0117] Figure 1 An embodiment of an accelerator primitive 100 of a spatial array of processing elements according to an embodiment of the present disclosure is shown. The accelerator primitive 100 can be part of a larger primitive. The accelerator primitive 100 runs one or more dataflow graphs. The dataflow Figure 1Generally, it can represent an explicit parallel program description, which appears in the compilation of sequential code. Some embodiments of this article (such as CSA) allow a data flow graph to be directly configured onto a CSA array, for example instead of being transformed into a sequential instruction stream. Some embodiments of this article allow memory access data flow operations (such as their types) to be performed by one or more processing elements (PEs) of a spatial array.
[0118] The deviation of the data flow graph from the sequential compilation process allows embodiments of CSA to support a familiar programming model and run existing high-performance computing (HPC) code directly (such as without using work sheets). The CSA processing element (PE) can be energy-efficient. Figure 1 In, the memory interface 102 can be coupled to a memory (such as Figure 2 the memory 202 in), to allow the accelerator primitive 100 to access (such as load and / or store) data from / to a (such as off-chip or off-system) memory. The illustrated accelerator primitive 100 is a heterogeneous array, which is composed of several types of PEs commonly coupled via an interconnection network 104. The accelerator primitive 100 can include one or more of integer arithmetic PEs, floating-point arithmetic PEs, communication circuits (such as network data flow endpoint circuits), and on-chip storage devices (such as part of a spatial array of processing elements 101). A data flow graph (such as a compiled data flow graph) can be overlaid on the accelerator primitive 100 for execution. In one embodiment, for a specific data flow graph, each PE only manipulates one or more (data flow) operations of the graph. The array of PEs can be heterogeneous, for example such that no PE supports a full CSA data flow architecture, and / or one or more PEs are programmed (such as customized) to only perform a few but efficient operations. Therefore, some embodiments of this article produce a processor or accelerator with an array of processing elements, which is computationally intensive compared to a roadmap architecture, but still achieves approximately an order of magnitude gain in terms of energy efficiency and performance compared to existing HPC.
[0119] Some embodiments of this article provide a performance increase from parallel execution within a (such as dense) spatial array of processing elements (such as CSA), where each PE utilized can, for example, perform its operation simultaneously when the input data is available. The efficiency increase can result from the efficiency of each PE, for example, where once each configuration (such as mapping) step and execution occur when local data arrives at the PE, the operation (such as behavior) of each PE is fixed, for example without considering other structural activities. In some embodiments, the PE is a (such as each single) data flow operator, such as a data flow operator that only operates on input data when (i) the input data has arrived at the data flow operator and (ii) there is space available for storing the output data (such as no operation occurs otherwise).
[0120] Some embodiments of the present disclosure include a spatial array of processing elements as an energy-efficient and high-performance way to accelerate user applications. In one embodiment, the (one or more) spatial arrays are configured via a serial process, where the latency of the configuration is fully exposed via a global reset. A part of this aspect may stem from the register transfer level (RTL) semantics of the array (e.g., a field programmable gate array (FPGA)). Programs for execution on the array (e.g., FPGA) may take the basic concept of a reset, where every part of the design is expected to be operable from a configuration reset. Some embodiments of the present disclosure provide a dataflow-style array where the PEs (e.g., all) conform to a flow controller microprotocol. The microprotocol may create the effect of distributed initialization. This microprotocol is capable of allowing, for example, pipelined configuration and extraction mechanisms with a regional (e.g., not the entire array) organization. Some embodiments of the present disclosure provide hazard detection and / or error recovery (e.g., manipulation) in a dataflow architecture.
[0121] Some embodiments of the present disclosure provide paradigm-breaking performance and significant improvements in energy efficiency across existing single-stream and parallel programs, e.g., all while preserving a familiar HPC programming model. Some embodiments of the present disclosure may target HPC such that floating-point energy efficiency is of great importance. Some embodiments of the present disclosure not only deliver competitive improvements in performance and reductions in energy, they also give these gains to existing HPC programs (which are written in mainstream HPC languages and used in mainstream HPC frameworks). Some embodiments of the architectures herein (e.g., considering compilation) provide several extensions in terms of direct support for the control-dataflow internal representations generated by modern compilers. Some embodiments of the present disclosure target a CSA dataflow compiler, e.g., which is capable of accepting C, C++, and Fortran programming languages for the CSA architecture.
[0122] Figure 2Shown is a hardware processor 200, according to an embodiment of the present disclosure, coupled to (e.g., connected to) a memory 202. In one embodiment, the hardware processor 200 and the memory 202 are a computing system 201. In certain embodiments, one or more of the accelerators are CSAs according to the present disclosure. In certain embodiments, one or more of the cores in the processor are those disclosed herein. The hardware processor 200 (e.g., each of its cores) may include a hardware decoder (e.g., decoding unit) and a hardware execution unit. The hardware processor 200 may include registers. Note that the drawings herein may not show all data communication couplings (e.g., connections). Those skilled in the art will understand that this does not affect the understanding of certain details in the drawings. Note that a single arrow in the drawings may not require one-way communication, e.g., it may indicate two-way communication (e.g., to / from that component or device). Note that a double arrow in the drawings may not require two-way communication, e.g., it may indicate one-way communication (e.g., to / from that component or device). Any one or all combinations of communication paths may be used in certain embodiments herein. According to an embodiment of the present disclosure, the shown hardware processor 200 includes a plurality of cores (0 to N, where N may be 1 or more) and hardware accelerators (0 to M, where M may be 1 or more). The hardware processor 200 (e.g., one or more of the accelerators and / or one or more of its cores) may be coupled to the memory 202 (e.g., data storage device) via (e.g., respective) memory interface circuits (0 to M, where M may be 1 or more). The memory interface circuit may be a Request Address File (RAF) circuit, as described below. The memory architecture herein (e.g., via the RAF) may manipulate memory dependencies, for example, via dependency tokens. In certain embodiments of the memory architecture, the compiler emits memory operations, which are configured onto a special memory interface circuit (e.g., the RAF). The spatial array (e.g., fabric) interface to the RAF may be channel-based. Certain embodiments herein extend the definition of memory operations and the implementation of the RAF to support program order descriptions. A load operation may accept an address stream of memory requests from a spatial array (e.g., fabric) and return a data stream when the requests are satisfied. A store operation may accept two streams, e.g., one for data and one for (e.g., destination) address. In one embodiment, each of these operations exactly corresponds to one memory operation in the source program. In one embodiment, individual operation channels are strongly ordered, but there is no implied order between channels.
[0123] (e.g., of a core) The hardware decoder may receive (e.g., a single) instruction (e.g., macro-instruction) and decode the instruction, e.g., into micro-instructions and / or micro-operations. (e.g., of a core) The hardware execution unit may run the decoded instruction (e.g., macro-instruction) to perform one or more operations.
[0124] The following Subsection 2 discloses embodiments of the CSA architecture. Specifically, new embodiments of integrating memory within a data flow execution model are disclosed. Subsection 3 studies the microarchitecture details of embodiments of the CSA. In one embodiment, the main goal of the CSA is to support programs generated by compilers. The following Subsection 4 examines embodiments of the CSA compilation tool chain. The advantages of embodiments of the CSA are compared with other architectures in the execution of the compiled code in Subsection 5. The performance of embodiments of the CSA microarchitecture is discussed in Subsection 6, other CSA details are discussed in Subsection 7, an example memory sort in an acceleration hardware (e.g., a spatial array of processing elements) is discussed in Subsection 8, and an overview is provided in Subsection 9.
[0125] 2. CSA Architecture
[0126] The goal of certain embodiments of the CSA is to run programs, such as those generated by compilers, quickly and efficiently. Certain embodiments of the CSA architecture provide programming abstractions that support the needs of compiler technologies and programming paradigms. Embodiments of the CSA run data flow graphs, such as program representations that are very similar to the internal representation (IR) of the compiler's own compiled program. In this model, a program is represented as a data flow graph, which consists of nodes (e.g., vertices) drawn from a set of data flow operators defined by the architecture (e.g., which include computational and control operations) and edges representing the passing of data between the data flow operators. Execution can be performed by injecting data flow tokens (e.g., which are or represent data values) into the data flow graph. Tokens can flow between each node (e.g., vertex) and be transformed at each node, e.g., to form a complete computation. A sample data flow graph and its deviation from high-level source code are shown in Figure 3A - 3C and Figure 4 an example of the execution of the data flow graph is shown.
[0127] By precisely providing the data flow graph execution support required by the compiler, embodiments of the CSA are configured for data flow graph execution. In one embodiment, the CSA is an accelerator (e.g., the accelerator in Figure 2 ), and it does not attempt to provide a part of the necessary but not commonly used mechanisms (e.g., system calls) available in a general-purpose processing core (e.g., the core in Figure 2 ). Thus, in this embodiment, the CSA is able to run many codes but not all codes. In exchange, the CSA gains significant performance and energy advantages. To achieve the acceleration of code written in common sequential languages, the embodiments herein also introduce several new architectural features to assist the compiler. One particular novelty is the handling of memory by the CSA, i.e., a topic that has been previously overlooked or inadequately addressed. Embodiments of the CSA are also unique in the use of data flow operators as its infrastructure interfaces, e.g., as opposed to look-up tables (LUTs).
[0128] Returning to the embodiments of the CSA, the data flow operators will be discussed next.
[0129] 2.1 Data flow operators
[0130] The key architectural interface of an embodiment of an accelerator (e.g., CSA) is the data flow operator, e.g., a direct representation as a node in a data flow graph. From an operational perspective, the data flow operator behaves in a streaming or data-driven manner. The data flow operator can run immediately when its incoming operands become available. The CSA data flow execution can (e.g., only) depend on the high-location domain state, e.g., giving rise to a highly scalable architecture with a distributed asynchronous execution model. The data flow operator can include arithmetic data flow operators, e.g., one or more of floating-point addition and multiplication, integer addition, subtraction and multiplication, various forms of comparison, logical operators, and shifts. However, embodiments of the CSA can also include an enrichment of control operators, which assist in the management of data flow tokens in the program graph. Examples of these include: a "pick" operator, e.g., which multiplexes two or more logical input channels into a single output channel; and a "switch" operator, e.g., which operates as a channel demultiplexer (e.g., outputs a single channel from two or more logical input channels). These operators can enable the compiler to implement control paradigms, e.g., conditional expressions and loops. Some embodiments of the CSA can include a limited set of data flow operators (e.g., for a smaller number of operations) to yield a dense and energy-efficient PE microarchitecture. Some embodiments can include data flow operators for complex arithmetic (which is common in HPC code). An example of a data flow operator is an ordinal data flow operator, e.g., to enable efficient control of a for loop (e.g., loop construct). An embodiment of the ordinal data flow operator that implements a loop introduces a feedback path between the conditional and post-condition update parts of the loop, e.g., the for loop terms are often related, e.g., an exit condition term (e.g., "M < i < N" or "i < N") is often followed by a decrement or increment term (e.g., "i++", etc., where i is the loop counter variable). In some embodiments, this can form a bottleneck in the performance of the ordinal data flow operator implementation, which is addressed by introducing a composite ordinal operation (e.g., which can perform the conditional and update of the for loop pattern in a single operation (e.g., a single cycle)). In one embodiment, a for loop includes one or more (e.g., all) of the following parts: initialization, condition, and afterthought. In one embodiment, the initialization declares any variables required (e.g., and assigns a value(s) to them). For example, if multiple variables are being used in the initialization part, the variables can be of the same type. In one embodiment, the condition checks a certain condition and exits the loop when "false". In one embodiment, the afterthought is only executed once at the end of each loop iteration and then repeated. The CSA data flow operator architecture is highly amenable to deployment-specific extensions. For example, in some embodiments, more complex mathematical data flow operators (e.g., trigonometric functions) can be included to accelerate certain math-intensive HPC workloads. Similarly, neural network tuning extensions can include data flow operators for vectorized low-precision arithmetic.
[0131] Certain embodiments of the present disclosure provide sequencer data flow operator architectures and sequencer microarchitectures such that, for example, the generation of control signals for (e.g., the most common) for loop constructs can achieve peak performance of one loop iteration per cycle (e.g., the cycle of an accelerator including the sequencer). Certain embodiments of the present disclosure can greatly improve the performance of many high-performance computing (HPC) applications. Certain embodiments of the sequencer data flow operator separate the generation of such loop control signals from the actual data flow tokens of the loop constructs themselves such that, for example, for many HPC applications, memory prefetching and / or data speculation (and associated energy waste) can be completely eliminated. Certain embodiments of the sequencer data flow operator can be formed by modifying one or more integer processing elements (PEs) and / or by adopting (e.g., minor) configuration changes and microarchitecture extensions, and the illustrated sequencer PEs can still operate as (e.g., basic) integer PEs. Full binary compatibility with (e.g., basic) integer PEs can also be achieved to minimize software engineering costs. Certain embodiments of the present disclosure can include a sequencer data flow operator (e.g., a circuit) that manipulates data (e.g., data tokens) such as 64-bit wide, 32-bit wide, etc. in a coarse-grained manner (e.g., as opposed to control tokens), and targets the highest achievable clock frequency (e.g., 1 - 1.5 GHz), while still using an energy-efficient circuit network topology / design.
[0132] Certain embodiments of the present disclosure include a sequencer data flow operator (e.g., a circuit) that minimizes overhead in terms of energy, area, throughput, and latency. Certain embodiments of the present disclosure include a sequencer data flow operator (e.g., a circuit) that minimizes the hardware resources utilized while achieving the highest possible performance.
[0133] Some embodiments of the sequencer data flow operator can generate a loop control signal at peak performance of one loop iteration per cycle (e.g., given no output data stream token backpressure), which can be up to 2 times (2X) to 3 times (3X) faster and / or at least 50% smaller compared to not using the sequencer data flow operator. Some embodiments of the sequencer data flow operator are significantly more energy efficient, e.g., because the communication between two adjacent PEs will be shorter and use dedicated wiring between them (e.g., without using an interconnection network or its channels). Some embodiments herein are directed to a sequencer data flow operator (e.g., a circuit) that takes as input a starting value, an ending value, and a stride (e.g., base value, limit, and stride respectively) and provides one or more outputs. In one embodiment, the sequencer data flow operator outputs a control signal (e.g., a control token), e.g., outputs a first indicator value (e.g., logic one) each time an output is sent and a second indicator value (e.g., logic zero) when the operation (e.g., for loop) is completed. In one embodiment, a compare data flow operator (e.g., less than, greater than, less than or equal to, or greater than or equal to) (e.g., the compare data flow operator of the sequencer) indicates when the operation (e.g., for loop) is to stop (e.g., based on the stride). In one embodiment (e.g., Figure 22 of), the sequencer data flow operator is formed from two processing elements, e.g., one processing element performs a stride (e.g., addition) operation and the other processing element performs a compare operation, e.g., such that the PEs are combined (e.g., along with additional circuitry and / or control signals) to form the sequencer data flow operator.
[0134] Figure 3A Shows a program source according to an embodiment of the present disclosure. The program source code includes a multiplication function (powY, e.g., where Y is the power of a value). Figure 3B Shows an embodiment according to the present disclosure, Figure 3AData flow diagram 300 of the program source. The data flow diagram 300 includes a pick node 304, a switch node 306, a multiplier node 308, and a sequencer node 310. Although the sequencer node 310 is shown as a single sequencer providing control signals (such as control tokens) to multiple nodes (such as the pick node 304 and the switch node 306), multiple sequencer nodes can be utilized (such as one sequencer node for each node that is sent the (one or more) control signals). The input "A" of the sequencer node 310 can be the number of iterations "n" or a value (such as a bit pattern) that causes the sequencer node 310 to perform the number of iterations "n". Optionally, one or more buffers can be included along one or more of the communication paths. The illustrated data flow diagram 300 can employ the pick node 304 to perform the operation of selecting the input X, multiplying X by Y (such as the multiplier node 308) "n" times, accumulating each iteration, and then outputting the result from the left output of the switch node 306. The sequencer node can provide control signals to cause these operations (such as pick and switch operations) to occur. Figure 3C An accelerator (such as a CSA) having multiple processing elements 301 configured to run Figure 3B a data flow diagram in accordance with an embodiment of the present disclosure is shown. More specifically, the data flow diagram 300 is mapped into an array of processing elements 301 (such as and the network (such as the interconnect) between them), such that each node of the data flow diagram 300 is represented as a data flow operator in the array of processing elements 301. For example, certain data flow operations can be implemented using processing elements, and / or certain data flow operations can be implemented using the communication network. In one embodiment, each coupling (such as a channel) (such as for controlling data (such as control tokens) and / or (such as separately) for input / output (such as payload) data (such as data flow tokens)) includes two paths, such as Figure 7A - 7B shown. The coupling can be as described below with reference to Figure 9 The forward path can transfer data (such as control data or input / output data) from the producer to the consumer. The multiplexer can be configured to, for example, Figure 7A direct data and valid bits from the producer to the consumer. In the case of multicast, the data will be directed to multiple consumer endpoints. The second part of this embodiment of the network is the flow control or backpressure path, which flows, for example, in Figure 7B the opposite direction of the forward data path and stalls the forward flow of data on the flow control or backpressure path until the data is used or there is space to store that data. In one embodiment, the signals include one or more of control signals (such as control tokens) from the sequencer data flow operator and / or input / output data signals (such as data flow tokens) from other data flow operators (such as pick operators and switch operators). For example, when the flow control or backpressure path (which, for example, in Figure 7Bflows in a direction opposite to the forward data path) halts the forward flow of the stalled data, e.g., when that forward data is being used or there is space to store that data, Figure 3C the forward flow of each admissible data (e.g., a control signal from the sequencer operator 310A (also referred to as the “sequencer data flow operator”) or an input / output data signal going to and / or from other operators) on the lines in. Thus, in some embodiments, each communication path can be stalled by a backpressure signal.
[0135] In one embodiment, one or more of the processing elements in the array of processing elements 301 access memory via the memory interface 302. In one embodiment, the pick node 304 of the data flow graph 300 thus corresponds to the pick operator 304A (e.g., represented thereby), the switch node 306 of the data flow graph 300 thus corresponds to the switch operator 306A (e.g., represented thereby), the multiplier node 308 of the data flow graph 300 thus corresponds to the multiplier operator 308A (e.g., represented thereby), and the sequencer node 310 of the data flow graph 300 thus corresponds to the sequencer operator 310A (e.g., the sequencer data flow operator) (e.g., represented thereby). Another processing element and / or flow control path network can provide control signals (e.g., control tokens) to the pick operator 304A and the switch operator 306A to perform Figure 3A the operations in. In the illustrated embodiment, the sequencer operator 310A provides control signals (e.g., control tokens) to the pick operator 304A and the switch operator 306A to perform Figure 3A the operations in. For example, if Y = 2, the variable X will be raised to the power of two “n” times, e.g., if X = 1, this will provide a square. In the illustrated embodiment, the path is configured (e.g., provided) from the right output of the switch operator 306A to the right input of the pick operator 304A, e.g., to iteratively receive the output from the multiplier operator 308A.
[0136] In one embodiment, the array of processing elements 301 (e.g., the sequencer operator 310A) is configured to run Figure 3B the data flow graph 300 before the execution begins. In one embodiment, the compiler performs the transformation from Figure 3A - 3B In one embodiment, the data flow graph nodes logically embed the data flow graph in the array of processing elements for the inputs in the array of processing elements, e.g., as further discussed below, such that the input / output paths are configured to produce the desired result.
[0137] 2.2 Latency-insensitive channels
[0138] Communication arcs are the second major component of the data flow graph. Some embodiments of the CSA describe these arcs as latency-insensitive channels, such as ordered, backpressure (e.g., not producing or sending an output until there is a place to store the output), point-to-point communication channels. Like data flow operators, latency-insensitive channels are substantially asynchronous, thus giving the freedom to compose many types of networks to implement the channels of a particular graph. Latency-insensitive channels can have arbitrarily long latencies and still reliably implement the CSA architecture. However, in some embodiments, there is a strong motivation to make the latency as small as possible in terms of performance and energy. Subsection 3.2 of this document discloses a network microarchitecture in which data flow graph channels are implemented with a latency of no more than one cycle in a pipelined manner. Embodiments of latency-insensitive channels provide a key abstraction layer that can be balanced using the CSA architecture in order to provide multiple runtime services to application programmers. For example, the CSA can balance the latency-insensitive channels in the implementation of the CSA configuration (loading a program onto the CSA array).
[0139] Figure 4 An example execution of a data flow graph 400 in accordance with an embodiment of the present disclosure is shown. The data flow graph 400 can be mapped onto multiple processing elements (e.g., as well as an interconnect network) such that each node (e.g., switch node, pick node, multiplier node, etc.) is represented as a data flow operator. In step 1, input values (e.g., for X being 1 in Figure 3B - 3C and for Y being 2 in Figure 3B - 3C ) can be loaded into the data flow graph 400 to perform a 1×2 multiplication operation “n” times (as controlled by sequencer node 410). One or more of the data input values can be static (e.g., constant) in operation (e.g., referring to Figure 3B - 3C , for X being 1 and for Y being 2), or updated during operation. In step 1, the sequencer node 410 is loaded with 2, e.g., which can indicate two iterations of the multiplication to be performed (e.g., for Figure 3A, n = 2). The sequencer node 410 can provide (e.g., preload) control signals that correspond to causing circuitry (e.g., the pick operator of pick node 404 and the switch operator of switch node 406) to perform multiplication, e.g., where the multiplier operator of multiplier node 408 outputs its result upon receiving an input. In step 2, the sequencer node 410 outputs zero to control the input of pick node 404 (e.g., the mux control signal) (e.g., sourcing an output from port "0" to it), and outputs zero to control the input of switch node 406 (e.g., the mux control signal) (e.g., providing its input from port "0" to a destination (e.g., a downstream processing element)). In step 3, the data value 1 is output from pick node 404 (e.g., and pick node 404 consumes its control signal "0") to multiplier node 408 for multiplication by the data value 2 in step 4. In step 4, the output of multiplier node 408 arrives at switch node 406, e.g., which causes switch node 406 to consume control signal "1" to output the value 2 from port "1" of switch node 406 in step 5. In step 5, the output of multiplier node 408 arrives at pick node 404 again (e.g., because 2 iterations (n = 2) are to be performed here), e.g., which causes pick node 404 to consume control signal "1" to output the value 2 from port "1" of pick node 404 in step 6. In step 6, the data value 2 is output from pick node 404 (e.g., and pick node 404 consumes its control signal "1") to multiplier node 408 for multiplication by the data value 2 in step 7. In step 7, the output of multiplier node 408 arrives at switch node 406, e.g., which causes switch node 406 to consume control signal "0" to output the value 4 from port "0" of switch node 406 in step 8. In step 8, the output of multiplier node 408 arrives at switch node 406 (e.g., because 2 iterations (n = 2) are to be performed here and n is zero at this time, so the operation ends), e.g., which causes switch node 406 to consume control signal "0" to output the value 4 from port "0" of switch node 406. The operation then completes. Thus, the CSA can be programmed accordingly such that the corresponding data flow operators of each node perform Figure 4 the operations in. Although the execution is serialized in this example, generally all data flow operations can run in parallel. The steps are in Figure 4used to distinguish data flow execution from any physical microarchitecture manifestation. In some embodiments, a downstream processing element sends a signal (or does not send a ready signal) (e.g., on a flow control path network) to the switching operator of the switching node 406 to stall the output from the switching node 406 (e.g., value 4), e.g., until the downstream processing element is ready for the output (e.g., has storage space). In some embodiments, the picking operator of the picking node 404 sends a signal (or does not send a ready signal) (e.g., on a flow control path network) to the upstream and downstream processing elements to stall the input (e.g., value 1) entering the picking node 404, e.g., until the processing element is ready for the input (e.g., has storage space). In some embodiments, the sequencer operator of the sequencer node 410 sends a signal (or does not send a ready signal) (e.g., on a flow control path network) to the upstream and downstream processing elements to stall the input (e.g., value 2) entering the sequencer node 410, e.g., until the processing element is ready for the input (e.g., has storage space). A spatial array (e.g., CSA) (e.g., PEs of the spatial array), a processor, or a system may include any of the disclosures herein, e.g., one or more PEs of a spatial array according to any of the architectures disclosed herein.
[0140] 2.3 Memory
[0141] Data flow architectures generally focus on communication and data manipulation, with less attention to state. However, enabling real-world software, especially programs written in legacy sequential languages, requires great attention to interfacing with memory. Some embodiments of CSA use architectural memory operations as the primary interface to a (large) stateful storage device. From the perspective of a data flow graph, memory operations are similar to other data flow operations, except that they have the side effect of updating shared storage. Specifically, memory operations in some embodiments herein have the same semantics as every other data flow operator, e.g., they "run" when their operands (e.g., addresses) are available and produce a response after some latency. Some embodiments herein explicitly separate operand inputs and result outputs, such that memory operators are naturally pipelined and have the possibility of generating many outstanding requests simultaneously, e.g., making them particularly well-suited to the latency and bandwidth characteristics of the memory subsystem. Embodiments of CSA provide basic memory operations, such as load (which takes an address channel and loads the value corresponding to the address into a response channel) and store. Embodiments of CSA may also provide more advanced operations, such as in-memory atomic and consistency operators. These operations may have semantics similar to their von Neumann counterparts. Embodiments of CSA may accelerate existing programs described in sequential languages (e.g., C and Fortran). The result of supporting these language models is to resolve program memory order, e.g., the serial ordering of memory operations typically prescribed by these languages.
[0142] Figure 5A Shows a program source (e.g., C code) 500 according to an embodiment of the present disclosure. According to the memory semantics of the C programming language, the memory copy (memcpy) should be serialized. However, if arrays A and B are known to be separate, memcpy can use an embodiment of CSA to serialize. Figure 5A Further shows the problem of program order. Generally, for example, for the same or different index values across a loop body, the compiler cannot prove that arrays A and B are different. This is called pointer or memory aliasing. Since compilers generate statically correct code, they are usually forced to serialize memory accesses. Typically, compilers for sequential von Neumann architectures use instruction scheduling as a natural means to enforce program order. However, embodiments of CSA do not have the concept of instruction or instruction-based program ordering as defined by the program counter. In some embodiments, the incoming dependence token (e.g., which does not contain architecturally visible information) is similar to all other data flow tokens, and memory operations may not execute until they receive a dependence token. In some embodiments, once its operation is visible to all logically subsequent dependent memory operations, a memory operation generates an outgoing dependence token. In some embodiments, dependence tokens are similar to other data flow tokens in the data flow graph. For example, since memory operations occur in a conditional context, dependence tokens can also be manipulated using the control operators described in Subsection 2.1, e.g., similar to any other token. Dependence tokens can have the effect of serializing memory accesses, e.g., providing the compiler with a means to define the order of memory accesses architecturally. Figure 5B Shows a program source (e.g., C code) 501 according to an embodiment of the present disclosure. The program source 501 can be a for loop construct for a memory copy operation to copy data from a vector "a" of "N" elements to a vector "b" of "N" elements.
[0143] 2.4 Runtime Services
[0144] The main architectural considerations for embodiments of the CSA relate to the actual execution of user-level programs, but it is also desirable to provide several support mechanisms that consolidate this execution. Chief among these are configuration (where data flow graphs are loaded into the CSA), extraction (where the state of the running graph is moved to memory), and exceptions (where mathematical, soft, and other types of errors in the structure may be detected and manipulated by external entities). Subsection 3.6 below discusses the nature of the latency-insensitive data flow architecture of embodiments of the CSA that produce these functions in a largely pipelined and efficient implementation. Conceptually, configuration can load the state of a data flow graph, for example, generally from memory, into the interconnect (and / or communication network) and processing elements (such as the fabric). During this step, all of the fabric in the CSA can be loaded with a new data flow graph and any data flow tokens that are alive in that graph, for example, as a result of a context switch. The latency-insensitive semantics of the CSA can permit distributed asynchronous initialization of the fabric. For example, when configuring PEs, they can immediately begin execution. Unconfigured PEs can backpressure their channels until they are configured, for example, to prevent communication between configured and unconfigured elements. CSA configuration can be divided into privileged and user-level states. This two-level division can enable the main configuration of the fabric to occur without invoking the operating system. During one embodiment of extraction, the logical view of the data flow graph is captured and committed to memory, for example, including all alive control and data flow tokens in the graph.
[0145] Extraction can also play a role in providing reliability guarantees through the creation of fabric checkpoints. Exceptions in the CSA are generally caused by the same events that cause exceptions in a processor, such as illegal operator arguments or reliability, availability, and serviceability (RAS) events. In some embodiments, exceptions are detected at the data flow operator level, for example, by checking argument values or through a modular arithmetic scheme. When an exception is detected, the data flow operator (such as a circuit) can pause and emit an exception message, for example, that contains the operation identifier and some details about the nature of the problem that has occurred. In one embodiment, the data flow operator will remain paused until it has been reconfigured. The exception message can then be passed to an associated processor (such as a core) for servicing, for example, which may include extracting the graph for software analysis.
[0146] 2.5 Primitives-level Architecture
[0147] Embodiments of the CSA computer architecture (e.g., for HPC and data center use) are tiled. Figure 6 and Figure 8 illustrates the primitives-level deployment of the CSA. Figure 8Shows a full primitive implementation of the CSA, which can be, for example, an accelerator for a processor with a core. The main advantage of this architecture is that it can reduce design risks, such as enabling the complete separation of the CSA and the core during manufacturing. In addition to allowing better component reuse, this also allows the design of components (such as the CSA cache) to consider only the CSA, for example, without having to incorporate the more stringent latency requirements of the core. Finally, the independent primitive allows the integration of the CSA with small or large cores. One embodiment of the CSA captures most vector parallel workloads, such that most vector-style workloads run directly on the CSA, but in some embodiments, vector-style instructions may be included in the core, for example, to support legacy binaries.
[0148] 3. Microarchitecture
[0149] In one embodiment, the goal of the CSA microarchitecture is to provide a high-quality implementation of each data flow operator specified by the CSA architecture. Embodiments of the CSA microarchitecture stipulate that each processing element (and / or communication network) of the microarchitecture corresponds to roughly one node (e.g., entity) in the architecture data flow graph. In one embodiment, the nodes in the data flow graph are distributed across multiple network data flow endpoint circuits. In some embodiments, this results in microarchitecture elements that are not only compact (thus resulting in a dense computing array), but also energy-efficient, for example, where the processing elements (PEs) are both simple and largely non-reused, such as running a single data flow operator of the CSA configuration (e.g., programming). To further reduce energy and implementation area, the CSA may include a configurable heterogeneous structure style, where each of its PEs implements only a subset of the data flow operators (e.g., having an independent subset of the data flow operators implemented by (one or more) network data flow endpoint circuits). Peripheral and support subsystems (such as the CSA cache) may be provided to support the distributed parallelism present in the main CSA processing structure itself. The implementation of the CSA microarchitecture can utilize the data flow and latency-insensitive communication abstractions present in the architecture. In some embodiments, there is a (substantially) one-to-one correspondence between the nodes in the compiler-generated graph and the data flow operators (e.g., data flow operator computing elements) in the CSA.
[0150] The following is a discussion of an example CSA, followed by a more detailed discussion of the microarchitecture. Some embodiments of this document provide a CSA that allows for easy compilation, for example, as opposed to existing FPGA compilers that manipulate a small subset of programming languages (such as C or C++) and require many hours to compile even small programs.
[0151] Some embodiments of the CSA architecture permit heterogeneous coarse-grained operations, such as double-precision floating point. Programs can be expressed with fewer coarse-grained operations, such that the disclosed compiler runs faster than traditional spatial compilers. Some embodiments include a structure having new processing elements that support sequential concepts, such as program-order memory access. Some embodiments implement hardware that supports coarse-grained data-flow style communication channels. This communication model is abstract and very close to the control-data-flow representation used by the compiler. Some embodiments herein include a network implementation that supports single-cycle latency communication, such as using (e.g., small) PEs that support single control-data-flow operations. In some embodiments, this not only improves energy efficiency and performance, it also simplifies compilation because the compiler makes a one-to-one mapping between high-level data-flow constructs and the structure. Thus, some embodiments herein simplify the task of compiling existing (e.g., C, C++, or Fortran) programs into CSA (e.g., structures).
[0152] Energy efficiency would be a primary concern in modern computer systems. Some embodiments herein provide new schemes for energy-efficient spatial architectures. In some embodiments, these architectures form a structure with a unique synthesis of a heterogeneous mix (and / or packet-switching communication network) of small energy-efficient data-flow oriented processing elements (PEs) and a lightweight circuit-switching communication network (e.g., an interconnect) with stable support for flow control, for example. Due to the energy advantages of each, the combination of these components can form a spatial accelerator (e.g., as part of a computer) suitable for running compiler-generated parallel programs in a highly energy-efficient manner. Since this structure is heterogeneous, some embodiments can be customized for different application domains by introducing new domain-specific PEs. For example, a structure for high-performance computing can include a certain customization for double-precision, fused multiply-add, while a structure for deep neural networks can include low-precision floating-point operations.
[0153] For example Figure 6 An embodiment of the spatial architecture scheme shown is a synthesis of lightweight processing elements (PEs) connected by a network between PEs. Generally, a PE can include a data-flow operator, such that once (e.g., all) input operands arrive at the data-flow operator, an operation (e.g., a micro-instruction or a set of micro-instructions) is run and the result is forwarded to a downstream operator. Thus, control, scheduling, and data storage can be distributed among the PEs, for example removing the overhead of the centralized structure that dominates traditional processors.
[0154] The program can be transformed into a data flow graph, which is mapped to the architecture by configuring the PEs and the network to express the control-data flow graph of the program. The communication channels can be flow-controlled and fully backpressure, such that for example a PE will stall when there is no data on a source communication channel (e.g., one or more sources) or the destination communication channel (e.g., one or more destinations) is full. In one embodiment, at runtime, data flows through the PEs and channels which have been configured to implement the operation (e.g., accelerate an algorithm). For example, data can be streamed from memory through the fabric and then output back to memory.
[0155] Embodiments of this architecture can achieve significant performance efficiency relative to traditional multi-core processors: the computations (e.g., in the form of PEs) can be simpler, more energy-efficient, and more plentiful than larger cores, and the communication can be direct and mostly short-range, e.g., as opposed to what occurs through a wide on-chip network as in a typical multi-core processor. Further, because embodiments of the architecture are extremely parallel, multiple aggressive circuit and device-level optimizations are possible without severely affecting throughput, e.g., low-leakage devices and low operating voltages. These low-level optimizations can achieve even greater performance benefits relative to traditional cores. The combination of efficiency at the architecture, circuit, and device levels in these embodiments is compelling. As transistor density continues to increase, embodiments of this architecture can achieve a larger effective area.
[0156] Embodiments herein provide a unique combination of data flow support and circuit switching in order to enable a fabric to be smaller, more energy-efficient, and provide higher aggregate performance compared to prior architectures. FPGAs are generally tuned to fine-grained bit manipulation, while embodiments herein are tuned to double-precision floating point operations present in HPC applications. Some embodiments herein may include an FPGA in addition to the CSA in accordance with the present disclosure.
[0157] Some embodiments herein combine a lightweight network with energy-efficient data flow processing elements (and / or communication networks) to form a high-throughput, low-latency, energy-efficient HPC fabric. This low-latency network enables the construction of processing elements (and / or communication networks) with less functionality, e.g., only one or two instructions and perhaps one architecturally visible register, since it is effective to combine multiple PEs together to form a complete program.
[0158] Relative to processor cores, embodiments of the CSA herein can provide greater computational density and energy efficiency. For example, when the PE is small (e.g., compared to a core), the CSA can perform more operations and has many more computational paradigms than a core, e.g., perhaps up to 16 times the number of FMAs of a vector processing unit (VPU). In some embodiments, the energy per operation is low in order to utilize all of these computational elements.
[0159] There are many energy advantages to embodiments of this data flow architecture. Parallelism is evident in the data flow graph, and embodiments of the CSA architecture do not expend or expend minimal energy to extract it, unlike out-of-order processors which must rediscover patterns each time an instruction is executed. Since in one embodiment each PE is responsible for a single operation, the register file and port count can be smaller, e.g., typically only one, and thus use less energy than their counterparts in a core. Some CSAs include many PEs, each of which holds in-progress program values, giving the aggregate effect of a large register file in a traditional architecture, which greatly reduces memory accesses. In embodiments where the memory is multi-port and distributed, the CSA can sustain more outstanding memory requests and utilize a larger bandwidth than a core. These advantages can be combined to yield an energy level per watt that is a small percentage of the cost of just the bare arithmetic circuitry. For example, in the case of integer multiplication, the CSA can consume no more than 25% of the energy compared to a base multiplication circuit. Relative to one embodiment of a core, integer operations in that CSA architecture consume less than 1 / 30 of the energy per integer operation.
[0160] From a programming perspective, the application-specific extensibility of embodiments of the CSA architecture yields significant advantages over vector processing units (VPUs). In traditional inflexible architectures, the number of functional units (e.g., floating point divide or various transcendental math functions) must be chosen at design time based on some projected use case. In embodiments of the CSA architecture, such functions can be configured (e.g., by the user rather than the manufacturer) into the architecture based on the requirements of each application. This can further increase application throughput. At the same time, the computational density of embodiments of the CSA is improved by avoiding hardening such functions and instead provisioning more instances of primitive functions (e.g., floating point multiplication). These advantages can be significant in HPC workloads, a portion of which spend over 75% of their floating point execution time in transcendental functions.
[0161] Certain embodiments of the CSA represent a significant advancement as a data flow-oriented spatial architecture. For example, the PEs of the present disclosure can be smaller but also more energy-efficient. These improvements can arise directly from the combination of data flow-oriented PEs with lightweight circuit-switching interconnects, e.g., which have a single-cycle latency, e.g., in contrast to packet-switching networks (e.g., which have at least 300% higher latency). Some embodiments of the PE support 32-bit or 64-bit operations. Some embodiments herein permit the introduction of new specialized PEs, e.g., for machine learning or security, rather than just homogeneous combinations. Some embodiments herein combine lightweight data flow-oriented processing elements with lightweight low-latency networks to form an energy-efficient computing fabric.
[0162] For some spatial architectures to succeed, programmers configure them with less workload, e.g., achieving significant power and performance advantages over sequential cores simultaneously. Some embodiments of the present disclosure provide a CSA (e.g., a spatial fabric) that is easy to program (e.g., via a compiler), power-efficient, and highly parallel. Some embodiments of the present disclosure provide a network (e.g., an interconnect) that achieves these three goals. From a programmability perspective, some embodiments of the network provide flow control channels, e.g., corresponding to the control-data flow graph (CDFG) model of execution used in a compiler. Some network embodiments utilize dedicated circuit-switched links, making program performance more amenable to being rolled out by humans and compilers because the performance is predictable. Some network embodiments provide high bandwidth and low latency. Some network embodiments (e.g., static circuit switching) provide a latency of 0 to 1 cycle (e.g., depending on the transmission distance). Some network embodiments provide high bandwidth by, e.g., laying out several networks in parallel and in lower-level metal. Some network embodiments communicate in lower-level metal and over short distances and are thus very energy-efficient.
[0163] Some embodiments of the network include architectural support for flow control. For example, in a spatial accelerator composed of small processing elements (PEs), communication latency and bandwidth can be critical to overall program performance. Some embodiments of the present disclosure provide a lightweight circuit-switched network that facilitates communication between PEs in a spatial processing array (e.g., Figure 6 the spatial array shown) and the microarchitectural control features required to support this network. Some embodiments of the network implement the formation of point-to-point flow control communication channels that support the communication of dataflow-oriented processing elements (PEs). In addition to point-to-point communication, some networks of the present disclosure also support multicast communication. The communication channels can be formed by statically configuring the network to form virtual circuits between PEs. The circuit-switching technology of the present disclosure can reduce communication latency and, correspondingly, minimize network buffering, e.g., resulting in high performance and high energy efficiency. In some embodiments of the network, the latency between PEs can be as low as zero cycles, meaning that the downstream PE can operate on the data in the cycle after it is generated. To achieve even higher bandwidth and accommodate more programs, multiple networks can be laid out in parallel, e.g., Figure 6 as shown.
[0164] Spatial architectures (e.g., Figure 6The spatial architecture shown) can be a composition of lightweight processing elements connected by a network between PEs (and / or a communication network). A program, regarded as a data flow graph, can be mapped to the architecture by configuring the PEs and the network. Generally, a PE can be configured as a data flow operator, and once (e.g., all) input operands arrive at the PE, an operation can occur and the result is forwarded to the intended downstream PE. PEs can communicate through dedicated virtual circuits formed by statically configuring a circuit-switched communication network. These virtual circuits can be flow-controlled and fully backpressure, e.g., causing a PE to stall when the source has no data or the destination is full. At runtime, data can flow through the PEs implementing the mapped algorithm. For example, data can be streamed from memory through the structure and then output to memory again. Embodiments of this architecture can achieve significant performance efficiency compared to traditional multi-core processors: e.g., where the computations in the form of PEs are simpler and more numerous than large cores, and the communication is direct, e.g., as opposed to the expansion of the memory system.
[0165] Figure 6 An accelerator primitive 600 including an array of processing elements (PEs) is shown in accordance with an embodiment of the present disclosure. The interconnect network is shown as a circuit-switched statically configured communication channel. For example, a set of channels is commonly coupled through switches (e.g., switch 610 in the first network and switch 611 in the second network). The first network and the second network can be independent or coupled together. For example, switch 610 can couple one or more of four data paths (612, 614, 616, 618) together, e.g., configured to perform operations according to a data flow graph. In one embodiment, the number of data paths is any number. Processing elements (e.g., processing element 604) can be as disclosed herein, e.g., in Figure 9 as disclosed. The accelerator primitive 600 includes a memory / cache hierarchy interface 602, e.g., to interface the accelerator primitive 600 with memory and / or cache. A data path (e.g., 618) can extend to another primitive or terminate at the edge of a primitive, e.g.. A processing element can include an input buffer (e.g., buffer 606) and an output buffer (e.g., buffer 608).
[0166] Operations can run based on the availability of their inputs and the state of the PE. A PE can obtain operands from an input channel and write results to an output channel, but can also use internal register states. Some embodiments herein include configurable data flow-friendly PEs. Figure 9A detailed block diagram of such a PE (integer PE) is shown. This PE consists of several I / O buffers, an ALU, storage registers, some instruction registers, and a scheduler. In each cycle, the scheduler can select an instruction for execution based on the availability of the input and output buffers and the status of the PE. The result of the operation can then be written to the output buffer or a register (e.g., local to the PE). The data written to the output buffer can be transferred to a downstream PE for further processing. A PE of this style can be extremely energy-efficient. For example, instead of reading data from a complex multi-port register file, the PE reads data from a register. Similarly, instructions can be stored directly in registers rather than in a virtualized instruction cache.
[0167] The instruction registers can be set during a special configuration step. During this step, in addition to the PE-to-PE network, auxiliary control wires and status can also be used to broadcast across several PEs including the fabric in the configuration. Due to parallelism, some embodiments of such a network can provide fast reconfiguration. For example, a primitive-sized fabric can be configured in less than about 10 microseconds.
[0168] Figure 9 An example configuration of a processing element is shown, for example where the architectural elements are sized minimally. In other embodiments, each of the components of the processing element is scaled individually to produce a new PE. For example, to handle more complex programs, a larger number of instructions executable by the PE can be introduced. A second dimension of configurability lies in the function of the PE arithmetic logic unit (ALU). Figure 9 In [the figure], an integer PE is shown, which can support addition, subtraction, and various logical operations. Other types of PEs can be created by substituting different types of functional units into the PE. For example, an integer multiplication PE may not have registers, a single instruction, and a single output buffer. Some embodiments of the PE decompose fused multiply-add (FMA) into independent but tightly coupled floating-point multiply and floating-point add units to improve support for multiply-accumulate workloads. The PE is further discussed below.
[0169] Figure 7A A configurable data path network 700 according to an embodiment of the present disclosure is shown (e.g., with reference to Figure 6 Network One or Network Two described). Network 700 includes a plurality of multiplexers (e.g., multiplexers 702, 704, 706), which can be configured (e.g., via their respective control signals) to connect one or more data paths (e.g., from PEs) together. Figure 7B A configurable flow control path network 701 according to an embodiment of the present disclosure is shown (e.g., with reference to Figure 6 Network One or Network Two described). The network can be a lightweight PE-to-PE network. Some embodiments of the network can be considered a collection of composable primitives of a distributed point-to-point data channel construction.Figure 7A A network is shown that has two enabled channels, namely the thick black line and the dashed black line. The thick black line channel is for multicast, for example, a single input is sent to two outputs. Note that the channels can intersect at certain points within a single network, even though dedicated circuit switching paths are formed between the channel endpoints. Additionally, this intersection can be made without introducing structural hazards between the two channels, enabling them to operate independently and at full bandwidth.
[0170] Implementing a distributed data channel can include Figure 7A - 7B the two paths shown. The forward or data path transfers data from the producer to the consumer. The multiplexer can be configured to direct, for example, data and valid bits from the producer to the consumer in Figure 7A . In the case of multicast, the data will be directed to multiple consumer endpoints. The second part of this embodiment of the network is the flow control or backpressure path, which flows, for example, in Figure 7B opposite to the forward data path. The consumer endpoints can assert the times when they are ready to accept new data. These signals can then be directed back to the producer using configurable logic combining (labeled as the (e.g., reflux) flow control function in Figure 7B ). In one embodiment, each flow control function circuit can be a plurality of switches (e.g., muxes), for example, similar to Figure 7A . The flow control path can manipulate the return control data from the consumer to the producer. Combining can enable multicast, for example, where each consumer is ready to receive data before the producer assumes the data has been received. In one embodiment, the PE is a PE having a data flow operator as its architectural interface. As a supplement or alternative, in one embodiment, the PE can be any kind of PE (e.g., in a fabric), non - restrictively, for example, a PE having an instruction pointer, a triggered instruction, or a state - machine - based architectural interface.
[0171] In addition to statically configuring the PEs, the network can also be statically configured. During the configuration step, configuration bits can be set in each network component. These bits control, for example, mux selection and flow control functions. The network can include multiple networks, such as a data path network and a flow control path network. The network or multiple networks can utilize paths of different widths (e.g., a first width and a narrower or wider width). In one embodiment, the data path network has a wider (e.g., bit - transfer) width than the flow control path network. In one embodiment, each of the first network and the second network includes its own data path network and flow control path network, such as data path network A and flow control path network A and a wider data path network B and flow control path network B.
[0172] Some embodiments of the network are bufferless, and data moves between producers and consumers in a single cycle. Some embodiments of the network are also unbounded, i.e., the network spans the entire fabric. In one embodiment, one PE communicates with any other PE in a single cycle. In one embodiment, to improve routing bandwidth, several networks can be arranged in parallel between rows of PEs.
[0173] Relative to FPGAs, some embodiments of the networks herein have three advantages: area, frequency, and program expressiveness. Some embodiments of the networks herein operate at a coarse-grain level, e.g., they reduce the number of configuration bits and thus reduce the area of the network. Some embodiments of the network also achieve area reduction by directly implementing flow control logic in the circuit (e.g., silicon). Some embodiments of hardened network implementations also enjoy a frequency advantage over FPGAs. Due to the area and frequency advantages, a power advantage may exist, where a lower voltage is used during throughput parity checks. Finally, some embodiments of the networks herein provide better high-level semantics than FPGA wires, especially relative to variable timing, and thus those some embodiments are easy to target by a compiler. Some embodiments of the networks herein can be considered a set of composable primitives for the construction of distributed point-to-point data channels.
[0174] In some embodiments, a multicast source may not assert that its data is valid unless it receives a ready signal from each sink. Thus, additional conjunction and control bits can be utilized in the multicast case.
[0175] Similar to some PEs, the network can be statically configured. During this step, configuration bits are set at each network component. These bits control, for example, mux selection and flow control functions. The forward path of our network requires some bits to swing its mux. In Figure 7A the example shown, four bits are required per hop: one bit is utilized for each of the east and west muxes, while two bits are utilized for the southward mux. In this embodiment, four bits can be used for the data path, but seven bits can be used for the flow control function (e.g., in a flow control path network). For example, if the CSA also utilizes the north-south direction, other embodiments may utilize more bits. The flow control function can use control bits for each direction from which flow control can originate. This enables static setting of the sensitivity of the flow control function. Table 1 below summarizes Figure 7B the Boolean algebra implementation of the flow control function of the network in
[0176] Table 1: Flow implementation
[0177]
[0178] For Figure 7BThe third process control box to the left, where EAST_WEST_SENSITIVE and NORTH_SOUTH_SENSITIVE are shown as being set to implement process control for bold lines and dashed channels, respectively.
[0179] Figure 8 Shown is a hardware processor primitive 800, including an accelerator 802, in accordance with an embodiment of the present disclosure. The accelerator 802 may be a CSA in accordance with the present disclosure. The primitive 800 includes a plurality of cache groups (e.g., cache group 808). A Request Address File (RAF) circuit 810 may be included, such as described below in Subsection 3.2. ODI may represent on-die interconnect, e.g., an interconnect that extends across the entire die and connects all primitives together. OTI may represent on-primitive interconnect, e.g., an interconnect that extends across a primitive, such as connecting cache groups on a primitive together.
[0180] 3.1 Processing Elements
[0181] In some embodiments, the CSA includes a heterogeneous PE array, where the fabric is composed of several types of PEs, each of which implements only a subset of the data flow operators. By way of example, Figure 9 Shown is a provisional implementation of a PE capable of implementing a wide set of integer and control operations. Other PEs (including those that support floating-point addition, floating-point multiplication, buffering, and certain control operations) may have a similar implementation style, e.g., having appropriate (data flow operator) circuitry in place of an ALU. The PEs of the CSA (e.g., data flow operators) may be configured (e.g., programmed) prior to the start of execution to implement a particular data flow operation from among the set supported by the PE. The configuration may include one or two control words that specify the operation code to control the ALU, direct various multiplexers within the PE, and drive data flow into and out of the PE channels. The data flow operators may be implemented by microcoding these configuration bits. Figure 9 The illustrated integer PE 900 in is organized as a single-stage logic pipeline that flows from top to bottom. Data enters the PE 900 from one of the local network sets, where it is recorded in an input buffer for subsequent operations. Each PE may support multiple wide data-oriented and narrow control-oriented channels. The number of preparatory channels may vary based on the PE functionality, but one embodiment of an integer-oriented PE has 2 wide and 1 - 2 narrow input and output channels. Although the integer PE is implemented as a single-cycle pipeline, other pipeline options may be utilized. For example, a multiplication PE may have multiple pipeline stages.
[0182] PE execution can be carried out in a data flow style. Based on the configured microcode, the scheduler can check the status of the PE input and output buffers, and when all the inputs of the configured operation have arrived and the output buffer of the operation is available, check the actual execution of the operation by a data flow operator (e.g., on an ALU). The resulting value can be placed in the configured output buffer. The transfer between the output buffer of one PE and the input buffer of another PE can be asynchronous when the buffer becomes available. In some embodiments, the PE is prepared such that at least one data flow operation is completed per cycle. Subsection 2 discusses data flow operators that include primitive operations (e.g., addition, exclusive OR, or pick). In some embodiments, the PE microarchitecture implements more than one data flow operator (e.g., a fused operator) within a single PE. This possibility occurs because different operators (e.g., arithmetic and control) can involve different paths within the PE. For example, in addition to several other useful fusion combinations, Figure 9 the PE shown can also fuse any arithmetic operation with a switch control operator. The energy, area, performance, and latency advantages of this capability are obvious. Through a small extension to the PE control path, more fusion combinations can be implemented in some embodiments. To manipulate a part of a more complex data flow operator (e.g., a floating-point fused multiply-add (FMA) and / or a loop control sequencer data flow operator), multiple PEs can be combined, e.g., instead of preparing a more complex single PE. In some embodiments, additional functional specific circuits (e.g., communication paths) are added between the combinable PEs. In one embodiment, the sequencer data flow operator implements a for loop control, and the combined path can be added between adjacent PEs to carry loop-related control information. For example, in the case where the combined behavior is not used for a specific data flow graph, such a PE combination can maintain fully pipelined behavior while preserving the utility of the basic PE. Some embodiments can provide advantages in terms of energy, area, performance, and latency. In one embodiment, more fusion combinations can be achieved through an extension to the PE control path. In one embodiment, the width of the processing element is 64 bits for heavy utilization of double-precision floating-point calculations in HPC and supports 64-bit memory addressing.
[0183] 3.2 Communication Network
[0184] Embodiments of the CSA microarchitecture provide a hierarchical structure for a network that together provides an implementation of an architectural abstraction of latency-insensitive channels across multiple communication scales. The lowest level of the CSA communication hierarchy can be the local network. The local network can be static circuit-switched, for example using configuration registers to swing the multiplexer(s) in the local network data path to form a fixed electrical path between communicating PEs. In one embodiment, the configuration of the local network is set once per data flow graph, for example at the same time as the PE configuration. In one embodiment, static circuit switching optimizes for energy, for example where the vast majority (perhaps greater than 95%) of the CSA communication traffic will cross the local network. A program can include items used in multiple expressions. To optimize for this situation, embodiments herein provide hardware support for multicast within the local network. A number of local networks can be combined together to form a routing channel, for example that is spread (as a grid) between rows and columns of PEs. As an optimization, a number of local networks can be included to carry control tokens. Compared to FPGA interconnects, the CSA local network can route at the granularity of data paths, and another difference can be the CSA's handling of control. One embodiment of the CSA local network is explicitly flow-controlled (e.g., backpressure). For example, for each forward data path and set of multiplexers, the CSA provides a backward flow control path that pairs physically with the forward data path group. The combination of the two microarchitecture paths can provide a low-latency, low-energy, small-area, point-to-point implementation of the latency-insensitive channel abstraction. In one embodiment, the flow control lines of the CSA are not visible to the user program, but they can be manipulated through the architecture in the services of the user program. For example, the exception handling mechanism described in Subsection 2.2 can be implemented by originating the flow control lines to a "nonexistent" state upon detection of an exception condition. This action can not only moderately stall those parts of the pipeline involved in the violating computation, but also save the machine state that caused the exception, for example for diagnostic analysis. The second network layer (e.g., the small backplane network) can be a shared packet-switching network. The small backplane network can include multiple distributed network controllers, network data flow endpoint circuits. The small backplane network (e.g. Figure 39The network shown schematically by the dashed box can provide more general long-range communication at the cost of, for example, latency, bandwidth, and energy. In some programs, most communication can occur on the local network, and thus the small backplane network provision will be significantly reduced by comparison. For example, each PE can be connected to multiple local networks, but the CSA will provision only one small backplane endpoint per logical neighborhood of the PE. Since the small backplane is effectively a shared network, each small backplane network can carry, for example, multiple logically independent channels and is provisioned with multiple virtual channels. In one embodiment, the main function of the small backplane network is to provide wide-ranging communication between PEs and between PEs and the memory. In addition to this capability, the small backplane can also include (one or more) network data flow endpoint circuits, for example, to perform certain data flow operations. In addition to this capability, the small backplane can also operate as a runtime support network, through which, for example, various services can access the complete fabric in a transparent manner to the user program. Through this capability, the small backplane endpoint can be used as a controller for its local neighborhood, for example, during CSA configuration. To form a channel across CSA primitives, three sub-channels and two local network channels (which carry traffic to / from a single channel in the small backplane network) can be utilized. In one embodiment, one small backplane channel is utilized, for example, one small backplane and two locals = a total of 3 network hops.
[0185] The composability of channels across network layers can be extended to higher-level network layers at the primitive, die, and fabric granularities.
[0186] Figure 9 A processing element 900 according to an embodiment of the present disclosure is shown. In one embodiment, the operation configuration register 919 is loaded during configuration (e.g., mapping) and specifies the (one or more) specific operations that this processing (e.g., computing) element is to perform. The register 920 activity can be controlled by that operation (the output of mux 916, e.g., controlled by scheduler 914). The scheduler 914 can schedule one or more operations of the processing element 900, for example, when input data and control inputs arrive. The control input buffer 922 is connected to the local network 902 (e.g., and the local network 902 can include a data path network as shown and as Figure 7A shown and as Figure 7Bin the process control path network), and is loaded with values upon arrival (e.g., the network has (one or more) data bits and (one or more) valid bits). The control output buffer 932, the data output buffer 934, and / or the data output buffer 936 can receive the output of the processing element 900, e.g., as controlled by an operation (the output of mux 916). Whenever the ALU 918 runs (also controlled by the output of mux 916), the status register 938 can be loaded. The data in the control input buffer 922 and the control output buffer 932 can be single bits. Mux 921 (e.g., operand A) and mux 923 (e.g., operand B) can originate inputs.
[0187] For example, assume the operation of this processing (e.g., computing) element is (or includes) Figure 3B the pick as described in. The processing element 900 then selects data from the data input buffer 924 or the data input buffer 926, e.g., to go to the data output buffer 934 (e.g., default) or the data output buffer 936. Thus, the control bit in 922 can indicate 0 when selecting from the data input buffer 924 or 1 when selecting from the data input buffer 926.
[0188] For example, assume the operation of this processing (e.g., computing) element is (or includes) Figure 3B the switch as described in. The processing element 900 outputs data from the data input buffer 924 (e.g., default) or the data input buffer 926 to the data output buffer 934 or the data output buffer 936, e.g.. Thus, the control bit in 922 can indicate 0 when outputting to the data output buffer 934 or 1 when outputting to the data output buffer 936.
[0189] Multiple networks (e.g., interconnections) can be connected to the processing element, e.g., (input) networks 902, 904, 906 and (output) networks 908, 910, 912. The connections can be switches, e.g., as referred to in Figure 7A and Figure 7B described. In one embodiment, each network includes two sub-networks (or two channels on the network), e.g., one for Figure 7A the data path network in and one for Figure 7B the process control (e.g., backpressure) path network in. As an example, the local network 902 (e.g., established as a control interconnection) is shown as switched (e.g., connected) to the control input buffer 922. In this embodiment, the data path (e.g., Figure 7AThe network in (e.g., a network) can carry a control input value (e.g., one or more bits) (e.g., a control token), and a flow control path (e.g., a network) can carry a backpressure signal (e.g., a backpressure or non-backpressure token) from the control input buffer 922, e.g., to indicate to an upstream producer (e.g., a PE) that a new control input value has not been loaded into (e.g., sent to) the control input buffer 922 until the backpressure signal indicates that there is space in the control input buffer 922 for the new control input value (e.g., from the control output buffer of the upstream producer). In one embodiment, the new control input value may not enter the control input buffer 922 until (i) the upstream producer receives a "space available" backpressure signal from the "control input" buffer 922 and (ii) the new control input value is sent, e.g., from the upstream producer, and this may stall the processing element 900 until that occurs (and space in the (one or more) target output buffers is available).
[0190] The data input buffer 924 and the data input buffer 926 may operate similarly. For example, the local network 904 (e.g., established as a data (as opposed to control) interconnect) is shown as being switched (e.g., connected) to the data input buffer 924. In this embodiment, the data path (e.g., Figure 7A the network in) can carry a data input value (e.g., one or more bits) (e.g., a data flow token), and a flow control path (e.g., a network) can carry a backpressure signal (e.g., a backpressure or non-backpressure token) from the data input buffer 924, e.g., to indicate to an upstream producer (e.g., a PE) that a new data input value has not been loaded into (e.g., sent to) the data input buffer 924 until the backpressure signal indicates that there is space in the data input buffer 924 for the new data input value (e.g., from the data output buffer of the upstream producer). In one embodiment, the new data input value may not enter the data input buffer 924 until (i) the upstream producer receives a "space available" backpressure signal from the "data input" buffer 924 and (ii) the new data input value is sent, e.g., from the upstream producer, and this may stall the processing element 900 until that occurs (and space in the (one or more) target output buffers is available). The control output value and / or the data output value may be stalled in their respective output buffers (e.g., 932, 934, 936) until the backpressure signal indicates that there is available space in the input buffer for the downstream (one or more) processing elements.
[0191] The processing element 900 may be stalled from execution until its operands (e.g., the control input value and its corresponding one or more data input values) are received and / or until there is space in the (one or more) output buffers of the processing element 900 for the data produced by the execution of the operation on those operands.
[0192] 3.3 Memory Interface
[0193] A Request Address File (RAF) circuit (simplified form shown in Figure 10 is responsible for running memory operations and acts as a mediator between the CSA fabric and the memory hierarchy. Thus, the main microarchitecture task of the RAF can be to rationalize out-of-order memory. Through this capability, the RAF circuit can be provisioned with a completion buffer, such as a queue-like structure, which reorders memory responses and returns them to the fabric in request order. A second main functionality of the RAF circuit can be to provide support in the form of address translation and page walkers. Incoming virtual addresses can be translated into physical addresses using a channel-associated Translation Lookaside Buffer (TLB). To provide sufficient memory bandwidth, each CSA primitive can include multiple RAF circuits. Similar to the various PEs of the fabric, the RAF circuits can operate in a dataflow style by checking the availability of input arguments and output buffers (if needed) before selecting which memory operations to run. However, unlike some PEs, the RAF circuits are multiplexed among several concurrent memory operations. The multiplexed RAF circuits can be used to minimize the area overhead of their various sub-components, such as shared Accelerator Cache Interface (ACI) ports (described in more detail in Subsection 3.4), shared virtual memory (SVM) support hardware, small backplane network interfaces, and other hardware management facilities. However, there are some program characteristics that can also drive this choice. In one embodiment, (e.g., valid) dataflow graphs poll memory in a shared virtual memory system. Memory latency-bound programs (similar to graph traversals) can saturate memory bandwidth with many independent memory operations due to memory-related control flow. Although each RAF can be multiplexed, the CSA can include multiple (e.g., between 8 and 32) RAFs at the primitive granularity to ensure sufficient cache bandwidth. The RAF can communicate with the rest of the fabric via a local network and a small backplane network. In the case of multiplexed RAFs, each RAF can be provisioned with several ports into the local network. These ports can be used as minimum latency high-determinacy paths to memory for latency-sensitive or high-bandwidth memory operations. Additionally, the RAF can be provisioned with small backplane network endpoints, such as those that provide memory access to runtime services and remote user-level memory accessors.
[0194] Figure 10Shows a Request Address File (RAF) circuit 1000 according to an embodiment of the present disclosure. In one embodiment, at configuration time, memory load and store operations in the data flow graph are specified in register 1010. The arcs to those memory operations in the data flow graph can then be connected to input queues 1022, 1024, and 1026. Thus, the arcs from those memory operations leave completion buffers 1028, 1030, or 1032. Correlation tokens (which can be a single bit) enter queues 1018 and 1020. The correlation tokens leave queue 1016. The correlation token counter 1014 can be a compact representation of the queue and keeps track of the number of correlation tokens for any given input queue. If the correlation token counter 1014 saturates, no additional correlation tokens can be generated for new memory operations. Accordingly, the memory ordering circuit (e.g., Figure 11 the RAF in
[0195] can stall scheduling new memory operations until the correlation token counter 1014 becomes unsaturated.
[0196] As an example of a load, an address enters queue 1022 and the scheduler 1012 matches it with the load in 1010. The completion buffer slot for this load is assigned in the order in which the addresses arrive. Assuming that this particular load in the figure does not specify a correlation, the address and the completion buffer slot are issued to the memory system by the scheduler (e.g., via memory command 1042). When the result returns to mux 1040 (schematically shown), it is stored in the specified completion buffer slot (e.g., as it carries the destination slot all the way through the memory system). The completion buffer sends the result back to the local network (e.g., local networks 1002, 1004, 1006, or 1008) in the order in which the addresses arrive.
[0197] 3.4 Cache
[0198] The data flow graph can be capable of generating a large number (e.g., word granularity) of requests in parallel. Thus, certain embodiments of the CSA provide sufficient bandwidth to the cache subsystem to service the CSA. A re-banked cache microarchitecture such as Figure 11 shown can be utilized. Figure 11Circuit 1100 is shown in accordance with an embodiment of the present disclosure and has a plurality of Request Address File (RAF) circuits (e.g., RAF circuit (1)) coupled between a plurality of accelerator primitives (1108, 1110, 1112, 1114) and a plurality of cache groups (e.g., cache group 1102). In one embodiment, the number of RAFs and cache groups may be in a ratio of 1:1 or 1:2. The cache groups may contain full cache lines (e.g., as opposed to sharding based on words), where each line has exactly one home in the cache. The cache lines may be mapped to the cache groups via a pseudo-random function. The CSA may integrate the SVM model with other tiling architectures. Some embodiments include an Accelerator Cache Interface (Interconnect) (ACI) network that connects the RAFs to the cache groups. This network may carry addresses and data between the RAFs and the cache. The topology of the ACI may be a cascaded crossbar, for example as a trade-off between latency and implementation complexity.
[0199] 3.5 Floating-Point Support
[0200] Some HPC applications are characterized by a need for significant floating-point bandwidth. To meet this need, embodiments of the CSA may be provisioned with multiple (e.g., between 128 and 256 each) floating-point addition and multiplication PEs, for example depending on the primitive configuration. The CSA may provide several other extended precision modes, for example to simplify math library implementation. The CSA floating-point PEs may support single and double precision, but lower precision PEs may support machine learning workloads. The CSA may provide an order of magnitude greater floating-point performance than a processor core. In one embodiment, in addition to increasing the floating-point bandwidth to power all the floating-point units, the energy consumed in floating-point operations is also reduced. For example, to reduce energy, the CSA may selectively gate the low-order bits of the floating-point multiplier array. In examining the behavior of floating-point arithmetic, the low-order bits of the multiplier array often do not affect the final rounded product. Figure 12 A floating-point multiplier 1200 is shown in accordance with an embodiment of the present disclosure and is divided into three regions (a result region, three potential carry regions (1202, 1204, 1206), and a gating region). In some embodiments, the carry regions may affect the result region, and the gating region may not affect the result region. Considering a g-bit gating region, the maximum carry can be:
[0201]
[0202] Given this maximum carry, if the result of the carry region is less than 2 c-g (where the carry region is c bits wide), the strobe region can be ignored because it does not affect the result region. Increasing g means that the strobe region will more likely be needed, while increasing c means that under random assumptions, the strobe region will not be used and can be disabled to avoid power consumption. In an embodiment of the CSA floating-point multiplication PE, a two-stage pipelined approach is utilized, where the carry region is determined first and then the strobe region is determined if it is found to affect the result. If more information related to the context of the multiplication is known, the CSA tunes the size of the strobe region more aggressively. In FMA, the multiplication result can be added to an accumulator, which is often much larger than either of the multiplicands. In this case, the adder exponent can be observed prior to multiplication, and the CSDA can adjust the strobe region accordingly. One embodiment of the CSA includes a scheme where a context value (which defines the minimum result of the computation) is provided to the associated multiplier to select the minimum energy strobe configuration.
[0203] 3.6 Runtime Services
[0204] In some embodiments, the CSA includes a heterogeneous and distributed structure, and thus the runtime service implementation adapts to several types of PEs in a parallel and distributed manner. Although the runtime services in the CSA can be critical, they can be infrequent relative to user-level computing. Thus, some implementations focus on overlaying services on hardware resources. To meet these goals, the CSA runtime services can be broadcast as a hierarchical structure, for example where each layer corresponds to a CSA network. At the primitive level, a single external-facing controller can accept or send commands to the associated cores with CSA primitives. The primitive-level controller can be used to coordinate regional controllers in the RAF, for example using the ACI network. The regional controllers can in turn coordinate local controllers at the termination of some small backplane networks (e.g., network data flow endpoint circuits). At the lowest level, service-specific microprotocols can run over the local network, for example during special modes controlled by the small backplane controller. The microprotocols can allow each PE (e.g., a PE class according to type) to interact with the runtime services according to its own needs. Thus, parallelism is implicit in this hierarchical organization, and operations at the lowest level can occur simultaneously. This parallelism can achieve the configuration of the CSA in the range of hundreds of nanoseconds to several microseconds, for example according to the configuration size and its position in the memory hierarchy. Thus, embodiments of the CSA balance the nature of the data flow graph to improve the implementation of each runtime service. A key observation is that the runtime services may only need to maintain a legal logical view of the data flow graph, e.g., a state that can be produced by some ordering of execution through the data flow operators. The services generally may not need to guarantee a temporal view of the data flow graph (e.g., the state of the data flow graph in the CSA at a particular point in time). This can allow the CSA to perform most runtime services in a distributed, pipelined, and parallel manner, e.g., as long as the services are organized to maintain a logical view of the data flow graph. The local configuration microprotocol can be a packet-based protocol overlaid on the local network. The configuration targets can be organized as a configuration chain, e.g., which is fixed in the microarchitecture. Structures (e.g., PEs) can be configured one at a time, for example using a single additional register per target, to achieve distributed coordination. To start the configuration, the controller can drive an out-of-band signal, which puts all the structure targets in the neighborhood into an unconfigured, suspended state and swings the multiplexers in the local network into a predefined conformation. When the structure (e.g., PE) targets are configured, i.e., they have fully received the configuration packet, they can set the configuration microprotocol register, thereby notifying the direct successor targets (e.g., PEs) about the configuration with which it can proceed using the next packet. There is no limit on the size of the configuration packet, and the packet can have a dynamically variable length. For example, a PE that configures a constant operand can have a configuration packet that is extended to include a constant field (e.g., Figure 3B - 3C the X and Y in Figure 13Shows the in - progress configuration of an accelerator 1300 having multiple processing elements (e.g., PE 1302, 1304, 1306, 1308) according to an embodiment of the present disclosure. Once configured, the PEs can operate subject to data - flow constraints. However, channels involving unconfigured PEs can be disabled by the micro - architecture, e.g., to prevent any undefined operations from occurring. These properties allow embodiments of the CSA to be initialized and run in a distributed fashion without any central control. From the unconfigured state, the configuration can occur completely in parallel, for example, in as little as 200 nanoseconds. However, due to the distributed initialization of embodiments of the CSA, the PEs can become active long before the entire structure is configured, e.g., sending requests to memory. The extraction can be performed in a manner almost identical to the configuration. The local network can be configured to extract data from one target at a time, and status bits are used to achieve distributed coordination. The CSA can organize the extraction to be lossless, i.e., when the extraction is complete, each extractable target has returned to its starting state. In this implementation, all the states in the target can cycle to an exit register (which is tied to the local network) in a scan - like fashion. However, in - place extraction can be achieved by introducing new paths at the register transfer level (RTL) or using existing circuitry to provide the same functionality at lower overhead. Similar to the configuration, the hierarchical extraction is implemented in parallel.
[0205] Figure 14 Shows a snapshot 1400 of an in - progress pipelined extraction according to an embodiment of the present disclosure. In some use cases of extraction (e.g., checkpoints), latency may not be an issue as long as the throughput of the structure is maintained. In these cases, the extraction can be organized in a pipelined fashion. Figure 14 This arrangement shown allows most of the structure to continue operating while disabling a narrow area for extraction. Configuration and extraction can be coordinated and composed to achieve pipelined context switching. Exceptions can be qualitatively different from configuration and extraction because they do not occur at the specified time; rather, they occur at any point during runtime anywhere in the structure. Thus, in one embodiment, the exception micro - protocol may not overlay the local network, which is occupied by the user program at runtime, and utilizes its own network. However, exceptions are inherently rare and are not sensitive to latency and bandwidth. Therefore, some embodiments of the CSA utilize a packet - switched network to forward exceptions to a local backplane termination, e.g., where they are forwarded up along a service hierarchy (e.g., in Figure 46 ). Packets in the local exception network can be extremely small. In many cases, a PE identifier (ID) of only two to eight bits is sufficient as a complete packet, e.g., because the CSA can create a unique exception identifier as the packet traverses the heterogeneous service hierarchy. This scheme can be desirable because it also reduces the area overhead of generating exceptions at each PE.
[0206] 4. Compilation
[0207] The ability to compile programs written in high-level languages onto a CSA is essential for industrial adoption. This subsection gives a high-level overview of the compilation strategy for embodiments of the CSA. First is a proposal for a CSA software framework that shows the expected properties of an ideal production-quality technology chain. Subsequently, the prototype compiler framework is discussed. Then, "control-to-dataflow transformation" is discussed, for example, to convert ordinary sequential control-flow code into CSA dataflow assembly code.
[0208] 4.1 Example Production Framework
[0209] Figure 15 Shows a compilation toolchain 1500 for an accelerator according to an embodiment of the present disclosure. This toolchain compiles high-level languages (such as C, C++, and Fortran) into a combination of host code (LLVM) intermediate representations (IRs) for specific regions to be accelerated. The CSA-specific part of this compilation toolchain takes the LLVM IR as its input, optimizes and compiles this IR into CSA assembly, for example, by adding appropriate buffering on latency-insensitive channels to obtain performance. The CSA assembly is then placed and routed onto the hardware fabric, and the PEs and network are configured for execution. In one embodiment, the toolchain supports just-in-time (JIT) CSA-specific compilation, combined with potential runtime feedback from actual execution. One of the key design features of this framework is the compilation of the (LLVM) IR of the CSA, rather than using a high-level language as input. While programs written in a high-level programming language specifically designed for the CSA may achieve maximum performance and / or energy efficiency, the adoption of new high-level languages or programming frameworks may be slow and is actually limited by the difficulty of converting existing codebases. Using the (LLVM) IR as input enables a wide range of existing programs to potentially run on the CSA, for example, without having to create a new language or significantly modify the front end of a new language that is desired to run on the CSA.
[0210] 4.2 Prototype Compiler
[0211] Figure 16Shows compiler 1600 of an accelerator according to an embodiment of the present disclosure. Compiler 1600 initially focused on ahead-of-time compilation of C and C++ via a (e.g., Clang) front-end. To compile (LLVM) IR, the compiler implements a CSA backend target within LLVM with three main levels. First, the CSA backend lowers the LLVM IR into target-specific machine instructions for sequential units, which implement most of the CSA operations combined with a traditional RISC-like control flow architecture (e.g., with branches and program counters). The sequential units in the toolchain can be useful assistants to the compiler and application developers, as it implements an incremental transformation of the program from control flow (CF) to data flow (DF), e.g., converting a segment of the code from control flow to data flow each time and verifying program correctness. The sequential units can also provide a model for manipulating code in a non-adaptive spatial array. Subsequently, the compiler converts these control flow instructions into data flow operators (e.g., code) of the CSA. This stage is described later in Subsection 4.3. The data flow operators (e.g., code) can have their sequences optimized, examples of which are described later in Subsection 4.4. Then, the CSA backend can run its own optimization passes on the data flow instructions. Finally, the compiler can dump the instructions in CSA assembly format. This assembly format is taken as input by a post-stage tool that places and routes the data flow instructions onto the actual CSA hardware.
[0212] 4.3 Control-to-Data Flow Transformation
[0213] A key part of the compiler can be implemented in a control-to-data flow transformation pass (or simply the data flow transformation pass). This pass takes a function represented in control flow form (e.g., a control flow graph (CFG) with sequential machine instructions operating on virtual registers) and converts it into a data flow function, which conceptually is a graph of data flow operations (instructions) connected by latency-insensitive channels (LICs). This subsection gives a high-level description of this pass, describing how it conceptually handles memory operations, branches, and loops in some embodiments.
[0214] Straight-Line Code
[0215] Figure 17A Shows sequential assembly code 1702 according to an embodiment of the present disclosure. Figure 17B Shows, according to an embodiment of the present disclosure, Figure 17A the data flow assembly code 1704 of the sequential assembly code 1702. Figure 17C Shows, according to an embodiment of the present disclosure, Figure 17B the data flow graph 1706 of the data flow assembly code 1704.
[0216] First, consider the simple case of converting straight-line sequential code into data flow. The data flow transformation pass can convert the basic blocks of sequential code (such as the code shown in Figure 17A ) into the CSA assembly code shown in Figure 17B . Conceptually, the CSA assembly representation in Figure 17B represents the data flow graph shown in Figure 17C . In this example, each sequential instruction is translated into machine CSA assembly. The.lic statement for data (for example) declares a latency-insensitive channel, which corresponds to a virtual register (such as Rdata) in the sequential code. In practice, the input to the data flow transformation pass can be in numbered virtual registers. However, for clarity, this subsection uses descriptive register names. Note that in this embodiment, load and store operations are supported in the CSA architecture, allowing more programs to run compared to architectures that only support pure data flow. Since the sequential code input to the compiler is in SSA (single static assignment) form, for a simple basic block, the control-data flow pass can convert each virtual register definition into the generation of a single value on a latency-insensitive channel. The SSA form allows multiple uses of a single definition of a virtual register (such as in Rdata2). To support this model, the CSA assembly code supports multiple uses of the same LIC (such as data2), where the simulator implicitly creates the necessary copies of the LIC. A key difference between sequential code and data flow code lies in the handling of memory operations. Figure 17A The code in is conceptually serial, which means that in the case where the addresses addr and addr3 overlap, the load32 (ld32) of addr3 should appear to occur after the st32 of addr.
[0217] Branch
[0218] To convert a program with multiple basic blocks and conditions into data flow, the compiler generates special data flow operators to replace branches. More specifically, the compiler uses a switch operator at the end of the basic blocks of the original CFG to direct the outgoing data, and a pick operator at the start of the basic blocks to select values from the appropriate incoming channels. As a specific example, consider the code in Figure 18A - 18C and the corresponding data flow graph, which conditionally calculates the value of y based on several inputs ai, x, and n. After computing the branch condition test, the data flow code uses a switch operator (such as see Figure 3B - 3C ) to direct the value in channel x to channel xF when the test is 0 or to channel xT when the test is 1. Similarly, the pick operator (such as see Figure 3B - 3C) It is used to send channel yF to y when the test is 0, or to send channel yT to y when the test is 1. In this example, as a result, even though the value of a is only used in the "true" branch of the condition, the CSA includes a switch operator that directs it to channel aT when the test is 1 and consumes (eats) the value when the test is 0. This latter case is expressed by setting the "false" output of the switch to %ign. Simply connecting channel a directly to the "true" path may not be correct because in the case where the execution actually takes the "false" path, this value of "a" will remain in the graph, resulting in an incorrect value of a for the next execution of the function. This example emphasizes the property of control equivalence, a key property in embodiments of correct data flow transformation.
[0219] Control equivalence: Consider a single-entry single-exit control flow graph G with two basic blocks A and B. A and B are control equivalent if all complete control flow paths through G visit A and B the same number of times.
[0220] LIC replacement: In control flow graph G, assume that an operation in basic block A defines virtual register x, and an operation in basic block B uses x. A correct control-to-data flow transformation can then use a latency-insensitive channel to replace x only if A and B are control equivalent. The control equivalence relation partitions the basic blocks of the CFG into strongly control-related regions. Figure 18A Shows C source code 1802 according to an embodiment of the present disclosure. Figure 18B Shows according to an embodiment of the present disclosure, Figure 18A the data flow assembly code 1804 of C source code 1802. Figure 18C Shows according to an embodiment of the present disclosure, Figure 18B the data flow graph 1806 of data flow assembly code 1804. In Figure 18A - 18C the example, the basic blocks before and after the condition are mutually control equivalent, but the basic blocks in the "true" and "false" paths are each in their own control-related regions. A correct algorithm for converting the CFG into a data flow is to make the compiler: (1) insert switches to compensate for mismatches in the execution frequencies of any values flowing between basic blocks that are not control equivalent; and (2) insert picks at the start of basic blocks to make a correct selection from any incoming values to the basic block. Generating the appropriate control signals for these picks and switches can be a key part of the data flow transformation.
[0221] Loop
[0222] Another important class of CFGs in data flow transformation is the CFG of single-entry single-exit loops, which is the common form of loops generated in (LLVM) IR. These loops can be almost aperiodic, except for a single back edge from the end of the loop back to the loop header block. The data flow transformation pass can use the same high-level strategies for loops as for branches. For example, it inserts switches at the end of the loop to direct values out of the loop (either exiting the loop or around the back edge to the start of the loop), and inserts picks at the start of the loop to select between the initial value entering the loop and the value experiencing the back edge. Figure 19A Shows C source code 1902 according to an embodiment of the present disclosure. Figure 19B Shows according to an embodiment of the present disclosure, Figure 19A the data flow assembly code 1904 of C source code 1902. Figure 19C Shows according to an embodiment of the present disclosure, Figure 19B the data flow graph 1906 of data flow assembly code 1904. Figure 19A - 19C Shows C and CSA assembly code for an example do-while loop (which sums the value of loop induction variable i) and the corresponding data flow graph. For each variable that conceptually loops around the loop (i and sum), this graph has a corresponding pick / switch pair that controls the flow of these values. Note that this example also uses pick / switch pairs to loop the value of n around the loop, even though n is a loop invariant. This repeated implementation of n converts the virtual register of n to a LIC because it matches the execution frequency between the conceptual definition of n outside the loop and one or more uses of n inside the loop. Generally, for correct data flow transformation, when a register is converted to a LIC, the register that migrates into the loop repeats for each iteration inside the loop body. Similarly, registers updated inside the loop and exiting the loop can be consumed, for example where a single final value is sent from the loop. Loops introduce wrinkles into the data flow transformation process, i.e., offsetting the control of the pick at the top of the loop and the switch at the bottom of the loop. For example, if Figure 18A the loop in
[0223] Sequence optimization
[0224] AlthoughFigure 19A the transformation of the code in Figure 19C the configuration to multiple processing elements to run
[0225] 1. Sequence: An embodiment of the sequence operation takes as input a triple of a base value, a limit, and a stride value, and uses those inputs to produce a value stream (equivalent to) a for loop. For example, if the base value is 10, the limit is 15, and the stride is 2, the seqlts32 operation produces a stream of three output values (i.e., 10; 12; 14;). It also produces a stream of 1; 1; 10 as control signals, which can be used, for example, to control other types of operations in the sequence family. The field in the operand of 32 can operate on 32 bits of data, for example, immediately. In another embodiment, the field is another numerical value, such as 64 instead of 32, and the operand can operate on 64 bits of data, for example, immediately.
[0226] 2. Stride: An embodiment of the stride operation takes as input a base value, a stride, and an input control stream of a control signal (ctl), and generates a corresponding linear sequence to match ctl. For example, for the stride32 operation, if the base value is 10, the stride is 1, and ctl is 1; 1; 1; 0, the output is 10; 11; 12. An embodiment of the stride operation can be considered a related sequence instruction that relies on the control flow of the sequence operation to determine the timing of the step, rather than making a comparison with the limit.
[0227] 3. Reduction: An embodiment of the reduction operation takes as input an initial value (init), a value stream in, and a stream of control signals (ctl), and outputs the sum of the initial value and the value stream. For example, redadd32 with init being 10, in being 3; 4; 2, and ctl being 1; 1; 1; 0 produces an output of 19.
[0228] 4. Repeat: An embodiment of the repeat operation repeats the input value according to the input control stream. For example, repeat32 with an input value of 42 and a control stream of 1; 1; 1; 0 will output three instances of 42.
[0229] 5. Onend: An embodiment of the onend operation conceptually matches the input values on the input stream in with the signals on the control signal (ctl) stream and returns a signal when all the matches are complete. For example, for a ctl input of 1; 1; 1; 0, the onend operation will match any three inputs on the value stream in and output an end signal when it reaches the 0 in ctl. In some embodiments, a sequence transformation pass in a compiler that runs after a data flow transformation to search for sequence candidates (e.g., pick and switch data flow operators (e.g., pair), which correspond to values that loop around a loop) will convert candidates for loop induction variables into sequence instructions and convert any remaining compatible candidates into relevant stride, repeat, or simplify operations.
[0230] Figure 20A Shows C source code 2002 according to an embodiment of the present disclosure. Figure 20B Shows, according to an embodiment of the present disclosure, Figure 20A the data flow assembly code 2004 of the C source code 2002. Figure 20C Shows, according to an embodiment of the present disclosure, Figure 20B the data flow graph 2006 of the data flow assembly code 2004. Figure 20A - 20C Shows an example of sequence optimization applied to a loop for computing a dot product. The seqlts64 operation can produce an output control flow of n 1s followed by a 0. Note that this example does not actually use the value of the induction variable i output by the sequence. Instead, this code uses the stride64 operation to stride across the addresses of x and y. Figure 20A The seqlts64 operation shown also produces two other control signal flow outputs, which are not used in this example (e.g., denoted by %ign). The inputs to the shown assembly code are n, x, and y, and the output is final_sum. The data flow graph 2006 can be mapped onto an array of processing elements (e.g., and the network (e.g., interconnect) between them), such that each node of the data flow graph 2006 is represented as a data flow operator in the array of processing elements (e.g., including a sequencer operator representing the sequencer node 2010).
[0231] Figure 21 Shows the implementation of an integer arithmetic / logic data flow operator 2101 on a processing element 2100 according to an embodiment of the present disclosure. In one embodiment, the integer arithmetic / logic data flow operator 2101 is an integer processing element, such as Figure 9 the integer processing element 900 in or other PEs. The operation selector can be a scheduler 2114, for example Figure 9The scheduler 914 or other PEs therein. In one embodiment, the operation configuration register 2109 is loaded during configuration (e.g., mapping) and specifies the particular operation(s) to be performed (e.g., performed by the ALU 2118) by this processing (e.g., computing) element. The scheduler 2114 (e.g., operation selector) may schedule one or more operations of the processing element 2100, for example, when the input data and control inputs arrive. Inputs and outputs (e.g., via one or more buffers) may be sent via a network (e.g., any of the networks described herein). The control input buffer 2122 may be connected to a local network (e.g., and the local network may include a data path network as shown in Figure 7A and a flow control path network as shown in Figure 7B ), and is loaded with values when they arrive (e.g., the network has one or more data bits and one or more valid bits). The control input buffer 2122 may be coupled to a zero generator 2125, for example, to add leading or trailing zeros to the value from the control input buffer 2122 to form the expected width (e.g., 64 bits) of the data item. The control output buffer 2132, data output buffer 2134, and / or data output buffer 2136 may receive the output of the processing element 2100, for example, as controlled by an operation (the output of the scheduler 2114). The data in the control input buffer 2122 and the control output buffer 2132 may be single bits. The mux 2121 (e.g., operand A) and mux 2123 (e.g., operand B) may source inputs.
[0232] For example, assume that the operation of this processing (e.g., computing) element is (or includes) Figure 3B the pick as described in
[0233] For example, assume that the operation of this processing (e.g., computing) element is (or includes) Figure 3B the switch as described in Figure 9 Figure 3B . The processing element 2100 outputs data, for example, from the data input buffer 2124 (e.g., default) or the data input buffer 2126 to the data output buffer 2134 or the data output buffer 2136. Thus, the control bit in 2122 may indicate 0 when outputting to the data output buffer 2134 or indicate 1 when outputting to the data output buffer 2136. Multiple networks (e.g., interconnections) may be connected to the processing element, such as an (input) network * (e.g., Figure 9networks 902, 904, 906 and (output) networks 908, 910, 912). The connection can be a switch, for example as referred to in Figure 7A and Figure 7B described. In one embodiment, each network includes two sub-networks (or two channels on the network), for example one for Figure 7A the data path network in Figure 7B and one for Figure 7A the flow control (e.g., backpressure) path network in
[0234] As an example, the local network can be switched (e.g., connected) to the control input buffer 2122 (e.g., as established for control interconnection). In this embodiment, the data path (e.g., Figure 7A the network in
[0234] can carry control input values (e.g., one or more bits) (e.g., control tokens), and the flow control path (e.g., the network) can carry a backpressure signal (e.g., backpressure or no-backpressure token) from the control input buffer 2122, for example to indicate to an upstream producer (e.g., a PE) that a new control input value has not been loaded into (e.g., sent to) the control input buffer 2122 until the backpressure signal indicates that there is space in the control input buffer 2122 for the new control input value (e.g., from the control output buffer of the upstream producer). In one embodiment, a new control input value may not enter the control input buffer 2122 until (i) the upstream producer receives a "space available" backpressure signal from the "control input" buffer 2122 and (ii) the new control input value is sent, e.g., from the upstream producer, and this can stall the processing element 2100 until that occurs (and space in the (one or more) target output buffers is available).
[0234] The data input buffer 2124 and the data input buffer 2126 can perform similarly. For example, the local network (e.g., as established for data (as opposed to control) interconnection) can be switched (e.g., connected) to the data input buffer 2124. In this embodiment, the data path (e.g., Figure 7AThe network) can carry data input values (e.g., one or more bits) (e.g., data stream tokens), and the flow control path (e.g., the network) can carry a backpressure signal (e.g., a backpressure or no-backpressure token) from the data input buffer 2124, e.g., to indicate to an upstream producer (e.g., a PE) that new data input values are not loaded into (e.g., sent to) the data input buffer 2124 until the backpressure signal indicates that there is space in the data input buffer 2124 for the new data input values (e.g., the data output buffer of the upstream producer). In one embodiment, the new data input values may not enter the data input buffer 2124 until (i) the upstream producer receives a "space available" backpressure signal from the "data input" buffer 2124 and (ii) the new data input values are sent, e.g., from the upstream producer, and this may stall the processing element 2100 until that occurs (and space in the (one or more) target output buffers is available). The control output values and / or data output values may stall in their respective output buffers (e.g., 2132, 2134, 2136) until the backpressure signal indicates that there is available space in the input buffer for the downstream (one or more) processing elements.
[0235] The processing element 2100 may stall from execution until its operands (e.g., control input values and their corresponding one or more data input values) are received and / or until there is space in the (one or more) output buffers of the processing element 2100 for the data generated by the execution of the operations on those operands. Some couplings (e.g., wires) are not shown in detail so as not to affect the understanding of some descriptions.
[0236] Although a heterogeneous CSA computing structure (e.g., different types of PEs) may be utilized (e.g., to optimize area / energy efficiency), there are (e.g., dark) circuits that exist but are not currently used (e.g., if the processing elements become too specialized) that can be harmful to the manufacturing cost and area / energy efficiency goals. In one embodiment, the sequencer data flow operator effectively supports sequence generation using two integer PEs with a (e.g., small) set of dedicated data / control wires connecting them, (e.g., a small amount of) additional control logic circuits, and / or storage devices. In one embodiment, each processing element forming the sequencer data flow operator operates in a first mode (e.g., as an independent (e.g., integer) PE) and a second mode (e.g., as a sequencer), e.g., operating in the first mode when it is not operating in the second mode.
[0237] The PEs may communicate using dedicated virtual circuits (formed by switching communication networks through static configuration circuits). Embodiments of these virtual circuits may be flow-controlled and fully backpressure, e.g., such that the PEs will stall when the source has no data or the destination is full.
[0238] Sequencer Data Flow Operator
[0239] Figure 22 Illustrated is the implementation of a sequencer data flow operator 2201 on processing elements (2200A, 2200B) in accordance with an embodiment of the present disclosure. In one embodiment, processing element 2200A performs arithmetic operations (such as addition or subtraction), and processing element 2220B performs comparison operations (e.g., to determine whether additional arithmetic operations should be triggered). This can be used in loop processing, where the number of iterations is determined by repeatedly incrementing and / or decrementing a base data value by a span data value until a specific threshold is reached or exceeded. The left portion (e.g., left half) (e.g., processing element 2200A) of the sequencer data flow operator 2201 has (e.g., a single) (e.g., 64-bit) register bank 2244, which is used to repeatedly accumulate span data (e.g., span data tokens) into base data (e.g., base data tokens). This can be referred to as a sequencer span PE (seqstr). The right portion (e.g., right half) (e.g., processing element 2200B) of the sequencer data flow operator 2201 has an ALU 2218B, which is used to perform comparison operations. This can be referred to as a sequencer comparison PE (seqcmp). The comparison result can be fed back (e.g., on data path 2241) from the sequencer comparison PE (seqcmp) (e.g., processing element 2200B) to the sequencer span PE (seqstr) (e.g., processing element 2200A), so that the two PEs jointly determine when the sequence generation ends (e.g., the sequencer comparison PE (seqcmp) (e.g., processing element 2200B) updates the sequencer span PE (seqstr) (e.g., processing element 2200A) when the end (e.g., limit or threshold) is reached).
[0240] In one embodiment, the data passed into the sequencer data flow operator 2201 includes a new span length, such that processing element 2200A performs the addition (or subtraction) of the span length to the total number of spans (e.g., iterations) thus far, and processing element 2200B performs the total number of spans (e.g., iterations) thus far and the total number of spans (e.g., iterations) to be performed (e.g., Figure 3A - 3CComparison of "n" or "A" in []. In one embodiment, the sequencer data flow operator 2201 (e.g., processing element 2200A) includes a sequencer span controller 2242, e.g., to track the arrival of the base value data token and the span value data token. When the base value data token has arrived, the sequencer span controller 2242 can immediately send a signal to the sequencer comparison PE (seqcmp) (e.g., processing element 2200B), such that the comparison operation can then begin. In addition to monitoring the base value data token arrival signal from the sequencer span controller 2242, the sequencer comparison controller 2240 can also monitor the arrival of the limit value data token to determine the time at which a valid comparison result can be generated. The sequencer span controller 2242 can then determine whether additional arithmetic operations (e.g., increment or decrement) should be triggered based on the actual value of the valid comparison result (e.g., a value of one indicates that additional arithmetic operations should be triggered, and a value of zero indicates that this particular sequence generation is complete). Additionally, the sequencer span controller 2242 can determine the (one or more) input operands for the additional arithmetic operations. For the first iteration, the base value data token can be the input operand. For all subsequent iterations, the register bank 2244 output can be the input operand. In one embodiment, the second input operand for the arithmetic operation can always be the span data token. The combination of the sequencer span controller 2242 and the sequencer comparison controller 2240 can generate a total of three control flows (or assertion flows) used in the loop processing. One is called the first flow. The start data token of the first flow can always be one, e.g., indicating that the 1st iteration of the loop can begin. All subsequent data tokens up to the Nth iteration of the loop can have a value of zero. As Figure 3C shown, the pick operator 304A can be controlled by the "first" flow generated by the sequencer data flow operator 310A. In the first iteration of the loop, Figure 3A the initial value of "res" in [](e.g., Figure 3C[[ X in []) will be the output of the pick operator 304A, which is fed to the multiplier operator 308A. (e.g., referring to , it can be seen that the inverse of the first flow is applied to the pick operator 404. In the first loop iteration, the value of one is passed to the multiplier node 408 at step 3. In the second loop iteration, the looped-back value of two is passed to the multiplier node 408 at step 6.)
[0241] The next control flow (or assertion flow) that the sequencer data flow operator can generate is called the last flow. For a loop with N iterations, the control data token associated with the Nth iteration has a value of one. The control data tokens associated with all previous iterations can have a value of zero. As shown, the switch operator 306A can be controlled by the last flow generated by the sequencer data flow operator 310A (e.g., referring to , the inverse of the last stream is applied to the switch node 406. In the first loop iteration, the output value of two loops back to the pick node 404 at step 5, which will become the data input for the second loop iteration. In the second and final loop iteration, the final output value of four is sent downstream at step 8 for further processing.)
[0242] The final control flow (or assert flow) that the sequencer data flow operator can generate is called the assert flow. For each iteration of the loop, a data token value of one can be generated. When the loop is completed, a data token value of zero can be generated. To accumulate the incremental values for each iteration of the loop and store the final accumulated value when the loop exits, a processing element can use a control flow similar to this. In one embodiment, it is incorrect to use the last stream for this use case when it is not desired to omit the final accumulation during the final iteration of the loop.
[0243] The sequencer comparison controller 2240 can cause the processing element 2200B to perform a comparison of the total number of spans (e.g., iterations) up to this point (e.g., stored in the register bank(s) 2244) with the total number of spans (e.g., iterations) to be performed (e.g., stored in the register bank(s) 2244) (e.g., "n" or "A" in). The sequencer data flow operator 2201 (e.g., processing element 2200A) can include a sequencer span controller 2242. The sequencer span controller 2242 can cause the processing element 2200A to perform an addition (or subtraction) of the span length (e.g., the increment per iteration) (e.g., in one embodiment, the span length is one unit (e.g., the numerical value one)) with the total number of spans (e.g., iterations) up to this point (e.g., "res" in). For each iteration of the operation (e.g., for loop), the sequencer data flow operator 2201 can output appropriate control signals (e.g., to the pick operator (e.g., implemented on its own PE) and / or the switch operator (e.g., implemented on its own PE)) (e.g., the control signals shown inside the circle in) (steps 1 - 8)) to cause each iteration of the total number of iterations to be executed. In one embodiment, the control signals are on a control data channel (e.g., narrower than the payload data) (e.g., using carried by the control input buffer 922 and / or the control output buffer 932 in). Another possible implementation of the sequencer data flow operator is to use a single integer PE that includes two ALUs (e.g., one for accumulation and another for comparison). The two ALUs can be pipelined (e.g., with additional pipeline hazard control circuitry) to maximize the circuit frequency, and / or the two ALUs can be placed in series within a single clock cycle, e.g., to simplify the controller. In one embodiment, the data passed into the sequencer data flow operator 2201 includes a new stride length, e.g., where the processing element 2200A performs the addition (or subtraction) of the stride length and the total number of spans (e.g., iterations) so far, and the processing element 2200B performs the comparison of the total number of spans (e.g., iterations) so far with the total number of spans (e.g., iterations) to be performed (e.g., "n" or "A" in).
[0244] As a supplement or alternative to forming the sequencer data flow operator, each of the processing elements 2200A and 2200B can operate as an integer PE.
[0245] In one embodiment, the operation configuration register 2109A is loaded during configuration (e.g., mapping) and specifies the particular operation(s) to be performed by this processing (e.g., computing) element. The scheduler 2114A (e.g., operation selector) can schedule one or more operations of the processing element 2100A, e.g., when input data and control inputs arrive. Inputs and outputs (e.g., via one or more buffers) can be sent via a network (e.g., any network described herein). The control input buffer 2122A can be connected to a local network (e.g., and the local network can include a data path network as shown and a in the flow control path network), and is loaded with values upon arrival (e.g., the network has (one or more) data bits and (one or more) valid bits). The control input buffer 2222A can be coupled to a zero generator 2225A, e.g., to add leading or trailing zeros to the value from the control input buffer 2222A to form the expected width of the data item (e.g., 64 bits). The control output buffer 2232A, the data output buffer 2234A, and / or the data output buffer 2236A can receive the output of the processing element 2200A, e.g., as controlled by an operation (the output of the scheduler 2214A). In one embodiment, the operation configuration register 2209A is loaded during configuration (e.g., mapping) and specifies the (one or more) specific operations (e.g., and if adjacent PEs 2200B are to be used for joint operations, such as sequence operations) to be performed by this processing (e.g., computing) element. The data in the control input buffer 2222A and the control output buffer 2232A can be single bits. The mux 2221A (e.g., operand A) and the mux 2223A (e.g., operand B) can source inputs.
[0246] For example, assume the operation of this processing (e.g., computing) element is (or includes) the pick as described in. The processing element 2200A then selects data from the data input buffer 2224A or the data input buffer 2226A, e.g., to go to the data output buffer 2234A (e.g., default) or the data output buffer 2236A. Thus, the control bit in 2222A can indicate 0 when selecting from the data input buffer 2224A or 1 when selecting from the data input buffer 2226A.
[0247] For example, assume the operation of this processing (e.g., computing) element is (or includes) the switch as described in. The processing element 2200A outputs data from the data input buffer 2224A (e.g., default) or the data input buffer 2226A to the data output buffer 2234A or the data output buffer 2236A, e.g.. Thus, the control bit in 2222A can indicate 0 when outputting to the data output buffer 2234A or 1 when outputting to the data output buffer 2236A.
[0248] Multiple networks (e.g., interconnections) can be connected to the processing element, e.g., (input) networks (e.g., networks 902, 904, 906 in and (output) networks 908, 910, 912). The connections can be switches, e.g., as referred to in and described. In one embodiment, each network includes two sub-networks (or two channels on the network), e.g., one for the data path network therein and a process control (e.g., backpressure) path network for therein. As an example, the local network may be switched (e.g., connected) to the control input buffer 2222A (e.g., as established for control interconnect). In this embodiment, the data path (e.g., the network therein) may carry control input values (e.g., one or more bits) (e.g., control tokens), and the process control path (e.g., network) may carry a backpressure signal (e.g., backpressure or no-backpressure token) from the control input buffer 2222A, e.g., to indicate to an upstream producer (e.g., PE) that a new control input value has not been loaded into (e.g., sent to) the control input buffer 2222A until the backpressure signal indicates that there is space in the control input buffer 2222A for the new control input value (e.g., from the control output buffer of the upstream producer). In one embodiment, the new control input value may not enter the control input buffer 2222A until (i) the upstream producer receives a "space available" backpressure signal from the "control input" buffer 2222A and (ii) the new control input value is sent, e.g., from the upstream producer, and this may stall the processing element 2200A until that occurs (and space in the (one or more) target output buffers is available).
[0249] The data input buffer 2224A and the data input buffer 2226A may perform similarly. For example, the local network (e.g., as established for data (as opposed to control) interconnect) may be switched (e.g., connected) to the data input buffer 2224A. In this embodiment, the data path (e.g., The network (e.g., in the network) can carry data input values (e.g., one or more bits) (e.g., data stream tokens), and the flow control path (e.g., the network) can carry a backpressure signal (e.g., backpressure or no backpressure token) from the data input buffer 2224A, e.g., to indicate to an upstream producer (e.g., a PE) that new data input values are not loaded into (e.g., sent to) the data input buffer 2224A until the backpressure signal indicates that there is space in the data input buffer 2224A for the new data input values (e.g., the data output buffer from the upstream producer). In one embodiment, the new data input values may not enter the data input buffer 2224A until (i) the upstream producer receives a "space available" backpressure signal from the "data input" buffer 2224A and (ii) the new data input values are sent, e.g., from the upstream producer, and this may stall the processing element 2200A until that occurs (and space in the (one or more) target output buffers is available). The control output values and / or data output values may be stalled in their respective output buffers (e.g., 2232A, 2234A, 2236A) until the backpressure signal indicates that there is available space in the input buffer for the downstream (one or more) processing elements.
[0250] The processing element 2200A can be stalled from execution until its operands (e.g., control input values and their corresponding one or more data input values) are received and / or until there is space in the (one or more) output buffers of the processing element 2200A for the data produced by the execution of the operations on those operands.
[0251] In one embodiment, the operation configuration register 2209B is loaded during configuration (e.g., mapping) and specifies the (one or more) specific operations to be performed by this processing (e.g., computing) element. The scheduler 2214B (e.g., operation selector) can schedule one or more operations of the processing element 2200A, e.g., when the input data and control inputs arrive. The inputs and outputs (e.g., via the (one or more) buffers) can be sent via a network (e.g., any network described herein). The control input buffer 2222B can be connected to a local network (e.g., and the local network can include a data path network as shown in and a in the process control path network), and is loaded with values upon arrival (e.g., the network has (one or more) data bits and (one or more) valid bits). The control input buffer 2222B can be coupled to a zero generator 2225B, e.g., to add leading or trailing zeros to the value from the control input buffer 2222B to form the expected width of the data item (e.g., 64 bits). The control output buffer 2232B, the data output buffer 2234B, and / or the data output buffer 2236B can receive the output of the processing element 2200B, e.g., as controlled by an operation (output of the scheduler 2214B). In one embodiment, the operation configuration register 2209B is loaded during configuration (e.g., mapping) and specifies the (one or more) specific operations to be performed by this processing (e.g., computing) element (e.g., and if adjacent PEs 2200B are to be used for joint operations, such as sequence operations). In one embodiment, the operation configuration register 2209A and the operation configuration register 2209B are loaded with data in the format described herein (e.g., as described in). The data in the control input buffer 2222B and the control output buffer 2232B can be single bits. The mux 2221B (e.g., operand A) and the mux 2223B (e.g., operand B) can source inputs.
[0252] For example, assume the operation of this processing (e.g., computing) element is (or includes) the pick as described in. The processing element 2200B then selects data from the data input buffer 2224B or the data input buffer 2226B, e.g., to go to the data output buffer 2234B (e.g., default) or the data output buffer 2236B. Thus, the control bit in 2222B can indicate 0 when selecting from the data input buffer 2224B or 1 when selecting from the data input buffer 2226B.
[0253] For example, assume the operation of this processing (e.g., computing) element is (or includes) Figure 3B the switch as described in. The processing element 2200B outputs data from the data input buffer 2224B (e.g., default) or the data input buffer 2226B to the data output buffer 2234B or the data output buffer 2236B, e.g.. Thus, the control bit in 2222B can indicate 0 when outputting to the data output buffer 2234B or 1 when outputting to the data output buffer 2236B.
[0254] Multiple networks (e.g., interconnections) can be connected to the processing element, e.g., (input) networks (e.g., Figure 9 networks 902, 904, 906 in) and (output) networks 908, 910, 912). The connections can be switches, e.g., with reference toFigure 7A and Figure 7B as described. In one embodiment, each network includes two sub - networks (or two channels on the network), for example, one for the Figure 7A data - path network in Figure 7B and one for the flow - control (e.g., back - pressure) path network in Figure 7A . As an example, the local network can be switched (e.g., connected) to the control input buffer 2222B (e.g., as established for control interconnection). In this embodiment, the data path (e.g., the network in Figure 7A ) can carry control input values (e.g., one or more bits) (e.g., control tokens), and the flow - control path (e.g., the network) can carry a back - pressure signal (e.g., back - pressure or no - back - pressure token) from the control input buffer 2222B, e.g., to indicate to an upstream producer (e.g., a PE) that a new control input value has not been loaded into (e.g., sent to) the control input buffer 2222B until the back - pressure signal indicates that there is space in the control input buffer 2222B for the new control input value (e.g., from the control output buffer of the upstream producer). In one embodiment, a new control input value may not enter the control input buffer 2222B until (i) the upstream producer receives a "space available" back - pressure signal from the "control input" buffer 2222B and (ii) the new control input value is sent, e.g., from the upstream producer, and this can stall the processing element 2200B until that occurs (and space in the (one or more) target output buffers is available).
[0255] The data input buffer 2224B and the data input buffer 2226B can operate similarly. For example, the local network (e.g., as established for data (as opposed to control) interconnection) can be switched (e.g., connected) to the data input buffer 2224B. In this embodiment, the data path (e.g., Figure 7AThe network (e.g., in the network) can carry data input values (e.g., one or more bits) (e.g., data stream tokens), and the flow control path (e.g., the network) can carry a backpressure signal (e.g., a backpressure or no-backpressure token) from the data input buffer 2224B, e.g., to indicate to an upstream producer (e.g., a PE) that new data input values are not loaded into (e.g., sent to) the data input buffer 2224B until the backpressure signal indicates that there is space in the data input buffer 2224B for the new data input values (e.g., the data output buffer of the upstream producer). In one embodiment, the new data input values may not enter the data input buffer 2224B until (i) the upstream producer receives a "space available" backpressure signal from the "data input" buffer 2224B and (ii) the new data input values are sent, e.g., from the upstream producer, and this may stall the processing element 2200B until that occurs (and space in the (one or more) target output buffers is available). The control output values and / or data output values may be stalled in their respective output buffers (e.g., 2232B, 2234B, 2236B) until the backpressure signal indicates that there is available space in the input buffer for the downstream (one or more) processing elements.
[0256] The processing element 2200B can be stalled from execution until its operands (e.g., control input values and their corresponding one or more data input values) are received and / or until there is space in the (one or more) output buffers of the processing element 2200B for the data produced by the execution of the operations on those operands.
[0257] In some embodiments, a processing element (PE) has one or more (e.g., two or three) operations that it can execute. For example, the PE can be configured based on the operations (e.g., operation values) input into the PE.
[0258] Figure 23 An example operation format 2300 for an integer arithmetic / logic data flow operator implementation on a processing element according to an embodiment of the present disclosure is shown. Although a 32-bit width of the operation value is shown, other bit widths are possible (e.g., 64 bits). According to the shown format, (e.g., lower) bits 20-0 (e.g., those 21 bits) are used to indicate to the processing element (e.g., the scheduler and / or controller) regarding a particular operation to be executed (e.g., and regarding which input(s) to use and / or to which output(s) to send the result). Other bits (e.g., bits 31-21) may be reserved for other purposes, e.g., filled with zeros when configuring the PE.
[0259] Figure 24An example arithmetic format 2400 of a sequencer data flow operator implementation on a processing element according to an embodiment of the present disclosure is shown. Although a 32-bit width of the operation value is shown, other bit widths are possible (e.g., 64 bits). According to the shown format, (e.g., lower) bits 20-0 (e.g., those 21 bits) are used to indicate to the processing element (e.g., scheduler and / or controller) regarding a particular operation to be performed (e.g., and regarding which input(s) to use and / or to which output(s) to send the result). Another bit or bits (e.g., other bits (e.g., bits 31-21), which are reserved for other uses according to Figure 23 the format 2300, e.g., padded with zeros when configuring the PE) can be used to switch between a first mode (e.g., as a stand-alone (e.g., integer) PE) and a second mode (e.g., as a sequencer), e.g., where the sequencer mode is one of the end bits. In one embodiment, the sequencer functionality is binary compatible with the integer PE by loading a "sequencer mode" bit in one of the (e.g., upper) bits of the configuration operation field, to save software engineering costs (e.g., based on the assumption that the configuration operation value is reported, it utilizes the (e.g., normal) data width of the CSA network (e.g., 32 bits or 64 bits), and the integer PE configuration uses less than the full data width (e.g., the configuration instructions for a basic integer PE can be only 21 bits wide)). In one embodiment, an operation configuration register (e.g., Figure 21 the operation configuration register 2109 of Figure 22 the operation configuration register 2209A and / or operation configuration register 2209B) is loaded during configuration (e.g., mapping) and specifies the particular operation(s) to be performed by this processing (e.g., computing) element, e.g., and couples two PEs together as a single sequencer data flow operator implementation. For example, when adjacent PEs both have their (one or more) sequencer mode bits set to, e.g., logic high (e.g., logic 1), the circuit between the two adjacent PEs (e.g., sequencer comparison data path 2243) can be enabled so that they work together on sequence operations. The sizes of the given fields are just examples (e.g., a 21-bit field for integer PE operations), and other sizes can be utilized in some embodiments. In one embodiment, only a subset of all the PEs in the array can include sequencer functionality.
[0260] Figure 25 An example arithmetic format 2500 of a sequencer data flow operator implementation on a processing element according to an embodiment of the present disclosure is shown. In one embodiment, the operation format 2500 is related to the sequencer span PE (seqstr) (e.g., Figure 22for use in conjunction with the processing element 2200A). Format 2500 includes using destination operand select bits (e.g., as existing in format 2300 or format 2400) (e.g., to route data to an output buffer) and / or source operand select bits (e.g., to route data from an input buffer) so as to allow the PE to source data from / to the buffer / PE. Another bit or bits (e.g., other bits (e.g., bits 30 - 21) reserved for other uses in format 2400 as per Figure 24 e.g., padding with zeros when configuring the PE) may be used to store additional destination operand select bits (e.g., due to the addition of the register bank(s) 2244) and / or additional source operand select bits (e.g., due to the addition of the register bank(s) 2244), e.g., to allow the PE to source data from / to the register bank(s) 2244. In one embodiment, format 2500 includes separating similarly typed fields (e.g., destination and source operand identification bits) grouped together (e.g., all input bits, all output bits, etc.) so as to keep the “integer PE configuration operation” format intact.
[0261] Figure 26 FIG. shows an example operation format 2600 for an sequencer data flow operator implementation on a processing element according to an embodiment of the present disclosure. Another possible alternative is to have reserved (e.g., spare) bits in the configuration bits (e.g., in bits 27 - 0). This may have the advantage of reducing software engineering costs for binary compatibility. Referring to Figure 22Sequencer data flow operator 2201 (e.g., one of the possible sequencer data flow operator implementations), in order to achieve a reasonable cycle time, the two ALUs used by sequencer data flow operator 2201 may not be in series on the sequencer comparison data path 2243 in the same clock cycle (e.g., the output of ALU 2218A in sequencer span (seqstr) processing element 2200A is first latched in a (e.g., 64-bit) register bank 2244 before being passed to sequencer comparison (seqcmp) processing element 2200B and, for example, input to ALU 2218B). Thus, in some embodiments, it is possible for the CSA to achieve the same frequency as the processor core (e.g., approximately 4 - 5 GHz). This may include, for example, programming the CSA to avoid pipeline hazards when backpressure occurs or when the input arrival time is arbitrarily delayed (caused by pipelining the two ALUs) in order to have correct functional behavior. The processing element may include a multiplier, a shifter, and / or some other dedicated ALU (e.g., in sequencer span (seqstr) processing element 2200A) if a particular application can utilize such a sequence generation algorithm. Similarly, if such a sequence generation algorithm becomes desirable for use in the CSA, the sequencer design can be extended to floating-point arithmetic / comparison or any other logical / arithmetic expression. In one embodiment, the sequencer can be self-cleaning by carefully aligning its control and internal reset signals to various controllers (e.g., finite state machine (FSM) and trigger control circuit). In other words, when the full sequence is generated based on the current set of three data input tokens (e.g., base value, span, and limit), all three data inputs (e.g., data tokens) can be fully dequeued, and thus the sequencer can accept a new set of data tokens to generate a new sequence. This would be useful for nested loops without reconfiguring the CSA (e.g., the interconnection of PEs and / or CSA).
[0262] Control paradigm
[0263] At the component level of individual processing, the data flow architecture used within a CSA can be very energy efficient when the circuit switches and performs useful computations / data transfers only when input data (e.g., one or more data tokens) is available and there is no backpressure on the corresponding output data (e.g., one or more data tokens). However, sequencer data flow operators can use more data input operands and can generate more output data operands (e.g., a token stream), where the corresponding data flow architecture controller / scheduler may be significantly more expensive in terms of its area / energy cost. Supporting more modes / functionality to meet the semantics of high-level programming constructs can also exacerbate this area / energy issue in some embodiments. While it is possible to expand the programmable state of the data flow architecture at the data flow operator level to achieve all the required functionality, some embodiments herein include a new control paradigm that uses the ability of having (e.g., small) embedded finite state machines (FSMs) to achieve the same set of functionality with lower energy / area cost and greater flexibility to expand the data flow PEs. To simplify implementation, some embodiments herein allow the PEs to partially exit the data flow mode and instead use one or more of the embedded state machines and later return to the full data flow style. This allows some embodiments to implement stateful functionality (e.g., a subset thereof) without being penalized by the overhead of a fully general solution. An additional advantage in some embodiments is that those embedded state machines can be largely separated from the main data flow architecture and allow the sequencer data flow operator to still operate as a (e.g., integer) PE, e.g., to maximize the effective silicon area utilization. As described below, the flexibility of this hybrid data flow / embedded state machine approach can also allow for an easy expansion of the microarchitecture with additional modes / functionality as needed. Some embodiments herein use embedded state machines to expand the data flow architecture, e.g., to allow more complex data flow operators (e.g., sequencers) to seamlessly transition between various control paradigms with greater flexibility and lower area / energy cost to achieve the same set of functionality.
[0264] Some embodiments herein utilize a single PE with an embedded state machine to distribute control as needed, and since each of the embedded state machines can be smaller (e.g., much smaller) (e.g., in terms of silicon area) compared to the independent operation of each including the state machine functionality, it allows for greater flexibility, lower energy / area cost, and better scalability for some (e.g., more complex) data flow operators.
[0265] Figure 27 FIG. 2700 shows a circuit implementing a sequencer data flow operator on multiple processing elements according to an embodiment of the present disclosure. As Figure 27 shown (e.g., showing Figure 22The sequencer span (portion of the seqstr processing element 2200A and portion of the sequencer compare (seqcmp) processing element 2200B (e.g., the last two digits of their common reference numeral)), the circuit 2700 will adapt such that, due to the LIC (latency-insensitive channel), the base value (e.g., starting value) data token and the span data token can arrive at any time and / or in any order. (e.g., Figure 22 Two (e.g., small and / or identical) finite state machines (FSMs) (2750, 2752) of the sequencer span (seqstr) processing element 2200A are used to track the arrival of those two data tokens (e.g., at input buffers 2724A and input buffer 2726A respectively, e.g., corresponding to Figure 22 input buffers 2224A and input buffer 2226A in). In one implementation, both FSMs 2750 and 2752 can each have only two states. One state is in_reset / invalid / data_token_has_not_arrived. The other state is out_of_reset / valid / data_token_has_arrived. Implementations with more states are possible in some embodiments. For example, if the arithmetic operations for the sequencer are power-consuming and / or are considered infrequent, power savings can be achieved by including states such as sleep state, wake-up state, full-power / active state, etc., to provide options for power gating and / or clock gating of the (e.g., arithmetic) circuits used inside the sequencer. AND logic gate 2756 can receive an input (e.g., logic one) from each of the FSMs (2750, 2752), which indicates the corresponding data token (e.g., base value (e.g., basic token)) in one of the buffers (2724A, 2726A) they each receive and the span value (e.g., data token) in the other buffer (2724A, 2726A) (e.g., indicating that the base and span data tokens have arrived) at a time. Data path 2758 (e.g., a single wire) can couple the output of the first AND logic gate 2756 to the second AND logic gate 2760. The second AND logic gate 2760 can also receive as an input from (e.g., Figure 22Output of the FSM 2754 of the sequencer compare (seqcmp) processing element 2200B. The FSM 2754 can receive inputs and indicate the time when a limit data token (e.g., a limit value (e.g., a limit token)) is in one of the buffers (2724B, 2726B) (e.g., either). In one implementation, the FSM 2754 can have only two states. One state is in_reset / invalid / data_token_has_not_arrived. The other state is out_of_reset / valid / data_token_has_arrived. Implementations with more states are possible in some embodiments. For example, states can be included such that the limit data token can arrive from the input buffer 2724B or 2726B to increase network routing flexibility. For example, states can be included that limit the limit data token to arriving only from one of the input buffers or a specific subset. If the dynamic reconfiguration time for changing that limit allows, some embodiments can have multiple loops of a sequencer that share loop control flow generation. By combining the outputs from the FSM 2750 and FSM 2752, this scheme can have the beneficial effect of reducing wire count (e.g., using 1 wire (e.g., data path 2758) between two adjacent PEs instead of 2 wires to signal the arrival of two data tokens). The FSM 2754 can track whether a "limit" data token has arrived (e.g., in either the input buffer 2724B or the input buffer 2726B), and a single "valid" signal (e.g., on data path 2762) can be used to signal to the seqstr controller 2742 and / or the seccmp controller 2740 about the ability to generate a valid comparison result (e.g., because "base", "span", and "limit" tokens have arrived). This can also create the flexibility to designate one or two (e.g., wide data) input buffers (e.g., corresponding channels) as possible receivers of the "limit" data token in the seqcmp PE, and by adding that functionality in the seqcmp PE, the complexity of the seqstr PE does not increase in some embodiments. Similarly, network channel binding can have different options (e.g., for base and span data tokens) on the seqstr Pe side without increasing seqcmp PE complexity.
[0266] Figure 28 Circuit 2800 is shown that supports a one-way mode for sequencer data flow operator implementation on a single processing element, in accordance with an embodiment of the present disclosure. As Figure 28 shown (e.g., shows Figure 22Part of the sequencer span (seqstr) processing element 2200A, e.g., which shares the last two digits in the reference numerals), to support the semantics of a do-while loop construct (e.g., in the C programming language), where the do-while loop will run at least one iteration of the loop regardless of whether the first comparison is successful or failed), the sequencer data flow operator supports a special mode called one_trip_mode. (e.g., small) FSM 2864 only enforces a "success" value for the comparison on the first iteration of the loop to support this functionality without touching the existing data flow architecture and / or the default mode sequencer controller. In one embodiment, FSM 2864 has two states. One state is in_reset / first_iteration_not_seen_yet, and the other state is out_of_reset_and_first_iteration_is_done. In one embodiment, FSM 2864 outputs a logic one (e.g., a voltage signal corresponding to logic one) until FSM 2864 sees the first loop iteration. That logic one hits an inverter (e.g., NOT) logic gate 2865 such that when the inverter logic gate 2868 receives a zero from FSM 2864 indicating that the first loop iteration is coming, the inverter logic gate 2865 outputs a logic one. If the one_trip_mode is enabled here (e.g., a one on the signal input 2867), then the AND logic gate 2866 will initially output a one, which will be output from the OR logic gate 2868 so that the (e.g., first) iteration of the loop is executed by, for example, the seqstr controller 2842 (e.g., corresponding to Figure 22 the seqstr controller 2242). Once the first iteration of the loop is complete, the combination of the inverter logic gate 2865 and the logic gate 2866 can ensure that additional loop iterations are not forced by FSM 2864 (e.g., the one_trip_mode circuit). Additionally, a signal (e.g., logic one) can be output from the sequencer comparison (seqcmp) processing element (e.g., Figure 22 on the data path 2241 of the processing element 2200B in Figure 22 to the OR logic gate 2868 so that another iteration of the loop is executed by, for example, the seqstr controller 2842 (e.g., corresponding to
[0267] Figure 29 Figure 29 Figure 22 the seqstr controller 2242). Although logic one and zero are discussed, other signals can be utilized, e.g., the inverses of the one and zero. Figure 22Part of the sequencer span (seqstr) processing element 2200A (e.g., sharing the last two digits in the reference numeral), circuit 2900 includes a simplified mode, e.g., to reconfigure the sequencer span (seqstr) processing element into a simplified operator. Given the semantics of a given simplification operation (e.g., the first in the control channel causes accumulation), thus (e.g., 64-bit) register file 2944 (e.g., Figure 22 register file 2244 in Figure 22 ALU 2218A in Figure 22 is a source operand of ALU 2918A (e.g., ALU 2218A in Figure 22 ), so the "base" value is pre-loaded into register file 2944. On the other hand, for loop constructs, it may not be necessary to pre-load the (e.g., 64-bit) register file 2944 because the first value stream data output token will originate directly from input data buffer 2926A (e.g., a channel). Input data buffer 2926A can be
[0268] Figure 30 Input data buffer 2224A or input data buffer 2226A in Figure 30 In some embodiments herein, the CSA does not require dedicated hardware for the simplified operator, but can re-use the sequencer span PE. Multiplexer 2970 can receive input signals to switch between the sequencer span mode (e.g., logic zero) and the simplified mode (e.g., logic zero). In the simplified mode, data (e.g., the base value) can be loaded from input data buffer 2926A into register file 2944 through multiplexer 2970. In the sequencer span mode, ALU 2918A can send data (e.g., as Figure 9 ALU 2218A sends data to register file 2244 in Figure 21 to register file 2944 through multiplexer 2970. Figure 22 Figure 22 Circuit 3000 showing the sequencer mode switched to the sequencer data flow operator implementation on a single processing element according to an embodiment of the present disclosure. As Figure 30 shown (e.g., showing part of the sequencer comparison (seqcmp) processing element 2200B, e.g., sharing the last two digits in the reference numeral), circuit 3000 saves energy costs (and deviates from the data flow architecture) because once the seqcmp PE is configured, the comparison operation code feeding ALU 3018B (e.g., from scheduler 3014) is statically presented to ALU 3018B (e.g., switched via multiplexer 3072). In one embodiment, the sequencer mode signal comes from the PE configuration register and / or the scheduler (e.g., in Figure 9 , Figure 21 or Figure 22In one embodiment (where multiple operations are possible in a single processing element), when it is not possible to statically expose multiple ALU opcodes to a single ALU, MUX 3072 can be used. In one embodiment, this has an energy advantage over data flow architectures because the only input that toggles is the "value" stream (e.g., which is the base value, base value + stride, base value + 2×stride, etc.), so the data change entropy is low because only a certain (e.g., the low-order bits in a 32-bit or 64-bit value) value is expected to change during each loop iteration. In a data flow architecture, the ALU opcode transitions from 0 to its correct value in the same cycle when a data token is provided to the ALU (e.g., triggering a CSA operation), but this can waste energy (due to additional bit toggling) and can also affect the cycle time.
[0269] Figure 31 Circuit 3100 is shown that switches between an active mode and a deactive mode of selective dequeueing of a sequencer data flow operator implementation on a single processing element, in accordance with an embodiment of the present disclosure. By using the underlying mechanisms of the data flow architecture and circuit to enqueue / dequeue data tokens, the dequeueing of three input data tokens can be fully user programmable. This has the additional beneficial effect of reducing area / energy costs. For example, for an algorithm like a merge sort of 256 elements, the stride can initially be 128 to divide the list into 2, then it is desired that the stride be 64 to divide the list into 4, and then it is desired that the stride be 32 to divide the list into 8, and so on. In all those recursive operations, the only new data token to be provided is the stride token. The base and limit tokens can remain in place to avoid wasting processing elements to repeatedly create duplicate loops that generate those tokens while the merge sort is running. Another example is, for example, a bubble sort for each loop iteration where the highest value is "pushed up" to the top of the memory array, changing the upper bound address for the next loop iteration (e.g., the base address and stride data tokens for the bubble sort address sweep do not change in the next iteration).
[0270] Sequencer Stride PE with Single PE Mode
[0271] In some embodiments, multiple (e.g., two) processing elements (e.g., sequencer stride (seqstr) processing element 2200A and sequencer compare (seqcmp) processing element 2200B) that work in cascade are used to form a sequencer data flow operator, e.g., for generating loop construction related data tokens (e.g., "value" stream, "first" stream, "last" stream, and "assert" stream). In certain embodiments, generating the "first" stream, "last" stream, and "assert" stream from a two-PE sequencer data flow operator can be redundant. Certain embodiments herein provide for a stride PE (e.g., Figure 22Expansion of the sequencer span (seqstr) processing element 2200A) in [the context], which allows the PE to operate in a single PE mode. This can provide even greater efficiency while retaining the flexibility to support multiple (e.g., three) basic data flow operator modes (e.g., basic integer PE mode, reduced operator mode, and sequencer mode). This expansion can reduce the structural area and energy required for implementation routines (e.g., Figure 5A or Figure 5B the memcpy code (routine) in [the context]) by approximately 20%. Some embodiments herein provide a sequencer span PE in single PE mode for use, for example, in any case where (e.g., loop) control decision flow can be shared between two or more sequence generation algorithms, thus significantly reducing energy usage and freeing up valuable real estate for other CSA data flow operators. Some embodiments herein allow the reuse of the sequencer compare (seqcmp) processing element (e.g., processing element 2200B, which accompanies the sequencer span (seqstr) processing element 2200A) in integer PE mode. In some embodiments, for example, in contrast to using two PE sequencer data flow operators to generate any loop construct sequence, a sequencer span PE in single PE mode can be used for sequencing operations. In certain embodiments, the sequencer compare (seqcmp) processing element of the sequencer data flow operator can be, for example, freed up and reused in integer PE mode, or clock gated and / or power gated to save energy.
[0272] In single PE mode, a sequencer span (seqstr) processing element (e.g., Figure 22 the seqstr PE2200A) of [the context] can be used without its accompanying sequencer compare (seqcmp) processing element (e.g., Figure 22 the seqcmp 2200B) of [the context] to generate an additional "value" flow when another full sequencer (e.g., seqstr Pe and seqcmp PE pair) can provide the correct "decision" flow. For example, when calculating a dot product, at least 2 arrays of the same size are iterated through. When passing through a memory copy loop, in some embodiments, each source address should have a corresponding destination address. Consider the following matrix multiplication code example.
[0273] Figure 32 Shows an example of matrix multiplication code 3200 according to an embodiment of the present disclosure. Figure 33A - 33B Shows a first sequencer data flow operator implementation on multiple processing elements that generate Figure 32 A[i][k] and B[k][j] of the matrix multiplication according to an embodiment of the present disclosure.
[0274] As can be seen from Figure 33A - 33BIt can be seen that the shown sequencer implementation for generating the address sequences of A[i][k] and B[k][j] utilizes two full-size sequencer data flow operators (3301, 3303) (e.g., two pairs of sequencer stride (seqstr) processing elements with their accompanying sequencer compare (seqcmp) processing elements, i.e., four PEs). It should be noted that the stride sizes of array A (stride size = 8) and array B (stride size = c2×8) can be different (e.g., as long as c2>1).
[0275] Certain embodiments of this disclosure can avoid utilizing two sequencer data flow operators. In one sequencer, the code run can reuse the controls from the sequencer, but it is not desirable to occupy two PEs. A single sequencer compare PE can issue its compare signals to multiple (e.g., seqstr) PEs on the array. Therefore, instead of just one seqstr and seqcmp pair of PEs as shown above Figure 22 there can be multiple seqstr PEs (e.g., Figure 22 the sequencer stride (seqstr) processing element 2200A) and one seqcmp PE that passes signals to multiple seqstr PEs.
[0276] Figure 34 Shown is a second optimized sequencer data flow operator implementation 3400 on multiple processing elements (two PEs in 3401 and one PE in 3405) that generate Figure 32 A[i][k] and B[k][j] for matrix multiplication according to an embodiment of the present disclosure. As seen in Figure 34 the optimized sequencer implementation for generating the address sequences of A[i][k] and B[k][j] uses only one full-size sequencer data flow operator 34701 and one sequencer stride PE (e.g., i.e., three PEs).
[0277] Figure 35 Shown is a sequencer data flow operator implementation 3500 on multiple processing elements (two PEs in 3501 and one PE in 3505) that transform a sparse memory access pattern into a dense memory access pattern according to an embodiment of the present disclosure. Also note that in embodiments where each seqstr Pe receives its own stride size data token, the embodiments herein can include the option of using different stride sizes to obtain the necessary new data layout (which is most beneficial from the perspective of energy / access time for future processing).
[0278] Figure 36FIG. 3600 is a flowchart showing an embodiment in accordance with the present disclosure. The illustrated process 3600 includes: decoding an instruction into a decoded instruction by a decoder of a core of a processor (3602); running the decoded instruction by an execution unit of the core of the processor to perform a first operation (3604); receiving an input of a data flow graph including a plurality of nodes forming a loop structure (3606); overlaying the data flow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements, wherein each node represents a data flow operator among the plurality of processing elements controlled by an sequencer data flow operator of the plurality of processing elements (3608); and generating control signals for at least one data flow operator among the plurality of processing elements for each of the data flow operators of the plurality of processing elements and the sequencer data flow operator that reach the plurality of processing elements through corresponding incoming operand sets, and performing a second operation of the data flow graph by using the interconnection network and the plurality of processing elements (3610).
[0279] Figure 37 FIG. 3701 is a flowchart showing an embodiment in accordance with the present disclosure. The illustrated process 3701 includes: receiving an input of a data flow graph including a plurality of nodes (3703); and overlaying the data flow graph onto a plurality of processing elements of the processor, a data path network between the plurality of processing elements, and a flow control path network between the plurality of processing elements, wherein each node represents a data flow operator among the plurality of processing elements (3705).
[0280] In one embodiment, the core writes a command to a memory queue, and the CSA (e.g., a plurality of processing elements) monitors the memory queue and starts running when reading the command. In one embodiment, the core runs a first part of a program, and the CSA (e.g., a plurality of processing elements) runs a second part of the program. In one embodiment, the code performs another task while the CSA runs its operation.
[0281] 5. CSA Advantages
[0282] In certain embodiments, the CSA architecture and microarchitecture provide profound energy, performance, and usability advantages over roadmap processor architectures and FPGAs. In this section, these architectures are compared with embodiments of the CSA, and the advantages of the CSA over each in accelerating parallel data flow graphs are emphasized.
[0283] 5.1 Processor
[0284] Figure 38 FIG. 3800 shows throughput versus energy per operation in an embodiment in accordance with the present disclosure. As Figure 38As shown, small cores are generally more energy-efficient than large cores, and in some workloads, this advantage can be translated into absolute performance through higher core counts. The CSA microarchitecture follows these observations and removes (e.g., most of) the energy-consuming control structures associated with the von Neumann architecture, including most of the instruction-side microarchitecture. By removing these overheads and implementing a simple single-operation PE, embodiments of CSA result in a dense and efficient spatial array. Different from small cores, which are typically fully serial, CSA can, for example, combine its PEs together via a circuit-switched local network to form an explicitly parallel aggregated data flow graph. The result is performance not only in parallel applications but also in serial applications. Different from cores, which can pay a high price for performance in terms of area and energy, CSA is already parallel in its native execution model. In some embodiments, CSA uses speculation to increase performance, for example, and it does not need to repeatedly re-extract parallelism from sequential program representations, thus avoiding two of the major energy taxes in the von Neumann architecture. Most of the structures in embodiments of CSA are distributed, small, and energy-efficient, as opposed to the centralized, bulky, and energy-consuming structures present in cores. Consider the case of registers in CSA: each PE can have several (e.g., 10 or fewer) storage registers. Individually, these registers can be more efficient than traditional register files. Taken together, these registers can provide the effect of a large in-structure register file. Thus, embodiments of CSA avoid most of the stack overflows and fills caused by traditional architectures, while using much less energy per state access. Of course, applications can still access memory. In embodiments of CSA, memory access requests and responses are architecturally separated, enabling the workload to maintain more outstanding memory accesses per unit area and energy. This property results in sufficiently high performance for cache-limited workloads and reduces the area and energy required to saturate the main memory in memory-limited workloads. Embodiments of CSA demonstrate a new form of energy efficiency that is unique to non-von Neumann architectures. As a result of (e.g., most of) the PEs running a single operation (e.g., instruction), the operand entropy is reduced. In the case of incremental operations, each execution can result in a large number of circuit-level toggles and very little energy consumption, i.e., the case investigated in Subsection 6.2. In contrast, the von Neumann architecture causes a large number of bit transitions through reuse. The asynchronous style of embodiments of CSA also enables microarchitecture optimizations, such as the floating-point optimizations described in Subsection 3.5, which are difficult to achieve in a strictly scheduled core pipeline. Since the PEs can be relatively simple and the behavior in a specific data flow graph is statically known, clock gating and power gating techniques can be applied more effectively than in coarser architectures. The graph execution style, small size, and extensibility of embodiments of CSA PE and network together enable the expression of many kinds of parallelism: instruction, data, pipeline, vector, memory, thread, and task parallelism can all be achieved.For example, in embodiments of the CSA, one application can use the arithmetic units to provide a high degree of address bandwidth, while another application can use those same units for computing. In many cases, multiple parallelisms can be combined to achieve even greater performance. Many key HPC operations can be replicated and pipelined, resulting in an order-of-magnitude performance gain. In contrast, von Neumann-style cores are typically optimized for one style of parallelism carefully chosen by the designer, resulting in a failure to capture all important application kernels. Because embodiments of the CSA exhibit and facilitate many forms of parallelism, it does not require a particular form of parallelism, or worse, the presence of a particular subroutine in the application, in order to benefit from the CSA. Many applications, including single-stream applications, can obtain performance and energy benefits from embodiments of the CSA, such as when compiled without modification. This overturns the long-standing trend that requires a large amount of work from programmers to obtain sufficient performance gains in single-stream applications. In fact, in some applications, embodiments of the CSA obtain greater performance from functionally equivalent but less "modern" code than from its complex contemporary cousin, which is painfully targeted at vector instructions.
[0285] 5.2 Comparison of CSA Embodiments and FPGAs
[0286] The choice of data flow operators for the infrastructure that is an embodiment of CSA differentiates those CSAs from FPGAs, and specifically, CSA is an excellent accelerator for HPC data flow graphs arising from traditional programming languages. The data flow operators are substantially asynchronous. This enables embodiments of CSA to not only have great freedom in implementation in the microarchitecture, but also enables them to simply and concisely adapt to abstract architecture concepts. For example, embodiments of CSA naturally adapt to many memory microarchitectures with a simple load-store interface, which is substantially asynchronous. Only the FPGA DRAM controller needs to be examined to understand the difference in complexity. Embodiments of CSA also balance the asynchrony to provide faster and more full-featured runtime services (such as configuration and extraction), which are considered to be 4 to 6 orders of magnitude faster than FPGAs. By narrowing the architecture interface, embodiments of CSA provide control over most of the timing paths at the microarchitecture level. This allows embodiments of CSA to operate at much higher frequencies than the more general control mechanisms provided in FPGAs. Similarly, clocks and resets, which can be the architectural basis of FPGAs, are microarchitectures in CSA, for example eliminating the need to support them as programmable entities. The data flow operators can be mostly coarse-grained. By only dealing with coarse operators, embodiments of CSA improve the density of the structure and its energy consumption: CSA runs operations directly rather than using look-up tables to simulate them. A second result of the coarseness is the simplification of placement and routing problems. The CSA data flow graph is many orders of magnitude smaller than the FPGA netlist, and the placement and routing times are proportionally reduced in embodiments of CSA. The significant differences between embodiments of CSA and FPGAs make CSA an excellent accelerator for data flow graphs arising from traditional programming languages, for example.
[0287] 6. Evaluation
[0288] CSA is a new computer architecture that has the potential to provide significant performance and energy benefits relative to roadmap processors. Consider the case of a single-stride address that computes a walk across an array. This case can be important in HPC applications, for example, where it spends a large amount of integer workload in computing address offsets. In address calculations and especially in stride address calculations, one argument is constant while the other changes only slightly per computation. Therefore, only a small number of bits rotate per cycle in most cases. In fact, it can be shown that using a derivation similar to the limit on floating-point carry bits described in Subsection 3.5, less than two bits of the input rotate per computation on average for stride calculations, reducing energy by 50% for a random rotation distribution. If a time-multiplexing approach is used, much of this energy savings may be lost. In one embodiment, CSA achieves roughly 3x energy efficiency over the core while producing an 8x performance gain. The parallelism gain achieved through embodiments of CSA can lead to reduced program run times, thereby producing a commensurate reduction in leakage energy. At the PE level, embodiments of CSA are extremely energy-efficient. A second important issue for CSA is whether it consumes reasonable energy at the primitive level. Since embodiments of CSA are able to implement every floating-point PE in the structure every cycle, it serves as a reasonable upper bound on energy and power consumption, for example, such that most of the energy goes into floating-point multiply and add.
[0289] 7. Other CSA Details
[0290] This subsection discusses other details of configuration and exception handling.
[0291] 7.1 Configuring the CSA Microarchitecture
[0292] This subsection discloses examples of how to configure CSA (e.g., the architecture), how to implement this configuration quickly, and how to minimize the resource overhead of the configuration. Quickly configuring the architecture can be of paramount importance in accelerating small parts of larger algorithms and thus in broadening the applicability of CSA. This subsection further discloses features that allow embodiments of CSA to be programmed with configurations of different lengths.
[0293] Embodiments of CSA (e.g., the architecture) differ from traditional cores in that they utilize a configuration step where a (e.g., large) portion of the architecture is loaded with a program configuration before program execution. The advantage of static configuration can be that very little energy is spent on configuration at runtime, for example, as opposed to sequential cores that spend energy fetching configuration information (instructions) almost every cycle. A previous drawback of configuration was that it was a coarse-grained step with potentially large latency, which imposed a lower bound on the size of programs that could be accelerated in the architecture due to the cost of context switching. This disclosure describes a scalable microarchitecture for quickly configuring a spatial array in a distributed manner, for example, that avoids the previous drawbacks.
[0294] As described above, the CSA may include lightweight processing elements connected via a network among PEs. A program, regarded as a control-data flow graph, is mapped onto the architecture by configuring configurable structural elements (CFEs), such as PEs and an interconnect (structural) network. Generally, a PE can be configured as a data flow operator, and once all input operands arrive at the PE, an operation occurs, and the result is forwarded to another or multiple PEs for consumption or output. PEs can communicate via dedicated virtual circuits formed by statically configuring a circuit-switched communication network. These virtual circuits can be flow-controlled and fully backpressure, such that a PE will stall when the source has no data or the destination is full. At runtime, data can flow through the PEs implementing the mapped algorithm. For example, data can be streamed from memory through the fabric and then output to memory again. This spatial architecture can achieve significant performance efficiency compared to traditional multi-core processors: the computations in the form of PEs are simpler and more numerous than large cores, and the communication can be direct, as opposed to the scaling of the memory system.
[0295] Embodiments of the CSA may not utilize packet switching (e.g., software-controlled), such as packet switching that requires a large amount of software assistance to implement, which slows down configuration. Embodiments of the CSA include out-of-band signaling and fixed configuration topologies in the network (e.g., only 2 - 3 bits according to the supported feature set) to avoid the need for a large amount of software support.
[0296] A key difference between embodiments of the CSA and the way used in FPGAs is that the CSA approach can use wide data words, is distributed, and includes a mechanism to directly fetch program data from memory. Embodiments of the CSA may not utilize JTAG-style single-bit communication for the benefit of area efficiency, e.g., because that may require milliseconds to fully configure a large FPGA fabric.
[0297] Embodiments of the CSA include a distributed configuration protocol and the microarchitecture that supports this protocol. Initially, the configuration state can reside in memory. Multiple (e.g., distributed) local configuration controllers (blocks) (LCCs) can stream parts of the overall program into local regions of the spatial fabric, for example, using a combination of a small set of control signals and the fabric. State elements can be used at each CFE to form a configuration chain, e.g., allowing individual CFEs to program themselves without global addressing.
[0298] Embodiments of CSA include specific hardware support for the formation of configuration chains, such as software that does not dynamically establish these chains at the cost of increased configuration time. Embodiments of CSA are not fully packet switched and include additional out-of-band control wires (e.g., control is sent not through a data path that requires additional cycles to gate and re-serialize this information). Embodiments of CSA reduce configuration latency (e.g., by at least 1 / 2) by fixing the configuration order and by providing explicit out-of-band control, without significantly increasing network complexity.
[0299] Embodiments of CSA do not use a serial mechanism for configuration, where data is streamed bit by bit into the fabric using a JTAG-like protocol. Embodiments of CSA utilize a coarse-grained fabric approach. In some embodiments, adding a few control wires or status elements to a 64- or 32-bit oriented CSA fabric has a lower cost than adding those same control mechanisms to a 4- or 6-bit fabric.
[0300] Figure 39 An accelerator primitive 3900 is shown in accordance with an embodiment of the present disclosure, including an array of processing elements (PEs) and local configuration controllers (3902, 3906). Each PE, each network controller (e.g., a network data flow endpoint circuit), and each switch can be configurable fabric elements (CFEs), e.g., which are configured (e.g., programmed) through an embodiment of the CSA architecture.
[0301] Embodiments of CSA include hardware that provides efficient distributed low-latency configuration of a heterogeneous spatial fabric. This can be achieved in four techniques. First, a hardware entity, a local configuration controller (LCC), is utilized, e.g., Figure 39 - 41 The LCC can fetch a configuration information stream from (e.g., virtual) memory. Second, a configuration data path can be included that is as wide as the native width of the PE fabric and can overlay the PE fabric. Third, new control signals can be received into the PE fabric that organize the configuration process. Fourth, status elements can be located (e.g., in registers) at each configurable endpoint that track the status of adjacent CFEs, allowing each CFE to configure itself explicitly without additional control signals. These four microarchitecture features can allow CSA to configure its CFE chains. To achieve low configuration latency, this can be partitioned by building many LCC and CFE chains. At configuration time, these can operate independently to load the fabric in parallel, e.g., greatly reducing latency. Due to these combinations, a fabric configured using an embodiment of the CSA architecture can be fully configured (e.g., in hundreds of nanoseconds). Details of the operation of the various components of an embodiment of a CSA configuration network are disclosed below.
[0302] Figure 40A - 40CShows a local configuration controller 4002 that configures a data path network according to an embodiment of the present disclosure. The shown network includes a plurality of multiplexers (e.g., multiplexers 4006, 4008, 4010), which are configurable (e.g., via their respective control signals) to connect one or more data paths (e.g., from a PE) together. Figure 40A Shows a network 4000 (e.g., fabric) that is configured (e.g., set) for a certain previous operation or program. Figure 40B Shows a local configuration controller 4002 (e.g., including network interface circuitry 4004 to send and / or receive signals) strobbing configuration signals, and the local network is set to a default configuration (e.g., as shown), which allows the LCC to send configuration data to all configurable fabric elements (CFEs) (e.g., muxes). Figure 40C Shows the LCC strobbing configuration information across the network and configuring the CFEs in a predetermined (e.g., silicon-defined) sequence. In one embodiment, when the CFEs are configured, they can immediately start operating. In another embodiment, the CFEs wait to start operating until the fabric is fully configured (e.g., as signaled by the configuration terminals (e.g., Figure 42 configuration terminal 4204 and configuration terminal 4208) of each local configuration controller). In one embodiment, the LCC obtains control of the network fabric by sending a special message or driving a signal. It then strobes the configuration data (e.g., for a number of cycles) into the CFEs in the fabric. In these figures, the multiplexer network is an analog of the "switch" shown in some figures (e.g., Figure 6 ).
[0303] Local configuration controller
[0304] Figure 41 Shows a (e.g., local) configuration controller 4102 according to an embodiment of the present disclosure. A local configuration controller (LCC) can be a hardware entity that is responsible for loading the local part of a fabric program (e.g., in a subset of primitives, etc.), interpreting these program parts, and then loading these program parts into the fabric by driving appropriate protocols on various configuration wires. With this ability, the LCC can be a dedicated sequential microcontroller.
[0305] The LCC operation can start when receiving a pointer to a code segment. Depending on the LCB microarchitecture, this pointer (e.g., stored in pointer register 4106) can arrive via a network (e.g., from within the CSA (fabric) itself) or via a memory system access to the LCC. When receiving such a pointer, the LCC optionally flushes the relevant state from parts of the fabric for context storage and then proceeds to immediately reconfigure the parts of the fabric for which it is responsible. The program loaded by the LCC can be a combination of fabric configuration data and control commands for the LCC (e.g., lightly encoded by it). When the LCC streams through the program section, it can interpret the program as a command stream and perform appropriate encoded actions to configure (e.g., load) the fabric.
[0306] Two different microarchitectures of the LCC are shown in Figure 39 For example, one or both of them are used in the CSA. The first places the LCC 3902 at the memory interface. In this case, the LCC can make direct requests to the memory system to load data. In the second case, the LCC 3906 is placed on the memory network, where it can only indirectly request the memory. In both cases, the logical operations of the LCB remain unchanged. In one embodiment, the LCC is notified about the program to be loaded, for example, through a set of control status registers (which are visible to the OS, for example) that will be used to inform the individual LCC about the new program pointer, etc.
[0307] Additional out-of-band control channels (e.g., wires)
[0308] In some embodiments, the configuration relies on 2 - 8 additional out-of-band control channels to improve the configuration speed, as defined below. For example, the configuration controller 4102 can include the following control channels, such as the CFG_START control channel 4108, the CFG_VALID control channel 4110, and the CFG_DONE control channel 4112, examples of each of which are discussed in Table 2 below.
[0309] Table 2: Control Channels
[0310]
[0311] Generally, the manipulation of configuration information can be left to the implementer of the specific CFE. For example, the selectable function CFE may have provisions for setting registers using existing data paths, while the fixed function CFE may simply set the configuration registers.
[0312] Due to the long wire delay when programming a large set of CFE, the CFG_VALID signal can be regarded as the clock / latch enable for the CFE component. Since this signal is used as a clock, in one embodiment, the duty cycle of the line is at most 50%. As a result, the configuration throughput is roughly halved. Optionally, a second CFG_VALID signal can be added to enable continuous programming.
[0313] In one embodiment, only CFG_START is strictly transmitted on a separate coupling (e.g., a wire), and for example, CFG_VALID and DFG_DONE can be overlaid on other network couplings.
[0314] Reuse of network resources
[0315] To reduce the configuration overhead, some embodiments of CSA utilize the existing network infrastructure to transmit configuration data. The LCC can use the chip-level memory hierarchy and the fabric-level communication network to move data from the storage device to the fabric. Therefore, in some embodiments of CSA, the configuration infrastructure adds no more than 2% to the overall fabric area and power.
[0316] The reuse of network resources in some embodiments of CSA can enable the network to have some hardware support for the configuration mechanism. The circuit-switching network of the embodiments of CSA enables the LCC to set its multiplexer in a specific configuration manner when the 'CFG_START' signal is asserted. The packet-switching network does not require expansion, but the LCC endpoints (e.g., configuration terminals) use specific addresses in the packet-switching network. Network reuse is optional, and some embodiments may find a dedicated configuration bus more convenient.
[0317] Per CFE state
[0318] Each CFE can hold a bit indicating whether it has been configured (see, for example, Figure 13 ). This bit can be de-asserted when the configuration start signal is driven, and then asserted when a specific CFE has been configured. In a configuration protocol, the CFEs are arranged to form a chain, where the CFE configuration status bits determine the topology of the chain. A CFE can read the configuration status bit of the adjacent CFE. If this adjacent CFE is configured and the current CFE is not configured, the CFE can determine any current configuration data targeted at the current CFE. When the 'CFG_DONE' signal is asserted, the CFE can set its configuration bit, for example, to cause the upstream CFE to be configured. As a base case of the configuration process, the configuration terminal that asserts that it is configured (e.g., Figure 39 the configuration terminal 3904 of the LCC 3902 or the configuration terminal 3908 of the LCC 3906 in
[0319] Inside the CFE, this bit can be used to drive the flow control ready signal. For example, when the configuration bit is de-asserted, the network control signal can be automatically clamped to a value that prevents data flow, and inside the PE, no operations or other actions will be scheduled.
[0320] Handling high-latency configuration paths
[0321] An embodiment of the LCC can drive signals over long distances, e.g., through many multiplexers and with many loads. Thus, it may be difficult to get the signal to reach a distant CFE within a short clock cycle. In some embodiments, the configuration signal is at a portion (e.g., a small portion) of the primary (e.g., CSA) clock frequency to ensure digital timing regularity during configuration. Clock division can be used in the out-of-band signaling protocol and does not require any modification of the primary clock tree.
[0322] Ensuring consistent fabric behavior during configuration Since some configuration schemes are distributed and have non-deterministic timing due to program and memory effects, different parts of the fabric can be configured at different times. Thus, some embodiments of the CSA provide mechanisms to prevent inconsistent operation between configured and unconfigured CFEs. In general, consistency is regarded as a property required of the CFE and maintained by the CFE itself, e.g., using internal CFE state. For example, when the CFE is in the unconfigured state, it can claim that its input buffer is full and its output is invalid. When configured, these values will be set to the true state of the buffer. Since enough of the fabric has been configured, these techniques can allow it to start operating. For example, if long-latency memory requests are issued early, this has the further effect of reducing context switching.
[0323] Variable-width configuration
[0324] Different CFEs can have different configuration word widths. For smaller CFE configuration words, the implementer can balance the latency by fairly assigning the CFE configuration load across network wires. To balance the load on the network wires, one option is to assign the configuration bits to different parts of the network wire to limit the network latency on any one wire. Wide data words can be manipulated using serialization / deserialization techniques. These decisions can be made on a per-fabric basis to optimize the behavior of a particular CSA (e.g., fabric). The network controller (e.g., one or more of network controller 3910 and network controller 3912) can communicate with each domain (e.g., subset) of the CSA (e.g., fabric), e.g., to send configuration information to one or more LCCs. The network controller can be part of a communication network (e.g., separate from the circuit-switching network). The network controller can include network data flow endpoint circuitry.
[0325] 7.2 Microarchitecture for low-latency configuration of CSA and timely fetch of configuration data of CSA
[0326] Embodiments of the CSA can be energy-saving and high-performance means for accelerating user applications. When considering whether a program (such as its data flow graph) can be successfully accelerated by an accelerator, both the time to configure the accelerator and the time to run the program can be considered. If the running time is short, the configuration time can play a major role in determining successful acceleration. Therefore, in some embodiments, to maximize the domain of accelerable programs, the configuration time is made as short as possible. One or more configuration caches can be included in the CSA, for example such that high-bandwidth low-latency storage enables fast reconfiguration. A description of several embodiments of the configuration cache follows.
[0327] In one embodiment, during configuration, the configuration hardware (such as the LCC) optionally accesses the configuration cache to obtain new configuration information. The configuration cache can operate as a traditional address-based cache or work in OS management mode, where the configuration is stored in a local address space and addressed by referring to that address space. If the configuration state is in the cache, in some embodiments no request is made to the backing store. In some embodiments, this configuration cache is separate from any (such as low-level) shared cache in the memory hierarchy.
[0328] Figure 42 An accelerator primitive 4200 is shown in accordance with an embodiment of the present disclosure, including an array of processing elements, a configuration cache (such as 4218 or 4220), and a local configuration controller (such as 4202 or 4206). In one embodiment, the configuration cache 4214 coexists with the local configuration controller 4202. In one embodiment, the configuration cache 4218 is located in the configuration domain of the local configuration controller 4206, for example where the first domain ends at the configuration terminal 4204, and the second domain ends at the configuration terminal 4208). The configuration cache can allow the local configuration controller to refer to the configuration cache during configuration, for example to obtain the configuration state with a lower latency than referring to memory. The configuration cache (storage device) can be dedicated or can be accessed as a configuration mode of a storage element within the structure (such as the local cache 4216).
[0329] Cache Mode
[0330] Demand cache - In this mode, the configuration cache operates as a true cache. The configuration controller issues an address-based request, which is checked against the tags in the cache. Misses are loaded into the cache and can then be referenced again during future reprogramming.
[0331] Structure-Inside Storage Device (Register) Caching — In this mode, the cache is configured to receive references to the configuration sequence in its own small address space rather than the larger address space of the host. This can improve memory density since the portion of the cache used to store tags can instead be used to store configuration.
[0332] In some embodiments, the configuration cache may have configuration data preloaded into it, for example, by external guidance or internal guidance. This can allow for a reduction in the latency of the loader. Some embodiments herein provide an interface to the configuration cache that permits the loading of a new configuration state into the cache, for example, even if the configuration is already running in the structure. The initiation of this loading can be done from an internal or external source. Embodiments of the preloading mechanism further reduce latency by removing the latency of cache loading from the configuration path.
[0333] Prefetch Mode
[0334] Explicit Prefetch — The configuration path is augmented with a new command, ConfigurationCachePrefetch. Instead of programming the structure, this command simply causes the relevant program configuration to be loaded into the configuration cache without programming the structure. Since this mechanism rides on top of the existing configuration infrastructure, it is exposed both within and outside the structure, for example, to cores and other entities accessing the memory space.
[0335] Implicit Prefetch — The global configuration controller may maintain a prefetch predictor and use this to initiate an explicit prefetch of the configuration cache in an automated manner.
[0336] 7.3 Hardware for Fast Reconfiguration of CSA in Response to Exceptions
[0337] Some embodiments of CSA (e.g., a spatial structure) include a large number of instructions and configuration states, for example, which are mainly static during the operation of the CSA. Thus, the configuration state can be vulnerable to soft errors. Fast and error-free recovery from these soft errors is critical for the long-term reliability and performance of the spatial system.
[0338] Certain embodiments of the present disclosure provide a fast configuration recovery loop, e.g., where configuration errors are detected and parts of the fabric are reconfigured immediately. Certain embodiments of the present disclosure include a configuration controller, e.g., having reliability, availability, and serviceability (RAS) reprogramming features. Certain embodiments of the CSA include circuitry for high-speed configuration, error reporting, and parity within the fabric. Using a combination of these three features and an optional configuration cache, the configuration / exception handling circuitry can recover from soft configuration errors. When detected, the soft errors can be communicated to the configuration cache, which initiates immediate reconfiguration of the fabric (e.g., that part of it). Certain embodiments provide dedicated reconfiguration circuitry, e.g., which is faster than any solution implemented indirectly within the fabric. In certain embodiments, the co-existing exception and configuration circuitry cooperate to reload the fabric upon configuration error detection.
[0339] Figure 43 An accelerator primitive 4300 is shown in accordance with an embodiment of the present disclosure, including an array of processing elements and a configuration and exception handling controller (4302, 4306) having reconfiguration circuitry (4318, 4322). In one embodiment, when a PE detects a configuration error via its local RAS feature, it sends a (e.g., configuration error or reconfiguration error) message to the configuration and exception handling controller (e.g., 4302 or 4306) via its exception generator. Upon receipt of this message, the configuration and exception handling controller (e.g., 4302 or 4306) initiates the co-existing reconfiguration circuitry (e.g., 4318 and / or 4322) to reload the configuration state. The configuration microarchitecture proceeds and reloads (e.g., only) the configuration state, and in certain embodiments only the configuration state of the PE reporting the RAS error. Upon completion of the reconfiguration, the fabric can resume normal operation. To reduce latency, the configuration state used by the configuration and exception handling controller (e.g., 4302 or 4306) can originate from the configuration cache. As a base case of the configuration or reconfiguration process, the configuration terminal (e.g., Figure 43 configuration terminal 4304 of configuration and exception handling controller 4302 or configuration terminal 4308 of configuration and exception handling controller 4306 in
[0340] Figure 44 A reconfiguration circuit 4418 is shown in accordance with an embodiment of the present disclosure. The reconfiguration circuit 4418 includes a configuration state register 4420 to store the configuration state (or a pointer thereto).
[0341] CSA fabric hardware that initiates reconfiguration
[0342] Some portions of an application for a CSA (such as a spatial array) may run infrequently or may be mutually exclusive with other portions of a program. To save area, to improve performance and / or reduce power, it may be useful to time-multiplex portions of the spatial structure between several different portions of a program data flow graph. Some embodiments herein include an interface through which a CSA (such as via a spatial program) may request that a portion of the structure be reprogrammed. This may enable the CSA to change itself dynamically in accordance with a dynamic control flow. Some embodiments herein allow the structure to initiate a reconfiguration (such as a reprogramming). Some embodiments herein provide a set of interfaces for triggering a configuration from within the structure. In some embodiments, a PE issues a reconfiguration request based on a certain determination in a program data flow graph. This request may be propagated through a network to a new configuration interface where it triggers a reconfiguration. Once the reconfiguration is complete, optionally a message notifying about the completion may be returned. Thus, some embodiments of a CSA provide a program (such as a data flow graph) guided reconfiguration capability.
[0343] Figure 45 An accelerator primitive 4500 is shown in accordance with an embodiment of the present disclosure, including an array of processing elements and a configuration and exception handling controller 4506 having a reconfiguration circuit 4518. Here, a portion of the structure issues a request for a (re)configuration to a configuration domain such as the configuration and exception handling controller 4506 and / or the reconfiguration circuit 4518. The domain (re)configures itself, and when the request is satisfied, the configuration and exception handling controller 4506 and / or the reconfiguration circuit 4518 issues a response to the structure to notify the structure about the (re)configuration completion. In one embodiment, the configuration and exception handling controller 4506 and / or the reconfiguration circuit 4518 disables communication during the time when the (re)configuration is in progress, so that there are no consistency issues during program operation.
[0344] Configuration mode
[0345] Configuration according to address - In this mode, the structure makes a direct request to load configuration data from a specific address.
[0346] Configuration according to reference - In this mode, the structure makes a request to load a new configuration, for example, according to a predetermined reference ID. This may simplify the determination of the code to be loaded, since the location of the code is abstracted.
[0347] Configuring multiple domains
[0348] A CSA may include an advanced configuration controller to support a multicast mechanism for broadcasting (such as via a network shown by a dashed box) configuration requests to multiple (such as distributed or local) configuration controllers. This may enable a single configuration request to be replicated across a larger portion of the structure, for example, triggering a widespread reconfiguration.
[0349] 7.5 Abnormal Aggregator
[0350] Some embodiments of the CSA may also encounter anomalies (e.g., abnormal conditions), such as floating-point underflow. When these conditions occur, a special handler may be called to correct the program or terminate the program. Some embodiments herein provide a system-level architecture for manipulating anomalies in the spatial structure. Since some spatial structures emphasize area efficiency, the embodiments herein minimize the total area while providing a general anomaly mechanism. Some embodiments herein provide small-area components for signaling anomaly conditions occurring within the CSA (e.g., a spatial array). Some embodiments herein provide an interface and signaling protocol for transmitting such anomalies and PE-level anomaly semantics. Some embodiments herein are dedicated anomaly handling capabilities, for example, and do not require explicit manipulation by the programmer.
[0351] One embodiment of the CSA anomaly architecture consists of four parts, for example Figure 46 - 47 as shown. These parts can be arranged in a hierarchical structure where anomalies flow from the producer and ultimately all the way to the primitive-level anomaly aggregator (e.g., the handler), which can interface with an anomaly service routine of a core, for example. The four parts can be:
[0352] 1. PE Anomaly Generator
[0353] 2. Local Anomaly Network
[0354] 3. Small Backplane Anomaly Aggregator
[0355] 4. Primitive-Level Anomaly Aggregator
[0356] Figure 46 An accelerator primitive 4600 is shown in accordance with an embodiment of the present disclosure, including an array of processing elements and a small backplane anomaly aggregator 4602 coupled to a primitive-level anomaly aggregator 4604. Figure 47 A processing element 4700 having an anomaly generator 4744 is shown in accordance with an embodiment of the present disclosure.
[0357] PE Anomaly Generator
[0358] The processing element 4700 may include Figure 9 processing elements 900, for example, where there are similar labels for similar components (e.g., local network 902 and local network 4702). An additional network 4713 (e.g., a channel) may be the anomaly network. The Pe may implement an interface to the anomaly network (e.g., Figure 47 the anomaly network 4713 (e.g., a channel)) of Figure 47Shows the microarchitecture of such an interface, where the PE has an exception generator 4744 (e.g., initiating an exception finite state machine (FSM) 4740 to gate exception groups (e.g., BOXID 4742) onto the exception network). The BOXID 4742 can be a unique identifier of an exception generation entity (e.g., a PE or a box) within the local exception network. When an exception is detected, the exception generator 4744 senses the exception network and gates out the BOXID when it finds the network to be idle. Exceptions can be caused by many conditions, non - restrictively such as arithmetic errors, failed ECC checks on states, etc. However, it can also be that an exception data stream operation is introduced, where there is a concept of a support structure similar to a breakpoint.
[0359] The initiation of an exception can occur explicitly through the execution of an instruction provided by a programmer, or implicitly when a hardened error condition (e.g., floating - point underflow) is detected. Upon an exception, the PE 4700 can enter a waiting state, where it waits to be serviced by a final exception handler, e.g., external to the PE4700. The content of the exception group depends on the implementation of the specific PE, as described below.
[0360] The local exception network (e.g., local) routes exception groups from the PE 4700 to the small backplane exception network. The exception network (e.g., 4713) can be, for example, a serial packet - switching network of a subset of PEs, which consists of (e.g., a single) control wire and one or more data wires (e.g., organized in a ring or tree topology). Each PE can have a (e.g., local) termination station in the exception network (e.g., a ring), where it can arbitrate to inject a message into the exception network.
[0361] The PE endpoint that needs to inject an exception group can observe its local exception network exit point. If the control signal indicates busy, the PE waits to start injecting its group. If the network is not busy, i.e., the downstream termination station has no packet to forward, the PE will proceed to start injecting.
[0362] Network packets can have variable or fixed lengths. Each packet can start with a fixed - length header field that identifies the source PE of the packet. This can be followed by a variable number of PE - specific fields that contain information, such as error codes, data values, or other useful status information.
[0363] Small backplane exception aggregator
[0364] The small backplane exception aggregator 4604 is responsible for assembling local exception networks into larger packets and sending them to the primitive-level exception aggregator 4602. The small backplane exception aggregator 4604 can pre-consider its own unique ID for local exception packets, for example, to ensure that exception messages are unambiguous. The small backplane exception aggregator 4604 can interface with special exception-only virtual channels in the small backplane network, for example, to ensure deadlock-free exceptions.
[0365] The small backplane exception aggregator 4604 may also be able to directly service certain classes of exceptions. For example, configuration requests from the fabric can be serviced from the small backplane network using a cache local to the small backplane network termination station.
[0366] Primitive-level exception aggregator
[0367] The final level of the exception system is the primitive-level exception aggregator 4602. The primitive-level exception aggregator 4602 is responsible for collecting exceptions from various small backplane-level exception aggregators (such as 4604) and forwarding them to the appropriate service hardware (such as cores). Thus, the primitive-level exception aggregator 4602 may include some internal tables and controllers to associate specific messages with handler routines. These tables can be indexed directly or using a small state machine to direct specific exceptions.
[0368] Similar to the small backplane exception aggregator, the primitive-level exception aggregator can service some exception requests. For example, it can initiate reprogramming of most of the PE fabric in response to a specific exception.
[0369] 7.6 Extraction controller
[0370] Certain embodiments of the CSA include one or more extraction controllers to extract data from the fabric. Embodiments are discussed below for how to implement this extraction quickly and how to minimize the resource overhead of data extraction. Data extraction can be used for critical tasks such as exception handling and context switching. Certain embodiments herein extract data from a heterogeneous spatial fabric by introducing features that allow a variable and dynamically variable number of extractable fabric elements (EFEs) (such as PEs, network controllers, and / or switches) with state to be extracted.
[0371] Embodiments of the CSA include a distributed data extraction protocol and the microarchitecture to support this protocol. Certain embodiments of the CSA include multiple local extraction controllers (LECs) that stream program data from a local region of the spatial fabric using a (e.g., small) set of control signals and the combination of fabric-provided networks. State elements can be used at each extractable fabric element (EFE) to form an extraction chain, for example, allowing individual EFEs to extract themselves without global addressing.
[0372] Embodiments of the CSA do not use a local network to extract program data. Embodiments of the CSA include specific hardware support (e.g., an extraction controller) for the formation of extraction chains, such as not relying on software that dynamically builds these chains at the cost of increased extraction time. Embodiments of the CSA are not fully packet-switched and include additional out-of-band control wires (e.g., control is sent not through the data path that requires extra cycles to gate and re-serialize this information). Embodiments of the CSA reduce extraction latency (e.g., by at least 1 / 2) by fixing the extraction order and by providing explicit out-of-band control, without significantly increasing network complexity.
[0373] Embodiments of the CSA do not use a serial mechanism for data extraction, where data is streamed bit-by-bit from the fabric using a JTAG-like protocol. Embodiments of the CSA utilize a coarse-grained fabric approach. In some embodiments, adding a few control wires or status elements to a 64- or 32-bit oriented CSA fabric has a lower cost than adding those same control mechanisms to a 4- or 6-bit fabric.
[0374] Figure 48 An accelerator primitive 4800 is shown in accordance with an embodiment of the present disclosure, including a processing element array and local extraction controllers (4802, 4806). Each PE, each network controller, and each switch can be an extractable fabric element (EFE), e.g., configured (e.g., programmed) by an embodiment of the CSA architecture.
[0375] Embodiments of the CSA include hardware that provides efficient distributed low-latency extraction from a heterogeneous fabric. This can be achieved in four techniques. First, utilize a hardware entity, a local extraction controller (LEC), such as Figure 48 - 50 The LEC can accept commands from a host (e.g., a processor core), such as to extract a data stream from the fabric array, and write this data back to virtual memory for the host to inspect. Second, an extraction data path can be included that is as wide as the native width of the PE fabric and can overlay the PE fabric. Third, new control signals can be received into the PE fabric that organize the extraction process. Fourth, status elements can be located (e.g., in registers) at each configurable endpoint that track the status of adjacent CFEs, allowing each EFE to explicitly derive its status without additional control signals. These four microarchitecture features can allow the CSA to extract data from EFE chains. To achieve low data extraction latency, some embodiments can partition the extraction problem by including multiple (e.g., many) LECs and EFE chains in the fabric. During extraction, these chains can operate independently to extract data from the fabric in parallel, e.g., greatly reducing latency. Due to these combinations, the CSA can perform a full state dump (e.g., in hundreds of nanoseconds).
[0376] Figure 49A - 49CShows a local extraction controller 4902 that configures a data path network according to an embodiment of the present disclosure. The shown network includes a plurality of multiplexers (e.g., multiplexers 4906, 4908, 4910), which are configurable (e.g., via their respective control signals) to connect one or more data paths (e.g., from PEs) together. Figure 49A Shows a network 4900 (e.g., fabric) that is configured (e.g., set) for some previous operation or program. Figure 49B Shows the local extraction controller 4902 (e.g., including network interface circuitry 4904 to send and / or receive signals) gating an extraction signal and all PEs controlled by the LEC entering the extraction mode. The last PE in the extraction chain (or the extraction terminal) may take control of the extraction channel (e.g., bus) and send data in accordance with (1) a signal from the LEC or (2) an internally generated signal (e.g., from the PE). Once complete, the PE may set its completion flag, e.g., enabling the next PE to extract its data. Figure 49C Shows that the farthest PE has completed the extraction process and has thus set one or more extraction status bits, e.g., it swings the mux into an adjacent network so that the next PE can start the extraction process. The extracted PE may resume normal operation. In some embodiments, the PE may remain disabled until another action is taken. In these figures, the multiplexer network is an analog of the "switch" shown in some figures (e.g., Figure 6 )
[0377] The following subsections describe the operation of the various components of an embodiment of the extraction network.
[0378] Local configuration controller
[0379] Figure 50 Shows an extraction controller 5002 according to an embodiment of the present disclosure. The local extraction controller (LEC) may be a hardware entity that is responsible for accepting extraction commands, coordinating the extraction process with the EFE, and / or storing the extracted data, e.g., into virtual memory. With this capability, the LEC may be a dedicated sequential microcontroller.
[0380] LEC operation may start upon receiving a pointer to a receive buffer (e.g., in virtual memory) where the fabric state will be written and an optional command that controls how much of the fabric will be extracted. Depending on the LEC microarchitecture, this pointer (e.g., stored in pointer register 5004) may arrive via the network or via access to the LEC's memory system. When it receives such a pointer (e.g., command), the LEC proceeds to extract the state from the portion of the fabric for which it is responsible. The LEC may stream this extracted data from the fabric into a buffer provided by an external calling program.
[0381] Two different microarchitectures of the LEC are in Figure 48Shown in. First, place the LEC 4802 at the memory interface. In this case, the LEC can directly request the memory system to write the extracted data. In the second case, place the LEC 4806 on the memory network, where it can only indirectly request the memory. In both cases, the logical operations of the LEC can remain unchanged. In one embodiment, the LEC is notified about the data expected to be extracted from the fabric, for example, through a set of control status registers (which will be used to notify individual LECs about new commands) (e.g., visible to the OS).
[0382] Additional out-of-band control channels (e.g., wires)
[0383] In some embodiments, the extraction relies on 2 - 8 additional out-of-band signals to improve the configuration speed, as defined below. Signals driven by the LEC can be labeled as LEC. Signals driven by the EFE (e.g., PE) can be labeled as EFE. The extraction controller 5002 can include the following control channels, such as the LEC_EXTRACT control channel 5106, the LEC_START control channel 5008, the LEC_STROBE control channel 5010, and the EFE_COMPLETE control channel 5012, examples of each are discussed in Table 3 below.
[0384] Table 3: Extraction Channels
[0385]
[0386] Generally, the manipulation of the extraction can be left to the implementer of the specific EFE. For example, a selectable function EFE may have provisions for dumping registers using the existing data path, while a fixed function EFE may simply have a multiplexer.
[0387] Due to the long wire delay when programming a large set of EFEs, the LEC_STROBE signal can be regarded as the clock / latch enable for the EFE component. Since this signal is used as a clock, in one embodiment, the duty cycle of the line is at most 50%. Therefore, the extraction throughput is roughly halved. Optionally, a second LEC_STROBE signal can be added to achieve continuous extraction.
[0388] In one embodiment, only the LEC_START is strictly transmitted on a separate coupling (e.g., wire), for example, other control channels can be overlaid on an existing network (e.g., wire).
[0389] Reuse of network resources
[0390] To reduce the overhead of data extraction, some embodiments of the CSA utilize the existing network infrastructure to transfer the extracted data. The LEC may utilize the chip-level memory hierarchy and the fabric-level communication network to move data from the fabric to the storage device. Thus, in some embodiments of the CSA, the extraction infrastructure adds no more than 2% to the overall fabric area and power.
[0391] The reuse of network resources in some embodiments of the CSA may enable the network to have some hardware support for the extraction protocol. The circuit-switching network of some embodiments of the CSA causes the LEC to set its multiplexer in a specific configured manner when the 'LEC_START' signal is asserted. The packet-switching network does not require expansion, but the LEC endpoints (e.g., extraction terminals) use specific addresses in the packet-switching network. Network reuse is optional, and some embodiments may find a dedicated configuration bus more convenient.
[0392] Per EFE state
[0393] Each EFE may maintain a bit indicating whether it has exported its state. This bit may be de-asserted when the extraction start signal is driven, and then asserted when a particular EFE completes extraction. In one extraction protocol, the EFEs are arranged in a chain, where the EFE extraction status bits determine the topology of the chain. An EFE may read the extraction status bit of the adjacent EFE. If this adjacent EFE has its extraction bit set while the current EFE does not, the EFE may determine that it owns the extraction bus. When an EFE dumps its last data value, it may drive the 'EFE_DONE' signal and set its extraction bit, e.g., enabling the upstream EFE to configure for extraction. The network adjacent to the EFE may observe this signal and also adjust its state to manipulate the transition. As a basic case of the extraction process, the extraction terminal that asserts extraction completion (e.g., Figure 39 extraction terminal 4804 of LEC 4802 or extraction terminal 4808 of LEC 4806 in
[0394] Within the EFE, this bit may be used to drive the flow control ready signal. For example, when the extraction bit is de-asserted, the network control signals may be automatically clamped to values that prevent data flow, while within the PE, operations or actions will not be scheduled.
[0395] Handling high-latency paths
[0396] One embodiment of an LEC can drive signals over long distances, such as through many multiplexers and with many loads. Consequently, it can be difficult to get the signal to the remote EFE within a short clock cycle. In some embodiments, the extracted signal is at a fraction (e.g., a small fraction) of the primary (e.g., CSA) clock frequency to ensure regular digital timing at the time of extraction. Clock division can be used for out-of-band signaling protocols and does not require any modification to the primary clock tree.
[0397] Ensuring consistent structure behavior during extraction Because certain extraction schemes are distributed and have non-deterministic timing due to program and memory effects, different members of a structure may be under extraction at different times. While driving LEC_EXTRACT, all network flow control signals may be driven to a logic low, for example, thereby freezing the operation of a specific segment of a structure.
[0398] The extraction process can be lossless. Therefore, once the extraction has completed, the set of PEs can be considered operational. Extensions to the extraction protocol can allow PEs to be optionally disabled after extraction. Alternatively, in an embodiment, starting configuration during the extraction process will have a similar effect.
[0399] Single PE extraction
[0400] In some cases, it may be advantageous to extract a single PE. In this case, the optional address signal can be driven as part of the start of the extraction process. This allows the PE to be directly enabled for extraction. Once the PE has been extracted, the extraction process can be stopped by lowering the LEC_EXTRACT signal. In this way, a single PE can be selectively extracted, for example, by a local extraction controller.
[0401] Controlling extraction backpressure
[0402] In embodiments where the LEC writes the extracted data to memory (e.g., for post-processing, such as by software), it may be subject to limited memory bandwidth. In the event that the LEC exhausts its buffer capacity or anticipates that it will exhaust its buffer capacity, it may stop strobing the LEC_STROBE signal until the buffering problem has been resolved.
[0403] It should be noted that in some of the drawings (e.g. Figure 39 、 Figure 42 、 Figure 43 、 Figure 45 、 Figure 46 and Figure 48 ), communications are schematically illustrated. In some embodiments, those communications may occur via a (e.g., interconnected) network.
[0404] 7.7 Flowchart
[0405] Figure 51FIG. 5100 is a flowchart illustrating an embodiment in accordance with the present disclosure. The illustrated process 5100 includes: decoding an instruction into a decoded instruction using a decoder of a core of a processor (5102); running the decoded instruction using an execution unit of the core of the processor to perform a first operation (5104); receiving an input of a data flow graph including a plurality of nodes (5106); overlaying the data flow graph onto an array of processing elements of the processor, wherein each node is represented as a data flow operator in the array of processing elements (5108); and performing a second operation of the data flow graph using the array of processing elements when an incoming operand set arrives at the array of processing elements (5110).
[0406] Figure 52 FIG. 5200 is a flowchart illustrating an embodiment in accordance with the present disclosure. The illustrated process 5200 includes: decoding an instruction into a decoded instruction using a decoder of a core of a processor (5202); running the decoded instruction using an execution unit of the core of the processor to perform a first operation (5204); receiving an input of a data flow graph including a plurality of nodes (5206); overlaying the data flow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements of the processor, wherein each node is represented as a data flow operator in the plurality of processing elements (5208); and performing a second operation of the data flow graph using the interconnection network and the plurality of processing elements when an incoming operand set arrives at the plurality of processing elements (5210).
[0407] 8. Example Memory Ordering in Accelerator Hardware (e.g., in a Spatial Array of Processing Elements)
[0408] Figure 53A FIG. 5300 is a block diagram of a system 5300 in accordance with an embodiment of the present disclosure that employs a memory ordering circuit 5305 between an insertion memory subsystem 5310 and accelerator hardware 5302. The memory subsystem 5310 may include known memory components including caches, memories, and one or more memory controllers associated with a processor-based architecture. The accelerator hardware 5302 may be a coarse-grained spatial architecture that consists of lightweight processing elements (or other types of processing components) connected by an inter-processor element (PE) network or another type of inter-component network.
[0409] In one embodiment, a program that is considered to control the data flow graph is mapped onto the spatial architecture by configuring the PEs and the communication network. Generally, the PEs are configured as data flow operators, similar to functional units in a processor: once the input operands arrive at a PE, an operation occurs and the result is forwarded downstream in a pipelined fashion to the downstream PEs. The data flow operators (or other types of operators) may optionally consume incoming data on a per-operator basis. For example, simple operators such as those that manipulate unconditional evaluations of arithmetic expressions often consume all of the incoming data. However, it is sometimes useful for an operator to maintain state, such as in an accumulation.
[0410] The PEs communicate using dedicated virtual circuits formed by statically configuring circuit-switching communication networks. These virtual circuits are flow-controlled and fully backpressure, such that the PEs will stall when the source has no data or the destination is full. At runtime, data flows through the PEs implementing a mapping algorithm according to a data flow graph, also referred to herein as a subroutine. For example, data may flow from memory through acceleration hardware 5302 and then be output back to memory. This architecture can achieve significant performance efficiency relative to traditional multi-core processors: the computations in the form of PEs are simpler and more numerous than the large cores, and the communication is direct, as opposed to the expansion of the memory subsystem 5310. However, memory system parallelism helps support parallel PE computations. If memory accesses are serialized, high parallelism may not be achievable. To promote parallelism of memory accesses, the disclosed memory sorting circuit 5305 includes a memory sorting architecture and a microarchitecture, as will be described in detail. In one embodiment, the memory sorting circuit 5305 is a request address file circuit (or "RAF") or other memory request circuit.
[0411] Figure 53B is a system 5300 according to an embodiment of the present disclosure that instead employs multiple memory sorting circuits 5305 Figure 53A block diagram. Each memory sorting circuit 5305 can be used as an interface between a memory subsystem 5310 and a portion of the acceleration hardware 5302 (e.g., a spatial array or primitive of processing elements). The memory subsystem 5310 may include multiple cache levels 12 (e.g., Figure 53B cache levels 12A, 12B, 12C, and 12D in the embodiment of
[0412] Each memory ordering circuit 5305 can accept read and write requests for the memory subsystem 5310. Requests from the acceleration hardware 5302 arrive at the memory ordering circuit 5305 in independent channels at each node of the data flow graph that initiates read and write accesses, also referred to herein as load or store accesses. Buffering is provided such that the processing of loads returns the requested data to the acceleration hardware 5302 in the order in which it was requested. In other words, iteration six data is returned before iteration seven data, and so on. Additionally, note that the request channels from the memory ordering circuit 5305 to a particular cache bank can be implemented as ordered channels, and any first request left before a second request will arrive at the cache bank before the second request.
[0413] Figure 54 FIG. 5400 is a block diagram showing the general functionality of memory operations into / out of the acceleration hardware 5302 according to an embodiment of the present disclosure. Operations occurring from the top of the acceleration hardware 5302 are understood to be to / from the memory of the memory subsystem 5310. Note that two load requests occur, followed by corresponding load responses. While the acceleration hardware 5302 is performing processing on the data from the load responses, a third load request and response occur, which triggers additional acceleration hardware processing. The results of the acceleration hardware processing of these three load operations are then passed into the load operations, and thus the final result is stored back into memory.
[0414] By considering this sequence of operations, it is apparent that a spatial array maps more naturally to channels. Additionally, the acceleration hardware 5302 is latency insensitive with respect to request and response channels and the inherent parallel processing that can occur. The acceleration hardware can also decouple the execution of a program from the implementation of the memory subsystem 5310 ( Figure 53A ) because interfacing with the memory occurs at discrete times separate from the multiple processing steps performed by the acceleration hardware 5302. For example, a load request to the memory and a load response from the memory are independent actions and can be scheduled differently depending on different circumstances of the correlation flow of the memory operations. A spatial structure such as the use of processing instructions facilitates this spatial separation and decoupling of load requests and load responses.
[0415] Figure 55FIG. 5500 is a block diagram showing the spatial correlation flow of a store operation 5501 in accordance with an embodiment of the present disclosure. The store operation is mentioned by way of example because the same flow may apply to a load operation (but without incoming data) or to other operators (such as a fence). A fence is an ordering operation of the memory subsystem that ensures that all previous memory operations of a certain type (e.g., all stores or all loads) have been completed. The store operation 5501 may receive an address 5502 (of the memory) and data 5504 received from the acceleration hardware 5302. The store operation 5501 may also receive an incoming correlation token 5508, and in response to the availability of these three items, the store operation 5501 may generate an outgoing correlation token 5512. The incoming correlation token (which may be, for example, the initial correlation token of a program) may be provided according to a configuration provided by the compiler of the program, or may be provided by the execution of memory-mapped input / output (I / O). Alternatively, if the program has already run, the incoming correlation token 5508 may be received from the acceleration hardware 5302, for example, associated with a previous memory operation (on which the store operation 5501 depends). The outgoing correlation token 5512 may be generated based on the address 5502 and data 5504 required by subsequent memory operations of the program.
[0416] Figure 56 is a detailed block diagram of a memory ordering circuit 5305 in accordance with an embodiment of the present disclosure. Figure 53A The memory ordering circuit 5305 may be coupled to the unordered memory subsystem 5310 discussed above and may include a cache 12 and a memory 18, as well as an associated unordered memory controller. The memory ordering circuit 5305 may include or be coupled to a communication network interface 20, which may be an inter-primitive or intra-primitive network interface and may be a circuit-switched network interface (as shown) and thus includes a circuit-switched interconnect. As an alternative or in addition, the communication network interface 20 may include a packet-switched interconnect.
[0417] The memory ordering circuit 5305 may also include, but is not limited to, a memory interface 5610, an operation queue 5612, an input queue 5616, a completion queue 5620, an operation configuration data structure 5624, and an operation manager circuit 5630, which may also include a scheduler circuit 5632 and an execution circuit 5634. In one embodiment, the memory interface 5610 may be circuit-switched, and in another embodiment, the memory interface 5610 may be packet-switched, or both may be present simultaneously. The operation queue 5612 may buffer memory operations (with corresponding arguments) for requests to be processed and may thus correspond to the addresses and data entering the input queue 5616.
[0418] More specifically, the input queue 5616 can be an aggregation of at least the following: a load address queue, a store address queue, a store data queue, and a dependency queue. When the input queue 5616 is implemented as an aggregation, the memory ordering circuit 5305 can provide sharing of the logical queues with additional control logic in order to logically separate the queues, which are separate channels with the memory ordering circuit. This can maximize the use of the input queue, but can also require additional complexity and space in the logic circuitry to manage the logical separation of the aggregated queues. Alternatively, as will be described with reference to Figure 57 the input queue 5616 can be implemented in an isolated manner, where there are independent hardware queues for each. Whether aggregated ( Figure 56 ) or disaggregated ( Figure 57 ), the implementation is substantially the same for facilitating the present disclosure, where the former uses additional logic to logically separate the queues within a single shared hardware queue.
[0419] When shared, the input queue 5616 and the completion queue 5620 can be implemented as fixed-size circular buffers. A circular buffer is an efficient implementation of a circular queue with first-in first-out (FIFO) data characteristics. Thus, these queues can enhance the semantic order of the program that requests memory operations. In one embodiment, the circular buffer (e.g., the store address queue) can have entries corresponding to the entries flowing through an associated queue (e.g., the store data queue or the dependency queue) at the same rate. In this way, the store address can be kept associated with the corresponding store data.
[0420] More specifically, the load address queue can buffer the incoming addresses of the memory 18 (from which data is retrieved). The store address queue can buffer the incoming addresses of the memory 18 (to which data is written), which are buffered in the store data queue. The dependency queue can buffer dependency tokens associated with the addresses of the load address queue and the store address queue. Each queue (represented as a separate channel) can be implemented with a fixed or dynamic number of entries. When fixed, the more entries available, the more efficient complex loop processing can be. However, having too many entries costs more area and energy to implement. In some cases, such as with an aggregated architecture, the disclosed input queue 5616 can share queue slots. The use of the slots in the queue can be statically allocated.
[0421] The completion queue 5620 can be a separate collection of queues to buffer the data received from the memory in response to the memory commands issued by the load operations. The completion queue 5620 can be used to hold the load operations that have been scheduled but for which data has not yet been received (and thus not yet completed). Thus, the completion queue 5620 can be used to reorder data and the operation flow.
[0422] The operation manager circuit 5630 (which will be described with reference to Figure 57 to 13More detailed description) can provide logic for scheduling and running queued memory operations when considering correlation tokens used to provide correct ordering of memory operations. The operation manager 5630 can access the operation configuration data structure 5624 to determine which queues are grouped together to form a given memory operation. For example, the operation configuration data structure 5624 can include specific correlation counters (or queues), input queues, output queues, and completion queues all grouped together for a particular memory operation. Since each successive memory operation can be assigned a different set of queues, access to the changing queues can be interleaved across subroutines of the memory operation. Knowing all these queues, the operation manager circuit 5630 can interface with the operation queue 5612, the input queue(s) 5616, the completion queue(s) 5620, and the memory subsystem 5310 to initially issue a memory operation to the memory subsystem 5310 when the successive memory operation becomes "executable", and then to complete the memory operation with some acknowledgement from the memory subsystem. This acknowledgement can be, for example, data in response to a load operation command or an acknowledgement of data stored in memory in response to a store operation command.
[0423] Figure 57 is in accordance with an embodiment of the present disclosure, Figure 53A Flowchart of the microarchitecture 5700 of the memory ordering circuit 5305. The memory subsystem 5310 can allow illegal execution of a program where the ordering of memory operations is incorrect due to the semantics of the C language (and other object-oriented programming languages). The microarchitecture 5700 can enhance the ordering of memory operations (the sequence of loads / stores to / from memory) such that the results of the instructions run by the accelerator hardware 5302 are correctly ordered. Multiple local networks 50 are shown to represent a part of the accelerator hardware 5302 coupled to the microarchitecture 5700.
[0424] From an architectural perspective, there are at least two goals: first, to correctly run general-order code, and second, to achieve high performance in the memory operations performed by the microarchitecture 5700. To ensure program correctness, the compiler expresses the correlation between store operations and load operations in a certain way as an array p, which is expressed via correlation tokens, as will be described. To improve performance, the microarchitecture 5700 looks up and issues as many load commands for the arrays as are legal for the program order in parallel.
[0425] In one embodiment, the microarchitecture 5700 can include the above-referenced Figure 56The operation queue 5612, input queue 5616, completion queue 5620, and operation manager circuit 5630, where individual queues may be referred to as channels. The microarchitecture 5700 may also include a plurality of dependency token counters 5714 (e.g., one per input queue), a set of dependency queues 5718 (e.g., one per input queue), an address multiplexer 5732, a store data multiplexer 5734, a completion queue index multiplexer 5736, and a load data multiplexer 5738. In one embodiment, the operation manager circuit 430 may direct these various multiplexers to generate memory commands 5750 (for transmission to the memory subsystem 5310) and receive responses to load commands from the memory subsystem 5310, as will be described.
[0426] As described, the input queue 5616 may include a load address queue 5722, a store address queue 5724, and a store data queue 5726. (The small numbers 0, 1, 2 are channel labels and will be referred to later in Figure 60 and Figure 63A and will be referred to. In various embodiments, these input queues may be multiplied to obtain additional channels for additional parallelization of the memory operation processing. Each dependency queue 5718 may be associated with one of the input queues 5616. More specifically, the dependency queue 5718 labeled B0 may be associated with the load address queue 5722, and the dependency queue labeled B1 may be associated with the store address queue 5724. If additional channels of the input queue 5616 are provided, the dependency queues 5718 may include additional corresponding channels.
[0427] In one embodiment, the completion queue 5620 may include a set of output buffers 5744 and 5746 for receiving load data from the memory subsystem 5310 and the completion queue 5742 to buffer the addresses and data of the load operations according to the indexes maintained by the operation manager circuit 5630. The operation manager circuit 5630 is capable of managing the indexes to ensure the orderly execution of the load operations and identifying the data received in the output buffers 5744 and 5746, which may be moved to the scheduled load operations in the completion queue 5742.
[0428] More specifically, because the memory subsystem 5310 is unordered, but the acceleration hardware 5302 completes operations in order, the microarchitecture 5700 can reorder memory operations by using the completion queue 5742. Three different sub-operations can be performed relative to the completion queue 5742, namely, allocation, enqueueing, and dequeueing. For allocation, the operation manager circuit 5630 can allocate an index to the completion queue 5742 in the next sequential slot of the completion queue. The operation manager circuit can provide this index to the memory subsystem 5310, which can then know the slot to which to write data for a load operation. For enqueueing, the memory subsystem 5310 can write data as an entry to the next sequential slot of the index in the completion queue 5742 (e.g., random access memory (RAM)), thereby setting the status bit of the entry to valid. For dequeueing, the operation manager circuit 5630 can provide the data stored in this next sequential slot to complete the load operation, thereby setting the status bit of the entry to invalid. The invalid entry can then be available for new allocation.
[0429] In one embodiment, the status signal 5648 can represent the status of the input queue 5616, the completion queue 5620, the dependency queue 5718, and the dependency token counter 5714. These statuses can include, for example, input status, output status, and control status, which can represent the presence or absence of dependency tokens associated with an input or output. The input status can include the presence or absence of an address, and the output status can include the presence or absence of a stored value and available completion buffer slots. The dependency token counter 5714 can be a compact representation of a queue and track the number of dependency tokens for any given input queue. If the dependency token counter 5714 saturates, no additional dependency tokens can be generated for new memory operations. Accordingly, the memory ordering circuit can stall scheduling new memory operations until the dependency token counter 5714 becomes unsaturated.
[0430] Still referring to Figure 58 , Figure 58 is a block diagram of an executable determiner circuit 5800 according to an embodiment of the present disclosure. The memory ordering circuit 5305 can employ different kinds of memory operations (e.g., load and store) to establish: ldNo[d,x]result.outN,addr.in64, order.in0,order.out0
[0431] stNo[d,x]addr.in64,data.inN, order.in0,order.out0
[0432] The executable determiner circuit 5800 can be integrated as part of the scheduler circuit 5632, and it can perform logical operations to determine whether a given memory operation is executable and thus ready to be issued to the memory. A memory operation can be run when the queue corresponding to its memory arguments has data and the associated dependency token exists. These memory arguments can include, for example, an input queue identifier 5810 (indicating the channel of the input queue 5616), an output queue identifier 5820 (indicating the channel of the completion queue 5620), a dependency queue identifier 5830 (e.g., which dependency queue or counter should be referenced), and an operation type indicator 5840 (e.g., a load operation or a store operation). Fields (e.g., of a memory request) can contain, for example, in the above format, which stores one or more bits to indicate the use of hazard checking hardware.
[0433] These memory arguments can be queued in the operation queue 5612 and used to schedule the issuance of memory operations associated with incoming addresses and data from the memory and the acceleration hardware 5302. (See Figure 59 .) The incoming status signal 5648 can be logically combined with these identifiers, and then the results can be added (e.g., via the AND gate 5850) to output an executable signal, e.g., which is asserted when the memory operation is executable. The incoming status signal 5648 can include the input status 5812 of the input queue identifier 5810, the output status 5822 of the output queue identifier 5820, and the control status 5832 of the dependency queue identifier 5830 (associated with the dependency token).
[0434] For a load operation, and by way of example, the memory ordering circuit 5305 can issue a load command when the load operation has space for an address (input status) and a load result in the buffered completion queue 5742 (output status). Similarly, the memory ordering circuit 5305 can issue a store command for a store operation when the store operation has an address and a data value (input status). Accordingly, the status signal 5648 can convey the empty (or full) level of the queue to which the status signal pertains. The operation type can then specify whether the logic generates an executable signal based on which address and data should be available.
[0435] To achieve coherent ordering, the scheduler circuit 5632 may extend the memory operation to include a coherence token as emphasized above in the example load and store operations. The control state 5832 may indicate whether the coherence token is available within the coherence queue identified by the coherence queue identifier 5830, which may be either the coherence queue 5718 (for incoming memory operations) or the coherence token counter 5714 (for completed memory operations). In this context, a relevant memory operation requires an additional ordering token to be issued and generates an additional ordering token upon completion of the memory operation, where completion means that the data resulting from the memory operation becomes available for subsequent memory operations of the program.
[0436] In one embodiment, further referring to Figure 57 , the operation manager circuit 5630 may direct the address multiplexer 5732 to select an address argument buffered within the load address queue 5722 or the store address queue 5724, depending on whether the load operation or the store operation is currently scheduled for execution. If it is a store operation, the operation manager circuit 5630 may also direct the store data multiplexer 5734 to select the corresponding data from the store data queue 5726. The operation manager circuit 5630 may also direct the completion queue index multiplexer 5736 to retrieve the load operation entry within the completion queue 5620 (indexed according to queue status and / or program order) to complete the load operation. The operation manager circuit 5630 may also direct the load data multiplexer 5738 to select the data received in the completion queue 5620 from the memory subsystem 5310 for the load operation waiting for completion. Thus, the operation manager circuit 5630 may direct the selection of inputs for forming the memory command 5750 (e.g., a load command or a store command) or the execution circuit 5634 to wait for the completion of the memory operation.
[0437] Figure 59 is a block diagram of an execution circuit 5634 according to an embodiment of the present disclosure, which may include a priority encoder 5906 and a selection circuit 5908 and generates output control lines 5910. In one embodiment, the execution circuit 5634 may access queued memory operations (in the operation queue 5612), which are determined to be executable ( Figure 58 ). The execution circuit 5634 may also receive schedules 5904A, 5904B, 5904C of queued memory operations (which have been queued and are also shown as ready to be issued to the memory). Thus, the priority encoder 5906 may receive the identification codes of the executable memory operations that have been scheduled and run certain rules (or follow a particular logic) to select the memory operation with the highest priority for execution from among those incoming memory operations. The priority encoder 5906 may output a selector signal 5907, which identifies the scheduled memory operation with the highest priority and thus has been selected.
[0438] For example, the priority encoder 5906 can be a circuit (such as a state machine or a simpler converter) that compresses multiple binary inputs into a smaller number of outputs, including possibly only one output. The output of the priority encoder is the binary representation of the original value of zeros starting from the most significant input bit. Thus, in one example, memory operation 0 ("zero"), memory operation one ("1"), and memory operation two ("2") are executable and scheduled, corresponding to 5904A, 5904B, and 5904C respectively. The priority encoder 5906 can be configured to output a selection signal 5907 to the selection circuit 5908, indicating memory operation zero as the memory operation with the highest priority. The selection circuit 5908 can be a multiplexer in one embodiment and is configured to output a selection (such as for memory operation zero) to the control line 5910 in response to the selector signal from the priority encoder 5906 (and indicating the selection of the memory operation with the highest priority) as a control signal. This control signal can go to the address multiplexer 5732, the store data multiplexer 5734, the completion queue index multiplexer 5736, and / or the load data multiplexer 5738, as referred to Figure 57 as described, to load the memory command 5750, which is then issued (sent) to the memory subsystem 5310. The transmission of the memory command can be understood as the issuance of the memory operation to the memory subsystem 5310.
[0439] Figure 60 is a block diagram of an exemplary load operation 6000 in logical and binary form according to an embodiment of the present disclosure. Also referring to Figure 58 , the logical representation of the load operation 6000 can include channel zero ("0") as the input queue identifier 5810 (corresponding to the load address queue 5722) and completion channel one ("1") as the output queue identifier 5820 (corresponding to the output buffer 5744). The correlation queue identifier 5830 can include two identifiers, namely channel B0 for the incoming correlation token (corresponding to the first of the correlation queue 5718) and counter C0 for the outgoing correlation token. The operation type 5840 has an indication of "load", which may also be a numerical indicator to indicate that the memory operation is a load operation. Below, the logical representation of the logical memory operation is a binary representation for illustration purposes, for example, where a load is indicated by "00". Figure 60 The load operation of Figure 62A can be extended to include other configurations (such as store operations (
[0440] )) or other types of memory operations (such as fences). Figure 61A - 61B , Figure 62A - 62B andFigure 63A - 63G Shown using a simplified example. For this example, the following code includes an array p that is accessed via indices i and i + 2:
[0441] for(i){
[0442] temp = p[i];
[0443] p[i + 2] = temp;
[0444] }
[0445] For this example, assume the array p contains 0, 1, 2, 3, 4, 5, 6, and at the end of the loop execution, the array p will contain 0, 1, 0, 1, 0, 1, 0. This code can be transformed by loop unrolling as shown in Figure 61A and Figure 61B Address dependencies are shown by the arrows in Figure 61A , where in each case, a load operation is dependent on a store operation to the same address. For example, for the first of these dependencies, a store (e.g., write) to p[2] needs to occur before a load (e.g., read) from p[2], and for the second of these dependencies, a store to p[3] needs to occur before a load from p[3], and so on. Since the compiler is conservative, the compiler annotates the dependencies between the two memory operations (i.e., load p[i] and store p[i + 2]). Note that reads and writes only sometimes conflict. The microarchitecture 5700 is designed to extract memory-level parallelism, where memory operations can be moved forward when there are no conflicts to the same address. This is especially true for load operations, which exhibit latency in code execution while waiting for a previous dependent store operation to complete. In the example code of Figure 61B , safe reordering is shown by the arrows on the left side of the unrolled code.
[0446] Refer to Figure 62A - 62B and Figure 63A - 63G to discuss the way in which the microarchitecture can perform this reordering. Note that this way is not the best possible, as the microarchitecture 5700 may not send memory commands to memory every cycle. However, with minimal hardware, the microarchitecture supports the dependency flow by running memory operations when the operands (e.g., the address and data for a store or the address for a load) and the dependency tokens are available.
[0447] Figure 62A is a block diagram of exemplary memory arguments for a load operation 6202 and a store operation 6204 in accordance with an embodiment of the present disclosure. These memory arguments, etc., are for Figure 60As described above and will not be elaborated here. However, note that the store operation 6204 does not have an indicator for the output queue identifier because no data is output to the acceleration hardware 5302. Instead, the store address in channel 1 of the input queue 5616 and the data in channel 2 as identified by the input queue identifier memory argument are scheduled in the memory command for transfer to the memory subsystem 5310 to complete the store operation 6204. In addition, both the input channel and the output channel of the correlation queue are implemented using counters. Since, as Figure 61A and Figure 61B shown, the load operation and the store operation are independent, the counter can cycle between the load operation and the store operation within the code flow.
[0448] Figure 62B is a block diagram showing the flow of the load operation and the store operation (e.g., Figure 57 the load operation 6202 and the store 6204 operation) of the memory ordering circuit microarchitecture 5700 according to an embodiment of the present disclosure through Figure 61A . For the sake of simplicity of illustration, not all components are shown, but reference can be made to the additional components shown in Figure 57 . Various ovals indicating "load" for the load operation 6202 and "store" for the store operation 6204 are overlaid on a part of the components of the microarchitecture 5700 as an indication of how the various channels serving as queues for the memory operations are queued and ordered through the microarchitecture 5700.
[0449] Figure 63A 、 Figure 63B 、 Figure 63C 、 Figure 63D 、 Figure 63E 、 Figure 63F 、 Figure 63G and Figure 63H are block diagrams showing the functional flow of the load operation and the store operation of the queues of the microarchitecture according to an embodiment of the present disclosure through Figure 62B . The Figure 61A and Figure 61B of each figure may correspond to the next cycle of processing performed by the microarchitecture 5700. The italicized values are the incoming values (entering the queue), and the bold values are the outgoing values (leaving the queue). All other values with normal font are the reserved values already present in the queue.
[0450] Figure 63AAmong them, the address p[0] enters the load address queue 5722, and the address p[2] enters the store address queue 5724, starting the control flow process. Note that the counter C0 for the dependency input of the load address queue is "1", and the counter C1 for the dependency output is zero. In contrast, the "1" of C0 indicates the dependency output value of the store operation. This indicates the incoming dependency of the load operation of p[0] and the outgoing dependency of the store operation of p[2]. However, these values are not yet active, but will become active in the Figure 63B in this way.
[0451] Figure 63B Among them, the address p[0] is in bold to indicate that it is outgoing in this cycle. The new address p[1] enters the load address queue, and the new address p[3] enters the store address queue. The zero ("0") fetch bit in the completion queue 5742 is also incoming, indicating that any data present for that index entry is invalid. As described, the values of the counters C0 and C1 are shown as incoming and are thus active in this cycle.
[0452] Figure 63C Among them, the outgoing address p[0] then leaves the load address queue, and the new address p[2] enters the load address queue. And the data ("0") enters the completion queue of the address p[0]. The validity bit is set to "1" to indicate that the data in the completion queue is valid. In addition, the new address p[4] enters the store address queue. The value of the counter C0 is shown as outgoing, and the value of the counter C1 is shown as incoming. The "1" value of C1 indicates the incoming dependency of the store operation for the address p[4].
[0453] Note that the address p[2] of the latest load operation and the value to be stored that first needs to be stored by the store operation passing through the address p[2] are at the top of the store address queue. Later, the index entry in the completion queue of the load operation from the address p[2] can remain buffered until the data from the store operation for the address p[2] is completed (see Figure 63F - 63H ).
[0454] Figure 63D Among them, the data ("0") leaves the completion queue of the address p[0], which is thus sent to the acceleration hardware 5302. In addition, the new address p[3] enters the load address queue, and the new address p[5] enters the store address queue. The values of the counters C0 and C1 remain unchanged.
[0455] Figure 63E Among them, the value ("0") of the address p[2] enters the store data queue, while the new address p[4] enters the load address queue, and the new address p[6] enters the store address queue. The counter values of C0 and C1 remain unchanged.
[0456] Figure 63F In this case, both the address p[2] in the data storage queue and the value of the address p[2] in the storage address queue (“0”) are out-of-play values. Similarly, the value of counter C1 is shown as out-of-play while the value of counter C0 remains unchanged. In addition, the new address p[5] enters the load address queue, and the new address p[7] enters the storage address queue.
[0457] Figure 63G In this case, the value (“0”) enters to indicate that the index value within the completion queue 5742 is invalid. The address p[1] is bolded to indicate that it leaves the load address queue, and the new address p[6] enters the load address queue. The new address p[8] also enters the storage address queue. The value of counter C0 enters as “1” corresponding to the in-play relevance of the load operation for address p[6] and the out-of-play relevance of the store operation for address p[8]. The value of counter C1 is “0” at this time and is shown as out-of-play.
[0458] Figure 63H In this case, the data value of “1” enters the completion queue 5742, and the validity bit also enters as “1” indicating that the buffered data is valid. This is the data required to complete the load operation for p[2]. Remember that this data must first be stored at address p[2], which occurs in Figure 63F In this case, the “0” value of counter C0 is out-of-play, and the “1” value of counter C1 is in-play. In addition, the new address p[7] enters the load address queue, and the new address p[9] enters the storage address queue.
[0459] In this embodiment, the process of running Figure 61A and Figure 61B can continue with the bounce correlation tokens between “0” and “1” for the load and store operations. This is due to the close correlation between p[i] and p[i + 2]. Another code with less frequent correlations can generate correlation tokens at a slower rate and thus reset counters C0 and C1 at a slower rate, resulting in the generation of higher-value tokens (corresponding to other semantically separated memory operations).
[0460] Figure 64 is a flowchart of a method 6400 for ordering memory operations between an acceleration hardware and an out-of-order memory subsystem in accordance with an embodiment of the present disclosure. The method 6400 can be performed by a system that can include hardware (e.g., circuitry, dedicated logic, and / or programmable logic), software (e.g., instructions executable on a computer system to perform hardware emulation), or a combination thereof. In an illustrative example, the method 6400 can be performed by the memory ordering circuit 5305 and various sub-components of the memory ordering circuit 5305.
[0461] More specifically, referring to Figure 64, method 6400 may begin with the memory ordering circuit queuing (6410) a memory operation in an operation queue of the memory ordering circuit. The memory operation and control arguments may form the queued memory operation, where the memory operation and control arguments are mapped to certain queues within the memory ordering circuit, as previously described. The memory ordering circuit may operate to issue the memory operation to the memory in association with the acceleration hardware to ensure that the memory operation is completed in program order. Method 6400 may continue with the memory ordering circuit receiving (6420) in a set of input queues from the acceleration hardware an address of memory associated with a second memory operation of the memory operation. In one embodiment, the load address queue of the set of input queues is the channel that receives the address. In another embodiment, the store address queue of the set of input queues is the channel that receives the address. Method 6400 may continue with the memory ordering circuit receiving from the acceleration hardware a correlation token associated with the address, where the correlati...
Claims
1. A processor, comprising: a core having a decoder that decodes instructions into decoded instructions and an execution unit that runs the decoded instructions to perform a first operation; a plurality of processing elements; and an interconnection network between the plurality of processing elements, the interconnection network to receive an input of a data flow graph comprising a plurality of nodes forming a loop construct, wherein the data flow graph is to be overlaid onto the interconnection network and the plurality of processing elements, and wherein each node represents a data flow operator among the plurality of processing elements controlled by an sequencer data flow operator of the plurality of processing elements, wherein the interconnection network and the plurality of processing elements are to perform a second operation when an incoming operand set arrives at each of the data flow operators of the plurality of processing elements and the sequencer data flow operator generates control signals for a first data flow operator representing a first node of the data flow graph and a second data flow operator representing a second node of the data flow graph.
2. The processor according to claim 1, wherein, The first data flow operator is a first processing element among the plurality of processing elements.
3. The processor according to claim 1, wherein, The second data flow operator is a second processing element among the plurality of processing elements.
4. The processor according to claim 1, wherein The first data flow operator representing the first node is a pick operator.
5. The processor according to claim 4, wherein, The second data flow operator representing the second node is a switch operator.
6. The processor according to claim 1, wherein, The sequencer data flow operator generates the control signals for the first data flow operator representing the first node and the second data flow operator representing the second node to perform a loop iteration of the loop construct in a single cycle of the processing elements.
7. The processor according to any one of claims 1-6, wherein, The sequencer data flow operator generates a next set of control signals for the loop iteration when receiving both a base data token and a stride data token.
8. A method, comprising: decoding instructions into decoded instructions using a decoder of a core of a processor; running the decoded instructions using an execution unit of the core of the processor to perform a first operation; receiving an input of a data flow graph comprising a plurality of nodes forming a loop construct; overlaying the data flow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements of the processor, wherein each node represents a data flow operator among the plurality of processing elements controlled by an sequencer data flow operator of the plurality of processing elements; and performing a second operation of the data flow graph using the interconnection network and the plurality of processing elements by an incoming operand set arriving at each of the data flow operators of the plurality of processing elements and the sequencer data flow operator generating control signals for a first data flow operator representing a first node of the data flow graph and a second data flow operator representing a second node of the data flow graph.
9. The method according to claim 8, wherein The first data flow operator is a first processing element among the plurality of processing elements.
10. The method according to claim 8, wherein, The second data flow operator is a second processing element among the plurality of processing elements.
11. The method according to claim 8, wherein, The first data flow operator representing the first node is a pick operator.
12. The method according to claim 11, wherein, The second data flow operator representing the second node is a switch operator.
13. The method according to claim 8, wherein, The sequencer data flow operator generates the control signals for the first data flow operator representing the first node and the second data flow operator representing the second node to execute the loop iteration of the loop construct in a single cycle of the processing element.
14. The method according to any one of claims 8 - 13, further comprising: The sequencer data flow operator generates the next set of control signals for the loop iteration when receiving both a base data token and a span data token.
15. A non-transitory machine-readable medium storing code that, when run by a machine, causes the machine to perform a method comprising the following steps: Decoding an instruction into a decoded instruction using a decoder of a core of a processor; Running the decoded instruction using an execution unit of the core of the processor to perform a first operation; Receiving an input of a data flow graph comprising a plurality of nodes forming a loop construct; Overlaying the data flow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements of the processor, wherein each node is represented as a data flow operator among the plurality of processing elements controlled by a sequencer data flow operator of the plurality of processing elements; And Performing a second operation of the data flow graph using the interconnection network and the plurality of processing elements, wherein a data flow operator of each of the data flow operators reaching the plurality of processing elements by a corresponding set of incoming operands and the sequencer data flow operator generates control signals for a first data flow operator representing a first node of the data flow graph and a second data flow operator representing a second node of the data flow graph.
16. The non-transitory machine-readable medium according to claim 15, wherein, The first data flow operator is a first processing element among the plurality of processing elements.
17. The non-transitory machine-readable medium according to claim 15, wherein, The second data flow operator is a second processing element among the plurality of processing elements.
18. The non-transitory machine-readable medium according to claim 15, wherein, The first data flow operator representing the first node is a pick operator.
19. The non-transitory machine-readable medium of claim 18, wherein, The second data flow operator representing the second node is a switch operator.
20. The non-transitory machine-readable medium according to claim 15, wherein, The sequencer data flow operator generates the control signals for the first data flow operator representing the first node and the second data flow operator representing the second node to execute the loop iteration of the loop construct in a single cycle of the processing element.
21. The non-transitory machine-readable medium according to any one of claims 15-20, wherein, The method further includes: the sequencer data flow operator generating the next set of control signals for the loop iteration when receiving both a base data token and a span data token.
22. A processor comprising: A core having a decoder that decodes an instruction into a decoded instruction and an execution unit that runs the decoded instruction to perform a first operation; And A component for receiving an input of a data flow graph comprising a plurality of nodes forming a loop construct, wherein the data flow graph is to be overlaid onto the component, and wherein each node is represented as a data flow operator controlled by a sequencer data flow operator, wherein the component is to perform a second operation when a set of incoming operands reaches the component and the sequencer data flow operator generates control signals for a first data flow operator representing a first node of the data flow graph and a second data flow operator representing a second node of the data flow graph.
23. The processor according to claim 22, wherein, The first data flow operator is a first processing element among a plurality of processing elements.
24. The processor according to claim 22, wherein, The second data flow operator is the second processing element among the plurality of processing elements.
25. The processor according to claim 22, wherein, The first data flow operator representing the first node is a pick operator.
26. The processor according to claim 25, wherein, The second data flow operator representing the second node is a switch operator.
27. The processor according to claim 23, wherein, The sequencer data flow operator generates the control signals for the first data flow operator representing the first node and the second data flow operator representing the second node to perform a loop iteration of the loop construct in a single cycle of the plurality of processing elements.
28. The processor according to any one of claims 22-27, further comprising the sequencer data flow operator generating a next set of control signals for loop iterations when receiving both a base data token and a span data token.
29. A computer program product storing machine-executable code that, when run by a machine, causes the machine to perform the method according to any one of claims 8-14.
Citation Information
Patent Citations
Programmable logic integrated circuit for digital algorithmic functions
US20080218203A1