Processor, method and system for configurable spatial accelerator with memory system performance, power reduction and atomic support features
By using a configurable space accelerator (CSA), which includes a processing element array with lightweight backpressure network connections, solving the problem that traditional processors are difficult to achieve high throughput and low energy consumption in high-performance computing, achieving significant improvements in high energy efficiency and performance.
Patent Information
- Application Number
- CN201810685952.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-07-01
- Filing Date
- 2018-06-28
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2038-06-28
AI Technical Summary
Existing processors are difficult to achieve both high throughput and low energy consumption in high-performance computing. The traditional von Neumann architecture has encountered difficulties in improving program execution performance and energy efficiency, especially under the goal of 10 billion secondary computing.
The configurable space accelerator (CSA) is adopted, which includes an array of processing components with a lightweight backpressure network connection, which reduces control overhead and achieves efficient parallel computing by directly executing data flow graphs.
High-performance computing is achieved on low-complexity and energy-efficient processing components, significantly improving the energy efficiency and performance of computing-intensive tasks, supporting programming models of mainstream HPC programs, and reducing energy consumption.
Smart Images

Figure CN109213523B_ABST
Abstract
Description
[0001] Statement Regarding Federally Funded Research and Development
[0002] This invention was made with Government support under Contract No. H98230A-13-D-0124 awarded by the Department of Defense. The Government has certain rights in this invention. Technical Field
[0003] The present disclosure relates generally to electronics, and more particularly, embodiments of the present disclosure relate to configurable spatial accelerators. Background Art
[0004] A processor or set of processors executes instructions from an instruction set (e.g., an instruction set architecture (ISA)). An instruction set is the part of a computer's architecture that involves programming and generally includes native data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I / O). It should be noted that the term "instruction" is generally used herein to refer to either macroinstructions (e.g., instructions provided to a processor for execution) or microinstructions (e.g., instructions generated by a processor's decoder decoding a macroinstruction). BRIEF DESCRIPTION OF THE DRAWINGS
[0005] The present disclosure is illustrated by way of example and not limitation in the accompanying figures and in which like reference numerals indicate similar elements and in which:
[0006] Figure 1 An accelerator sheet according to an embodiment of the present disclosure is illustrated.
[0007] Figure 2 Illustrated is a hardware processor coupled to a memory according to an embodiment of the present disclosure.
[0008] Figure 3A A program source according to an embodiment of the present disclosure is illustrated.
[0009] Figure 3B FIG. 1 illustrates an embodiment according to the present disclosure Figure 3A The data flow graph of the program source in .
[0010] Figure 3C The accelerator according to an embodiment of the present disclosure is shown. Figure 3B A data flow graph of multiple processing elements.
[0011] Figure 4 Illustrate an example execution of a data flow graph according to an embodiment of the present disclosure.
[0012] Figure 5 A program source according to an embodiment of the present disclosure is illustrated.
[0013] Figure 6An accelerator slice including an array of processing elements is illustrated according to an embodiment of the present disclosure.
[0014] Figure 7A Illustrated is a configurable data path network according to an embodiment of the present disclosure.
[0015] Figure 7B Illustrated is a configurable flow control path network according to an embodiment of the present disclosure.
[0016] Figure 8 Illustrated is a hardware processor slice including an accelerator according to an embodiment of the present disclosure.
[0017] Figure 9 Illustrated are processing elements according to an embodiment of the present disclosure.
[0018] Figure 10 Illustrated is a request address file (RAF) circuit according to an embodiment of the present disclosure.
[0019] Figure 11A Illustrated are multiple request address file (RAF) circuits coupled between multiple accelerator slices and multiple cache banks according to an embodiment of the present disclosure.
[0020] Figures 11B-11D Regarding a memory subsystem for efficient memory fetching according to an embodiment of the present disclosure.
[0021] Figure 11E-11H Regarding a storage management unit for processing storage to a memory in a spatial computing structure according to an embodiment of the present disclosure.
[0022] Figure 11I-11J FIG. 1 illustrates an architecture for atomic operations according to an embodiment of the present disclosure.
[0023] Figure 12 The diagram illustrates a floating-point multiplier partitioned into three regions (a result region, three potential carry regions, and a gate region) according to an embodiment of the present disclosure.
[0024] Figure 13 An in-flight configuration of an accelerator having multiple processing elements according to an embodiment of the present disclosure is illustrated.
[0025] Figure 14 Illustrate a snapshot of an in-flight pipelined extraction according to an embodiment of the present disclosure.
[0026] Figure 15 FIG. 4 illustrates a compilation tool chain for an accelerator according to an embodiment of the present disclosure.
[0027] Figure 16 FIG. 1 illustrates a compiler for an accelerator according to an embodiment of the present disclosure.
[0028] Figure 17A Serialized assembly code according to an embodiment of the present disclosure is illustrated.
[0029] Figure 17B FIG. 1 shows an embodiment of the present disclosure. Figure 17A Data flow assembly code for serialization assembly code.
[0030] Figure 17C FIG. 1 illustrates an embodiment of the present disclosure for an accelerator. Figure 17B The data flow diagram of the data flow assembly code.
[0031] Figure 18A FIG. 1 illustrates C source code according to an embodiment of the present disclosure.
[0032] Figure 18B FIG. 1 shows an embodiment of the present disclosure. Figure 18A The data flow of C source code is compiled into assembly code.
[0033] Figure 18C FIG. 1 illustrates an embodiment of the present disclosure for an accelerator. Figure 18B The data flow diagram of the data flow assembly code.
[0034] Figure 19A FIG. 1 illustrates C source code according to an embodiment of the present disclosure.
[0035] Figure 19B FIG. 1 shows an embodiment of the present disclosure. Figure 19A The data flow of C source code is compiled into assembly code.
[0036] Figure 19C FIG. 1 illustrates an embodiment of the present disclosure for an accelerator. Figure 19B The data flow diagram of the data flow assembly code.
[0037] Figure 20A A flowchart according to an embodiment of the present disclosure is illustrated.
[0038] Figure 20B A flowchart according to an embodiment of the present disclosure is illustrated.
[0039] Figure 21 Graph illustrating throughput versus energy per operation according to an embodiment of the present disclosure.
[0040] Figure 22 An accelerator slice according to an embodiment of the present disclosure is illustrated, the accelerator slice including an array of processing elements and a local configuration controller.
[0041] Figures 23A-23C Illustrated is a local configuration controller configuring a data path network according to an embodiment of the present disclosure.
[0042] Figure 24 A configuration controller according to an embodiment of the present disclosure is illustrated.
[0043] Figure 25 An accelerator slice is illustrated that includes an array of processing elements, a configuration cache, and a local configuration controller according to an embodiment of the present disclosure.
[0044] Figure 26 An accelerator slice including an array of processing elements and a configuration and exception handling controller with reconfiguration circuitry is illustrated according to an embodiment of the present disclosure.
[0045] Figure 27 A reconfiguration circuit according to an embodiment of the present disclosure is illustrated.
[0046] Figure 28 An accelerator slice including an array of processing elements and a configuration and exception handling controller with reconfiguration circuitry is illustrated according to an embodiment of the present disclosure.
[0047] Figure 29 An accelerator slice according to an embodiment of the present disclosure is illustrated, the accelerator slice including an array of processing elements and a mezzanine anomaly aggregator coupled to a chip-level anomaly aggregator.
[0048] Figure 30 Illustrated is a processing element with an exception generator according to an embodiment of the present disclosure.
[0049] Figure 31 An accelerator slice is illustrated, comprising an array of processing elements and a local fetch controller, according to an embodiment of the present disclosure.
[0050] Figures 32A-32C Illustrated is a local extraction controller configuring a data path network according to an embodiment of the present disclosure.
[0051] Figure 33 An extraction controller according to an embodiment of the present disclosure is shown.
[0052] Figure 34 A flowchart according to an embodiment of the present disclosure is shown.
[0053] Figure 35 A flowchart according to an embodiment of the present disclosure is shown.
[0054] Figure 36A is a block diagram illustrating a general vector friendly instruction format and class A instruction templates thereof according to an embodiment of the present disclosure.
[0055] Figure 36B is a block diagram illustrating a generic vector friendly instruction format and its class B instruction template according to an embodiment of the present disclosure.
[0056] Figure 37A is a diagram illustrating a method for Figure 36A and Figure 36B Block diagram of the fields of the generic vector friendly instruction format in .
[0057] Figure 37B is a diagram illustrating a structure of a complete opcode field according to one embodiment of the present disclosure. Figure 37A Block diagram of the fields of the specialized vector friendly instruction format in .
[0058] Figure 37C FIG. 1 is a diagram illustrating a register index field according to an embodiment of the present disclosure. Figure 37A Block diagram of the fields of the specialized vector friendly instruction format in .
[0059] Figure 37D FIG. 3 is a diagram illustrating the structure of the extended operation field 3650 according to one embodiment of the present disclosure. Figure 37A Block diagram of the fields of the specialized vector friendly instruction format in .
[0060] Figure 38 is a block diagram of a register architecture according to one embodiment of the present disclosure.
[0061] Figure 39A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register-renaming, out-of-order issue / execution pipeline according to embodiments of the present disclosure.
[0062] Figure 39B is a block diagram illustrating both an exemplary embodiment of an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor according to embodiments of the present disclosure.
[0063] Figure 40A is a block diagram of a single processor core and its connections to the on-die interconnect network and its local subset of the Level 2 (L2) cache according to an embodiment of the present disclosure.
[0064] Figure 40B According to an embodiment of the present disclosure Figure 40A An expanded view of a portion of a processor core in FIG.
[0065] Figure 41 is a block diagram of a processor that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to an embodiment of the present disclosure.
[0066] Figure 42 is a block diagram of a system according to one embodiment of the present disclosure.
[0067] Figure 43is a block diagram of a more specific exemplary system according to an embodiment of the present disclosure.
[0068] Figure 44 Shown is a block diagram of a second, more specific exemplary system according to an embodiment of the present disclosure.
[0069] Figure 45 Shown is a block diagram of a system on a chip (SoC) according to an embodiment of the present disclosure.
[0070] Figure 46 is a block diagram illustrating converting binary instructions in a source instruction set into binary instructions in a target instruction set using a software instruction converter according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0071] In the following description, a number of specific details are set forth. However, it should be understood that embodiments of the present disclosure may be implemented without these specific details. In other instances, well-known circuits, structures, and technologies are not shown in detail to avoid obscuring the understanding of this description.
[0072] References in the specification to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment necessarily includes that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an embodiment, it is understood that it is within the knowledge of those skilled in the art to be able to affect such feature, structure, or characteristic in conjunction with other embodiments, whether or not explicitly described.
[0073] A processor (e.g., having one or more cores) may execute instructions (e.g., an instruction thread) to operate on data, for example, to perform arithmetic, logical, or other functions. For example, software may request an operation, and a hardware processor (e.g., one or more cores of the hardware processor) may perform the operation in response to the request. A non-limiting example of an operation is a blend operation that inputs multiple vector elements and outputs a vector having the blended multiple elements. In some embodiments, multiple operations are performed using the execution of a single instruction.
[0074] For example, exascale performance as defined by the U.S. Department of Energy may require system-level floating-point performance exceeding 10 Mbps within a given (e.g., 20 MW) power budget. 18
[0014] Certain embodiments herein relate to a configurable spatial accelerator (CSA) for high performance computing (HPC). Certain embodiments of the CSA are directed to direct execution of data flow graphs to achieve a computationally intensive yet energy-efficient spatial microarchitecture that far exceeds conventional roadmap architectures. The following includes a description of the architectural concepts of embodiments of the CSA and certain features thereof. As with any revolutionary architecture, programmability can be a risk. To mitigate this issue, embodiments of the CSA architecture have been co-designed with a compilation toolchain (also discussed below).
[0075] 1. Introduction
[0076] Exascale computing goals may require massive system-level floating-point performance (e.g., 1 ExaFLOP) within a dramatic power budget (e.g., 20MW). However, it has become difficult to simultaneously improve the performance and energy efficiency of program execution using classical von Neumann architectures: out-of-order scheduling, simultaneous multi-threaded operation, complex register files, and other structures provide performance, but at a high energy cost. Certain embodiments herein achieve both performance and energy requirements. Exascale computing power-performance goals may require both high throughput and low energy consumption for each operation. Certain embodiments herein provide this by providing a large number of low-complexity, energy-efficient processing (e.g., compute) elements that significantly eliminate the control overhead of previous processor designs. Guided by this observation, certain embodiments herein include a configurable spatial accelerator (CSA) that includes, for example, an array of processing elements (PEs) connected by a lightweight back-pressure network. An example of a CSA slice is in Figure 1 Some embodiments of processing (e.g., compute) elements are dataflow operators, e.g., multiple dataflow operators that only process input data when both (i) the input data has arrived at the dataflow operator and (ii) there is space available to store output data (e.g., otherwise no processing is occurring). Some embodiments (e.g., accelerators or CSAs) do not utilize triggered instructions.
[0077] Figure 1 The accelerator slice 100 according to an embodiment of the present disclosure is shown. The accelerator slice 100 may be part of a larger slice. The accelerator slice 100 executes one or more data flow graphs. Figure 1Generally can refer to the explicit parallel program description that appears at compile time for serialized code. Certain embodiments herein (e.g., CSA) allow data flow graphs to be configured directly onto CSA arrays, e.g., rather than being transformed into a serialized instruction stream. The deviation of data flow graphs from the serialized compilation flow allows embodiments of CSA to support familiar programming models and directly execute existing high performance computing (HPC) code (e.g., without using work tables). CSA processing elements (PEs) can be energy efficient. In Figure 1 , the memory interface 102 may be coupled to a memory (eg, Figure 2 Memory 202 in the chip allows the accelerator slice 100 to access (e.g., load and / or store) data to (e.g., off-die) memory. The depicted accelerator slice 100 is composed of a heterogeneous array of several types of PEs coupled together via an interconnect network 104. The accelerator slice 100 may include one or more of the following: integer arithmetic PEs, floating-point arithmetic PEs, communication circuitry, and fabric storage. A dataflow graph (e.g., a compiled dataflow graph) may be overlaid on the accelerator slice 100 for execution. In one embodiment, for a particular dataflow graph, each PE only handles one or two operations in the graph. The PE array may be heterogeneous, for example, such that no PE supports the full CSA dataflow architecture and / or one or more PEs are programmed (e.g., customized) to perform only a few, but highly efficient, operations. Certain embodiments herein thus implement accelerators having an array of processing elements that is more computationally intensive than roadmap architectures, and achieves approximately orders of magnitude gains in energy efficiency and performance relative to existing HPC offerings.
[0078] Performance gains can come from (e.g., dense) parallel execution within a CSA, where each PE can execute concurrently if input data is available. Efficiency gains can come from the efficiency of each PE, for example, where the operations (e.g., behavior) of each PE are fixed once for each configuration (e.g., mapping) step, and execution occurs when local data arrives at the PE (e.g., without regard to other architectural activity). In some embodiments, the PEs are dataflow operators (e.g., each PE is a single dataflow operator), e.g., a dataflow operator that only processes input data when both (i) input data has arrived at the dataflow operator and (ii) there is space available to store output data (e.g., otherwise no processing is occurring). These properties enable embodiments to provide paradigm-shifting levels of performance and significant energy efficiency improvements across a wide class of existing single-stream and parallel programs (e.g., all programs), while maintaining a familiar HPC programming model. Certain embodiments of the CSA may be targeted at HPC, where floating-point energy efficiency is paramount. Certain embodiments of the CSA not only achieve impressive performance improvements and energy reductions, but also transfer these gains to existing HPC programs written in mainstream HPC languages and for mainstream HPC frameworks. Certain embodiments of the CSA architecture (e.g., with compilation in mind) provide several extensions in direct support for the internal representation of control dataflow generated by modern compilers. Certain embodiments herein relate to CSA dataflow compilers (e.g., which can accept C, C++, and Fortran programming languages) for targeting the CSA architecture.
[0079] Section 2 below discloses an embodiment of the CSA architecture. Specifically, it discloses a novel embodiment for integrating memory within a dataflow execution model. Section 3 explores the microarchitectural details of an embodiment of the CSA. In one embodiment, the primary purpose of the CSA is to support compiler-generated programs. Section 4 below examines an embodiment of a CSA compilation toolchain. In Section 5, the advantages of an embodiment of the CSA are compared to other architectures in the execution of compiled code. Finally, Section 6 discusses the performance of an embodiment of the CSA microarchitecture, Section 7 discusses further CSA details, and Section 8 provides a summary.
[0080] 2. Architecture
[0081] Certain embodiments of CSAs aim to execute programs (e.g., programs produced by a compiler) quickly and efficiently. Certain embodiments of the CSA architecture provide programming abstractions that support the needs of compiler technology and programming paradigms. Embodiments of CSAs implement dataflow graphs, e.g., a representation of a program that is much like the compiler's own internal representation (IR) of a compiled program. In this model, a program is represented as a dataflow graph consisting of nodes (e.g., vertices) drawn from a set of architecturally defined dataflow operators (e.g., covering both computational and control operations) and edges representing the transfer of data between the dataflow operators. Execution can proceed by injecting dataflow tokens (e.g., as data values or dataflow tokens representing data values) into the dataflow graph. Tokens can flow between them and can be transformed at each node (e.g., vertex) to form, for example, a complete computation. In Figure 3A - a sample data flow graph and its deviation from the high-level source code is shown in FIG3C , and Figure 5 An example of the execution of a data flow graph is shown.
[0082] An embodiment of a CSA configures data flow graph execution by providing exactly the data flow graph execution support required by the compiler. In one embodiment, the CSA is an accelerator (e.g., Figure 2 accelerators in ), and it does not seek to provide Figure 2
[0014] The CSA utilizes some of the necessary but infrequently used mechanisms (such as system calls) available on the cores in the CSA. Thus, in this embodiment, the CSA can execute much code, but not all code. In exchange, the CSA gains significant performance and energy advantages. To achieve acceleration of code written in commonly used serializing languages, the embodiments herein also introduce several novel architectural features to assist the compiler. One particular novelty is the CSA's handling of memory, a topic that has previously been overlooked or poorly addressed. Embodiments of the CSA are also unique in using dataflow operators (e.g., as opposed to lookup tables (LUTs)) as their basic architectural interface.
[0083] Figure 2The diagram illustrates a hardware processor 200 coupled to (e.g., connected to) memory 202 in accordance with an embodiment of the present disclosure. In one embodiment, the hardware processor 200 and memory 202 are a computing system 201. In certain embodiments, one or more of the accelerators are CSAs in accordance with the present disclosure. In certain embodiments, one or more of the cores in the processor are those disclosed herein. The hardware processor 200 (e.g., each of its cores) may include a hardware decoder (e.g., a decode unit) and a hardware execution unit. The hardware processor 200 may include registers. Note that the figures herein may not depict all data communication couplings (e.g., connections). Those skilled in the art will appreciate that this is to avoid obscuring certain details in the figures. Note that a bidirectional arrow in a figure may not require bidirectional communication; for example, it may indicate unidirectional communication (e.g., to or from that component or device). In certain embodiments herein, any or all combinations of communication paths may be utilized. The depicted hardware processor 200 includes multiple cores (0 to N, where N may be 1 or greater) and multiple hardware accelerators (0 to M, where M may be 1 or greater). Hardware processor 200 (e.g., its accelerator(s) and / or core(s)) may be coupled to memory 202 (e.g., a data storage device). A hardware decoder (e.g., of a core) may receive (e.g., a single) instruction (e.g., a macroinstruction) and decode the instruction into, for example, microinstructions and / or micro-operations. A hardware execution unit (e.g., of a core) may execute the decoded instruction (e.g., a macroinstruction) to perform one or more operations. Returning to the embodiment of the CSA, dataflow operators are discussed below.
[0084] 2.1 Data Flow Operators
[0085] A key architectural interface of embodiments of an accelerator (e.g., a CSA) is dataflow operators, e.g., as direct representations of nodes in a dataflow graph. From an operational perspective, dataflow operators behave in a streaming or data-driven manner. A dataflow operator can be executed as soon as its incoming operands are available. CSA dataflow execution can rely (e.g., solely) on highly localized state, leading to a highly scalable architecture with, for example, a distributed asynchronous execution model. Dataflow operators can include arithmetic dataflow operators, e.g., one or more of the following: floating-point addition and multiplication; integer addition, subtraction, and multiplication; various forms of comparisons, logical operators, and shifts. However, embodiments of a CSA may also include a rich set of control operators that assist in the management of dataflow tokens in a program graph. Examples of these control operators include a "pick" operator (e.g., which multiplexes two or more logical input channels into a single output channel) and a "switch" operator (e.g., which operates as a channel demultiplexer) (e.g., outputting a single channel from two or more logical input channels). These operators enable the compiler to implement control paradigms (such as conditional expressions). Certain embodiments of the CSA may include a limited set of dataflow operators (e.g., relative to a small number of operations) to achieve a dense and energy-efficient PE microarchitecture. Certain embodiments may include dataflow operators for complex operations commonly found in HPC code. The CSA dataflow operator architecture is highly adaptable to deployment-specific extensions. For example, more complex mathematical dataflow operators (e.g., trigonometric functions) may be included in certain embodiments to accelerate certain math-intensive HPC workloads. Similarly, neural network-tuned extensions may include dataflow operators for vectorized, low-precision arithmetic.
[0086] Figure 3A The program source according to an embodiment of the present disclosure is shown in FIG. The program source code includes a multiplication function (func). Figure 3B FIG. 1 shows an embodiment of the present disclosure. Figure 3A 3. Dataflow graph 300 for a program source. Dataflow graph 300 includes a pick node 304, a switch node 306, and a multiplication node 308. Buffers may optionally be included along one or more of the communication paths. The depicted dataflow graph 300 may perform the following operations: select input X using pick node 304, multiply X by Y (e.g., multiplication node 308), and then output the output result from the left side of switch node 306. Figure 3C3B . More specifically, dataflow graph 300 is overlaid onto an array of processing elements 301 (and, for example, a plurality of (e.g., interconnect) networks therebetween) such that, for example, each node in dataflow graph 300 is represented as a dataflow operator in the array of processing elements 301. In one embodiment, one or more of the processing elements in the array of processing elements 301 are used to access memory via memory interface 302. In one embodiment, pick node 304 of dataflow graph 300 corresponds to pick operator 304A (e.g., represented by pick operator 304A), switch node 306 of dataflow graph 300 corresponds to switch operator 306A (e.g., represented by switch operator 306A), and multiplier node 308 of dataflow graph 300 corresponds to multiplier operator 308A (e.g., represented by multiplier operator 308A). Another processing element and / or a flow control path network may provide a control signal (eg, a control token) to the pick operator 304A and the switch operator 306A to perform Figure 3A In one embodiment, the array of processing elements 301 is configured to perform operations before execution begins. Figure 3B In one embodiment, the compiler performs the data flow diagram 300 from FIG. 3A to Figure 3B In one embodiment, the input of a dataflow graph node into the array of processing elements logically embeds the dataflow graph into the array of processing elements (e.g., as discussed further below) such that the input / output paths are configured to produce the desired results.
[0087] 2.2 Waiting time insensitive channels
[0088] Communication arcs (arcs) are the second main component of a dataflow graph. Certain embodiments of a CSA describe these arcs as latency-insensitive channels, such as ordered, backpressured (e.g., output is not generated or sent until there is space to store it), point-to-point communication channels. Like dataflow operators, latency-insensitive channels are fundamentally asynchronous, giving the freedom to combine many types of networks to implement channels for a particular graph. Latency-insensitive channels can have arbitrarily long latencies and still faithfully implement the CSA architecture. However, in certain embodiments, there are strong performance and energy incentives to minimize latency. Section 3.2 herein discloses a network microarchitecture in which dataflow graph channels are implemented in a pipelined fashion with a latency of no more than one cycle. Embodiments of latency-insensitive channels provide a key abstraction layer that can be leveraged with the CSA architecture to provide many runtime services to application programmers. For example, a CSA can leverage latency-insensitive channels when implementing CSA configuration (loading a program onto a CSA array).
[0089] Figure 4 4. An example execution of a data flow diagram 400 according to an embodiment of the present disclosure is illustrated. In step 1, an input value (e.g., Figure 3B 1 of X in and for Figure 3B 2 of Y in ) can be loaded into the data flow graph 400 to perform a 1*2 multiplication operation. One or more of the data input values can be static (e.g., constant) in the operation (e.g., referring to Figure 3B , X is 1 and Y is 2) or is updated during the operation. In step 2, the processing element or other circuit (e.g., on the flow control path network) outputs 0 to the control input (e.g., mux control signal) of the pick node 404 (e.g., to obtain "1" from the port as a source to its output), and outputs 0 to control the input (e.g., mux control signal) of the switch node 406 (e.g., to cause its input to be provided outwardly from port "0" to a destination (e.g., a downstream processing element)). In step 3, the data value 1 is output from the pick node 404 (and, for example, consumes its control signal "0" at the pick node 404) to the multiplier node 408 to be multiplied with the data value 2 in step 4. In step 4, the output of the multiplier node 408 reaches the switch node 406, for example, which causes the switch node 406 to consume the control signal "0" to output the value 2 from the port "0" of the switch node 406 in step 5. The operation is then completed. The CSA can therefore be programmed accordingly so that the corresponding data flow operator of each node performs Figure 4 Although the execution is serialized in this example, in principle all data flow operations can be performed in parallel. Figure 4In one embodiment, the downstream processing element is operable to send a ready signal (or not send a ready signal) to the switch 406 (e.g., on a flow control path network) to stall output from the switch 406 until the downstream processing element is ready for output (e.g., has storage space).
[0090] 2.3 Memory
[0091] Dataflow architectures generally focus on communication and data manipulation, with less attention paid to state. However, enabling real-world software, especially programs written in traditional sequential languages, requires significant attention to interfacing with memory. Certain embodiments of CSAs use architectural memory operations as their primary interface to (e.g., large) stateful stores. From a dataflow graph perspective, memory operations are similar to other dataflow operations, except that they have the side effect of updating shared storage. Specifically, the memory operations of certain embodiments herein have the same semantics as every other dataflow operator; for example, they "execute" when their operands (e.g., addresses) are available and a response is generated after some latency. Certain embodiments herein explicitly decouple operand inputs from result outputs, making memory operators inherently pipelined and capable of generating many simultaneous pending requests, making them highly adaptable to the latency and bandwidth characteristics of a memory subsystem, for example. Embodiments of CSAs provide basic memory operations such as loads and stores, which take an address channel and fill a response channel with the value corresponding to that address. Embodiments of CSAs also provide more advanced operations, such as in-memory atomics and consistency operators. These operations can have semantics similar to their von Neumann equivalents. Embodiments of CSA can accelerate existing programs written in sequential languages such as C and Fortran. Support for these language models results in addressing program memory order, e.g., serial ordering of memory operations typically specified by these languages.
[0092] Figure 5 A program source (e.g., C code) 500 according to an embodiment of the present disclosure is shown. According to the memory semantics of the C programming language, memory copy (memcpy) should be serialized. However, if it is known that array A and array B are disjoint, memcpy can be parallelized using an embodiment of CSA. Figure 5The problem of program order is further illustrated. In general, compilers cannot prove that array A is different from array B, for example, whether for the same index value or for different index values across loop bodies. This is called pointer or memory aliasing. Because compilers are used to generate statically correct code, they are often forced to serialize memory accesses. Typically, compilers for serialized von Neumann architectures use instruction ordering as a natural means of enforcing program order. However, embodiments of CSA do not have the concept of instruction ordering defined by the program counter or instruction-based program ordering. In some embodiments, incoming dependency tokens (e.g., which do not contain architecturally visible information) are like all other dataflow tokens, and memory operations cannot be executed until they receive a dependency token. In some embodiments, once the operation of a memory operation is visible to logically subsequent dependent memory operations, these memory operations generate outgoing dependency tokens. In some embodiments, dependency tokens are similar to other dataflow tokens in a dataflow graph. For example, because memory operations occur in a conditional context, dependency tokens can also be manipulated using the control operators described in Section 2.1 (e.g., like any other token). Dependency tokens may have the effect of serializing memory accesses, thereby providing a compiler with a means to architecturally define the order of memory accesses, for example.
[0093] 2.4 Runtime Services
[0094] The primary architectural considerations for embodiments of the CSA concern the actual execution of user-level programs, but it is also desirable to provide several supporting mechanisms that underpin this execution. Chief among these are configuration (where dataflow graphs are loaded into the CSA), fetching (where the state of the execution graph is moved to memory), and exceptions (where mathematical, soft, and other types of errors in the fabric can be detected and handled by external entities). Section 3.6 below discusses the latency-insensitive dataflow architecture properties of embodiments of the CSA for efficient, highly pipelined implementations of these functions. Conceptually, configuration loads the state of a dataflow graph (e.g., typically from memory) into the interconnect and processing elements (e.g., fabrics). During this step, all fabrics in the CSA can be loaded with the new dataflow graph, and any dataflow tokens surviving in that graph, for example, as a result of a context switch. The latency-insensitive semantics of the CSA permit distributed asynchronous initialization of fabrics; for example, PEs can begin execution immediately upon configuration. Unconfigured PEs can backpressure their channels until they are configured, thereby, for example, preventing communication between configured and unconfigured elements. The CAS configuration can be partitioned into privileged and user-level states. Such two-level partitioning allows the main configuration of the fabric to occur without invoking the operating system. In one embodiment of extraction, a logical diagram of a data flow graph is captured and committed to memory, e.g., including all live control and data flow tokens and states in the graph.
[0095] Extraction can also play a role in providing reliability assurance by creating structural checkpoints. Exceptions in a CSA can generally be caused by the same events that cause exceptions in a processor, such as illegal operator arguments or reliability, availability, and durability (RAS) events. In some embodiments, exceptions are detected at the level of the dataflow operator (e.g., checking argument values) or through a modular arithmetic scheme. Upon detecting an exception, the dataflow operator (e.g., a circuit) can stop and emit an exception message that, for example, contains both an operation identifier and some details about the nature of the problem that occurred. In some embodiments, the dataflow operator will remain stopped until it has been reconfigured. Subsequently, the exception message can be passed to the associated processor (e.g., a core) for servicing (e.g., which may include an extraction graph for software analysis).
[0096] 2.5 Chip-Level Architecture
[0097] Embodiments of a CSA computer architecture (eg, for HPC and datacenter use) are sharded. Figure 6 and Figure 8 Shows the slice-level deployment of CSA. Figure 8A full-slice implementation of a CSA is shown, which can be, for example, an accelerator for a processor with a core. A major advantage of this architecture can be reduced design risk, for example, allowing the CSA to be completely decoupled from the core at fabrication time. In addition to allowing better component reuse, this can also allow components (like the CSA cache) to be considered CSA-only, rather than, for example, needing to incorporate more stringent latency requirements for the core. Ultimately, separate slices can allow the CSA to be integrated with either small or large cores. One embodiment of the CSA captures most vector-parallel workloads, allowing most vector-based workloads to run directly on the CSA, but in some embodiments, vector-based instructions in the core can be included, for example, to support legacy binaries.
[0098] 3. Microarchitecture
[0099] In one embodiment, the CSA microarchitecture aims to provide a high-quality implementation of each dataflow operator specified by the CAS architecture. Embodiments of the CSA microarchitecture provide that each processing element of the microarchitecture corresponds to approximately one node (e.g., entity) in the architecture's dataflow graph. In certain embodiments, this results in an architectural element that is both compact, resulting in a dense computational array, and energy-efficient, for example, where processing elements (PEs) are both simple and highly unmultiplexed (e.g., configured (e.g., programmed) to perform a single dataflow operation). To further reduce energy and implementation area, the CSA may include a configurable heterogeneous architecture where each PE implements only a subset of the dataflow operators. Peripheral and support subsystems (such as CSA caches) may be provided to support the distributed parallelism inherent in the main CSA processing structure itself. Implementations of the CSA microarchitecture may implement dataflow and latency-insensitive communication abstractions present in the architecture. In certain embodiments, there is (e.g., substantially) a one-to-one correspondence between nodes in the compiler-generated graph and dataflow operators (e.g., dataflow operator compute elements) in the CSA.
[0100] The following is a discussion of an example CSA, followed by a more detailed discussion of the microarchitecture. Certain embodiments herein provide a CSA that allows for easy compilation, for example, in contrast to existing FPGA compilers, which handle a small subset of programming languages (e.g., C or C++) and can take many hours to compile even small programs.
[0101] Certain embodiments of the CSA architecture allow for heterogeneous coarse-grained operations such as double-precision floating point. Programs can be expressed in terms of less coarse-grained operations, for example, allowing the disclosed compiler to run faster than traditional spatial compilers. Certain embodiments include an architecture with new processing elements to support serialization concepts such as program-ordered memory access. Certain embodiments implement hardware to support coarse-grained dataflow-type communication channels. This communication model is abstract and closely resembles the control dataflow representation used by compilers. Certain embodiments herein include a network implementation that supports single-cycle latency communication, for example, utilizing (e.g., small) PEs that support a single control dataflow operation. In certain embodiments, this not only improves energy efficiency and performance, but also simplifies compilation because the compiler performs a one-to-one mapping between high-level dataflow constructs and structures. Certain embodiments herein therefore simplify the task of compiling existing (e.g., C, C++, or Fortran) programs to CSA (e.g., structures).
[0102] Energy efficiency can be a primary consideration in modern computer systems. Certain embodiments herein provide new models of energy-efficient spatial architectures. In certain embodiments, these architectures form a structure having a unique composition of a heterogeneous mix of small, energy-efficient, dataflow-oriented processing elements (PEs) and a lightweight circuit-switched communication network (e.g., interconnect), for example, with enhanced support for flow control. Due to the energy advantages of each, the combination of these two components can form a spatial accelerator (e.g., as part of a computer) suitable for executing compiler-generated parallel programs in an extremely energy-efficient manner. Because the structure is heterogeneous, certain embodiments can be customized for different application domains by introducing new domain-specific PEs. For example, a structure for high-performance computing may include some customization for double-precision, fused multiply-add, while a structure for deep neural networks may include low-precision floating-point operations.
[0103] Spatial architecture patterns (e.g. Figure 6 A processor (exemplified in FIG) is composed of lightweight processing elements (PEs) connected by a network of PEs. Generally speaking, a PE may include a dataflow operator, for example, where once all input operands arrive at the dataflow operator, an operation (e.g., a microinstruction or a set of microinstructions) is performed and the result is forwarded to a downstream operator. Thus, control, scheduling, and data storage can be distributed across multiple PEs, for example, removing the overhead of the centralized architecture that dominates classical processors.
[0104] A program can be converted into a data flow graph by configuring the PEs and the network to express the control data flow graph of the program, which is mapped onto the architecture. Communication channels can be flow controlled and fully backpressured, so that, for example, if the source communication channel has no data or the destination communication channel is full, the PE will stop. In one embodiment, at runtime, data flows through the PEs and channels that have been configured to implement the operation (e.g., the accelerated algorithm). For example, data can flow in from memory through the fabric and then out back to memory.
[0105] Embodiments of such an architecture can achieve remarkable performance efficiencies relative to traditional multi-core processors: computations (e.g., in the form of PEs) can be simpler, more energy-efficient, and richer than in larger cores, and communications can be direct and primarily short-range, as opposed to, for example, being conducted over a wide, full-chip network as in a typical multi-core processor. Furthermore, because embodiments of the architecture are extremely parallel, many powerful circuit and device-level optimizations are possible without severely impacting throughput, such as low-leakage devices and low operating voltages. These lower-level optimizations can achieve even greater performance advantages over traditional cores. The combination of efficiencies yielded by these embodiments at the architectural, circuit, and device levels is compelling. As transistor density continues to increase, embodiments of the architecture can achieve even larger effective areas.
[0106] The embodiments herein provide a unique combination of data flow support and circuit switching, enabling fabrics to be smaller, more energy efficient, and provide higher aggregate performance than previous architectures. FPGAs are generally tuned for fine-grained bit manipulation, while the embodiments herein are tuned for double-precision floating-point operations found in HPC applications. Certain embodiments herein may include an FPGA in addition to a CSA according to the present disclosure.
[0107] Certain embodiments herein combine a lightweight network with energy-efficient dataflow processing elements to form a high-throughput, low-latency, energy-efficient HPC fabric. The low-latency network allows the creation of processing elements with less functionality (e.g., only one or two instructions and perhaps only one architecturally visible register, as it is efficient to aggregate multiple PEs together to form a complete program).
[0108] Relative to processor cores, CSA embodiments herein can provide greater computational density and energy efficiency. For example, when the PEs (e.g., compared to the cores) are very small, the CSA can perform many more operations than the cores and can have much more computational parallelism than the cores, for example, perhaps up to 16 times the number of FMAs as a vector processing unit (VPU). To utilize all of these computational elements, in some embodiments, the energy per operation is very low.
[0109] The energy advantages of embodiments of the dataflow architecture of the present application are numerous. Parallelism is explicit in the dataflow graph, and embodiments of the CSA architecture expend no or minimal energy to extract this parallelism, unlike, for example, out-of-order processors, which must rediscover parallelism each time an instruction is executed. In one embodiment, because each PE is responsible for a single operation, the register file and port count can be small, often just one, and thus use less energy than their counterparts in the core. Some CSAs include many PEs, each of which holds live program values, giving the aggregation effect of a large register file in traditional architectures, which significantly reduces memory accesses. In embodiments where memory is multi-ported and distributed, a CSA can maintain many more outstanding memory requests and utilize more bandwidth than a core. These advantages combine to achieve energy-per-watt levels that are only a small percentage of the cost of bare arithmetic circuitry. For example, in the case of integer multiplication, a CSA can consume no more than 25% more energy than the underlying multiplication circuitry. Relative to one embodiment of the core, integer operations in that CSA structure consume less than 1 / 30th the energy per integer operation.
[0110] From a programming perspective, the application-specific adaptability of embodiments of the CSA architecture offers significant advantages over vector processing units (VPUs). In traditional, inflexible architectures, the number of functional units, such as floating-point division or various transcendental math functions, must be chosen at design time based on a number of desired use cases. In embodiments of the CSA architecture, such functions can be configured into the architecture (e.g., by the user rather than the manufacturer) based on the requirements of each application. Application throughput can thereby be further increased. At the same time, by avoiding the need to harden such functions and instead providing more instances of primitive functions like floating-point multiplication, the computational density of embodiments of the CSA is improved. These advantages can be significant in HPC workloads, some of which spend 75% of their floating-point execution time in transcendental functions.
[0111] Certain embodiments of CSA represent significant advances in dataflow-oriented spatial architectures, e.g., the PEs of the present disclosure can be smaller and more energy efficient. These improvements can be directly derived from the combination of dataflow-oriented PEs with lightweight, circuit-switched interconnects, e.g., having single-cycle latency, as opposed to packet-switched networks (e.g., having at least 300% higher latency). Certain embodiments of the PEs support 32-bit or 64-bit operations. Certain embodiments herein allow the introduction of new application-specific PEs, e.g., for machine learning or security, and not just homogeneous combinations. Certain embodiments herein combine lightweight, dataflow-oriented processing elements with lightweight, low-latency networks to form energy-efficient computing structures.
[0112] To enable certain spatial architectures to succeed, programmers will have to spend relatively little effort to configure them, for example, while simultaneously achieving significant power and performance advantages over serialized cores. Certain embodiments herein provide CSAs (e.g., spatial structures) that are easy to program (e.g., by a compiler), highly power-efficient, and highly parallel. Certain embodiments herein provide (e.g., interconnected) networks that achieve these three goals. From a programmability perspective, certain network embodiments provide flow-controlled channels that correspond, for example, to the control data flow graph (CDFG) model of execution used in compilers. Certain network embodiments utilize dedicated circuit-switched links, making program performance easier to deduce by both humans and compilers because performance is predictable. Certain network embodiments provide both high bandwidth and low latency. Certain network embodiments (e.g., static, circuit-switched) provide latency of 0 to 1 cycle (e.g., depending on the transmission distance). Certain network embodiments provide high bandwidth by arranging several networks in parallel (and, for example, in low-level metal). Certain network embodiments communicate in low-level metal and over short distances, and therefore have very high power efficiency.
[0113] Certain embodiments of the network include architectural support for flow control. For example, in a spatial accelerator composed of small processing elements (PEs), communication latency and bandwidth can be critical to overall program performance. Certain embodiments herein provide a lightweight, circuit-switched network that facilitates spatial processing arrays (such as, Figure 6 ) in a spatial array shown in . Certain embodiments of the network implement the construction of point-to-point, flow-controlled communication channels that support communication of processing elements (PEs) oriented towards data flows. In addition to point-to-point communication, certain networks herein also support multicast communication. Communication channels can be formed by statically configuring the network to form virtual circuits between PEs. The circuit switching technology herein can reduce communication latency and correspondingly minimize network buffering, thereby, for example, producing both high performance and high energy efficiency. In certain embodiments of the network, inter-PE latency can be as low as zero cycles, which means that downstream PEs can operate on the data within the cycle after the data is generated. In order to obtain even higher bandwidth, and to allow more programs, multiple networks can be arranged in parallel, for example, as Figure 6 As shown in .
[0114] Spatial architecture (such as Figure 6The spatial architecture shown in ( ) can be composed of lightweight processing elements connected by a network between PEs. Programs viewed as data flow graphs can be mapped onto the architecture by configuring the PEs and the network. In general, PEs can be configured as data flow operators, and once all input operands arrive at the PE, some operations can then occur and the results are forwarded to the required downstream PEs. The PEs can communicate via dedicated virtual circuits, which are formed by statically configuring a circuit-switched communication network. These virtual circuits can be flow controlled and fully backpressured, so that, for example, if the source has no data or the destination is full, the PE will stop. At runtime, data can flow through the PEs that implement the mapped algorithm. For example, data can flow from memory through the fabric and then out back to memory. Embodiments of this architecture can achieve superior performance efficiency relative to traditional multi-core processors: for example, where, as opposed to an extended memory system, computation in the form of PEs is simpler and more numerous than with larger cores, and communication is direct.
[0115] Figure 6 An accelerator slice 600 is illustrated according to an embodiment of the present disclosure, the accelerator slice 600 including an array of processing elements (PEs). The interconnection network is depicted as circuit-switched, statically configured communication channels. For example, a collection of channels are coupled together by switching devices (e.g., switching devices 610 in a first network and switching devices 620 in a second network). The first network and the second network can be separate or can be coupled together. For example, the switching device 610 can couple together one or more of the four data paths 612, 614, 616, 618, e.g., configured to perform operations according to a data flow graph. In one embodiment, the number of data paths is arbitrarily large. The processing elements (e.g., processing element 604) can be as disclosed herein, e.g., Figure 9 As in . Accelerator slice 600 includes a memory / cache hierarchy interface 602 to, for example, interface accelerator slice 600 with memory and / or cache. Data paths (e.g., 618) may extend to another slice or may terminate, for example, at the edge of a slice. Processing elements may include input buffers (e.g., buffer 606) and output buffers (e.g., buffer 608).
[0116] Operations can be performed based on the availability of inputs to those operations and the state of the PE. The PE can fetch operands from input channels and write results to output channels, but can also use internal register state. Certain embodiments herein include configurable dataflow-friendly PEs. Figure 9A detailed block diagram of one such PE is shown: an integer PE. This PE consists of several I / O buffers, an ALU, storage registers, some instruction registers, and a scheduler. At each cycle, the scheduler can select an instruction for execution based on the availability of input and output buffers and the state of the PE. The result of the operation is then written to the output register, or to a register (e.g., local to the PE). The data written to the output buffer can be transferred to a downstream PE for further processing. This PE style can be extremely energy efficient, for example, the PE reads data from registers instead of from a complex multi-port register set. Similarly, instructions can be stored directly in registers rather than in a virtualized instruction cache.
[0117] The instruction register can be set during a special configuration step. During this step, in addition to the inter-PE network, auxiliary control lines and status can also be used to flow the configuration across the several PEs comprising the fabric. As a result of the parallelism, certain embodiments of such a network can provide fast reconfiguration; for example, a chip-sized fabric can be configured in less than approximately 10 microseconds.
[0118] Figure 9 represents an example configuration of a processing element, for example, where all architectural elements are sized to their minimum. In other embodiments, each of the multiple components of a processing element is independently scaled to produce a new PE. For example, to handle more complex programs, a greater number of instructions that can be executed by a PE may be introduced. The second dimension of configurability is the functionality of the PE arithmetic logic unit (ALU). In Figure 9 In [1], integer PEs are depicted as supporting addition, subtraction, and various logical operations. Other types of PEs can be created by substituting different types of functional units into PEs. For example, an integer multiplication PE may have no registers, a single instruction, and a single output buffer. Certain embodiments of PEs deconstruct fused multiply-add (FMA) into separate but tightly coupled floating-point multiplication and floating-point addition units to improve support for multiply-add-heavy workloads. PEs are discussed further below.
[0119] Figure 7A The diagram according to the embodiment of the present disclosure (for example, referring to Figure 6 Configurable data path network 700 (e.g., network 1 or network 2 discussed). Network 700 includes a plurality of multiplexers (e.g., multiplexers 702, 704, 706) that can be configured (e.g., via their respective control signals) to connect one or more data paths (e.g., from PEs) together. Figure 7B The diagram according to the embodiment of the present disclosure (for example, referring to Figure 6Configurable flow control path network 7801 (in network 1 or network 2 as discussed). The network can be a lightweight PE-to-PE network. Certain embodiments of the network can be viewed as a collection of building blocks for constructing distributed point-to-point data channels. Figure 7A shows a network with two channels (bold black line and dotted black line) enabled. The bold black channel is multicast, for example, a single input is sent to two outputs. Note that even if dedicated circuit-switched paths are formed between the channel endpoints, the channels can cross at some points within a single network. Furthermore, this crossing does not introduce structural hazards between the two channels, allowing each to operate independently and at full bandwidth.
[0120] Implementing a distributed data channel may include Figure 7A-7B The two paths shown in . The forward or data path carries data from the producer to the consumer. The multiplexer can be configured to direct data and valid bits from the producer to the consumer, for example, Figure 7A In the case of multicast, the data will be directed to multiple consumer endpoints. The second part of this embodiment of the network is the flow control or backpressure path, which flows in the opposite direction of the forward data path, such as Figure 7B As shown in . Consumer endpoints can assert when they are ready to accept new data. Configurable logic can then be used (in Figure 7B In one embodiment, each flow control function circuit can be a plurality of switching devices (e.g., a plurality of muxes), for example, with Figure 7A Similarly. The flow control path can handle the return of control data from the consumer to the producer. The junction can enable multicasting, for example, where each consumer is ready to receive data before the producer assumes that the data has been received. In one embodiment, the PE is a PE having a data flow operator as its architectural interface. Additionally or alternatively, in one embodiment, the PE can be any type of PE (e.g., in a fabric), such as, but not limited to, a PE having an instruction pointer, triggered instructions, or a state machine-based architectural interface.
[0121] In addition to, for example, PE being statically configured, the network may also be statically configured. During this configuration step, configuration bits may be set at each network component. These bits control, for example, mux selection and flow control functions. The network may include multiple networks, for example, a data path network and a flow control path network. A network or multiple networks may utilize paths of different widths (for example, a first width and a narrower or wider width). In one embodiment, the data path network has a width (for example, bit transmission) wider than the width of the flow control path network. In one embodiment, each of the first network and the second network includes its own data path network and flow control path network, for example, data path network A and flow control path network A and a wider data path network B and flow control path network B.
[0122] Some embodiments of the network are unbuffered, and data is designed to move between producers and consumers in a single cycle. Some embodiments of the network are also unbounded, that is, the network spans the entire fabric. In one embodiment, a PE is designed to communicate with any other PE in a single cycle. In one embodiment, to improve routing bandwidth, several networks can be arranged in parallel between rows of PEs.
[0123] Relative to FPGAs, certain embodiments of the networks herein have three advantages: area, frequency, and program expression. Certain embodiments of the networks herein operate at a coarse granularity, which, for example, reduces the number of configuration bits and thereby reduces the area of the network. Certain embodiments of the networks also achieve area reduction by directly implementing flow control logic in the circuit (e.g., silicon). Certain embodiments of the enhanced network implementation also enjoy frequency advantages relative to FPGAs. Due to the area and frequency advantages, power advantages may exist when lower voltages are used at throughput parity. Finally, certain embodiments of the networks provide better high-level semantics than FPGA lines, especially with respect to variable timing, and therefore, those embodiments are more easily targeted by compilers. Certain embodiments of the networks herein can be viewed as a collection of constituent primitives for building distributed point-to-point data channels.
[0124] In some embodiments, a multicast source may not be able to assert that its data is valid unless it receives a ready signal from each sink. Therefore, in the multicast case, additional binding and control bits may be utilized.
[0125] Like some PEs, the network can be statically configured. During this step, configuration bits are set at each network component. These bits control functions such as mux selection and flow control. The forward path of the network of this application requires some bits to enable the mux of the forward path to swing. Figure 7AIn the example shown in FIG, four bits are required per hop: one bit is used for each of the east and west muxes, while two bits are used for the south mux. In this embodiment, four bits are available for the data path, but seven bits are available for the flow control function (e.g., in a flow control path network). Other embodiments may utilize more bits if, for example, the CSA further utilizes a north-south direction. The flow control function may use a control bit for each direction from which flow control may come. This allows for statically setting the sensitivity of the flow control function. Table 1 below summarizes the bits used for Figure 7B In the Boolean algebraic implementation of the network flow control function in FIG, the configuration bits are capitalized. In this example, seven bits are utilized.
[0126] Table 1: Stream implementation methods
[0127]
[0128] For from Figure 7B The third flow control box from the left in FIG, EAST_WEST_SENSITIVE and NORTH_SOUTH_SENSITIVE are depicted as being set to implement flow control for the bold line channel and the dotted line channel, respectively.
[0129] Figure 8
[0066] The diagram illustrates a hardware processor slice 800 including an accelerator 802 according to an embodiment of the present disclosure. The accelerator 802 may be a CSA according to the present disclosure. The slice 800 includes a plurality of cache blocks (e.g., cache block 808). A request address file (RAF) circuit 810 may be included, for example, as discussed below in Section 3.2. ODI may refer to an on-die interconnect, e.g., an interconnect that extends across the entire die, connecting all slices. OTI may refer to an on-chip interconnect (e.g., that extends across the slice, e.g., connecting cache blocks on a slice together).
[0130] 3.1 Processing Elements
[0131] In some embodiments, a CSA comprises an array of heterogeneous PEs, where the structure consists of several types of PEs, each of which implements only a subset of the data flow operators. As an example, Figure 9A tentative implementation of a PE capable of implementing a broad set of integer and control operations is shown. Other PEs (including those supporting floating-point addition, floating-point multiplication, buffering, and certain control operations) may also have similar implementation styles, for example, replacing the ALU with appropriate (dataflow operator) circuitry. Before execution begins, a CSA's PEs (e.g., dataflow operators) may be configured (e.g., programmed) to implement specific dataflow operations from the set supported by the PE. The configuration may include one or two control words that specify opcodes that control the ALU, direct various multiplexers within the PE, and drive dataflow into and out of the PE channels. The dataflow operators may be implemented by microcoding these configuration bits. Figure 9 The integer PE 900 depicted in FIG is organized into a single-stage logical pipeline flowing from top to bottom. Data enters PE 900 from one of a set of local networks, where it is stored in input buffers for subsequent operations. Each PE can support multiple wide data-oriented channels and narrow control-oriented channels. The number of channels provided can vary based on the functionality of the PE, but one embodiment of an integer-oriented PE has two wide and one to two narrow input and output channels. Although the integer PE is implemented as a single-cycle pipeline, other pipeline options are possible. For example, a multiplication PE can have multiple pipeline stages.
[0132] PE execution can proceed in a data flow style. Based on the configuration microcode, the scheduler can check the status of the PE entry and exit buffers and arrange for the actual execution of the operation by the data operator (e.g., on the ALU) when all inputs for the configured operation have arrived and the exit buffer of the operation is available. The resulting value can be placed in the configured exit buffer. When the buffer becomes available, the transfer between the exit buffer of one PE and the entry buffer of another PE can occur asynchronously. In some embodiments, the PE is provided so that at least one data flow operation is completed for each cycle. Section 2 discusses data flow operators covering primitive operations (such as add, exclusive OR (xor), or select). Certain embodiments can provide advantages in energy, area, performance, and latency. In one embodiment, more fusion combinations can be enabled by extending the PE control path. In one embodiment, the width of the processing element is 64 bits, for example, for high utilization of double-precision floating-point calculations in HPC and for supporting 64-bit memory addressing.
[0133] 3.2 Communication Network
[0134] Embodiments of the CSA microarchitecture provide a hierarchy of multiple networks that together provide an architectural abstraction for implementing latency-insensitive channels across multiple communication scales. The lowest level of the CSA communication hierarchy can be a local network. The local network can be statically circuit-switched, for example, using configuration registers to oscillate multiplexer(s) in the local network data path to form a fixed electrical path between communicating PEs. In one embodiment, the configuration of the local network is set once for each dataflow graph (e.g., at the same time as PE configuration). In one embodiment, static circuit switching is optimized for energy, for example, where the majority (perhaps greater than 95%) of CSA communication traffic will traverse the local network. Programs can include terms used in multiple expressions. To optimize for this situation, embodiments herein provide hardware support for multicast within the local network. Several local networks can be aggregated to form routing channels, which are, for example, spread across rows and columns of PEs (in a mesh). As an optimization, several local networks can be included to carry control tokens. Compared to the FPGA interconnect, the CSA local network can be routed at the granularity of a datapath, and another difference may be the CSA's handling of control. One embodiment of the CSA local network is explicitly flow controlled (e.g., backpressure). For example, for each forward datapath and set of multiplexers, the CSA is used to provide a backward flow control path that is physically paired with the forward datapath. The combination of the two microarchitectural paths may provide a low latency, low energy, small area, point-to-point implementation of a latency-insensitive channel abstraction. In one embodiment, the flow control lines of the CSA are not visible to the user program, but these flow control lines may be manipulated by the architecture that maintains the user program. For example, the exception handling mechanism described in Section 2.2 may be implemented by pulling the flow control lines to a "non-existent" state upon detection of an exception condition. This action may not only gently stop those parts of the pipeline involved in the offending computation, but may also preserve the machine state prior to the exception for, for example, diagnostic analysis. The second network layer (e.g., a mezzanine network) may be a shared packet-switched network. The mezzanine network (e.g., composed of Figure 22The mezzanine network (schematically indicated by the dashed box in the figure) can provide more general long-distance communication at the expense of latency, bandwidth, and energy. In well-routed applications, most communication can occur on the local network, so the mezzanine network provisioning will be significantly reduced in comparison. For example, each PE can be connected to multiple local networks, but the CSA will only provision one mezzanine endpoint for each logical neighborhood of the PE. Because the mezzanine is effectively a shared network, each mezzanine network can carry multiple logically independent channels and be provisioned, for example, as multiple virtual channels. In one embodiment, the main function of the mezzanine network is to provide wide-range communication between PEs and between PEs and storage. In addition to this capability, the mezzanine can also operate as a runtime support network, for example, through which various services can access the entire structure in a user-program transparent manner. In this capability, the mezzanine endpoint can act as a controller for its local neighborhood, for example during CSA configuration. To form a channel across a CSA slice, three subchannels and two local network channels (which carry traffic to and from a single channel in the mezzanine network) can be utilized. In one embodiment, one mezzanine channel is utilized, for example, one mezzanine and two native = 3 network hops total.
[0135] The composability of channels across network layers is extended to higher-level network layers at inter-chip, inter-die, and fabric granularity.
[0136] Figure 9 9. The diagram illustrates a processing element 900 according to an embodiment of the present disclosure. In one embodiment, operation configuration registers 919 are loaded during configuration (e.g., mapping) and specify a particular operation (or operations) to be performed by the process (e.g., compute element). The activity of registers 920 may be controlled by that operation (e.g., the output of mux 916, e.g., controlled by scheduler 914). For example, as input data and control inputs arrive, scheduler 914 may schedule one or more operations of processing element 900. Control input buffers 922 are connected to local network 902 (e.g., and local network 902 may include a data path network such as that shown in FIG. 7A and a control input buffer such as that shown in FIG. 7B). Figure 7B916). The control input buffer 922 and the control output buffer 932 may be loaded with a value when a value arrives (e.g., the network has data bit(s) and valid bit(s). The control output buffer 932, data output buffer 934, and / or data output buffer 936 may receive the output of the processing element 900 (e.g., as controlled by the operation (output of mux 916)). The status register 938 may be loaded each time the ALU 918 executes (also controlled by the output of mux 916). The data in the control input buffer 922 and the control output buffer 932 may be a single bit. The mux 921 (e.g., operand A) and the mux 923 (e.g., operand B) may serve as sources of input.
[0137] For example, suppose the operation of the processing (e.g., computing) element is (or includes) Figure 3B 922. The processing element 900 is then used to select data from either data input buffer 924 or data input buffer 926, for example, to go to data output buffer 934 (e.g., by default) or data output buffer 936. Thus, if selecting from data input buffer 924, the control bit in 922 may indicate a 0, or if selecting from data input buffer 926, the control bit in 922 may indicate a 1.
[0138] For example, suppose the operation of the processing (e.g., computing) element is (or includes) Figure 3B The processing element 900 is used to output data from the data input buffer 924 (e.g., by default) or the data input buffer 926 to the data output buffer 934 or the data output buffer 936. Therefore, if the output is to the data output buffer 934, the control bit in 922 may indicate 0, or if the output is to the data output buffer 936, the control bit in 922 may indicate 1.
[0139] A plurality of networks (eg, interconnects) (eg, (input) networks 902, 904, 906 and (output) networks 908, 910, 912) may be connected to the processing element. The connections may be, for example, with reference to FIG. 7A and FIG. Figure 7B In one embodiment, each network includes two sub-networks (or two channels on the network), for example, one for Figure 7A The data path network in Figure 7B As an example, local network 902 (e.g., established as a control interconnect) is depicted as being switched (e.g., connected) to control input buffer 922. In this embodiment, the data path (e.g., Figure 7A922 (e.g., a network in the flow control path (e.g., a network in the flow control path (e.g., a network) may carry a control input value (e.g., one or more bits) (e.g., a control token), and the flow control path (e.g., a network) may carry a backpressure signal (e.g., a backpressure or no-backpressure token) from the control input buffer 922 to indicate, for example, to an upstream producer (e.g., a PE) that a new control input value (e.g., from a control output buffer of the upstream producer) will not be loaded into (e.g., sent to) the control input buffer 922 until the backpressure signal indicates that there is room in the control input buffer 922 for the new control input value. In one embodiment, the new control input value may not enter the control input buffer 922 until both (i) the upstream producer receives a “space available” backpressure signal from the “control input” buffer 922; and (ii) the new control input value is sent from the upstream producer, and this may stall the processing element 900 until that occurs (and space is available in the target, output buffer(s)).
[0140] Data input buffer 924 and data input buffer 926 can be implemented in a similar manner, for example, local network 904 (e.g., established as a data (as opposed to control) interconnect) is depicted as being switched (e.g., connected) to data input buffer 924. In this embodiment, the data path (e.g., Figure 7A , a network in the data input buffer 924 may carry a data input value (e.g., one or more bits) (e.g., a data flow token), and the flow control path (e.g., a network) may carry a backpressure signal (e.g., a backpressure or no-backpressure token) from the data input buffer 924 to indicate, for example, to an upstream producer (e.g., a PE) that a new data input value (e.g., from a data output buffer of the upstream producer) will not be loaded into (e.g., sent to) the data input buffer 924 until the backpressure signal indicates that there is room in the data input buffer 924 for the new data input value. In one embodiment, the new data input value may not enter the data input buffer 924 until both (i) the upstream producer receives a “space available” backpressure signal from the “data input” buffer 924; and (ii) the new data input value is sent from the upstream producer, for example, and this may stall the processing element 900 until that occurs (and space is available in the target, output buffer(s)). Control output values and / or data outputs may be stalled in their respective output buffers (eg, 932 , 934 , 936 ) until a backpressure signal indicates that there is available space in an input buffer for the downstream processing element(s).
[0141] Processing element 900 may stall execution until its operands (e.g., a control input value and one or more corresponding data input values for the control input value) are received and / or until there is room in the output buffer(s) of processing element 900 for data to be produced by performing operations on those operands.
[0142] 3.3 Memory Interface
[0143] Request Address File (RAF) circuit (in Figure 10A RAF (a simplified version of which is shown in Figure 1) may be responsible for executing memory operations and acting as an intermediary between the CSA fabric and the memory hierarchy. Thus, the primary microarchitectural role of the RAF may be to rationalize the out-of-order memory subsystem with the in-order semantics of the CSA fabric. In this capacity, the RAF circuitry may be provided with completion buffers (e.g., queue-like structures) that reorder memory responses and return these memory requests to the fabric in the order in which they were requested. A second primary function of the RAF circuitry may be to provide support in the form of address translation and page walkers. Incoming virtual addresses are translated into physical addresses using channel-associated translation lookaside buffers (TLBs). To provide ample memory bandwidth, each CSA slice may include multiple RAF circuits. Like the fabric's various PEs, the RAF circuitry may operate in a dataflow fashion by checking the availability of input arguments and output buffers (if needed) before selecting a memory operation to execute. However, unlike some PEs, the RAF circuitry is multiplexed between several co-located memory operations. The reused RAF circuit can be used to minimize the area overhead of its various subcomponents, such as shared accelerator cache interface (ACI) ports (described in more detail in Section 3.4), shared virtual memory (SVM) support hardware, mezzanine network interfaces, and other hardware management facilities. However, there are some program characteristics that also promote this choice. In one embodiment, (e.g., efficient) data flow graphs are used to round-robin the memory in the shared virtual memory system. Memory latency-constrained programs (like graph traversals) can utilize many separate memory operations to saturate the memory bandwidth due to the control flow that depends on the memory. Although each RAF can be reused, CAS can include multiple (e.g., between 8 and 32) RAFs of slice granularity to ensure sufficient cache bandwidth. RAFs can communicate with the rest of the structure via both the local network and the mezzanine network. When RAFs are reused, each RAF can be supplied to the local network together with several ports. These ports can serve as the lowest latency, highly deterministic path to the memory for use by latency-sensitive or high-bandwidth memory operations. Additionally, the RAF may be provisioned with a mezzanine network endpoint that provides memory access to runtime services and remote user-level memory accessors, for example.
[0144] Figure 10The diagram illustrates a request address file (RAF) circuit 1000 according to an embodiment of the present disclosure. In one embodiment, during configuration, memory load and store operations already in the data flow graph are specified in register 1010. Arcs to those memory operations in the data flow graph can then be connected to input queues 1022, 1024, and 1026. Arcs from those memory operations are then used to leave completion buffers 1028, 1030, or 1032. Dependency tokens (which can be multiple individual bits) arrive in queues 1018 and 1020. Dependency tokens will leave from queue 1016. Dependency token counter 1014 can be a compact representation of the queues and can track the number of dependency tokens for any given input queue. If dependency token counter 1014 is saturated, no additional dependency tokens can be generated for new memory operations. Accordingly, the memory ordering circuit (e.g., RAF in FIG. 11 ) stops scheduling new memory operations until dependency token counter 1014 becomes unsaturated.
[0145] As an example of a load, an address arrives in queue 1022, and scheduler 1012 matches queue 1022 with the load in 1010. Completion buffer slots for the loads are assigned in the order the addresses arrive. Assuming that this particular load in the figure has no specified dependencies, the address and completion buffer slot are dispatched to the memory system by the scheduler (e.g., via memory command 1042). When the result returns to mux 1040 (shown schematically), the result is stored in its designated completion buffer slot (e.g., because the result carries the target slot all the way through the memory system). The completion buffer sends the results back to the local network (e.g., local network 1002, 1004, 1006, or 1008) in the order the addresses arrived.
[0146] Stores can be simpler, except that both the address and the data must arrive before any operation can be dispatched to the memory system.
[0147] 3.4 Cache
[0148] The data flow graph may be able to generate a large number of (e.g., word-granular) requests in parallel. Therefore, some embodiments of the CSA provide sufficient bandwidth to the cache subsystem to maintain the CSA. Figure 11A ). Figure 11AA circuit 1100 according to an embodiment of the present disclosure is shown, having multiple request address file (RAF) circuits (e.g., RAF circuit 1) coupled between multiple accelerator slices 1108, 1110, 1112, 1114 and multiple cache blocks (e.g., cache block 1102). In one embodiment, the number of RAFs and cache blocks can be in a 1:1 or 1:2 ratio. A cache block can include a complete cache line (e.g., as opposed to word-by-word sharing), and each line has a specific home in the cache. Cache lines can be mapped to cache blocks via a pseudorandom function. The CSA can employ an SVM model for integration with other sharding architectures. Certain embodiments include an accelerator cache interconnect (ACI) network connecting the RAFs to the cache blocks. The network can carry addresses and data between the RAFs and the cache. The ACI topology can be a cascaded crossbar switch, for example, as a trade-off between latency and implementation complexity.
[0149] Figures 11B-11D Regarding a memory subsystem for efficient memory fetching according to an embodiment of the present invention. In an accelerator, memory flows can often be calculated well in advance of data usage and are often strided. This capability is limited by bottlenecks at the level close to where requests originate, where response buffers are allocated, and where pending misses are tracked. Thus, although memory accesses can be calculated, they cannot be issued. Therefore, it is expected that embodiments of the present invention having a programmable memory subsystem architecture can enable an accelerator to maintain high memory bandwidth with very little area consumption by allowing the accelerator to issue memory requests more directly to a level closer to main memory. Embodiments can eliminate the need for area-consuming cache hierarchies for a wide range of accelerators and have the potential to improve accelerator figures of merit (including raw performance and performance per unit area).
[0150] The spatial architecture may assume that memory fetches will be pushed in a purely demand-driven manner by the accelerator structure itself. Multi-level prefetching according to embodiments of the present invention improves on this approach by allowing the programmer to expose well-formed request streams to the memory hierarchy. Embodiments can improve the handling of these streams by carefully issuing requests to different levels of the memory hierarchy.
[0151] In current accelerator designs, programmers can often expose fetch patterns to memory in the form of relatively regular stride access patterns. If a given accelerator exhibits such a pattern, it is possible to employ principles from hardware prefetching to significantly reduce apparent memory latency. By carefully coordinating fetches from different levels of the memory hierarchy, it is possible to achieve perfect L1 cache performance for such streaming workloads. Because the fetch pattern can be directly exposed by the accelerator programmer, the overhead area of orchestrated prefetching is minimized compared to traditional prefetching schemes. Furthermore, because the fetch pattern is specified by the programmer, the overhead associated with traditional speculative prefetching schemes, such as incorrect learning and mistimed performance, can be largely avoided.
[0152] Figure 11B A memory subsystem utilizing multi-level memory streaming (MLMS) technology according to an embodiment of the present invention is shown. The memory subsystem includes a streamer element (SE) and a reorder buffer (ROB). The SE allows accelerator designers to specify a stylized fetch mode, such as single-dimensional or multi-dimensional streaming fetch mode.
[0153] Memory requests can be initiated by the SE or directly by the accelerator. Regardless of their origin, memory requests are injected into the ROB, which is dedicated to sorting requests returned by the non-uniform memory hierarchy. Traditionally, all memory requests are injected into the ROB. If the ROB becomes full, for example due to long request latencies, no new requests can be issued and the accelerator can stall. MLMS extends the existing cache hierarchy to allow fetch requests that load data but do not return a response to the requester to be issued to all levels of the memory hierarchy. By carefully coordinating fetch, read, and write requests between memory hierarchies, MLMS can improve accelerator performance.
[0154] In an embodiment, MLMS introduces new fetch request paths into the memory hierarchy. For each MLMS-enabled memory tier, MLMS adds paths from the accelerator to that tier. Depending on the implementation of the specific memory system tiers, the implementation of these paths may include discrete routing or extensions of the network packet format. For example, MLMS-enabled tier 1 paths may be routed, while MLMS-enabled tier 3 may augment the uncore packet format with new message types. Discrete and network implementations of MLMS may be combined in a single memory hierarchy. Unlike demand requests, these paths bypass tiers close to the originating request to improve latency and reduce hardware complexity.
[0155] At each MLMS-enabled memory level, a priority arbiter is introduced to choose between fetch requests and demand requests. Demand requests are given priority.
[0156] Embodiments may include two methods of exposing the MLMS interface to hardware designers.
[0157] The first approach is MLMS augmented streaming. In accelerator architectures that utilize explicitly programmed SEs, MLMS adds tracking registers in addition to the baseline demand flow. These tracking registers automatically issue fetch requests ahead of the demand load / store flow based on the fetch advance distance. This fetch advance distance can be fixed or programmable by the programmer at execution time. Figure 11C This extension is shown in [1]. Additional registers are added to the streamer logic to perform fetches ahead of demand streams. These fetches have the effect of pulling the required data closer to the accelerator, reducing latency for demand access to that data. Unlike existing prefetchers, the augmented MLMS streamer requires no learning mechanism and does not require a fully associative structure, thus reducing area overhead while providing excellent performance and scalability.
[0158] The second approach is direct accelerator fetch. For programmable accelerators like ASICs, MLMS extends the memory interface to accept fetch requests. The accelerator can issue fetch requests directly to each level of the memory hierarchy. This is achieved by extending the memory interface control bits to specify the target level of the cache. Figure 11D An example of such encoding is shown in . The accelerator hardware may optionally include a fetch state machine that performs the same functions as an MLMS-enabled streamer. However, an ASIC-type implementation may also include more advanced fetch schemes, for example, one that handles runtime data dependencies.
[0159] Figure 11E-11H Regarding a storage management unit (SMU) for handling stores to memory in a spatial computing architecture according to an embodiment of the present invention. Writes to cache-coherent memory spaces may use messages to preserve cache coherence properties and expensive buffering to support long-term possession to obtain ownership. Use of an SMU according to an embodiment of the present invention may eliminate some of those messages and utilize the cache itself for buffering.
[0160] Memory accesses from a graph executing on an accelerator built with a spatial compute fabric are weakly consistent in that no observer can make assumptions about the order of those references. Figure 11EThis is illustrated as the execution of the graph continuing along the graph's arcs as data is available and without being subject to certain program ordering constraints, as expressed in conventional von Neumann architectures. However, transferring control from the CPU proxy to the accelerator or vice versa requires that memory accesses occurring before the control transfer be completed, and that memory accesses after the control transfer not be initiated until the control transfer has occurred. This allows the communication occurring through memory between these two entities to be understood and replicated.
[0161] During execution on the spatial computing fabric, the interface between the spatial computing fabric and the cache / memory system occurs through a number of ports, each of which is represented by a Figure 11F RAF control as shown in . Figure 11F The RAF in the RAF can service ordered memory accesses from up to N channels of compute boxes in the space structure. For read accesses on each channel, the return value on the associated return channel occurs in the same order as the access arrived at the RAF. For store accesses, the address in the RAF is matched with the ordered data channel containing the store data. Write accesses are sent to the cache / memory system in the order they arrive.
[0162] Logically, each channel is given a slice of internal storage in the RAF, which is treated as a ring buffer. Even if memory read responses from the cache / memory hierarchy are not returned out of order in the order they were requested, the buffering in the ring buffer allows the tail pointer to return the responses back into the spatial structure in the order they were requested.
[0163] Figure 11G The relationship between the RAF and the cache is shown. To keep energy to a minimum, the goal can be to send only one message from the RAF to the corresponding cache block for each store. To hit a row that is already owned, the row is simply updated with new data from that store.
[0164] However, it may be desirable to maintain high bandwidth in the presence of a cache miss. A cache in a cache coherence protocol that supports single-writer semantics (writes can only be made to a unique and single copy of the data) will issue a command to acquire sole ownership (RFO - Read for Ownership) of the line to be written. However, since graphs executed on spatial computation fabrics are considered weakly consistent, it may be possible to wait on a miss and not notify the cache coherence system immediately, but instead fetch an empty cache entry to hold the new version of the data for the line being written and only update the portion represented by that store.
[0165] At each cache block, the bookkeeping of the storage is done in a hardware structure called SMU, such as Figure 11H . The SMU observes up to N rows at a time. In an embodiment, communication with the cache coherence system for a row in the SMU may not occur immediately, but will wait until the row is around and is the Nth oldest such row in the SMU, or until the last store to the row occurs that represents a change to all data in the row. If all slices of the row have been written, a full-block-write is issued to the cache coherence system. If not all slices have been written, a masked write is issued to the cache coherence system, and the agent (e.g., memory controller) will get the rest of the row, merge the data, and store the latest row.
[0166] An alternative embodiment that can use more energy but still preserve a high bandwidth memory flow includes issuing an RFO immediately upon receiving a write to a newly managed line (typically, this is what a conventional pipeline does). In the event that a complete line is subsequently written, filling the RFO is unnecessary, and furthermore, the line now resides in the cache rather than just being streamed to memory. Another embodiment would have an RFO predictor and / or timeout mechanism that indicates when to check whether an RFO is needed. The RFO predictor can use channel attributes, history, or other indicators to decide whether an RFO should be issued. Other alternatives include having a threshold position (e.g., N / 2 by age) and at that threshold position, deciding to issue an RFO. It may be desirable to avoid delaying until the line has become the Nth line by age before deciding that an RFO needs to be issued, which would create an unnecessary latency bubble in the memory flow.
[0167] After control is transferred back from the spatial computation fabric to the CPU agent, all pending writes must be in the cache coherent system. This means that any lines still described by the SMU should be sent to the cache coherent system, either by using a full block write or a masked write, or by obtaining ownership via an RFO, and merged with the return fill into a valid cache line. This also creates unnecessary latency bubbles; mechanisms such as the RFO predictor can reduce the number of lines that are not in the cache coherent system when a control transfer is to be performed.
[0168] Figure 11IThe diagram illustrates an architecture for low-latency atomic operations according to an embodiment of the present invention. Spatial architectures may implement only portions of a program, implying that these architectures must typically communicate with a core or another spatial architecture. Traditionally, such communications have occurred through shared memory using special operations such as atomic operations. However, throughput-oriented architectures or architectures running at slow frequencies (like FPGAs) experience long wait times when performing such operations directly. Therefore, the use of a special microcontroller that is capable of performing atomic operations around a spatial architecture according to an embodiment of the present invention may be desired. The controller may accept commands from the architecture, execute them at high speed, and return the results of the operations to the architecture.
[0169] Figure 11I The system-level architecture of the programmable atomic infrastructure is shown. A programmable microcontroller (MCU) is introduced at the fabric memory interface. The microcontroller includes a basic ALU for performing atomic operations and an extended interface to the cache memory, which may include metadata bits describing cache line status, as well as a FIFO-based interface to the fabric. The basic ALU is used to perform atomic operations, and the extended interface to the cache memory may include metadata bits describing cache line status. Storage for microcontroller instructions that can be linked into the cache memory hierarchy is also included.
[0170] In addition to issuing encoded commands directly to the microcontroller, the structure can also issue references to stored routines in the MCU. These references are converted into instruction pointers for the MCU. This reduces the control burden within the space structure because commands can be encoded with only a few bits.
[0171] At runtime, the fabric can push commands into the MCU by signaling the MCU FIFO. The MCU switches to the command and begins executing the set of instructions associated with the command. The MCU can service a single command, or it can have support for simultaneous execution of commands in the form of simultaneous multithreading (SMT).
[0172] Atomic operations may fail, so the MCU can allow the operation to be retried, which can speed up the completion of the operation. The MCU can also return a status indicator to indicate to the structure whether the operation succeeded or failed.
[0173] Figure 11J Another architecture for implementing atomic operations according to an embodiment of the present invention is shown. It includes a memory interface that exposes some behaviors (particularly consistency) of the memory hierarchy to the structure. By making these behaviors visible to the structure, the structure can implement various atomic operations. Figure 11JThe system-level architecture of the atomic interface is shown. At the beginning of an operation, the structure injects a specially tagged load instruction. The load sets a trace bit at the cache for the cache line associated with the load. This bit can be encoded in other cache states to save implementation area. The load then returns the data to the structure following the normal path. When the atomic operation has completed, the structure issues the annotated store operation. If the cache retains ownership, a success indicator is returned. Otherwise, a failure indicator is returned and the store is aborted.
[0174] 3.5 floating point support
[0175] Certain HPC applications are characterized by their need for significant floating-point bandwidth. To meet this need, embodiments of the CSA can be provisioned with multiple (e.g., each with between 128 and 256) floating-point addition and multiplication processors, depending on the slice configuration. The CSA can offer several other extended-precision modes, for example to simplify math library implementations. CSA floating-point processors can support both single- and double-precision, but lower-precision processors can support machine learning workloads. The CSA can offer floating-point performance that is an order of magnitude higher than that of the processing cores. In one embodiment, in addition to increasing floating-point bandwidth, the energy consumed in floating-point operations is reduced to drive all floating-point units. For example, to reduce energy, the CSA can selectively gate the low-order bits of the floating-point multiplier array. When examining the behavior of floating-point arithmetic, the low-order bits of the multiplication array may not often affect the final rounded product. Figure 12 The diagram shows a floating-point multiplier 1200 partitioned into three regions (a result region, three potential carry regions 1202, 1204, 1206, and a gate region) according to an embodiment of the present disclosure. In some embodiments, the carry region may affect the result region, while the gate region is less likely to affect the result region. Considering a g-bit gate region, the maximum carry may be:
[0176]
[0177] Given this maximum carry, if the result of the carry area is less than 2 cIf the carry region is c bits wide (where the carry region is c bits wide), the gated region can be ignored because it does not affect the result region. Increasing g means it is more likely that the gated region will be needed, while increasing c means that, under random assumptions, the gated region will not be used and can be disabled to avoid energy consumption. In an embodiment of the CSA floating-point multiplication PE, a two-stage pipelined approach is utilized, where the carry region is first determined, and then, if the gated region is found to affect the result, the gated region is determined. If more information about the context of the multiplication is known, the CSA adjusts the size of the gated region more aggressively. In FMA, the multiplication result may be added to an accumulator, which is often much larger than either of the multiplicands. In this case, the addend exponent can be observed in advance of the multiplication, and the CSDA can adjust the gated region accordingly. One embodiment of the CSA includes a scheme in which a context value (which constrains the minimum result of the computation) is provided to the relevant multipliers to select the lowest-energy gating configuration.
[0178] 3.6 Runtime Services
[0179] In certain embodiments, the CSA comprises a heterogeneous distributed architecture, and therefore, the runtime service implementation is designed to accommodate several types of PEs in a parallel, distributed manner. While runtime services in the CSA may be critical, they may be less frequent than user-level computations. Therefore, certain embodiments focus on overlaying services onto hardware resources. To achieve these goals, CSA runtime services may be structured as a hierarchy, with each layer corresponding to a CSA network, for example. At the slice level, a single externally-facing controller may accept or send service commands to cores associated with a CSA slice. The slice-level controller may serve to coordinate with regional controllers at the RAF (e.g., using the ACI network). The regional controllers may, in turn, coordinate with local controllers at certain mezzanine network stations. At the lowest level, service-specific microprotocols may be executed on the local network (e.g., during special modes controlled by the mezzanine controller). The microprotocols may allow each PE (e.g., PE classes divided by type) to interact with runtime services according to its needs. Consequently, parallelism is implicit in this hierarchical organization, and operations at the lowest levels can occur simultaneously. This parallelism can enable configuration of CSA slices in a range of hundreds of nanoseconds to a few microseconds, depending, for example, on the configured size of the CSA slices and their location in the memory hierarchy. Thus, embodiments of the CSA leverage the properties of dataflow graphs to improve the implementation of each runtime service. A key observation is that runtime services may only need to maintain a valid logical view of the dataflow graph (e.g., the state that can be generated by a certain ordering of the execution of dataflow operators). Services generally may not need to guarantee a temporal view of the dataflow graph (e.g., the state of the dataflow graph in the CSA at a given moment). For example, assuming that services are arranged to maintain a logical view of the dataflow graph, this can allow the CSA to perform most runtime services in a distributed, pipelined, parallel manner. The local configuration microprotocol can be a packet-based protocol overlaid on a local network. Configuration targets can be organized into configuration chains, e.g., fixed in the microarchitecture. Structural (e.g., PE) targets can be configured one at a time, e.g., using a single additional register per target to achieve distributed coordination. To begin configuration, the controller may drive an out-of-band signal that places all fabric targets within its neighborhood into an unconfigured, paused state and swings multiplexers in the local network to a predefined configuration. When fabric (e.g., PE) targets are configured (i.e., they have fully received their configuration packet), they may set their configuration micro-protocol registers, thereby notifying the next target (e.g., PE) that it may continue configuration using subsequent packets. There is no limit on the size of a configuration packet, and packets may have dynamically variable lengths. For example, a PE that configures a constant operand may have a length set to include a constant field (e.g., Figure 3B-3C The configuration grouping of X and Y in . Figure 13 The diagram illustrates the on-the-fly configuration of an accelerator 1300 having multiple processing elements (e.g., PEs 1302, 1304, 1306, 1308) according to an embodiment of the present disclosure. Once configured, the PEs can execute subject to data flow constraints. However, channels involving unconfigured PEs can be disabled by the microarchitecture, for example, to prevent any undefined operations from occurring. These properties allow embodiments of the CSA to initialize and execute in a distributed manner, without any centralized control. From an unconfigured state, configuration can occur entirely in parallel (e.g., perhaps in as little as 200 nanoseconds). However, due to the distributed initialization of embodiments of the CSA, PEs may become active, for example, sending requests to memory long before the entire structure is configured. Extraction can proceed in much the same manner as configuration. Local networks can be observed to extract data from one target at a time, and extract state bits used to implement distributed coordination. The CSA can arrange for extraction to be non-destructive, i.e., upon completion of the extraction, each extractable target has returned to its starting state. In this implementation, all states in the target are propagated to egress registers connected to the local network in a scan-like fashion. However, in-place extraction can be implemented by introducing new paths in the register transfer level (RTL) or using existing lines to provide the same functionality with lower overhead. Similar configurations and hierarchical extraction are implemented in parallel.
[0180] Figure 14 1400 illustrates a snapshot of an in-flight pipelined extraction according to an embodiment of the present disclosure. In some use cases for extraction (such as checkpointing), latency may not be a concern as long as fabric throughput can be maintained. In these cases, extraction can be arranged in a pipelined manner. Figure 14The arrangement shown in FIGURE 2 permits most of the fabric to continue executing, while narrow regions are disabled from fetching. Configuration and fetching can be coordinated and composed to achieve pipelined context switching. Qualitatively, exceptions can differ from configuration and fetching in that, rather than occurring at a specific time, they can occur anywhere in the fabric at any time during runtime. Therefore, in one embodiment, the exception microprotocol may not be overlaid on the local network and utilize its own network, which is occupied by the user program at runtime. However, exceptions are inherently rare and insensitive to latency and bandwidth. Therefore, certain embodiments of the CSA utilize a packet-switched network to carry exceptions to a local mezzanine station, for example, where they are forwarded further up the service hierarchy (e.g., as shown in FIGURE 29). Packets in the local exception network can be extremely small. In many cases, only 2 to 8 bits of PE identification (ID) are sufficient for a complete packet, for example because the CSA can create unique exception identifiers as the packet traverses the exception service hierarchy. Such an approach can be desirable because it reduces the area overhead of generating exceptions at each PE.
[0181] 4. Compile
[0182] Compiling programs written in high-level languages onto CSAs may be necessary for industrial applications. This section provides a high-level overview of the compilation strategy for an embodiment of a CSA. First, a CSA software framework is proposed that illustrates the desired properties of an ideal production-quality toolchain. Next, a prototype compiler framework is discussed. This is followed by a discussion of "control-data flow conversion," which is used, for example, to convert ordinary serialized control flow code into CSA data flow assembly code.
[0183] 4.1 Example Production Framework
[0184] Figure 15The diagram illustrates a compilation toolchain 1500 for an accelerator according to an embodiment of the present disclosure. The toolchain compiles high-level languages (such as C, C++, and Fortran) into a combination of (LLVM) intermediate representations (IR) of the main code for the specific area to be accelerated. The CSA-specific portion of the compilation toolchain takes LLVM IR as its input, optimizes and compiles it into CSA assembly, for example, adding appropriate buffering for performance on latency-insensitive channels. It then places and routes the CSA assembly onto the hardware architecture and configures the PEs and network for execution. In one embodiment, the toolchain supports CSA-specific compilation as just-in-time (JIT) compilation, thereby incorporating potential runtime feedback from actual execution. One of the key design features of the framework is compiling (LLVM) IR to obtain CSA, rather than using a higher-level language as input. While programs written in high-level programming languages specifically designed for CSAs can achieve the highest performance and / or energy efficiency, adopting new high-level languages or programming frameworks can be slow and limited in practice due to the difficulty of converting existing code bases. Using (LLVM) IR as input enables a wide range of existing programs to potentially execute on the CSA, e.g. without the need to create a new language, nor to significantly modify the front end of a new language that one wants to run on the CSA.
[0185] 4.2 Prototype Compiler
[0186] Figure 16 The figure illustrates a compiler 1600 for an accelerator according to an embodiment of the present disclosure. The compiler 1600 initially focuses on ahead-of-time compilation of C or C++ via a front-end (e.g., Clang). To compile (LLVM) IR, the compiler implements the CSA backend target within LLVM using three main stages. First, the CSA backend reduces the LLVM IR to target-specific machine instructions for a serialization unit, which implements most CSA operations as well as a traditional RISC-like control flow architecture (e.g., using branches and a program counter). The serialization unit in the toolchain can serve as a useful aid for both the compiler and application developers, as it allows for incremental transformation from control flow (CF) to data flow (DF), for example, converting a code segment at a certain moment from control flow to data flow and verifying program correctness. The serialization unit also provides a model for handling code that does not fit in a spatial array. The compiler then converts this control flow into data flow operators (e.g., code) for the CSA. This stage is described later in Section 4.3. The CSA backend can then run its own optimization rounds on the data flow operations. Finally, the compiler can dump instructions in CSA assembly format. This assembly format is taken as input to subsequent tools that place and route data flow operations on the actual CSA hardware.
[0187] 4.3 Control to Data Flow Conversion
[0188] The key part of the compiler can be implemented in the control-dataflow conversion pass (or simply the dataflow conversion pass). This pass takes a function expressed in control flow form, such as a control flow graph (CFG) with serialized machine instructions that operate on virtual registers, and converts it into a dataflow function, which is conceptually a graph of dataflow operations (instructions) connected by latency-insensitive channels (LICs). This section gives a high-level description of this pass, describing how it conceptually handles memory operations, branches, and loops in some embodiments.
[0189] Straight Line Code
[0190] Figure 17A Serialized assembly code 1702 is illustrated according to an embodiment of the present disclosure. Figure 17B FIG. 1 shows an embodiment of the present disclosure. Figure 17A The data flow assembly code 1704 of the serialization assembly code 1702. Figure 17C FIG. 1 illustrates a method for an accelerator according to an embodiment of the present disclosure. Figure 17B A data flow graph 1706 of the data flow assembly code 1704 is shown.
[0191] First, consider the simple case of converting a straight-line serialization code into a data flow. The data flow conversion pass can convert a basic serialization code block (such as Figure 17A ) is converted to the CSA assembly code shown in FIG17B . Conceptually, Figure 17B The CSA assembly representation in Figure 17C. In this example, each serializing instruction is converted to matching CSA assembly. (For example, the .lic declaration for data declares a latency-insensitive lane corresponding to a virtual register (e.g., Rdata) in the serializing code. In practice, the input to the dataflow conversion pass can be in numbered virtual registers. However, for clarity, this section uses descriptive register names. Note that in this embodiment, load and store operations are supported in the CSA architecture, allowing for many more program runs than architectures that only support pure dataflow. Because the serializing code input to the compiler is in SSA (single static assignment) form, for simple basic blocks, the control-dataflow pass can convert each virtual register definition into the generation of a single value on the latency-insensitive lane. SSA form allows multiple uses of a single definition of a virtual register (such as in Rdata2). To support this model, the CSA assembly code supports multiple uses of the same LIC (e.g., data2), and the simulator implicitly creates the necessary copies of the LIC. A key difference between serialization code and dataflow code is the handling of memory operations. Figure 17A The code in is conceptually serial, meaning that in the case where the addresses of addr and addr3 overlap, the load32 (ld32) of addr3 should appear to occur after the st32 of addr.
[0192] branch
[0193] To convert a program with multiple basic blocks and conditional statements into dataflow, the compiler generates special dataflow operators to replace branches. More specifically, the compiler uses switch operators to direct outgoing data at the end of a basic block in the original CFG, and uses pick operators to select values from the appropriate incoming channels at the beginning of a basic block. As a concrete example, consider Figures 18A-18C The code and corresponding data flow graph in , which conditionally computes the value of y based on the following inputs: a, i, x, and n. After computing the branch conditional test, the data flow code uses the switch operator (see, for example, Figure 3B-3C ) to: if the test is 0, direct the value in channel x to channel xF, or if the test is 1, direct the value in channel x to channel xT. Similarly, the pick operator (see, for example, Figure 3B-3C) is used to send channel yF to y if the test is 0, or to send channel yT to y if the test is 1. In this example, it is demonstrated that even if the value of a is only used in the true branch of the conditional statement, the CSA will include a switch operator that directs the value of a to channel aT when the test is 1, and consumes (devours) the value when the test is 0. The latter case is expressed by setting the false output of the switch to %ign. Simply connecting the channel directly to the true path may not be correct because, in the event that execution actually takes the false path, the value of "a" will be left in the graph, resulting in an incorrect value of a for the next execution of the function. This example highlights the property of control equivalence, a key property in embodiments of correct dataflow transformations.
[0194] Control Equivalence : Consider a single-entry, single-exit control flow graph G with two basic blocks A and B. If all completion control flow paths through G visit A and B the same number of times, then A and B are control equivalent.
[0195] LIC replacement In a control flow graph G, suppose an operation in basic block A defines a virtual register x and an operation in basic block B uses x. Then a correct control-data flow transformation can replace x with a latency-insensitive channel only if A and B are control equivalent. The control equivalence relation partitions the basic blocks of the CFG into regions of strong control dependencies. Figure 18A Illustrated is C source code 1802 according to an embodiment of the present disclosure. Figure 18B FIG. 1 shows an embodiment of the present disclosure. Figure 18A The data flow assembly code 1804 of the C source code 1802. Figure 18C A data flow diagram 1806 is shown for the data flow assembly code 1804 of FIG. 18B according to an embodiment of the present disclosure. Figures 18A-18C In the example, the basic blocks before and after the conditional statement are control-equivalent to each other, but the basic blocks in the true and false paths are each located in their control dependency regions. A correct algorithm for converting a CFG to dataflow is for the compiler to: (1) insert switches to compensate for the mismatch in execution frequency for any values flowing between basic blocks that are not control-equivalent; and (2) insert pickers at the beginning of the basic block to correctly select from any incoming value to the basic block. Generating appropriate control signals for these pickers and switches can be a key part of the dataflow conversion.
[0196] cycle
[0197] Another important class of CFGs in dataflow transformations is the CFG for single-entry, single-exit loops, which is a common form of loop generated in (LLVM) IR. These loops can be almost acyclic except for a single backedge from the end of the loop back to the loop header block. Dataflow transformation passes can use the same high-level strategies to transform loops as for branches, for example, dataflow transformation passes insert switches at the end of the loop to direct values out of the loop (either out of the loop exit or around the backedge to the beginning of the loop), and insert pickers at the beginning of the loop to select between the initial value entering the loop and the value arriving via the backedge. Figure 19A Illustrated is C source code 1902 according to an embodiment of the present disclosure. Figure 19B FIG. 1 shows an embodiment of the present disclosure. Figure 19A C source code 1904 and data flow assembly code 1902. Figure 19C FIG. 1 shows an embodiment of the present disclosure. Figure 19B Data flow diagram 1900 of data flow assembly code 1904. Figure 19A - Figure 19C The C and CSA assembly code for an example do-while loop that adds up the values of the loop induction variable i, and the corresponding data flow diagram are shown. For each variable (i and sum) that conceptually loops around the loop, the diagram has a corresponding pick / switch pair that controls the flow of these values. Note that even though n is a loop invariant, this example uses a pick / switch pair to loop the value of n around the loop. This duplication of n enables the virtual register for n to be converted into the LIC because it matches the execution frequency between the conceptual definition of n outside the loop and the one or more uses of n inside the loop. Generally speaking, to achieve correct data flow conversion, when registers are converted into the LIC, registers that live in the loop will be repeated once for each iteration inside the loop body. Similarly, registers that are updated within the loop and live out of the loop will be consumed (e.g., with a single final value that is sent out of the loop). Looping introduces folds into the data flow conversion process, that is, the control for the pick at the top of the loop and the switch at the bottom of the loop are offset. For example, if Figure 18A Executes three iterations and exits, then the control for the picker should be 0, 1, 1, and the control for the switch should be 1, 1, 0. This control is achieved by starting the picker channel with an initial extra 0 when the function begins at loop 0 (which is specified in the assembly by the directives .value 0 and .avail 0), and then copying the output switch into the picker. Note that the last 0 in the switch restores the final 0 to the picker, ensuring that the final state of the dataflow graph matches its initial state.
[0198] Figure 20A 2000 according to an embodiment of the present disclosure. The depicted flow 2000 includes: decoding an instruction into a decoded instruction 2002 using a decoder of a core of a processor; executing the decoded instruction to perform a first operation 2004 using an execution unit of the core of the processor; receiving an input of a dataflow graph comprising a plurality of nodes 2006; overlaying the dataflow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements 2008; and performing a second operation 2010 of the dataflow graph using the interconnection network and the plurality of processing elements when a corresponding incoming operand set arrives at each of the dataflow operators of the plurality of processing elements.
[0199] Figure 20B 2001 illustrates a flow chart according to an embodiment of the present disclosure. The depicted flow includes receiving an input of a dataflow graph comprising a plurality of nodes 2003; overlaying the dataflow graph onto a plurality of processing elements of a processor, a data path network between the plurality of processing elements, and a flow control path network between the plurality of processing elements, with each node represented as a dataflow operator 2005 in the plurality of processing elements.
[0200] In one embodiment, the core writes commands to a memory queue, and the CSA (e.g., multiple processing elements) monitors the memory queue and begins execution when the command is read. In one embodiment, the core executes a first portion of a program, and the CSA (e.g., multiple processing elements) executes a second portion of the program. In one embodiment, while the CSA is executing operations, the core performs other work.
[0201] 5. CSA Advantages
[0202] In certain embodiments, CSA architectures and microarchitectures offer profound energy, performance, and usability advantages over roadmap processor architectures and FPGAs. In this section, these architectures are compared with embodiments of CSAs, highlighting the superiority of CSAs over each in accelerating parallel dataflow graphs.
[0203] 5.1 Processor
[0204] Figure 21 Graph 2100 illustrating throughput versus energy per operation according to an embodiment of the present disclosure. Figure 21As shown in [1], small cores are generally more energy efficient than large cores, and in some workloads, this advantage can be translated into absolute performance through higher core counts. The CSA microarchitecture follows these observations to their conclusion and removes (e.g., most) of the energy-hungry control structures associated with the von Neumann architecture (including most in the instruction-side microarchitecture). By removing this overhead and implementing simple single-operation PEs, embodiments of the CSA achieve dense, space-efficient arrays. Unlike small cores, which are typically very serial, a CSA can cluster its PEs together, for example, via a circuit-switched local network, to form an explicitly parallel aggregate dataflow graph. The result is performance not only in parallel applications, but also in serial applications. Unlike cores, which are expensive in terms of area and energy, CSAs are already parallel in their native execution model. In certain embodiments, CSAs neither require speculation to improve performance nor repeatedly re-extract parallelism from a serialized program representation, thereby avoiding two of the main energy burdens of the von Neumann architecture. Most structures in CSA embodiments are distributed, small, and energy-efficient, in contrast to the centralized, bulky, and energy-hungry structures found in cores. Consider the case of registers in a CSA: Each PE may have a few (e.g., 10 or fewer) storage registers. Individually, these registers can be more efficient than a traditional register bank. When aggregated, these registers can provide the effect of a register bank in a larger architecture. As a result, CSA embodiments avoid most of the stack overflows and fill-ups caused by classic architectures, while using much less energy for each state access. Of course, applications can still access memory. In CSA embodiments, memory access requests and responses are architecturally decoupled, allowing workloads to maintain many more pending memory accesses per unit of area and energy. This property enables significantly higher performance for cache-bound workloads and reduces the area and energy required to saturate main memory in memory-bound workloads. CSA embodiments expose new forms of energy efficiency that are unique to non-von Neumann architectures. One consequence of executing a single operation (e.g., instruction) at (e.g., most) PEs is reduced operand entropy. In the case of incremental operation, each execution results in a small number of circuit-level switches and very little energy consumption, a situation examined in detail in Section 6.2. In contrast, von Neumann is multiplexed, resulting in a large number of bit transitions. The asynchronous style of CSA embodiments also enables microarchitectural optimizations, such as the floating-point optimizations described in Section 3.5, that are difficult to implement in a tightly scheduled core pipeline. Because PEs can be relatively simple and their behavior in a particular dataflow graph is statically known, clock gating and power gating techniques can be employed more efficiently than in coarser architectures.The graph execution style, small size, and scalability of embodiments of CSAs, PEs, and networks collectively enable the expression of many types of parallelism: instruction, data, pipeline, vector, memory, thread, and task parallelism can all be implemented. For example, in a CSA embodiment, one application can use the arithmetic units to provide high levels of address bandwidth, while another application can use those same units for computation. In many cases, multiple types of parallelism can be combined to achieve even higher performance. Many key HPC operations can be both replicated and pipelined, resulting in performance gains of multiple orders of magnitude. In contrast, von Neumann cores are typically optimized for a single parallelism style carefully chosen by the architect, resulting in an inability to capture all important application kernels. Precisely because CSA embodiments expose and facilitate many forms of parallelism, they do not mandate that specific forms of parallelism, or worse, specific subroutines, exist in an application to benefit from the CSA. For example, many applications (including single-stream applications) can gain both performance and energy benefits from CSA embodiments even when compiled without modification. This goes against the long-standing trend of requiring significant programmer effort to achieve significant performance gains in single-stream applications. In fact, in some applications, embodiments of the CSA achieve more performance from functionally equivalent, but less "modern," code than from its complex contemporary counterparts that have been hardened to target vector instructions.
[0205] 5.2 Comparison between CSA Implementation and FPGA
[0206] The choice of dataflow operators as the fundamental architecture of CSA embodiments distinguishes those CSAs from FPGAs, specifically as superior accelerators for HPC dataflow graphs generated from traditional programming languages. Dataflow operators are fundamentally asynchronous. This allows CSA embodiments not only to have implementation freedom in the microarchitecture, but also to adapt CSA embodiments to abstract architectural concepts simply and elegantly. For example, CSA embodiments naturally adapt to many memory microarchitectures, which are fundamentally asynchronous, using a simple load-store interface. One need only examine FPGA DRAM controllers to appreciate the difference in replication. CSA embodiments also exploit asynchrony to provide faster and more fully featured runtime services like configuration and extraction, believed to be 4-6 orders of magnitude faster than FPGA blocks. By narrowing the architectural interface, CSA embodiments provide control over most timing paths at the microarchitecture level. This allows CSA embodiments to operate at a much higher frequency than the more general control mechanisms provided in FPGAs. Similarly, clocks and resets, which may be architecturally fundamental to an FPGA, are microarchitectural in a CSA, eliminating, for example, the need to support clocks and resets as programmable entities. Dataflow operators can be coarse-grained for the most part. By performing processing only in coarse operators, embodiments of a CSA improve both the density of the structure and its energy consumption. The CSA performs operations directly rather than emulating them using lookup tables. A second consequence of coarseness is that the placement and routing problem is simplified. CSA dataflow graphs are many orders of magnitude smaller than FPGA netlists, and in embodiments of a CSA, placement and routing times are correspondingly reduced. The significant differences between embodiments of a CSA and an FPGA make the CSA superior as an accelerator, for example, for dataflows generated from traditional programming languages.
[0207] 6. Evaluation
[0208] CSAs are a novel computer architecture that offer significant performance and energy advantages over roadmap processors. Consider the case of computing a single-stride address for a walk across an array. This case can be important in HPC applications (e.g., where a large amount of integer work is spent computing address offsets). In address computations, especially strided address computations, one argument is constant for each computation, and the other argument changes only slightly. Therefore, in most cases, only a few bits switch per cycle. In fact, using a derivation similar to the constraint on the floating-point carry bit described in Section 3.5, it can be shown that for strided computations, on average, fewer than two input bits switch per computation, resulting in a 50% energy reduction for a random switching distribution. Many of these energy savings are lost if time multiplexing is used. In one embodiment, CSAs achieve approximately 3x (three times) the energy efficiency of a core while achieving an 8x (eight times) performance gain. The parallelism gains achieved by CSA embodiments result in reduced program runtimes, resulting in correspondingly significant leakage energy reductions. At the PE level, CSA embodiments are extremely energy efficient. The second important question for CSA is whether it uses a reasonable amount of energy at the chip level. Since CSA embodiments can exercise every floating-point PE in the fabric every cycle, it serves as a reasonable upper bound on energy and power consumption, for example, making most of the energy go into floating-point multiplication and addition.
[0209] 7. Future CSA Details
[0210] This section discusses further details of configuration and exception handling.
[0211] 7.1 Microarchitecture for Configuring CSA
[0212] This section discloses examples of how to configure a CSA (e.g., a structure), how to quickly implement that configuration, and how to minimize the resource overhead of the configuration. Quickly configuring structures is crucial for accelerating small parts of larger algorithms and, therefore, for broadening the applicability of CSAs. This section further discusses features that allow embodiments of a CSA to be programmed with configurations of varying lengths.
[0213] Embodiments of a CSA (e.g., a fabric) may differ from conventional cores in that they may utilize a configuration step in which (e.g., large) portions of the fabric are loaded in advance with program configuration, prior to program execution. An advantage of static configuration may be that very little energy is expended at runtime while configuring, in contrast, for example, to a serializing core that expends energy fetching configuration information (instructions) nearly every cycle. A previous disadvantage of configuration was that it was a coarse-grained step with potentially long latencies that set a lower bound on the size of programs that could be accelerated in the fabric due to the cost of context switches. The present disclosure describes a scalable microarchitecture for rapidly configuring spatial arrays in a distributed manner that, for example, avoids the previous disadvantages.
[0214] As discussed above, a CSA can include lightweight processing elements connected by an inter-PE network. By configuring configurable fabric elements (CFEs) (e.g., PEs and an interconnect (fabric) network), a program, viewed as a control-data flow graph, is then mapped onto the architecture. Generally speaking, a PE can be configured as a dataflow operator, and once all input operands arrive at a PE, some operation occurs, and the result is forwarded to another PE or PEs for consumption or output. PEs can communicate via dedicated virtual circuits, formed by statically configuring a circuit-switched communication network. These virtual circuits can be flow-controlled and fully backpressured, so that, for example, if the source has no data or the destination is full, the PE will stall. At runtime, data can flow through the PEs implementing the mapped algorithm. For example, data can flow from memory through the fabric and then outbound back to memory. This type of spatial architecture can achieve superior performance efficiency relative to traditional multi-core processors: in contrast to extended memory systems, computation can be simpler and more numerous in PE form than in larger cores, and communication can be direct.
[0215] Embodiments of a CSA may not utilize (e.g., software-controlled) packet switching (e.g., packet switching that requires significant software assistance to implement), which slows down configuration. Embodiments of a CSA include out-of-band signaling in the network (e.g., only 2-3 bits of out-of-band signaling, depending on the supported feature set) and a fixed configuration topology to avoid the need for extensive software support.
[0216] A key difference between embodiments of a CSA and the approach used in FPGAs is that the CSA approach can use wide data words, is distributed, and includes mechanisms for fetching program data directly from memory. Embodiments of a CSA can avoid utilizing JTAG-type single-bit communication for area efficiency, for example because that can require several milliseconds to fully configure a large FPGA fabric.
[0217] Embodiments of the CSA include a distributed configuration protocol and a microarchitecture to support this protocol. Initially, the configuration state may reside in memory. Multiple (e.g., distributed) local configuration controllers (LCCs) can stream portions of the overall program to their local regions in the spatial fabric, for example, using a combination of a small set of control signals and a fabric-provided network. State elements can be used at each CFE to form a configuration chain, for example, allowing each CFE to self-program without global addressing.
[0218] Embodiments of the CSA include specific hardware support for forming configuration chains, e.g., rather than software that dynamically builds these chains at the expense of increased configuration time. Embodiments of the CSA are not purely packet-switched and do include additional out-of-band control lines (e.g., control is not sent over the data path, requiring additional cycles to gate and reserialize this information). Embodiments of the CSA reduce configuration latency (e.g., by at least half) by fixing configuration ordering and by providing explicit out-of-band control, while not significantly increasing network complexity.
[0219] Embodiments of the CSA do not use serial configuration for configurations where data is streamed into the fabric bit by bit using a JTAG-like protocol. Embodiments of the CSA utilize a coarse-grained fabric approach. In some embodiments, adding some control lines or status elements to a 64-bit or 32-bit oriented CSA fabric is less costly than adding those same control mechanisms to a 4-bit or 6-bit fabric.
[0220] Figure 22 An accelerator slice 2200 according to an embodiment of the present disclosure is illustrated, comprising an array of processing elements (PEs) and local configuration controllers 2202, 2206. Each PE, each network controller, and each switch may be a configurable fabric element (CFE), for example, configured (e.g., programmed) by an embodiment of a CSA architecture.
[0221] An embodiment of a CSA includes hardware that provides efficient, distributed, low-latency configuration of heterogeneous spatial structures. This can be achieved based on four techniques. First, by utilizing a hardware entity (a local configuration controller (LCC)), such as Figure 22-24As shown in . LCCs can fetch a stream of configuration information from (e.g., virtual) memory. Second, a configuration data path can be included, for example, that is as wide as the native width of the PE structure and can be overlaid on top of the PE structure. Third, new control signals can be received into the PE structure that orchestrate the configuration process. Fourth, a state element can be located (e.g., in a register) at each configurable endpoint that tracks the state of adjacent CFEs, allowing each CFE to explicitly self-configure without the need for additional control signals. These four microarchitectural features allow a CSA to configure its chain of CFEs. To achieve low configuration latency, the configuration can be partitioned by establishing many LCCs and CFE chains. During configuration, these can operate independently to load structures in parallel, thereby dynamically reducing latency, for example. As a result of these combinations, structures configured using embodiments of the CSA architecture can be fully configured (e.g., in a few hundred nanoseconds). Detailed operation of various components of embodiments of the CSA configuration network is disclosed below.
[0222] Figures 23A-23C A local configuration controller 2302 is illustrated configuring a datapath network according to an embodiment of the present disclosure. The depicted network includes a plurality of multiplexers (e.g., multiplexers 2306, 2308, 2310) that can be configured (e.g., via their respective control signals) to connect one or more datapaths (e.g., from PEs) together. Figure 23A Figure 23 shows a network 2300 (e.g., a fabric) configured (e.g., set) for some prior operation or program. Figure 23 shows a local configuration controller 2302 (e.g., including a network interface circuit 2304 for sending and / or receiving signals) strobing configuration signals, and the local network is set to a default configuration (e.g., as depicted in the figure) that allows the LLC to send configuration data to all configurable fabric elements (CFEs) (e.g., muxes). Figure 23C The LCC is shown strobing configuration information across the network to configure the CFEs in a predetermined (e.g., silicon-defined) sequence. In one embodiment, when the CFEs are configured, they can begin operation immediately. In another embodiment, the CFEs wait to begin operation until the fabric has been fully configured (e.g., for each local configuration controller, by a configuration terminator (e.g., Figure 25 In one embodiment, the LCC gains control of the network fabric by sending a special message or driving a signal. It then gates the configuration data to the CFEs in the fabric (e.g., over a period of many cycles). In these figures, the multiplexer network is not identical to the multiplexer network in some figures (e.g., Figure 6 ) is similar to the “switching device” shown in ).
[0223] Local Configuration Controller
[0224] Figure 24 The diagram illustrates a (e.g., local) configuration controller 2402 according to an embodiment of the present disclosure. A local configuration controller (LCC) can be a hardware entity responsible for loading local portions of a fabric program (e.g., within a subset of a slice or elsewhere), interpreting these program portions, and subsequently loading these program portions into the fabric by driving appropriate protocols over various configuration lines. In this capacity, an LLC can be a dedicated serialized microcontroller.
[0225] When the LLC operation receives a pointer to a code segment, it can begin. Depending on the LCC microarchitecture, this pointer (e.g., stored in pointer register 2406) comes to the LCC either over the network (e.g., from within the CSA (structure) itself) or through a memory system access. When the LLC receives such a pointer, it can optionally drain the relevant state from its portion of the structure used for context storage and then proceed to immediately reconfigure the portion of the structure that the LLC is responsible for. The program loaded by the LCC can be a combination of configuration data for the structure and control commands for the LCC, for example, the configuration data and the control commands are lightly encoded. When the LLC has the program portion streamed in, it can interpret the program as a command stream and perform the appropriate encoded actions to configure (e.g., load) the structure.
[0226] exist Figure 22 2 shows two different microarchitectures for LCCs, one or both of which may be used in a CSA, for example. The first microarchitecture places LCC 2202 at the memory interface. In this case, the LCC can make direct requests to the memory system to load data. In the second case, LCC 2206 is placed on a memory network where it can only make requests to memory indirectly. In both cases, the logical operation of the LCC remains unchanged. In one embodiment, the LCC is notified of the program to be loaded, for example, by a set of control status registers (e.g., visible to the OS), which are used to notify each LCC of a new program pointer, etc.
[0227] Additional out-of-band control channels (e.g., wires)
[0228] In some embodiments, configuration relies on 2-8 additional out-of-band control channels to improve configuration speed, as defined below. For example, the configuration controller 2402 may include the following control channels: e.g., a CFG_START control channel 2408, a CFG_VALID control channel 2410, and a CFG_DONE control channel 2412, examples of each of which are discussed below in Table 2.
[0229] Table 2: Control channels
[0230]
[0231] In general, the handling of configuration information may be left to the implementer of a particular CFE. For example, a selectable-function CFE may have provisions to set registers using existing data paths, whereas a fixed-function CFE may simply set configuration registers.
[0232] Due to the long line delays when programming large sets of CFEs, the CFG_VALID signal can be considered a clock / latch enable for the CFE components. Because this signal is used as a clock, in one embodiment, the line's duty cycle is at most 50%. As a result, configuration throughput is approximately halved. Optionally, a second CFG_VALID signal can be added to allow continuous programming.
[0233] In one embodiment, only CFG_START is strictly passed on standalone coupling devices (eg, wires), for example, CFG_VALID and CFG_DONE may be overlaid on top of other network coupling devices.
[0234] Reuse of network resources
[0235] To reduce configuration overhead, certain embodiments of the CSA leverage existing network infrastructure to transfer configuration data. LCCs can leverage both the chip-level memory hierarchy and the fabric-level communication network to move data from storage to the fabric. As a result, in certain embodiments of the CSA, the configuration infrastructure adds no more than 2% to the total fabric area and power.
[0236] Reuse of network resources in certain embodiments of the CSA can enable networks with some hardware support for configuration mechanisms. Circuit-switched networks of embodiments of the CSA cause the LCC to set up their multiplexers in a specific manner for configuration when the 'CFG_START' signal is asserted. Packet-switched networks do not require extensions, but LCC endpoints (e.g., configuration terminators) use specific addresses in packet-switched networks. Network reuse is optional, and some embodiments may find a dedicated configuration bus more convenient.
[0237] Each CFE status
[0238] Each CFE may maintain a bit indicating whether it has been configured (see, for example, Figure 13). This bit may be de-asserted when the configuration start signal is driven, and subsequently asserted once a particular CFE has been configured. In one configuration protocol, the CFEs are arranged in a chain, and the CFE and configuration status bits determine the topology of the chain. A CFE may read the configuration status bits of an immediately adjacent CFE. If the adjacent CFE is configured and the current CFE is not configured, then the CFE and any current configuration data is determined to be for the current CFE. When the 'CFG_DONE' signal is asserted, the CFE may set its configuration bits to, for example, enable the upstream CFE to be configured. As a base case for the configuration process, the CFE asserts its configured configuration terminator (e.g., in Figure 22 A configuration terminator 2204 for LCC 2202 or a configuration terminator 2208 for LCC 2206) may be included at the end of the chain.
[0239] Within the CFE, this bit can be used to drive flow control ready signals. For example, when the configuration bit is deasserted, network control signals can be automatically clamped to values that prevent data flow, and within the PE, no operations or other actions will be scheduled.
[0240] Handling high-latency configuration paths
[0241] An embodiment of an LCC may, for example, drive signals over long distances through many multiplexers and utilize numerous loads. Therefore, it may be difficult for signals to reach the remote CFE within a short clock cycle. In some embodiments, configuration signals are at a certain frequency division (e.g., fractional) of the master (e.g., CSA) clock signal to ensure digital timing compliance during configuration. Clock division can be used in out-of-band signaling protocols and does not require any modifications to the master clock tree.
[0242] Ensure consistent structural behavior during configuration
[0243] Because some configuration schemes are distributed and because program and memory effects have non-deterministic timing, different parts of the fabric may be configured at different times. As a result, certain embodiments of the CSA provide mechanisms for preventing inconsistent operation between configured and unconfigured CFEs. In general, consistency is considered a property that is required and maintained by the CFE itself, for example using internal CFE state. For example, when a CFE is in an unconfigured state, it may declare its input buffers to be full and its outputs to be invalid. When configured, these values will be set to the true states of the buffers. As enough parts of the fabric come out of configuration, these techniques can allow the fabric to begin operation. This has the effect of further reducing context switch latency, for example, if long latency memory requests are issued early.
[0244] Variable width configuration
[0245] Different CFEs may have different configuration word widths. For smaller CFE configuration words, the implementer can balance latency by fairly assigning CFE configuration loading across network lines. To balance loading on network lines, one option is to assign configuration bits to different portions of the network lines to limit the net latency on any one line. Wide data words can be handled by using serialization / deserialization techniques. These decisions can be taken on a fabric-by-fabric basis to optimize the behavior of a particular CSA (e.g., fabric). A network controller (e.g., one or more of network controller 2210 and network controller 2212) can communicate with each domain (e.g., subset) of a CSA (e.g., fabric) to, for example, send configuration information to one or more LCCs.
[0246] 7.2 Microarchitecture for Low-Latency Configuration of CSAs and Timely Fetching of CSA Configuration Data
[0247] Embodiments of a CSA can be an energy-efficient and high-performance means of accelerating user applications. When considering whether a program (e.g., a data flow graph of a program) can be successfully accelerated by an accelerator, both the time used to configure the accelerator and the time used to run the program can be considered. If the run time is short, the configuration time will play a large role in determining successful acceleration. Therefore, in order to maximize the domain of accelerable programs, in some embodiments, the configuration time is made as short as possible. One or more configuration caches can be included in the CSA, for example to enable high-bandwidth, low-latency storage to achieve rapid reconfiguration. What follows is a description of several embodiments of the configuration cache.
[0248] In one embodiment, during configuration, the configuration hardware (e.g., LCC) may optionally access a configuration cache to obtain new configuration information. The configuration cache may operate as either a traditional address-based cache or in an OS-managed mode where the configuration is stored in a local address space and addressed by referencing that address space. If the configuration state is in the cache, then in some embodiments, no request to the backing store will be made. In some embodiments, the configuration cache is separate from any (e.g., lower-level) shared caches in the memory hierarchy.
[0249] Figure 25An accelerator slice 2500 according to an embodiment of the present disclosure is illustrated, comprising an array of processing elements, a configuration cache (e.g., 2518 or 2520), and a local configuration controller (e.g., 2502 or 2506). In one embodiment, configuration cache 2514 is co-located with local configuration controller 2502. In one embodiment, configuration cache 2518 is located within a configuration domain of local configuration controller 2506, e.g., a first domain terminates at configuration terminator 2504 and a second domain terminates at configuration terminator 2508. The configuration cache allows the local configuration controller to reference the configuration cache during configuration, e.g., to obtain configuration state with lower latency than referencing memory. The configuration cache (storage) can be either dedicated or accessible as a configuration mode within a storage element (e.g., local cache 2516) within the fabric.
[0250] Cache Mode
[0251] 1. Demand Caching - In this mode, the configuration cache operates as a true cache. The configuration controller issues address-based requests, which are checked against the tags in the cache. Misses can be loaded into the cache and can then be referenced during future reprogramming.
[0252] 2. In-Fabric Storage (Scratchpad) Cache - In this mode, the configuration cache receives references to configuration sequences in its own small address space rather than the host's larger address space. This can improve memory density because the portion of the cache used to store tags can instead be used to store configurations.
[0253] In some embodiments, the configuration cache may have configuration data preloaded therein (e.g., via external or internal instructions). This may allow for a reduction in latency for loading programs. Certain embodiments herein provide a structure for accessing the configuration cache that, for example, allows new configuration states to be loaded into the cache even while a configuration is already running within the structure. The initiation of this loading may occur from an internal or external source. Embodiments of the preloading mechanism further reduce latency by removing latency from cache loads of the configuration path.
[0254] Prefetch Mode
[0255] 1. Explicit Prefetch – The configuration path is augmented with a new command, ConfigurationCachePrefetch. Instead of programming the structure, this command simply causes the relevant program configuration to be loaded into the configuration cache without programming the structure. Because this mechanism piggybacks on the existing configuration infrastructure, it is exposed both within the structure and externally to the core and other entities accessing memory space.
[0256] 2. Implicit prefetching - The global configuration controller may maintain a prefetch predictor and use it to initiate (eg, in an automated manner) explicit prefetches of the configuration cache.
[0257] 7.3 Hardware for Rapid Reconfiguration of CSA in Response to Exceptions
[0258] Certain embodiments of CSAs (e.g., space structures) include a large number of instructions and configuration states, for example, which are largely static during operation of the CSA. Consequently, the configuration states may be susceptible to soft errors. Rapid and error-free recovery from these soft errors may be critical to the long-term reliability and performance of space systems.
[0259] Certain embodiments herein provide a fast configuration recovery cycle, e.g., in which a configuration error is detected and portions of a structure are immediately reconfigured. Certain embodiments herein include, for example, a configuration controller with a reliability, availability, and durability (RAS) reprogramming feature. Certain embodiments of the CSA include circuitry for high-speed configuration, error reporting, and parity checking within a spatial structure. Using a combination of these three features and an optional configuration cache, the configuration / exception handling circuitry can recover from soft errors in the configuration. When detected, a soft error can be transferred to a configuration cache, which initiates an immediate reconfiguration of the structure (e.g., that portion of the structure). Certain embodiments provide dedicated reconfiguration circuitry, e.g., that is faster than any solution that would be implemented indirectly in the structure. In certain embodiments, the exception and configuration circuitry located together collaborate to reload the structure upon a configuration error detection.
[0260] Figure 26The diagram illustrates an accelerator slice 2600 according to an embodiment of the present disclosure, comprising an array of processing elements and configuration and exception handling controllers 2602 and 2606 with reconfiguration circuitry 2618 and 2622. In one embodiment, when a PE detects a configuration error through its RAS feature, it sends a message (e.g., a configuration error or reconfiguration error) to the configuration and exception handling controller (e.g., 2602 or 2606) via its exception generator. Upon receiving this message, the configuration and exception handling controller (e.g., 2602 or 2606) activates co-located reconfiguration circuitry (e.g., 2618 or 2622, respectively) to reload the configuration state. The microarchitecture continues to configure and (e.g., only) reloads the configuration state, and in some embodiments, only reloads the configuration state for the PE reporting the RAS error. After reconfiguration is complete, the architecture can resume normal operation. To reduce latency, the configuration state used by the configuration and exception handling controller (e.g., 2602 or 2606) can be sourced from a configuration cache. As a base case of the configuration or reconfiguration process, the configuration terminator (e.g., Figure 26 A configuration terminator 2604 for the configuration and exception handling controller 2602 or a configuration terminator 2608 for the configuration and exception handling controller 2606) may be included at the end of the chain.
[0261] Figure 27 Reconfiguration circuitry 2718 is illustrated in accordance with an embodiment of the present disclosure. Reconfiguration circuitry 2718 includes a configuration state register 2720 for storing a configuration state (or a pointer to the configuration state).
[0262] 7.4 Hardware for fabric-initiated reconfiguration of CSA
[0263] Some parts of an application for a CSA (e.g., a spatial array) may be run infrequently or may be mutually exclusive with respect to other parts of the program. To save area, to improve performance and / or to reduce power, it may be useful to time-multiplex multiple parts of a spatial structure between several different parts of a program data flow graph. Certain embodiments herein include an interface through which a CSA (e.g., via a spatial program) may request that part of the structure be reprogrammed. This may enable the CSA to dynamically change itself based on dynamic control flow. Certain embodiments herein may allow structure-initiated reconfiguration (e.g., reprogramming). Certain embodiments herein provide a set of interfaces for triggering configuration from within a structure. In some embodiments, a PE issues a reconfiguration request based on a decision in the program data flow graph. The request may travel over the network to our new configuration interface, where it triggers the reconfiguration. Once the reconfiguration is complete, a message notifying the completion may optionally be returned. Certain embodiments of the CSA therefore provide program (e.g., data flow graph)-guided reconfiguration capabilities.
[0264] Figure 28 The accelerator slice 2800 according to an embodiment of the present disclosure is shown, and includes an array of processing elements and a configuration and exception handling controller 2806 with reconfiguration circuitry 2818. Here, a portion of the fabric issues a request for (re)configuration to a configuration domain, such as the configuration and exception handling controller 2806 and / or the reconfiguration circuitry 2818. The domain (re)configures itself, and when the request has been satisfied, the configuration and exception handling controller 2806 and / or the reconfiguration circuitry 2818 issues a response to the fabric to notify it that the (re)configuration is complete. In one embodiment, the configuration and exception handling controller 2806 and / or the reconfiguration circuitry 2818 disables communication while (re)configuration is in progress, so that there are no consistency issues with the program during operation.
[0265] Configuration Mode
[0266] Configuration by Address - In this mode, the fabric makes a direct request to load configuration data from a specific address.
[0267] Configuration by reference - In this mode, the structure makes a request to load a new configuration, for example, by a predefined reference ID. This simplifies the determination of the code to load because the location of the code has been abstracted.
[0268] Configuring multiple domains
[0269] The CSA may include a higher-level configuration controller to support a multicast mechanism to broadcast configuration requests to multiple (e.g., distributed or local) configuration controllers (e.g., via a network indicated by a dashed box). This can enable a single configuration request to be replicated across multiple, larger portions of the fabric, for example, to trigger a wide reconfiguration.
[0270] 7.5 Exception Aggregator
[0271] Certain embodiments of the CSA may also experience exceptions (e.g., exceptional conditions), such as floating point underflow. When these conditions occur, special handling routines may be called to either correct the program or terminate it. Certain embodiments herein provide a system-level architecture for handling exceptions in a spatial structure. Because certain spatial structures emphasize area efficiency, embodiments herein minimize the total area while providing a general exception mechanism. Certain embodiments herein provide a low-area means for signaling exceptional conditions occurring in a CSA (e.g., a spatial array). Certain embodiments herein provide interfaces and signaling protocols for communicating such exceptions as well as PE-level exception semantics. Certain embodiments herein are dedicated exception handling capabilities and, for example, do not require explicit handling by the programmer.
[0272] One embodiment of the CSA exception architecture consists of four parts, such as Figure 29-30 These parts can be arranged in a hierarchy where exceptions flow from the producer and eventually flow up to a slice-level exception aggregator (e.g., a handler), which can meet with an exception maintainer, such as a core. The four parts can be:
[0273] 1.PE Exception Generator
[0274] 2. Local abnormal network
[0275] 3. Interlayer Abnormal Aggregator
[0276] 4. Slice-level exception aggregator
[0277] Figure 29 An accelerator slice 2900 is shown including an array of processing elements and a mezzanine anomaly aggregator 2902 coupled to a chip-level anomaly aggregator 2904 , according to an embodiment of the present disclosure. Figure 30 A processing element 3000 is shown with an exception generator 3044 according to an embodiment of the present disclosure.
[0278] PE exception generator
[0279] Processing element 3000 may include Figure 9The processing element 900, for example, similar numbers are similar components, such as local network 902 and local network 3002. The additional network 3013 (e.g., channel) can be an abnormal network. The PE can be implemented to the abnormal network (e.g., Figure 30 An interface on an abnormal network 3013 (eg, a channel). For example, Figure 30 The microarchitecture of such an interface is illustrated, wherein a PE has an exception generator 3044 (e.g., to initiate an exception finite state machine (FSM) 3040 to gate an exception packet (e.g., BOXID 3042) out onto an exception network). BOXID 3042 can be a unique identifier for an exception-generating entity (e.g., a PE or block) within a local exception network. When an exception is detected, exception generator 3044 senses the exception network and gates out the BOXID when the network is found to be idle. Exceptions can be caused by many conditions, such as, but not limited to, arithmetic errors, failed ECC checks on state, etc. However, it is also possible to introduce exceptional data flow operations using the idea of supporting constructs such as breakpoints.
[0280] Exceptions can be raised either explicitly through programmer-provided instructions or implicitly when a hardened error condition (e.g., floating-point underflow) is detected. When an exception occurs, the PE 3000 may enter a wait state, in which it waits for service by, for example, a final exception handler external to the PE 3000. The contents of the exception packet depend on the implementation of the particular PE, as described below.
[0281] Local abnormal network
[0282] The (e.g., local) anomaly network directs anomaly packets from PE 3000 to the mezzanine anomaly network. The anomaly network (e.g., 3013) can be a serial packet-switched network consisting of (e.g., a single control line) and one or more data lines, organized in a ring or tree topology, for example, for a subset of PEs. Each PE can have a (e.g., ring) station in the (e.g., local) anomaly network, at which, for example, the PE can arbitrate for injecting messages into the anomaly network.
[0283] PE endpoints that need to inject anomaly packets can observe their local anomaly network exit point. If the control signal indicates busy, the PE will wait to start injecting packets for that PE. If the network is not busy, that is, the downstream station has no packets to forward, the PE will proceed with the injection.
[0284] Network packets can be of variable or fixed length. Each packet can begin with a fixed-length header field that identifies the source PE of the packet. This header field can be followed by a variable number of PE-specific fields containing information such as error codes, data values, or other useful status information.
[0285] Interlayer Abnormal Aggregator
[0286] The interlayer anomaly aggregator 2904 is responsible for assembling local anomaly networks into larger packets and sending these larger packets to the slice-level anomaly aggregator 2902. The interlayer anomaly aggregator 2904 can prepend the local anomaly packet with its own unique ID, for example, to ensure that the anomaly message is unambiguous. The interlayer anomaly aggregator 2904 can interface with a special virtual channel in the interlayer network that is used only for anomalies, for example, to ensure that the anomaly is deadlock-free.
[0287] The mezzanine anomaly aggregator 2904 may also be able to maintain certain categories of anomalies directly. For example, configuration requests from the fabric may be distributed out of the mezzanine network using a cache local to the mezzanine network station.
[0288] Slice-level exception aggregator
[0289] The final level of the exception system is the chip-level exception aggregator 2902. The chip-level exception aggregator 2902 is responsible for collecting exceptions from the various mezzanine-level exception aggregators (e.g., 2904) and forwarding these exceptions to the appropriate maintenance hardware (e.g., core). Thus, the chip-level exception aggregator 2902 may include several internal tables and controllers for associating specific messages with handler routines. These tables can be indexed directly or with a small state machine to direct specific exceptions.
[0290] Like the mezzanine exception aggregator, the slice-level exception aggregator can service some exception requests. For example, it can initiate reprogramming of a large portion of the PE structure in response to a specific exception.
[0291] 7.6 Exception Controller
[0292] Certain embodiments of a CSA include extraction controller(s) for extracting data from a fabric. The following discusses embodiments of how to quickly implement this extraction and minimize the resource overhead of data extraction. Data extraction can be used for critical tasks such as exception handling and context switching. Certain embodiments herein extract data from a heterogeneous spatial fabric by introducing features that allow extractable fabric elements (EFEs) (e.g., PEs, network controllers, and / or switches) to have a variable and dynamically variable number of states to be extracted.
[0293] Embodiments of the CSA include a distributed data extraction protocol and a microarchitecture to support this protocol. Certain embodiments of the CSA include multiple local extraction controllers (LECs) that use a combination of a (e.g., small) set of control signals and a fabric-provided network to cause program data to flow from their local regions in the spatial fabric. State elements can be used at each extractable fabric element (EFE) to form an extraction chain, for example, allowing individual EFEs to extract themselves without requiring global addressing.
[0294] Embodiments of the CSA do not use the local network to extract program data. Embodiments of the CSA include, for example, specific hardware support (e.g., extraction controllers) for forming extraction chains and do not rely on software to dynamically establish these chains (e.g., at the expense of increased extraction time). Embodiments of the CSA are not purely packet-switched and do include additional out-of-band control lines (e.g., controls are not sent over the data path, requiring additional cycles to gate and reserialize this information). Embodiments of the CSA reduce extraction latency (e.g., by at least half) by fixing the extraction order and by providing explicit out-of-band control, while not significantly increasing network complexity.
[0295] Embodiments of the CSA do not use a serial mechanism for data extraction where data is streamed bit by bit from the fabric using a JTAG-like protocol. Embodiments of the CSA utilize a coarse-grained fabric approach. In some embodiments, adding some control lines or status elements to a 64-bit or 32-bit oriented CSA fabric is less costly than adding those same control mechanisms to a 4-bit or 6-bit fabric.
[0296] Figure 31 An accelerator slice 3100 according to an embodiment of the present disclosure is illustrated, comprising an array of processing elements and local extraction controllers 3102, 3106. Each PE, each network controller, and each switch may be an extractable fabric element (EFE), for example, configured (e.g., programmed) by an embodiment of the CSA architecture.
[0297] An embodiment of a CSA includes hardware that provides efficient, distributed, low-latency extraction of heterogeneous spatial structures. This can be achieved based on four techniques. First, by utilizing a hardware entity (a local extraction controller (LEC)), such as Figures 31-33As shown in [1], the LEC can accept commands from the host (e.g., a processor core), such as extracting a data stream from a spatial array and writing the data back to virtual memory for inspection by the host. Second, an extraction data path can be included, for example, one that is as wide as the native width of the PE structure and can be overlaid on top of the PE structure. Third, new control signals can be received into the PE structure to orchestrate the extraction process. Fourth, a state element can be located (e.g., in a register) at each configurable endpoint that tracks the state of adjacent EFEs, allowing each EFE to explicitly output its state without requiring additional control signals. These four microarchitectural features allow the CSA to extract data from a chain of EFEs. To achieve low data extraction latency, certain embodiments can partition the extraction problem by including multiple (e.g., many) LECs and EFE chains in the structure. During extraction, these chains can operate independently to extract data from the structure in parallel, thereby, for example, dramatically reducing latency. As a result of these features, the CSA can perform a complete state dump (e.g., in a few hundred nanoseconds).
[0298] Figures 32A-32C 32. A local extraction controller 3202 is shown configuring a datapath network according to an embodiment of the present disclosure. The depicted network includes a plurality of multiplexers (e.g., multiplexers 3206, 3208, 3210) that can be configured (e.g., via their respective control signals) to connect one or more datapaths (e.g., from PEs) together. Figure 32A Illustrated is a network 3200 (eg, structure) configured (eg, set up) for some prior operating procedures. Figure 32B The diagram shows a local extraction controller 3202 (e.g., including a network interface circuit 3204 for sending and / or receiving signals) strobing an extraction signal, and all PEs controlled by the LEC enter extraction mode. The last PE in the extraction chain (or extraction terminator) can master the extraction channel (e.g., bus) and send data based on (1) a signal from the LEC or (2) a signal generated internally (e.g., from the PE). Once completed, the PE can set its completion flag, for example, to enable the next PE to extract its data. Figure 32C The furthest PE is shown having completed the extraction process and, as a result, has set its one or more extraction status bits, e.g., which cause the mux to swing to the adjacent network to enable the next PE to begin the extraction process. The extracted PE may resume normal operation. In some embodiments, the PE may remain disabled until other action is taken. In these figures, the multiplexer network is not the same as in some figures (e.g., Figure 6 ) is similar to the “switching device” shown in ).
[0299] The next section describes the operation of the various components of an embodiment of the extraction network.
[0300] Local Extraction Controller
[0301] Figure 33 The figure shows an extraction controller 3302 according to an embodiment of the present disclosure. The local extraction controller (LEC) can be a hardware entity responsible for accepting extraction commands, coordinating the extraction process of the EFE, and / or storing the extracted data to, for example, virtual memory. In this capacity, the LEC can be a dedicated serialized microcontroller.
[0302] LEC operation can begin when the LEC receives a pointer to a buffer (e.g., in virtual memory) where the fabric state is to be written, and optionally a command controlling how much of the fabric state is to be extracted. Depending on the LEC microarchitecture, this pointer (e.g., stored in pointer register 3304) can come to the LEC either over the network or through a memory system access. When the LEC receives such a pointer (e.g., a command), it proceeds to extract the state from the portion of the fabric it is responsible for. The LEC can stream this extracted data out of the fabric into a buffer provided by the external caller.
[0303] exist Figure 31 Two different microarchitectures for LECs are shown in FIG. The first places LEC 3102 at the memory interface. In this case, the LEC can make direct requests to the memory system to write the data being fetched. In the second case, LEC 3106 is placed on a memory network, where LCC 3106 can only make requests to the memory indirectly. In both cases, the logical operation of the LEC may not change. In one embodiment, the LEC is informed of the desire to fetch data from the structure, for example, by a set of control status registers (e.g., visible to the OS), which will be used to notify each LEC of the new command.
[0304] Additional out-of-band control channels (e.g., wires)
[0305] In some embodiments, extraction relies on 2-8 additional out-of-band signals to improve configuration speed, as defined below. Signals driven by the LEC may be labeled LEC. Signals driven by the EFE (e.g., PE) may be labeled EFE. Configuration controller 3302 may include the following control channels, such as LEC_EXTRACT control channel 3406, LEC_START control channel 3308, LEC_STROBE control channel 3310, and EFE_COMPLETE control channel 3312, examples of each of which are discussed below in Table 3.
[0306] Table 3: Extraction channels
[0307]
[0308] In general, the handling of the extraction can be left to the implementer of the specific EFE. For example, a selectable function EFE may have provisions to use existing data paths to dump registers, while a fixed function EFE may just have multiplexers.
[0309] Due to the long line delays when programming a large set of EFEs, the LEC_STROBE signal can be considered a clock / latch enable for the EFE components. Because this signal is used as a clock, in one embodiment, the duty cycle of this line is at most 50%. As a result, the extraction throughput is approximately halved. Optionally, a second LEC_STROBE signal can be added to enable continuous extraction.
[0310] In one embodiment, only LEC_START is strictly communicated over a separate coupling device (eg, wire), for example, other control channels may be overlaid over an existing network (eg, wire).
[0311] Reuse of network resources
[0312] To reduce data fetch overhead, certain embodiments of the CSA leverage existing network infrastructure to deliver fetched data. LECs can leverage both the chip-level memory hierarchy and the fabric-level communication network to move data from the fabric to storage. As a result, in certain embodiments of the CSA, the fetch infrastructure contributes no more than 2% to the total fabric area and power.
[0313] Reuse of network resources in certain embodiments of the CSA may enable networks with some hardware support for the extraction protocol. Circuit-switched networks require certain embodiments of the CSA to have the LEC configure their multiplexers in a specific way when the "LEC_START" signal is asserted. Packet-switched networks do not require this extension, but LEC endpoints (e.g., extraction terminators) use specific addresses in packet-switched networks. Network reuse is optional, and some embodiments may find a dedicated configuration bus more convenient.
[0314] Each EFE status
[0315] Each EFE may maintain a bit that indicates whether it has output its status. This bit may be de-asserted when the extract start signal is driven, and then asserted once the particular EFE has completed the extract. In one extract protocol, the EFEs are arranged to form a chain, and the EFE extract status bits determine the topology of the chain. An EFE may read the extract status bits of the immediately adjacent EFEs. If the adjacent EFE has its extract bit set and the current EFE does not have its extract bit set, then the EFE may determine that it owns the extract bus. When an EFE dumps its last data value, it may drive the "EFE_DONE" signal and set its extract bit, thereby enabling, for example, an upstream EFE to be configured for extraction. The network adjacent to the EFE may observe this signal and also adjust its state to handle this transition. As a base case for the extract process, an extract terminator (e.g., in Figure 22 An extraction terminator 3104 for LEC 3102 or an extraction terminator 3108 for LEC 3106) may be included at the end of the chain.
[0316] Within the EFE, this bit can be used to drive flow control ready signals. For example, when the extract bit is deasserted, network control signals can be automatically clamped to values that prevent data flow, and within the PE, no operations or actions will be scheduled.
[0317] Handling high-latency paths
[0318] An embodiment of an LEC may, for example, drive signals over long distances through numerous multiplexers and utilize numerous loads. Consequently, it may be difficult for signals to reach the remote EFE within a short clock cycle. In some embodiments, the extracted signal is at a certain frequency division (e.g., fractional) of the master (e.g., CSA) clock signal to ensure digital timing compliance during extraction. Clock division can be used in out-of-band signaling protocols and does not require any modifications to the master clock tree.
[0319] Ensure consistent structural behavior during extraction
[0320] Because some extraction schemes are distributed and have non-deterministic timing due to program and memory effects, different members of the structure may be in the extraction state at different times. When LEC_EXTRACT is driven, all network flow control signals may be driven to logic low, for example, thereby freezing the operation of a particular segment of the structure.
[0321] The extraction process can be non-destructive. Thus, once extraction is complete, the set of PEs can be considered running. Extensions to the extraction protocol can allow PEs to be optionally disabled after extraction. Alternatively, in an embodiment, starting configuration during the extraction process will have a similar effect.
[0322] Single PE extraction
[0323] In some cases, extracting a single PE may be expedient. In this case, as part of the initialization of the extraction process, an optional address signal may be driven. This allows the extraction of that PE to be directly enabled. Once the PE has been extracted, the extraction process terminates with the LEC_EXTRACT signal falling. In this way, a single PE can be selectively extracted, for example, by a local extraction controller.
[0324] Disposal of extraction back pressure
[0325] In embodiments where the LEC writes the extracted data to memory (e.g., for post-processing, e.g., in software), it may be constrained by limited memory bandwidth. In the event that the LEC exhausts its buffer capacity or anticipates that it will exhaust its buffer capacity, the LEC may stop strobing LEC_STROBE until the buffering issue has been resolved.
[0326] Note that in some of the drawings (e.g. Figure 22 、 25 , 26, 28, 29 and 31), schematically illustrating communications. In some embodiments, those communications can occur through a (e.g., interconnected) network.
[0327] 7.7 Flowchart
[0328] Figure 34 3400 according to an embodiment of the present disclosure. The depicted flow 3400 includes: decoding an instruction into a decoded instruction 3402 using a decoder of a core of a processor; executing the decoded instruction to perform a first operation 3404 using an execution unit of the core of the processor; receiving an input 3406 of a dataflow graph comprising a plurality of nodes; overlaying the dataflow graph onto an array of processing elements of the processor, with each node represented as a dataflow operator 3408 in the array of processing elements; and performing a second operation 3410 of the dataflow graph using the array of processing elements when an incoming operand set arrives at the array of processing elements.
[0329] Figure 3535. The flowchart 3500 according to an embodiment of the present disclosure is illustrated. The depicted flowchart 3500 includes: decoding an instruction into a decoded instruction 3502 using a decoder of a core of a processor; executing the decoded instruction to perform a first operation 3504 using an execution unit of the core of the processor; receiving an input 3506 of a dataflow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements 3508; and performing a second operation 3510 of the dataflow graph using the interconnection network and the plurality of processing elements when an incoming operand set arrives at the plurality of processing elements.
[0330] 8. Summary
[0331] ExaFLOP-scale supercomputing can be a challenge in high-performance computing that may not be met by conventional von Neumann architectures. To achieve ExaFLOPs, embodiments of the CSA provide heterogeneous spatial arrays that are targeted for direct execution of dataflow graphs (e.g., generated by a compiler). In addition to laying out the architectural principles of embodiments of the CSA, embodiments of the CSA are described and evaluated above that demonstrate 10x (10 times) higher performance and energy than existing products. The code generated by the compiler can have significant performance and energy gains compared to roadmap architectures. As a heterogeneous parameterized architecture, embodiments of the CSA can be easily adapted to all computing use cases. For example, a mobile version of the CSA can be scaled to 32 bits, while an array focused on machine learning can feature a significant number of vectorized 8-bit multiplication units. The main advantages of embodiments of the CSA are high performance, extremely high energy efficiency, and features relevant to all forms of computing, from supercomputing and data centers to the Internet of Things.
[0332] In one embodiment, a processor includes: a plurality of processing elements; and an interconnect network between the plurality of processing elements, the interconnect network configured to receive input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the interconnect network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are configured to perform operations upon incoming operand sets arriving at the plurality of processing elements. The processor also includes a streamer element configured to prefetch the incoming operand sets from two or more levels of a memory system.
[0333] The streamer element may prefetch based on a programmable memory access pattern. The streamer element may include a plurality of tracking registers for prefetching ahead of a demand stream. The plurality of tracking registers may include an x-dimensional register for prefetching ahead of a first dimension in a multi-dimensional streaming fetch pattern. The plurality of tracking registers may include a y-dimensional register for prefetching ahead of a second dimension in a multi-dimensional streaming fetch pattern.
[0334] In an embodiment, a processor includes: a plurality of processing elements; an interconnection network between the plurality of processing elements, the interconnection network for receiving input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the interconnection network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are configured to perform operations when incoming operand sets arrive at the plurality of processing elements; and a memory management unit for processing storage of output data from the operations in a memory.
[0335] The memory management unit may include an address register for storing an address of a cache line in the address register, wherein the cache line is for storing a plurality of data values, at least two of the data values being from two different processing elements. The memory management unit is configured to track the storage of the plurality of data values in the cache line. The memory management unit also includes a plurality of mask bits for performing a masked write in response to determining that less than a complete cache line is to be stored in the memory.
[0336] In an embodiment, a processor includes: a plurality of processing elements; an interconnection network between the plurality of processing elements, the interconnection network for receiving an input of a dataflow graph including a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the interconnection network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are configured to perform a first operation when an incoming operand set arrives at the plurality of processing elements; and a microcontroller for performing a second operation, wherein the second operation is an atomic operation.
[0337] In an embodiment, a method includes: receiving an input of a dataflow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of a processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements; prefetching, by a streamer element, incoming operand sets from two or more levels of a memory system; and performing operations of the dataflow graph using the interconnection network and the plurality of processing elements when the incoming operand sets arrive at the plurality of processing elements.
[0338] The streamer element may prefetch based on a programmable memory access pattern. The streamer element may include a plurality of tracking registers for prefetching ahead of a demand stream. The plurality of tracking registers may include an x-dimensional register for prefetching ahead of a first dimension in a multi-dimensional streaming fetch pattern. The plurality of tracking registers may include a y-dimensional register for prefetching ahead of a second dimension in a multi-dimensional streaming fetch pattern.
[0339] In an embodiment, a method includes: receiving an input of a dataflow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of a processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements; performing operations of the dataflow graph using the interconnection network and the plurality of processing elements when an incoming operand set arrives at the plurality of processing elements; and processing, by a memory management unit, storing output data from the operations in a memory.
[0340] The memory management unit may include an address register for storing an address of a cache line in the address register, wherein the cache line is for storing a plurality of data values, at least two of the data values being from two different processing elements. Processing by the memory management unit may include tracking the storage of the plurality of data values in the cache line. The method may also include determining that less than a full cache line is to be stored in the memory; and using a plurality of mask bits in the memory management unit to perform a masked write.
[0341] In an embodiment, a method includes: receiving an input of a dataflow graph including a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of a processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements; performing a first operation of the dataflow graph using the interconnection network and the plurality of processing elements when an incoming operand set arrives at the plurality of processing elements; and performing a second operation by a microcontroller, wherein the second operation is an atomic operation.
[0342] In one embodiment, a processor includes: a plurality of processing elements; and an interconnect network between the plurality of processing elements, the interconnect network configured to receive input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the interconnect network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements is configured to perform an operation upon arrival of a corresponding incoming operand set at each of the dataflow operators of the plurality of processing elements. The processor further includes a streamer element configured to prefetch the incoming operand sets from two or more levels of a memory system.
[0343] When a backpressure signal from a downstream processing element indicates that storage in the downstream processing element is not available for an output of a processing element in the plurality of processing elements, the processing element may cease execution. The processor may include a flow control path network for carrying backpressure signals according to a dataflow graph. The dataflow token may cause an output from a dataflow operator receiving the dataflow token to be sent to an input buffer of a particular processing element in the plurality of processing elements. The operation may include a memory access, and the plurality of processing elements may include a memory access dataflow operator that will not perform the memory access until a memory dependency token is received from a logically preceding dataflow operator. The plurality of processing elements may include processing elements of a first type and processing elements of a second, different type.
[0344] In another embodiment, a method receives an input of a dataflow graph comprising a plurality of nodes; overlays the dataflow graph onto a plurality of processing elements of a processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements; and, when a corresponding incoming operand set arrives at each of the dataflow operators of the plurality of processing elements, performs an operation of the dataflow graph using the interconnection network and the plurality of processing elements. The processor further includes a streamer element for prefetching the incoming operand sets from two or more levels of a memory system.
[0345] The method may include halting execution by a processing element of a plurality of processing elements when a backpressure signal from a downstream processing element indicates that storage in the downstream processing element is not available for an output of the processing element. The method may include sending a backpressure signal on a flow control path according to a dataflow graph. The dataflow token may cause an output from a dataflow operator receiving the dataflow token to be sent to an input buffer of a particular processing element of the plurality of processing elements. The method may include not performing a storage access until a memory dependency token is received from a logically preceding dataflow operator, wherein the operation comprises a memory access and the plurality of processing elements comprises a memory access dataflow operator. The method may include providing a first type of processing element and a second, different type of processing element of the plurality of processing elements.
[0346] In yet another embodiment, an apparatus includes: a datapath network between a plurality of processing elements; and a flow control path network between the plurality of processing elements, wherein the datapath network and the flow control path network are configured to receive input of a dataflow graph comprising a plurality of nodes, the dataflow graph being configured to be overlaid onto the datapath network, the flow control path network, and the plurality of processing elements, with each node being represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements being configured to perform an operation upon arrival of a corresponding incoming operand set at each of the dataflow operators of the plurality of processing elements. The processor further includes a streamer element configured to prefetch the incoming operand sets from two or more levels of a memory system.
[0347] The flow control path network can carry backpressure signals to a plurality of dataflow operators according to a dataflow graph. Dataflow tokens sent to the dataflow operators on the datapath network can cause outputs from the dataflow operators to be sent on the datapath network to input buffers of specific processing elements among the plurality of processing elements. The datapath network can be a static circuit-switched network for carrying corresponding sets of input operands to each of the dataflow operators according to the dataflow graph. The flow control path network can transmit backpressure signals from downstream processing elements according to the dataflow graph to indicate that storage in the downstream processing element is unavailable for outputs of the processing element. At least one data path of the datapath network and at least one flow control path of the flow control path network can form a channelized circuit with backpressure control. The flow control path network can be serially pipelined at at least two processing elements among the plurality of processing elements.
[0348] In another embodiment, a method includes receiving an input of a dataflow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of a processor, a datapath network between the plurality of processing elements, and a flow control path network between the plurality of processing elements, wherein each node is represented as a dataflow operator in the plurality of processing elements. The method may include, based on the dataflow graph, using the flow control path network to carry backpressure signals to the plurality of dataflow operators. The method may include sending a dataflow token to the dataflow operator on the datapath network so that output from the dataflow operator is sent on the datapath network to an input buffer of a particular processing element in the plurality of processing elements. The method may include, based on the dataflow graph, setting a plurality of switches of the datapath network and / or a plurality of switches of the flow control path network to carry a corresponding set of input operands to each of the dataflow operators, wherein the datapath network is a static circuit-switched network. The method may include, based on the dataflow graph, using the flow control path network to transmit a backpressure signal from a downstream processing element to indicate that storage in the downstream processing element is unavailable for output from the processing element. The method may include forming a channelized circuit with backpressure control using at least one data path of a data path network and at least one flow control path of a flow control path network.
[0349] In yet another embodiment, a processor includes: a plurality of processing elements; and a network device between the plurality of processing elements, the network device configured to receive an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the network device and the plurality of processing elements, each node being represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements being configured to perform an operation upon arrival of a corresponding incoming operand set at each of the dataflow operators of the plurality of processing elements. The processor further includes a streamer element configured to prefetch the incoming operand sets from two or more levels of a memory system.
[0350] In another embodiment, an apparatus includes: a datapath device between a plurality of processing elements; and a flow control path device between the plurality of processing elements, wherein the datapath device and the flow control path device are configured to receive input of a dataflow graph comprising a plurality of nodes, the dataflow graph being configured to be overlaid onto the datapath device, the flow control path device, and the plurality of processing elements, with each node being represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements being configured to perform an operation upon arrival of a corresponding incoming operand set at each of the dataflow operators of the plurality of processing elements. The processor further includes a streamer element configured to prefetch the incoming operand sets from two or more levels of a memory system.
[0351] In one embodiment, a processor includes an array of processing elements configured to receive an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the array of processing elements and each node is represented as a dataflow operator in the array of processing elements, and the array of processing elements is configured to perform operations upon arrival of an incoming operand set at the array of processing elements. The processor further includes a streamer element configured to prefetch the incoming operand set from two or more levels of a memory system.
[0352] The array of processing elements may not perform an operation until an incoming operand set arrives at the array of processing elements and storage in the array of processing elements is available for an output of the operation. The array of processing elements may include a network (or channels) for carrying dataflow tokens and control tokens to a plurality of dataflow operators. The operation may include a memory access, and the array of processing elements may include a memory access dataflow operator that is configured not to perform a memory access until a memory dependency token is received from a logically preceding dataflow operator. Each processing element may only perform one or two operations of the dataflow graph.
[0353] In another embodiment, a method includes receiving an input of a dataflow graph comprising a plurality of nodes, overlaying the dataflow graph onto an array of processing elements of a processor, with each node represented as a dataflow operator in the array of processing elements, and executing operations of the dataflow graph using the array of processing elements when an incoming operand set arrives at the array of processing elements. The processor also includes a streamer element for prefetching the incoming operand set from two or more levels of a memory system.
[0354] The array of processing elements may not perform an operation until an incoming operand set arrives at the array of processing elements and storage in the array of processing elements is available for the output of a second operation. The array of processing elements may include a network that carries dataflow tokens and control tokens to a plurality of dataflow operators. The operation may include a memory access, and the array of processing elements includes a memory access dataflow operator that will not perform a memory access until a memory dependency token is received from a logically preceding dataflow operator. Each processing element may only perform one or two operations of the dataflow graph.
[0355] In yet another embodiment, a non-transitory machine-readable medium stores code that, when executed by a machine, causes the machine to perform a method comprising: receiving an input of a dataflow graph comprising a plurality of nodes; overlaying the dataflow graph onto an array of processing elements of a processor, with each node represented as a dataflow operator in the array of processing elements; and performing operations of the dataflow graph using the array of processing elements when an incoming operand set arrives at the array of processing elements. The processor further comprises a streamer element for prefetching the incoming operand set from two or more levels of a memory system.
[0356] The array of processing elements may not perform a second operation until an incoming operand set arrives at the array of processing elements and storage in the array of processing elements is available for the output of the second operation. The array of processing elements may include a network that carries dataflow tokens and control tokens to a plurality of dataflow operators. The operation may include a memory access, and the array of processing elements includes a memory access dataflow operator that will not perform a memory access until a memory dependency token is received from a logically preceding dataflow operator. Each processing element may only perform one or two operations of the dataflow graph.
[0357] In another embodiment, a processor includes: means for receiving an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid into the means and each node is represented as a dataflow operator in the means, and the means is configured to perform an operation upon arrival of an incoming operand set at the means. The processor also includes a streamer element configured to prefetch the incoming operand set from two or more levels of a memory system.
[0358] In one embodiment, a processor includes: a core having a decoder and an execution unit, the decoder for decoding an instruction into a decoded instruction, the execution unit for executing the decoded instruction to perform a first operation; a plurality of processing elements; and an interconnect network between the plurality of processing elements, the interconnect network for receiving input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is adapted to be overlaid onto the interconnect network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are adapted to perform a second operation upon arrival of an incoming operand set at the plurality of processing elements. The processor also includes a streamer element for prefetching the incoming operand set from two or more levels of a memory system.
[0359] The processor may further include a plurality of configuration controllers, each coupled to a respective subset of the plurality of processing elements, and each configured to load configuration information from storage and cause the respective subset of the plurality of processing elements to couple according to the configuration information. The processor may include a plurality of configuration caches, each coupled to a respective configuration cache to retrieve configuration information for the respective subset of the plurality of processing elements. The first operation performed by the execution unit may prefetch configuration information into each of the plurality of configuration caches. Each of the plurality of configuration controllers may include reconfiguration circuitry configured to: upon receiving a configuration error message from at least one processing element in the respective subset of the plurality of processing elements, cause reconfiguration of at least one processing element. Each of the plurality of configuration controllers may include reconfiguration circuitry configured to: upon receiving a reconfiguration request message, cause reconfiguration of the respective subset of the plurality of processing elements; and disable communication with the respective subset of the plurality of processing elements until reconfiguration is complete. The processor may include a plurality of exception aggregators, each coupled to a respective subset of the plurality of processing elements, to collect exceptions from the respective subset of the plurality of processing elements and forward the exceptions to the core for maintenance. The processor may include a plurality of extraction controllers, each extraction controller coupled to a respective subset of the plurality of processing elements, and each extraction controller to cause state data from the respective subset of the plurality of processing elements to be saved to the memory.
[0360] In another embodiment, a method includes: decoding an instruction into a decoded instruction using a decoder of a core of a processor; executing the decoded instruction to perform a first operation using an execution unit of the core of the processor; receiving an input of a dataflow graph including a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements; and performing a second operation of the dataflow graph using the interconnection network and the plurality of processing elements when an incoming operand set arrives at the plurality of processing elements. The processor also includes a streamer element to prefetch the incoming operand set from two or more levels of a memory system.
[0361] The method may include: loading configuration information for a respective subset of a plurality of processing elements from storage; and causing coupling for each respective subset of the plurality of processing elements based on the configuration information. The method may include: retrieving the configuration information for the respective subset of the plurality of processing elements from a respective configuration cache of a plurality of configuration caches. The first operation performed by the execution unit may be prefetching the configuration information into each of the plurality of configuration caches. The method may include: upon receiving a configuration error message from at least one processing element in the respective subset of the plurality of processing elements, causing reconfiguration of at least one processing element. The method may include: upon receiving a reconfiguration request message, causing reconfiguration of the respective subset of the plurality of processing elements; and disabling communication with the respective subset of the plurality of processing elements until the reconfiguration is complete. The method may include: collecting exceptions from the respective subset of the plurality of processing elements; and forwarding the exceptions to a core for maintenance. The method may include: causing state data from the respective subset of the plurality of processing elements to be saved to a memory.
[0362] In yet another embodiment, a non-transitory machine-readable medium stores code that, when executed by a machine, causes the machine to perform a method comprising: decoding an instruction into a decoded instruction using a decoder of a core of a processor; executing the decoded instruction to perform a first operation using an execution unit of the core of the processor; receiving an input of a dataflow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of the processor and an interconnection network between the plurality of processing elements of the processor, with each node represented as a dataflow operator in the plurality of processing elements; and performing a second operation of the dataflow graph using the interconnection network and the plurality of processing elements when an incoming operand set arrives at the plurality of processing elements. The processor also includes a streamer element to prefetch the incoming operand set from two or more levels of a memory system.
[0363] The method may include: loading configuration information for a respective subset of a plurality of processing elements from storage; and causing coupling for each respective subset of the plurality of processing elements based on the configuration information. The method may include: retrieving the configuration information for the respective subset of the plurality of processing elements from a respective configuration cache of a plurality of configuration caches. The first operation performed by the execution unit may be prefetching the configuration information into each of the plurality of configuration caches. The method may include: upon receiving a configuration error message from at least one processing element in the respective subset of processing elements, causing reconfiguration of the at least one processing element. The method may include: upon receiving a reconfiguration request message, causing reconfiguration of the respective subset of the plurality of processing elements; and disabling communication with the respective subset of the plurality of processing elements until the reconfiguration is complete. The method may include: collecting exceptions from the respective subset of the plurality of processing elements; and forwarding the exceptions to a core for maintenance. The method may include: causing state data from the respective subset of the plurality of processing elements to be saved to a memory.
[0364] In another embodiment, a processor includes: a plurality of processing elements; and a device between the plurality of processing elements, the device configured to receive an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the device and the plurality of processing elements, each node being represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements configured to perform operations upon arrival of an incoming operand set at the plurality of processing elements. The processor further includes a streamer element configured to prefetch the incoming operand set from two or more levels of a memory system.
[0365] In yet another embodiment, an apparatus includes a data storage device storing code that, when executed by a hardware processor, causes the hardware processor to perform the method disclosed herein. The apparatus may be as described in the "Detailed Description." The method may be as described in the "Detailed Description."
[0366] In another embodiment, a non-transitory machine-readable medium stores code that, when executed by a machine, causes the machine to perform a method including any of the methods disclosed herein.
[0367] An instruction set (e.g., for execution by a core) may include one or more instruction formats. A given instruction format may define various fields (e.g., number of bits, bit positions) to specify the operation to be performed (e.g., opcode), operand(s) on which the operation is to be performed, and / or other data field(s) (e.g., mask), among others. Some instruction formats are further decomposed through the definition of instruction templates (or subformats). For example, instruction templates for a given instruction format may be defined to have different subsets of the fields of the instruction format (the fields included are generally in the same order, but at least some fields have different bit positions because fewer fields are included), and / or to have given fields interpreted differently. Thus, each instruction of an ISA is expressed using a given instruction format (and, if defined, according to a given one of the instruction templates for that instruction format) and includes fields for specifying the operation and operands. For example, an exemplary ADD (addition) instruction has a specific opcode and instruction format, the specific instruction format including an opcode field for specifying the opcode and an operand field for selecting operands (source 1 / destination and source 2); and the ADD instruction, when present in an instruction stream, will have specific content in the operand field for selecting a specific operand. A set of SIMD extensions known as Advanced Vector Extensions (AVX) (AVX1 and AVX2) and utilizing a Vector Extensions (VEX) encoding scheme have been introduced and / or released (see, e.g., June 2016). 64 and IA-32 Architectures Software Developer's Manual; and see the February 2016 Architecture Instruction Set Extension Programming Reference).
[0368] Example instruction format
[0369] The embodiments of the instruction(s) described herein can be embodied in different formats. In addition, exemplary systems, architectures, and pipelines are described in detail below. The embodiments of the instruction(s) can be executed on such systems, architectures, and pipelines, but are not limited to those described in detail.
[0370] Generic vector-friendly instruction format
[0371] The vector friendly instruction format is an instruction format suitable for vector instructions (e.g., there are specific fields dedicated to vector operations). Although embodiments are described in which both vector and scalar operations are supported through the vector friendly instruction format, alternative embodiments use only vector operations through the vector friendly instruction format.
[0372] Figure 36A-Figure 36B is a block diagram illustrating a general vector friendly instruction format and instruction templates thereof according to an embodiment of the present disclosure. Figure 36Ais a block diagram illustrating a general vector friendly instruction format and its class A instruction template according to an embodiment of the present disclosure; and Figure 36B 36 is a block diagram illustrating a generic vector friendly instruction format and its class B instruction templates according to an embodiment of the present disclosure. Specifically, class A and class B instruction templates are defined for the generic vector friendly instruction format 3600, both of which include instruction templates with no memory access 3605 and instruction templates with memory access 3620. The term "generic" in the context of the vector friendly instruction format refers to an instruction format that is not tied to any particular instruction set.
[0373] Although embodiments of the present disclosure will be described in which the vector friendly instruction format supports: a 64-byte vector operand length (or size) with a 32-bit (4-byte) or 64-bit (8-byte) data element width (or size) (and thus, a 64-byte vector consists of 16 doubleword-sized elements, or alternatively, 8 quadword-sized elements); a 64-byte vector operand length (or size) with a 16-bit (2-byte) or 8-bit (1-byte) data element width (or size); a 32-byte vector operand length (or size) with a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte) data element width (or size); and a 32-byte vector operand length (or size) with a 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte) data element width (or size). or 8-bit (1 byte) data element width (or size); and 16-byte vector operand length (or size) with 32-bit (4-byte), 64-bit (8-byte), 16-bit (2-byte), or 8-bit (1-byte) data element width (or size); however, alternative embodiments may support larger, smaller, and / or different vector operand sizes (e.g., 256-byte vector operands) and larger, smaller, or different data element widths (e.g., 128-bit (16-byte) data element width).
[0374] Figure 36A The Class A instruction templates include: 1) within the instruction templates with no memory access 3605, there are shown instruction templates for full rounding control type operations 3610 without memory access and instruction templates for data transformation type operations 3615 without memory access; and 2) within the instruction templates with memory access 3620, there are shown instruction templates for timeliness of memory access 3625 and instruction templates for non-timeliness of memory access 3630. Figure 36B The Class B instruction templates include: 1) within the instruction template with no memory access 3605, an instruction template for a partial rounding control type operation 3612 with write mask control without memory access and an instruction template for a vsize type operation 3617 with write mask control without memory access; and 2) within the instruction template with memory access 3620, an instruction template for a write mask control 3627 with memory access.
[0375] The generic vector friendly instruction format 3600 includes the following listed in accordance with Figures 36A-36BThe following fields are in the order shown in the figure.
[0376] Format field 3640 - The specific value in this field (the instruction format identifier value) uniquely identifies the vector friendly instruction format, and thereby identifies that the instruction appears in the vector friendly instruction format in the instruction stream. Thus, this field is optional in the sense that it is not required for instruction sets that only have the general vector friendly instruction format.
[0377] Basic operation field 3642 - its content distinguishes different basic operations.
[0378] Register index field 3644 - its contents specify the location of the source or destination operand in a register or in memory, either directly or through address generation. These fields include a sufficient number of bits to select N registers from a PxQ (e.g., 32x512, 16x128, 32x1024, 64x1024) register file. Although in one embodiment N may be up to three source registers and one destination register, alternative embodiments may support more or fewer source and destination registers (e.g., up to two sources may be supported, one of which may also serve as a destination; up to three sources may be supported, one of which may also serve as a destination; up to two sources and one destination may be supported).
[0379] Modifier field 3646 - its content distinguishes instructions appearing in the generic vector instruction format that specify memory access from instructions appearing in the generic vector instruction format that do not specify memory access; that is, between instruction templates with no memory access 3605 and instruction templates with memory access 3620. Memory access operations read and / or write to the memory hierarchy (in some cases, using values in registers to specify the source and / or destination addresses), while non-memory access operations do not (e.g., the source and / or destination are registers). Although in one embodiment, this field also selects between three different ways to perform memory address calculations, alternative embodiments may support more, fewer, or different ways to perform memory address calculations.
[0380] Extended Operation Field 3650 - Its contents distinguish which of various operations are to be performed in addition to the base operation. This field is context-specific. In one embodiment of the present disclosure, this field is divided into a Class field 3668, an Alpha field 3652, and a Beta field 3654. Extended Operation Field 3650 allows for common groups of operations to be performed in a single instruction rather than two, three, or four instructions.
[0381] Scale field 3660 - its content allows for memory address generation (e.g., for use with (2 比例* The address of (index + base address) is generated by scaling the contents of the index field.
[0382] Displacement field 3662A - its contents are used as part of memory address generation (e.g., for use with (2 比例 *Address generation of (index + base address + displacement).
[0383] Displacement Factor field 3662B (note that the juxtaposition of displacement field 3662A directly over displacement factor field 3662B indicates that one or the other is used) - its contents are used as part of address generation; it specifies the displacement factor by which the size (N) of the memory access will be scaled - where N is the number of bytes in the memory access (e.g., for use with (2 比例 * index + base address + scaled displacement). Redundant low-order bits are ignored, and therefore the contents of the displacement factor field are multiplied by the total size of the memory operand (N) to generate the final displacement to be used in calculating the effective address. The value of N is determined by the processor hardware at runtime based on the full opcode field 3674 (described later herein) and the data manipulation field 3654C. The displacement field 3662A and the displacement factor field 3662B are optional in the sense that they are not used in instruction templates without memory access 3605 and / or different embodiments may implement only one or neither of the two.
[0384] Data element width field 3664 - its contents distinguish which of multiple data element widths will be used (in some embodiments for all instructions; in other embodiments for only some of the instructions). This field is optional in the sense that it is not required if only one data element width is supported and / or some aspect of the opcode is used to support the data element width.
[0385] Writemask field 3670—its contents control, on a per-data-element-position basis, whether the data element positions in the destination vector operand reflect the results of the base and augmented operations. Class A instruction templates support merge-writemasking, while class B instruction templates support both merge-writemasking and zero-writemasking. When merged, the vector mask allows any set of elements in the destination to be protected from updates during the execution of any operation (specified by the base and augmented operations); in another embodiment, the old value of each element of the destination where the corresponding mask bit has a value of 0 is preserved. Conversely, when zeroed, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base and augmented operations); in one embodiment, elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span from the first to the last element being modified); however, the modified elements do not necessarily have to be contiguous. Thus, writemask field 3670 allows for partial vector operations, including loads, stores, arithmetic, logical, and more. Although embodiments of the present disclosure are described in which the contents of the write mask field 3670 select one of a plurality of write mask registers that contains the write mask to be used (and thereby, the contents of the write mask field 3670 indirectly identify the masking to be performed), alternative embodiments alternatively or additionally allow the contents of the mask write field 3670 to directly specify the masking to be performed.
[0386] Immediate field 3672 - its contents allow specification of an immediate value. This field is optional in the sense that it is not present in implementations that do not support the generic vector friendly format for immediate values and is not present in instructions that do not use immediate values.
[0387] Class field 3668 - its content distinguishes between instructions of different classes. Figure 36A-Figure 36B , the content of this field selects between class A and class B instructions. Figure 36A-Figure 36B In , a rounded square is used to indicate that a specific value exists in a field (for example, in Figure 36A-Figure 36B Class A 3668A and Class B 3668B for class field 3668 respectively).
[0388] Class A instruction template
[0389] In the case of the class A non-memory access 3605 instruction templates, the α field 3652 is interpreted as an RS field 3652A whose content distinguishes which of the different extended operation types is to be performed (e.g., the instruction templates for the no-memory access rounding operation 3610 and the no-memory access data transformation operation 3615 specify rounding 3652A.1 and data transformation 3652A.2, respectively), while the β field 3654 distinguishes which of the specified types of operations is to be performed. In the no-memory access 3605 instruction templates, the scale field 3660, the displacement field 3662A, and the displacement scale field 3662B are not present.
[0390] Instruction templates with no memory access – fully rounded control type operations
[0391] In the instruction template for the no-memory-access, full-round-controlled operation 3610, the beta field 3654 is interpreted as a round control field 3654A that provides static rounding for its contents. Although in the described embodiment of the present disclosure, the round control field 3654A includes a suppress all floating-point exceptions (SAE) field 3656 and a round operation control field 3658, alternative embodiments may support both concepts, may encode both concepts into the same field, or may have only one or the other of these concepts / fields (e.g., may have only the round operation control field 3658).
[0392] SAE field 3656 - its content distinguishes whether exception event reporting is disabled; when the content of SAE field 3656 indicates that suppression is enabled, the given instruction does not report any kind of floating point exception flags and does not invoke any floating point exception handler.
[0393] Round operation control field 3658 - its contents distinguish which of a set of rounding operations is to be performed (e.g., round up, round down, round toward zero, and round to nearest). Thus, the round operation control field 3658 allows the rounding mode to be changed on an instruction-by-instruction basis. In one embodiment of the present disclosure in which the processor includes a control register for specifying the rounding mode, the contents of the round operation control field 3650 override the register value.
[0394] Instruction templates without memory access - data transformation operations
[0395] In the instruction template for the no-memory-access data transform type operation 3615, the beta field 3654 is interpreted as a data transform field 3654B, the contents of which distinguish which of multiple data transforms is to be performed (eg, no data transform, blend, broadcast).
[0396] In the case of an instruction template of type A memory access 3620, the alpha field 3652 is interpreted as an eviction hint field 3652B, the contents of which distinguish which of the eviction hints is to be used (in Figure 36A In the example, the instruction templates for memory access temporal 3625 and memory access non-temporal 3630 specify temporal 3652B.1 and non-temporal 3652B.2, respectively, and the beta field 3654 is interpreted as a data manipulation field 3654C, the content of which distinguishes which of a plurality of data manipulation operations (also called primitives) is to be performed (e.g., no manipulation, broadcast, upcast of the source, and downcast of the destination). The instruction template for memory access 3620 includes a scale field 3660 and optionally includes a displacement field 3662A or a displacement scale field 3662B.
[0397] Vector memory instructions use the translation support to perform vector loads from memory and vector stores to memory. Like normal vector instructions, vector memory instructions transfer data to and from memory in an element-wise manner, where the actual elements transferred are specified by the contents of the vector mask selected as the writemask.
[0398] Memory access instruction templates - time-sensitive
[0399] Temporal data is data that is likely to be reused quickly enough to benefit from cache operations. However, this is a hint, and different processors can implement it in different ways, including ignoring the hint completely.
[0400] Memory access instruction templates - non-temporal
[0401] Non-temporal data is data that is unlikely to be reused quickly enough to benefit from cache operations in the first level cache and should be given eviction priority. However, this is a hint, and different processors can implement it in different ways, including ignoring the hint completely.
[0402] Class B instruction template
[0403] In the case of class B instruction templates, the alpha field 3652 is interpreted as a write mask control (Z) field 3652C, the contents of which distinguish whether the write mask controlled by the write mask field 3670 should be merge or zero.
[0404] In the case of a Class B non-memory access 3605 instruction template, a portion of the β field 3654 is interpreted as an RL field 3657A, the contents of which distinguish which of the different extended operation types is to be performed (e.g., the instruction template for the no-memory-access write mask control partial rounding control type operation 3612 and the instruction template for the no-memory-access write mask control VSIZE type operation 3617 specify rounding 3657A.1 and vector length (VSIZE) 3657A.2, respectively), while the remainder of the β field 3654 distinguishes which of the specified types of operations is to be performed. In the no-memory-access 3605 instruction template, the scale field 3660, the displacement field 3662A, and the displacement scale field 3662B are not present.
[0405] In the instruction template for the write mask control partial round control type operation 3610 with no memory access, the remainder of the β field 3654 is interpreted as the round operation field 3659A, and exception event reporting is disabled (the given instruction does not report any kind of floating-point exception flag and does not invoke any floating-point exception handler).
[0406] Round Operation Control Field 3659A - As with the Round Operation Control Field 3658, its contents distinguish which of a set of rounding operations (e.g., round up, round down, round toward zero, and round to nearest) is to be performed. Thus, the Round Operation Control Field 3659A allows the rounding mode to be changed on an instruction-by-instruction basis. In one embodiment of the present disclosure in which the processor includes a control register for specifying the rounding mode, the contents of the Round Operation Control Field 3650 override the register value.
[0407] In the instruction template for the write mask control VSIZE type operation 3617 without memory access, the remainder of the β field 3654 is interpreted as a vector length field 3659B, the contents of which distinguish which of multiple data vector lengths is to be performed (e.g., 128 bytes, 256 bytes, or 512 bytes).
[0408] In the case of a memory access 3620 instruction template of type B, a portion of the beta field 3654 is interpreted as a broadcast field 3657B, the contents of which distinguish whether a broadcast-type data manipulation operation is to be performed, while the remainder of the beta field 3654 is interpreted as a vector length field 3659B. The memory access 3620 instruction template includes a scale field 3660 and optionally includes a displacement field 3662A or a displacement scale field 3662B.
[0409] For the generic vector friendly instruction format 3600, the full opcode field 3674 is shown to include a format field 3640, a base operation field 3642, and a data element width field 3664. Although one embodiment is shown in which the full opcode field 3674 includes all of these fields, in embodiments that do not support all of these fields, the full opcode field 3674 includes less than all of these fields. The full opcode field 3674 provides an operation code (opcode).
[0410] The augment operation field 3650, the data element width field 3664, and the write mask field 3670 allow these features to be specified on an instruction-by-instruction basis in the generic vector friendly instruction format.
[0411] The combination of the write mask field and the data element width field creates various types of instructions because these instructions allow the mask to be applied based on different data element widths.
[0412] The various instruction templates found within classes A and B are beneficial in different situations. In some embodiments of the present disclosure, different processors or different cores within a processor may support only class A, only class B, or both classes. For example, a high-performance general-purpose out-of-order core intended for general-purpose computing may support only class B, a core intended primarily for graphics and / or scientific (throughput) computing may support only class A, and a core intended for both general-purpose computing and graphics and / or scientific (throughput) computing may support both classes A and B (of course, cores with some mix of templates and instructions from both classes, but not all templates and instructions from both classes, are within the scope of this disclosure). Similarly, a single processor may include multiple cores, all of which support the same class, or different cores in which support different classes. For example, in a processor with separate graphics and general-purpose cores, one of the graphics cores intended primarily for graphics and / or scientific computing may support only class A, while one or more of the general-purpose cores may be a high-performance general-purpose core with out-of-order execution and register renaming intended for general-purpose computing that supports only class B. Another processor that does not have a separate graphics core may include one or more general-purpose in-order or out-of-order cores that support both class A and class B. Of course, in different embodiments of the present disclosure, features from one class may also be implemented in the other class. A program written in a high-level language will be made (e.g., just-in-time or statically compiled) into a variety of different executable forms, including: 1) a form with only instructions of the class(es) supported by the target processor for execution; or 2) a form with alternative routines written using different combinations of instructions from all classes and with control flow code that selects these routines to execute based on the instructions supported by the processor currently executing the code.
[0413] Exemplary dedicated vector friendly instruction format
[0414] Figure 37 is a block diagram illustrating an exemplary dedicated vector friendly instruction format according to an embodiment of the present disclosure. Figure 37 shows a dedicated vector friendly instruction format 3700, which specifies the position, size, interpretation and order of each field, as well as the values of some of those fields. In this sense, the dedicated vector friendly instruction format 3700 is dedicated. The dedicated vector friendly instruction format 3700 can be used to extend the x86 instruction set, and some of the fields are similar or identical to those used in the existing x86 instruction set and its extensions (e.g., AVX). The format remains consistent with the prefix encoding field, real opcode byte field, MOD R / M field, SIB field, displacement field, and immediate field of the existing x86 instruction set with the extension. The fields from Figure 36 are illustrated, and the fields from Figure 37 are mapped to the fields from Figure 36.
[0415] It should be understood that although embodiments of the present disclosure are described with reference to the specific vector friendly instruction format 3700 in the context of the general vector friendly instruction format 3600 for illustrative purposes, the present disclosure is not limited to the specific vector friendly instruction format 3700 unless otherwise stated. For example, the general vector friendly instruction format 3600 contemplates various possible sizes for various fields, while the specific vector friendly instruction format 3700 is illustrated as having fields of specific sizes. As a specific example, although the data element width field 3664 is illustrated as a one-bit field in the specific vector friendly instruction format 3700, the present disclosure is not limited thereto (i.e., the general vector friendly instruction format 3600 contemplates other sizes for the data element width field 3664).
[0416] The generic vector friendly instruction format 3600 includes the following listed in accordance with Figure 37A The following fields are in the order shown in the figure.
[0417] EVEX prefix (bytes 0-3) 3702 - encoded in four bytes.
[0418] Format field 3640 (EVEX byte 0, bits [7:0]) - The first byte (EVEX byte 0) is the format field 3640, and it contains 0x62 (a unique value used to distinguish the vector friendly instruction format in one embodiment of the present disclosure).
[0419] The second through fourth bytes (EVEX bytes 1-3) include a number of bit fields that provide specific capabilities.
[0420] REX field 3705 (EVEX byte 1, bits [7-5]) - consists of the EVEX.R bit field (EVEX byte 1, bits [7]–R), the EVEX.X bit field (EVEX byte 1, bits [6]–X), and (3657BEX byte 1, bits [5]–B). The EVEX.R, EVEX.X, and EVEX.B bit fields provide the same functionality as the corresponding VEX bit fields and are encoded using 1's complement form, i.e., ZMM0 is encoded as 1111B and ZMM15 is encoded as 0000B. The other fields of these instructions encode the lower three bits of the register index (rrr, xxx, and bbb) as known in the art, whereby Rrrr, Xxxx, and Bbbb can be formed by adding EVEX.R, EVEX.X, and EVEX.B.
[0421] REX' field 3610 - This is the first portion of the REX' field 3610 and is the EVEX.R' bit field (EVEX byte 1, bit [4] - R') used to encode the upper 16 or lower 16 registers of the extended 32 register set. In one embodiment of the present disclosure, this bit is stored in a bit-reversed format along with the other bits indicated below to distinguish it from the BOUND instruction (in 32-bit mode of the well-known x86) whose real opcode byte is 62, but does not accept a value of 11 in the MOD field in the MODR / M field (described below); alternative embodiments of the present disclosure do not store this indicated bit and the other indicated bits below in a reversed format. A value of 1 is used to encode the lower 16 registers. In other words, R'Rrrr is formed by combining EVEX.R', EVEX.R, and the other RRRs from the other fields.
[0422] Opcode map field 3715 (EVEX byte 1, bits [3:0] – mmmm) – its contents encode the implied leading opcode byte (0F, 0F 38, or 0F 3).
[0423] Data element width field 3664 (EVEX byte 2, bit [7] – W) – denoted by the notation EVEX.W. EVEX.W is used to define the granularity (size) of the data type (32-bit data element or 64-bit data element).
[0424] EVEX.vvvv 3720 (EVEX byte 2, bits [6:3]-vvvv) - The purpose of EVEX.vvvv may include the following: 1) EVEX.vvvv encodes the first source register operand specified in inverted (one's complement) form and is valid for instructions with two or more source operands; 2) EVEX.vvvv encodes the destination register operand specified in one's complement form for a particular vector displacement; or 3) EVEX.vvvv does not encode any operand; this field is reserved and should contain 1111b. Thus, EVEX.vvvv field 3720 encodes the four low-order bits of the first source register designator stored in inverted (one's complement) form. Depending on the instruction, additional different EVEX bit fields are used to extend the designator size to 32 registers.
[0425] EVEX.U 3668 Class field (EVEX byte 2, bit [2] - U) - If EVEX.U = 0, it indicates Class A or EVEX.U0; if EVEX.U = 1, it indicates Class B or EVEX.U1.
[0426] Prefix encoding field 3725 (EVEX byte 2, bits [1:0]-pp)—provides additional bits for the base operation field. In addition to supporting legacy SSE instructions in EVEX prefix format, this also has the benefit of compressing the SIMD prefix (the EVEX prefix only requires 2 bits, rather than a byte to express the SIMD prefix). In one embodiment, to support legacy SSE instructions using SIMD prefixes (66H, F2H, F3H) in both legacy and EVEX prefix formats, these legacy SIMD prefixes are encoded into the SIMD prefix encoding field; and at runtime, they are expanded into legacy SIMD prefixes before being provided to the decoder's PLA (thus, the PLA can execute these legacy instructions in both legacy and EVEX formats without modification). While newer instructions may use the contents of the EVEX prefix encoding field directly as an opcode extension, certain embodiments extend this in a similar manner for consistency, but allow for different meanings specified by these legacy SIMD prefixes. Alternative embodiments may redesign the PLA to support 2-bit SIMD prefix encodings, thereby eliminating the need for expansion.
[0427] Alpha field 3652 (EVEX byte 3, bit [7] - EH, also known as EVEX.EH, EVEX.rs, EVEX.RL, EVEX.WriteMaskControl, and EVEX.N; also illustrated as alpha) - As previously described, this field is context specific.
[0428] Beta field 3654 (EVEX byte 3, bits [6:4] - SSS, also known as EVEX.s 2-0 EVEX.r 2-0 , EVEX.rr1, EVEX.LL0, EVEX.LLB, also illustrated as βββ) - As mentioned earlier, this field is context-specific.
[0429] REX' field 3610 - This is the remainder of the REX' field and is the EVEX.V' bit field (EVEX byte 3, bit [3] - V') that can be used to encode the upper 16 or lower 16 registers of the extended 32 register set. This bit is stored in a bit-reversed format. A value of 1 is used to encode the lower 16 registers. In other words, V'VVVV is formed by combining EVEX.V', EVEX.vvvv.
[0430] Write mask field 3670 (EVEX byte 3, bits [2:0] - kkk) - its contents specify the index of a register in the write mask register, as previously described. In one embodiment of the present disclosure, a special value of EVEX.kkk = 000 has special behavior that implies no write mask is used for a particular instruction (this can be implemented in various ways, including using a write mask hardwired to all objects or hardware that bypasses the masking hardware).
[0431] The real opcode field 3730 (byte 4) is also called the opcode byte. A portion of the opcode is specified in this field.
[0432] The MOD R / M field 3740 (byte 5) includes a MOD field 3742, a Reg field 3744, and an R / M field 3746. As previously described, the contents of the MOD field 3742 distinguish memory access operations from non-memory access operations. The role of the Reg field 3744 can be summarized as follows: encoding a destination register operand or a source register operand; or being treated as an opcode extension and not used to encode any instruction operand. The role of the R / M field 3746 can include the following: encoding an instruction operand that references a memory address; or encoding a destination register operand or a source register operand.
[0433] Scale, Index, Base (SIB) Byte (Byte 6) - As previously mentioned, the contents of the scale field 3650 are used for memory address generation. SIB.xxx 3754 and SIB.bbb 3756 - The contents of these fields have been mentioned previously for register indices Xxxx and Bbbb.
[0434] Displacement field 3662A (bytes 7-10) - When the MOD field 3742 contains 10, bytes 7-10 are the displacement field 3662A, and it works the same as a traditional 32-bit displacement (disp32) and works at byte granularity.
[0435] Displacement Factor Field 3662B (Byte 7)—When MOD field 3742 contains 01, byte 7 is the displacement factor field 3662B. This field is located in the same location as the traditional x86 instruction set 8-bit displacement (disp8), which operates at byte granularity. Because disp8 is sign-extended, it can only address offsets between -128 and 127 bytes. In terms of a 64-byte cache line, disp8 uses 8 bits that can be set to only four truly useful values: 128, -64, 0, and 64. Because a larger range is often needed, disp32 is used; however, disp32 requires 4 bytes. In contrast to disp8 and disp32, displacement factor field 3662B is a reinterpretation of disp8. When displacement factor field 3662B is used, the actual displacement is determined by multiplying the contents of the displacement factor field by the size (N) of the memory operand being accessed. This type of displacement is referred to as disp8*N. This reduces the average instruction length (a single byte is used for the displacement, but with a much larger range). This type of compressed displacement is based on the assumption that the effective displacement is a multiple of the granularity of the memory access, and thus the redundant low-order bits of the address offset do not need to be encoded. In other words, the displacement factor field 3662B replaces the traditional x86 instruction set 8-bit displacement. Thus, the displacement factor field 3662B is encoded in the same manner as the x86 instruction set 8-bit displacement (thus, there is no change in the ModRM / SIB encoding rules), the only difference being that disp8 is overloaded to disp8*N. In other words, there is no change in the encoding rules or encoding length, only in the hardware's interpretation of the displacement value (which requires scaling the displacement by the size of the memory operand to obtain a byte-based address offset). The immediate field 3672 operates as previously described.
[0436] Full opcode field
[0437] Figure 37B 36 is a block diagram illustrating the fields of a specific vector friendly instruction format 3700 that make up a full opcode field 3674 according to one embodiment of the present disclosure. Specifically, the full opcode field 3674 includes a format field 3640, a base operation field 3642, and a data element width (W) field 3664. The base operation field 3642 includes a prefix encoding field 3725, an opcode map field 3715, and a real opcode field 3730.
[0438] Register index field
[0439] Figure 37C 37 is a block diagram illustrating fields of a specific vector friendly instruction format 3700 that make up the register index field 3644 according to one embodiment of the present disclosure. Specifically, the register index field 3644 includes a REX field 3705, a REX' field 3710, a MODR / M.reg field 3744, a MODR / Mr / m field 3746, a VVVV field 3720, a xxx field 3754, and a bbb field 3756.
[0440] Expand the operation field
[0441] Figure 37D 3652A. This is a block diagram illustrating the fields of the specific vector friendly instruction format 3700 that make up the extended operation field 3650 according to one embodiment of the present disclosure. When the class (U) field 3668 contains 0, it indicates EVEX.U0 (class A 3668A); when it contains 1, it indicates EVEX.U1 (class B 3668B). When U=0 and the MOD field 3742 contains 11 (indicating no memory access operation), the alpha field 3652 (EVEX byte 3, bit [7]-EH) is interpreted as the rs field 3652A. When the rs field 3652A contains 1 (round 3652A.1), the beta field 3654 (EVEX byte 3, bits [6:4]-SSS) is interpreted as the round control field 3654A. The round control field 3654A includes a one-bit SAE field 3656 and a two-bit round operation field 3658. When the rs field 3652A contains 0 (data transformation 3652A.2), the beta field 3654 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data transformation field 3654B. When U=0 and the MOD field 3742 contains 00, 01, or 10 (indicating a memory access operation), the alpha field 3652 (EVEX byte 3, bits [7]-EH) is interpreted as an eviction hint (EH) field 3652B, and the beta field 3654 (EVEX byte 3, bits [6:4]-SSS) is interpreted as a three-bit data manipulation field 3654C.
[0442] When U=1, the alpha field 3652 (EVEX byte 3, bit [7]–EH) is interpreted as the write mask control (Z) field 3652C. When U=1 and the MOD field 3742 contains 11 (indicating no memory access operation), a portion of the beta field 3654 (EVEX byte 3, bit [4]–S0) is interpreted as the RL field 3657A; when it contains 1 (rounded 3657A.1), the remainder of the beta field 3654 (EVEX byte 3, bits [6-5]–S0) is interpreted as the RL field 3657A. 2-1) is interpreted as the rounding operation field 3659A, and when the RL field 3657A contains 0 (VSIZE 3657.A2), the remainder of the beta field 3654 (EVEX byte 3, bits [6-5]-S 2-1 ) is interpreted as the vector length field 3659B (EVEX byte 3, bits [6-5]–L 1-0 When U=1 and the MOD field 3742 contains 00, 01, or 10 (indicating a memory access operation), the beta field 3654 (EVEX byte 3, bits [6:4]–SSS) is interpreted as the vector length field 3659B (EVEX byte 3, bits [6-5]–L 1-0 ) and the broadcast field 3657B (EVEX byte 3, bit [4]–B).
[0443] Exemplary Register Architecture
[0444] Figure 38 is a block diagram of a register architecture 3800 according to one embodiment of the present disclosure. In the illustrated embodiment, there are 32 512-bit wide vector registers 3810; these registers are referenced as zmm0 through zmm31. The lower-order 256 bits of the lower 16 zmm registers are overlaid on registers ymm0-16. The lower-order 128 bits of the lower 16 zmm registers (the lower-order 128 bits of the ymm registers) are overlaid on registers xmm0-15. The specific vector friendly instruction format 3700 operates on these overlaid register files, as illustrated in the following table.
[0445]
[0446] In other words, the vector length field 3659B selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the previous length, and instruction templates that do not have a vector length field 3659B operate on the maximum vector length. In addition, in one embodiment, the class B instruction templates of the dedicated vector friendly instruction format 3700 operate on packed or scalar single / double precision floating point data and packed or scalar integer data. Scalar operations are operations performed on the lowest-order data element position in the zmm / ymm / xmm register; depending on the embodiment, higher-order data element positions either remain the same as before the instruction or are reset to zero.
[0447] Write mask registers 3815 - In the illustrated embodiment, there are eight write mask registers (k0 through k7), each 64 bits in size. In an alternative embodiment, the write mask registers 3805 are 16 bits in size. As previously mentioned, in one embodiment of the present disclosure, vector mask register k0 cannot be used as a write mask; when the encoding that normally indicates k0 is used as a write mask, it selects a hardwired write mask of 0xFFFF, effectively disabling write masking for that instruction.
[0448] General purpose registers 3825 - In the embodiment shown, there are sixteen 64-bit general purpose registers that are used with existing x86 addressing modes to address memory operands. These registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.
[0449] A scalar floating-point stack register file (x87 stack) 3845 on which the MMX packed integer flat register file 3850 is overlaid - in the illustrated embodiment, the x87 stack is an eight-element stack used to perform scalar floating-point operations on 32 / 64 / 80-bit floating-point data using the x87 instruction set extension; while MMX registers are used to perform operations on 64-bit packed integer data, as well as to hold operands for some operations performed between MMX and XMM registers.
[0450] Alternative embodiments of the present disclosure may use wider or narrower registers. Additionally, alternative embodiments of the present disclosure may use more, fewer, or different register files and registers.
[0451] Exemplary Core Architectures, Processors, and Computer Architectures
[0452] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores may include: 1) a general-purpose in-order core intended for general-purpose computing; 2) a high-performance general-purpose out-of-order core intended for general-purpose computing; 3) a specialized core intended primarily for graphics and / or scientific (throughput) computing. Different processor implementations may include: 1) a CPU that includes one or more general-purpose in-order cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores intended for general-purpose computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors lead to different computer system architectures, which may include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case such a coprocessor is sometimes referred to as dedicated logic or as a dedicated core, such as integrated graphics and / or scientific (throughput) logic); and 4) a system on a chip, which may include the described CPU (sometimes referred to as application core(s) or application processor(s), the coprocessor described above, and additional functionality on the same die. An exemplary core architecture is described next, followed by a description of an exemplary processor and computer architecture.
[0453] Exemplary Core Architecture
[0454] In-order and out-of-order core block diagram
[0455] Figure 39A is a block diagram illustrating an exemplary in-order pipeline and an exemplary register-renaming out-of-order issue / execution pipeline according to various embodiments of the present disclosure. Figure 39B is a block diagram illustrating an exemplary embodiment of an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor according to various embodiments of the present disclosure. Figure 39A-Figure 39B The solid line boxes in illustrate the in-order pipeline and in-order core, while the optional addition of dashed line boxes illustrates the register renaming, out-of-order issue / execution pipeline and core. Considering that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0456] exist Figure 39A , the processor pipeline 3900 includes a fetch stage 3902, a length decode stage 3904, a decode stage 3906, an allocation stage 3908, a rename stage 3910, a schedule (also called a dispatch or issue) stage 3912, a register read / memory read stage 3914, an execute stage 3916, a write back / memory write stage 3918, an exception handling stage 3922, and a commit stage 3924.
[0457] Figure 39BA processor core 3990 is shown, comprising a front end unit 3930 coupled to an execution engine unit 3950, and both the front end unit 3930 and the execution engine unit 3950 are coupled to a memory unit 3970. The core 3990 may be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 3990 may be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and the like.
[0458] The front end unit 3930 includes a branch prediction unit 3932, which is coupled to an instruction cache unit 3934, which is coupled to an instruction translation lookaside buffer (TLB) 3936, which is coupled to an instruction fetch unit 3938, which is coupled to a decode unit 3940. The decode unit 3940 (or decoder or decode unit) can decode an instruction (e.g., a macroinstruction) and generate as output one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals decoded from, or otherwise reflecting, or derived from, the original instruction. The decode unit 3940 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLA), microcode read-only memories (ROMs), and the like. In one embodiment, core 3990 includes a microcode ROM or other medium that stores microcode for certain macroinstructions (e.g., in decode unit 3940 or otherwise within front end unit 3930). Decode unit 3940 is coupled to rename / allocator unit 3952 in execution engine unit 3950.
[0459] The execution engine unit 3950 includes a rename / allocator unit 3952, which is coupled to a retirement unit 3954 and a set of one or more scheduler units 3956. The scheduler unit(s) 3956 represent any number of different schedulers, including reservation stations, central instruction windows, and the like. The scheduler unit(s) 3956 are coupled to physical register file(s) units 3958. Each of the physical register file(s) units 3958 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating point, packed integers, packed floating point, vector integers, vector floating point, state (e.g., an instruction pointer, which is the address of the next instruction to be executed), and the like. In one embodiment, the physical register file(s) units 3958 include a vector register unit, a write mask register unit, and a scalar register unit. These register units may provide architectural vector registers, vector mask registers, and general purpose registers. Physical register file(s) units 3958 are overlapped by retirement units 3954 to illustrate various ways in which register renaming and out-of-order execution can be implemented (e.g., using reorder buffer(s) and retirement register file(s); using future files(s), history buffer(s), retirement register file(s); using register maps and register pools, etc.). Retirement units 3954 and physical register file(s) units 3958 are coupled to execution cluster(s) 3960. Execution cluster(s) 3960 include a set of one or more execution units 3962 and a set of one or more memory access units 3964. Execution units 3962 can perform various operations (e.g., shifts, additions, subtractions, multiplications) and can operate on various data types (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some embodiments may include multiple execution units dedicated to a particular function or set of functions, other embodiments may include only one execution unit or multiple execution units that all perform all functions. Scheduler unit(s) 3956, physical register file(s) 3958, and execution cluster(s) 3960 are shown as potentially multiple because some embodiments create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each with its own scheduler unit(s), physical register file(s) units, and / or execution cluster—and in the case of separate memory access pipelines, certain embodiments are implemented in which only the execution cluster for that pipeline has memory access unit(s) 3964). It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution, and the remaining pipelines may be in-order.
[0460] A set of memory access units 3964 is coupled to a memory unit 3970, which includes a data TLB unit 3972, which is coupled to a data cache unit 3974, which is coupled to a second level (L2) cache unit 3976. In one exemplary embodiment, the memory access unit 3964 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 3972 in the memory unit 3970. The instruction cache unit 3934 is also coupled to a second level (L2) cache unit 3976 in the memory unit 3970. The L2 cache unit 3976 is coupled to one or more other levels of cache and ultimately to main memory.
[0461] As an example, the exemplary register renaming out-of-order issue / execution core architecture may implement the pipeline 3900 as follows: 1) instruction fetch 3938 executes the fetch stage 3902 and the length decode stage 3904; 2) the decode unit 3940 executes the decode stage 3906; 3) the rename / allocator unit 3952 executes the allocate stage 3908 and the rename stage 3910; 4) (multiple) scheduler units 3956 execute the schedule stage 3912; 5) (multiple) physical register file units 3958 and memory units 3970 execute the register read / memory read stage 3914; the execution cluster 3960 executes the execute stage 3916; 6) the memory unit 3970 and the (multiple) physical register file units 3958 execute the write back / memory write stage 3918; 7) each unit may be involved in the exception handling stage 3922; and 8) the retirement unit 3954 and the (multiple) physical register file units 3958 execute the commit stage 3924.
[0462] Core 3990 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions that have been added with newer versions); the MIPS instruction set from MIPS Technologies, Inc. of Sunnyvale, California; the ARM instruction set from ARM Holdings, Inc. of Sunnyvale, California (with optional additional extensions such as NEON)), including the instruction(s) described herein. In one embodiment, core 3990 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.
[0463] It should be understood that a core may support multithreading (executing two or more sets of operations or threads in parallel) and that this multithreading may be accomplished in a variety of ways, including time-shared multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads that the physical core is simultaneously multithreading), or a combination thereof (e.g., time-shared fetch and decode and subsequent operations such as Hyper-Threading (Simultaneous Multithreading).
[0464] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. Although the illustrated embodiment of the processor also includes separate instruction and data cache units 3934 / 3974 and a shared L2 cache unit 3976, alternative embodiments may have a single internal cache for both instructions and data, such as, for example, a first level (L1) internal cache or multiple levels of internal cache. In some embodiments, the system may include a combination of internal caches and external caches external to the core and / or processor. Alternatively, all caches may be external to the core and / or processor.
[0465] Specific exemplary in-order core architecture
[0466] Figures 40A-40B A block diagram illustrating a more specific exemplary in-order core architecture is shown, which would be one logic block among several logic blocks in a chip (including other cores of the same and / or different types). Depending on the application, the logic block communicates with some fixed function logic, memory I / O interfaces, and other necessary I / O logic via a high-bandwidth interconnect network (e.g., a ring network).
[0467] Figure 40A 4004 and 4016. The block diagram of a single processor core and its connection to an on-die interconnect network 4002 and its local subset 4004 of a level 2 (L2) cache according to an embodiment of the present disclosure. In one embodiment, the instruction decode unit 4000 supports the x86 instruction set with the packed data instruction set extension. The L1 cache 4006 allows low-latency access to cache memory in the scalar and vector units. Although in one embodiment (to simplify the design), the scalar unit 4008 and the vector unit 4010 use separate register sets (scalar registers 4012 and vector registers 4014, respectively), and data transferred between these registers is written to memory and then read back from the level 1 (L1) cache 4006, alternative embodiments of the present disclosure may use a different approach (e.g., using a single register set or including a communication path that allows data to be transferred between the two register files without being written and read back).
[0468] The local subset 4004 of the L2 cache is part of the global L2 cache, which is divided into multiple separate local subsets, one for each processor core. Each processor core has a direct access path to its own local subset 4004 of the L2 cache. Data read by a processor core is stored in its L2 cache subset 4004 and can be quickly accessed in parallel with other processor cores accessing their own local L2 cache subsets. Data written by a processor core is stored in its own L2 cache subset 4004 and flushed from other subsets when necessary. The ring network ensures the consistency of shared data. The ring network is bidirectional to allow agents such as processor cores, L2 caches, and other logic blocks to communicate with each other within the chip. Each ring data path is 1012 bits wide in each direction.
[0469] Figure 40B According to an embodiment of the present disclosure Figure 40A An expanded view of a portion of a processor core in FIG. Figure 40B Includes the L1 data cache 4006A portion of the L1 cache 4004, as well as more details about the vector unit 4010 and vector registers 4014. Specifically, the vector unit 4010 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 4028) that executes one or more of integer, single-precision floating-point, and double-precision floating-point instructions. The VPU supports blending of register inputs via blend unit 4020, numerical conversion via numerical conversion units 4022A-B, and copying of memory inputs via copy unit 4024. Write mask register 4026 allows for predictive vector writes.
[0470] Figure 41 is a block diagram of a processor 4100 that may have more than one core, may have an integrated memory controller, and may have integrated graphics, according to an embodiment of the present disclosure. Figure 41 The solid line box in the figure illustrates a processor 4100 having a single core 4102A, a system agent 4110, a set 4116 of one or more bus controller units, while the optional addition of the dashed line box illustrates an alternative processor 4100 having multiple cores 4102A-N, a set 4114 of one or more integrated memory controller units in the system agent unit 4110, and dedicated logic 4108.
[0471] Thus, different implementations of processor 4100 may include: 1) a CPU, wherein specialized logic 4108 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores), and cores 4102A-N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of the two); 2) a coprocessor, wherein cores 4102A-N are a large number of specialized cores intended primarily for graphics and / or scientific (throughput); and 3) a coprocessor, wherein cores 4102A-N are a large number of general-purpose in-order cores. Thus, processor 4100 may be a general-purpose processor, a coprocessor, or a specialized processor, such as, for example, a network or communications processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput many-integrated-core (MIC) coprocessor (including 30 or more cores), an embedded processor, or the like. The processor may be implemented on one or more chips. Processor 4100 may be part of one or more substrates and / or may be implemented on one or more substrates using any of a variety of process technologies (such as, for example, BiCMOS, CMOS, or NMOS).
[0472] The memory hierarchy includes one or more cache levels within the core, a set of one or more shared cache units 4106, and external memory (not shown) coupled to a set of integrated memory controller units 4114. The set of shared cache units 4106 may include one or more intermediate levels of cache, such as a level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, a last level cache (LLC), and / or combinations thereof. While in one embodiment, a ring-based interconnect unit 4112 interconnects the integrated graphics logic 4108, the set of shared cache units 4106, and the system agent unit 4110 / integrated memory controller unit(s) 4114, alternative embodiments may use any number of well-known techniques to interconnect such units. In one embodiment, coherency is maintained between the one or more cache units 4106 and the cores 4102A-N.
[0473] In some embodiments, one or more cores 4102A-N may be multithreaded. System agent 4110 includes components that coordinate and operate cores 4102A-N. System agent unit 4110 may include, for example, a power control unit (PCU) and a display unit. The PCU may include, or may include, logic and components required to regulate the power state of cores 4102A-N and integrated graphics logic 4108. The display unit is used to drive one or more externally connected displays.
[0474] The cores 4102A-N may be homogeneous or heterogeneous with respect to the architectural instruction set; that is, two or more of the cores 4102A-N may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of the instruction set or a different instruction set.
[0475] Exemplary Computer Architecture
[0476] Figures 42-45 is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptops, desktops, handheld PCs, personal digital assistants, engineering workstations, servers, network appliances, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, a wide variety of systems or electronic devices that can include a processor and / or other execution logic as disclosed herein are generally suitable.
[0477] Now refer to Figure 42 , shown is a block diagram of a system 4200 according to one embodiment of the present disclosure. System 4200 may include one or more processors 4210, 4215 coupled to a controller hub 4220. In one embodiment, controller hub 4220 includes a graphics memory controller hub (GMCH) 4290 and an input / output hub (IOH) 4250 (which may be on separate chips); GMCH 4290 includes memory and a graphics controller, to which memory 4240 and coprocessor 4245 are coupled; and IOH 4250 couples input / output (I / O) devices 4260 to GMCH 4290. Alternatively, one or both of the memory and graphics controller are integrated within the processor (as described herein), with memory 4240 and coprocessor 4245 directly coupled to processor 4210, and controller hub 4220 and IOH 4250 being on a single chip. The memory 4240 may include a compiler module 4240A, which is used to store code, for example, which when executed causes the processor to perform any method of the present disclosure.
[0478] The optional addition of processor 4215 is Figure 42 Each processor 4210 , 4215 may include one or more of the processing cores described herein and may be a version of processor 4100 .
[0479] The memory 4240 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), or a combination of the two. For at least one embodiment, the controller hub 4220 communicates with the processor(s) 4210, 4215 via a multi-drop bus such as a front-side bus (FSB), a point-to-point interface such as a Quick Path Interconnect (QPI), or a similar connection 4295.
[0480] In one embodiment, coprocessor 4245 is a special purpose processor such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, etc. In one embodiment, controller hub 4220 may include an integrated graphics accelerator.
[0481] There may be various differences between the physical resources 4210 and 4215 in terms of a range of quality metrics including architectural, microarchitectural, thermal, and power consumption characteristics.
[0482] In one embodiment, processor 4210 executes instructions that control general types of data processing operations. Embedded within these instructions may be coprocessor instructions. Processor 4210 recognizes these coprocessor instructions as being of a type that should be executed by attached coprocessor 4245. Accordingly, processor 4210 issues these coprocessor instructions (or control signals representing coprocessor instructions) to coprocessor 4245 over a coprocessor bus or other interconnect. Coprocessor(s) 4245 accept and execute the received coprocessor instructions.
[0483] Now see Figure 43 , shown is a block diagram of a first more specific exemplary system 4300 according to an embodiment of the present disclosure. Figure 43 As shown in FIG, multiprocessor system 4300 is a point-to-point interconnect system and includes a first processor 4370 and a second processor 4380 coupled via a point-to-point interconnect 4350. Each of processors 4370 and 4380 may be a version of processor 4100. In one embodiment of the present disclosure, processors 4370 and 4380 are processors 4210 and 4215, respectively, and coprocessor 4338 is coprocessor 4245. In another embodiment, processors 4370 and 4380 are processor 4210 and coprocessor 4245, respectively.
[0484] Processors 4370 and 4380 are shown as including integrated memory controller (IMC) units 4372 and 4382, respectively. Processor 4370 also includes point-to-point (PP) interfaces 4376 and 4378 as part of its bus controller unit; similarly, second processor 4380 includes PP interfaces 4386 and 4388. Processors 4370, 4380 can exchange information via PP interface 4350 using point-to-point (PP) interface circuits 4378, 4388. Figure 43 As shown in FIG, IMCs 4372 and 4382 couple the processors to respective memories, namely, memory 4332 and memory 4334, which may be portions of main memory locally attached to the respective processors.
[0485] Processors 4370, 4380 may each exchange information with a chipset 4390 via respective PP interfaces 4352, 4354 using point-to-point interface circuits 4376, 4394, 4386, 4398. Chipset 4390 may optionally exchange information with a coprocessor 4338 via a high-performance interface 4339. In one embodiment, coprocessor 4338 is a special-purpose processor such as, for example, a high-throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, or the like.
[0486] A shared cache (not shown) may be included in either processor, or external to both processors but connected to the processors via the PP interconnect, such that if the processors are placed in a low power mode, local cache information of either or both processors may be stored in the shared cache.
[0487] Chipset 4390 may be coupled to first bus 4316 via interface 4396. In one embodiment, first bus 4316 may be a Peripheral Component Interconnect (PCI) bus or a bus such as PCI Express or another third generation I / O interconnect bus, although the scope of the present disclosure is not limited in this regard.
[0488] like Figure 43As shown in , various I / O devices 4314 may be coupled to the first bus 4316, along with a bus bridge 4318 that couples the first bus 4316 to a second bus 4320. In one embodiment, one or more additional processors 4315, such as a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (such as, for example, a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array, or any other processor, are coupled to the first bus 4316. In one embodiment, the second bus 4320 may be a low pin count (LPC) bus. In one embodiment, various devices may be coupled to the second bus 4320, including, for example, a keyboard and / or mouse 4322, communication devices 4327, and a storage unit 4328, such as a disk drive or other mass storage device, which may include instructions / code and data 4330. Additionally, an audio I / O 4324 may be coupled to the second bus 4320. Note that other architectures are possible. For example, instead of Figure 43 Instead of a point-to-point architecture, the system can implement a multi-drop bus or other such architecture.
[0489] Now refer to Figure 44 , shown is a block diagram of a second more specific exemplary system 4400 according to an embodiment of the present disclosure. Figure 43 and 44 Similar elements in the same reference numerals are used, and Figure 44 Omitted Figure 43 Some aspects of Figure 44 other aspects.
[0490] Figure 44 The illustrated processors 4370, 4380 may respectively include integrated memory and I / O control logic ("CL") 4372 and 4382. Thus, the CL 4372, 4382 includes an integrated memory controller unit and includes I / O control logic. Figure 44 The illustration shows not only memories 4332, 4334 coupled to CL 4372, 4382, but also I / O devices 4414 coupled to control logic 4372, 4382. Legacy I / O devices 4415 are coupled to chipset 4390.
[0491] Now refer to Figure 45 , shown is a block diagram of a SoC 4500 according to an embodiment of the present disclosure. Figure 41 Similar elements in the FIGURE 1 use similar reference numerals. In addition, the dashed boxes are optional features on more advanced SoCs. Figure 45In the embodiment, interconnect unit(s) 4502 are coupled to: application processor 4510, which includes a set of one or more cores 202A-N and a shared cache unit(s) 4106; system agent unit 4110; bus controller unit(s) 4116; integrated memory controller unit(s) 4114; one or more coprocessors 4520, which may include integrated graphics logic, image processors, audio processors, and video processors; static random access memory (SRAM) unit 4530; direct memory access (DMA) unit 4532; and display unit 4540 for coupling to one or more external displays. In one embodiment, coprocessor(s) 4520 include specialized processors such as, for example, a network or communication processor, a compression engine, a GPGPU, a high-throughput MIC processor, or an embedded processor, among others.
[0492] The various embodiments of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementations. The embodiments of the present disclosure may be implemented as a computer program or program code executed on a programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0493] Program code (such as Figure 43 The code 4330 illustrated in FIG. 4 is applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0494] The program code can be implemented in a high-level process-oriented programming language or an object-oriented programming language to communicate with the processing system. If necessary, the program code can also be implemented in assembly language or machine language. In fact, the mechanism described herein is not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0495] One or more aspects of at least one embodiment may be implemented as representative instructions stored on a machine-readable medium that represent various logic within a processor, which, when read by a machine, causes the machine to fabricate logic for performing the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible, machine-readable medium and supplied to various customers or manufacturing facilities to load into fabrication machines that actually manufacture the logic or processor.
[0496] Such machine-readable storage media may include, but are not limited to, a non-transitory, tangible arrangement of an article of manufacture manufactured or formed by a machine or apparatus, including a storage medium such as a hard disk; any other type of disk, including a floppy disk, an optical disk, a compact disk read only memory (CD-ROM), a compact disk rewritable (CD-RW), and a magneto-optical disk; a semiconductor device, such as a read-only memory (ROM), a random access memory (RAM) such as a dynamic random access memory (DRAM) and a static random access memory (SRAM), an erasable programmable read-only memory (EPROM), flash memory, an electrically erasable programmable read-only memory (EEPROM); a phase change memory (PCM); a magnetic or optical card; or any other type of medium suitable for storing electronic instructions.
[0497] Therefore, embodiments of the present disclosure also include non-transitory tangible machine-readable media containing instructions or containing design data, such as hardware description language (HDL), which defines the structures, circuits, devices, processors and / or system features described herein. These embodiments are also referred to as program products.
[0498] Simulation (including binary conversion, code deformation, etc.)
[0499] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may transform (e.g., using static binary transformation, dynamic binary transformation including dynamic compilation), morph, emulate, or otherwise convert an instruction into one or more other instructions to be processed by the core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on-processor, off-processor, or partially on-processor and partially off-processor.
[0500] Figure 46 1 is a block diagram illustrating a method for converting binary instructions in a source instruction set into binary instructions in a target instruction set using a software instruction converter according to an embodiment of the present disclosure. In the illustrated embodiment, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 46It is shown that an x86 compiler 4604 can be used to compile a program in a high-level language 4602 to generate x86 binary code 4606 that can be natively executed by a processor 4616 having at least one x86 instruction set core. A processor 4616 having at least one x86 instruction set core represents any processor that performs substantially the same functionality as an Intel processor having at least one x86 instruction set core by compatibly executing or otherwise performing: 1) an essential portion of the instruction set of the Intel x86 instruction set core, or 2) an object code version of an application or other software that is targeted to run on an Intel processor having at least one x86 instruction set core to achieve substantially the same results as an Intel processor having at least one x86 instruction set core. The x86 compiler 4604 represents a compiler that can be used to generate x86 binary code 4606 (e.g., object code) that can be executed on a processor 4616 having at least one x86 instruction set core with or without additional linking processing. Similarly, Figure 46 An alternative instruction set compiler 4608 is shown as being used to compile a program in a high-level language 4602 to generate alternative instruction set binary code 4610 that can be natively executed by a processor 4614 that does not have at least one x86 instruction set core (e.g., a processor having a core that executes the MIPS instruction set of MIPS Technologies, Inc. of Sunnyvale, California, and / or the ARM instruction set of ARM Holdings, Inc. of Sunnyvale, California). An instruction converter 4612 is used to convert x86 binary code 4606 into code that can be natively executed by a processor 4614 that does not have an x86 instruction set core. This converted code is unlikely to be identical to the alternative instruction set binary code 4610 because an instruction converter capable of doing so would be difficult to manufacture; however, the converted code will perform general operations and be composed of instructions from the alternative instruction set. Thus, instruction converter 4612 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have an x86 instruction set processor or core to execute x86 binary code 4606, through emulation, simulation, or any other process.
Claims
1. A processor for data flow graph processing, comprising: multiple processing elements; an interconnection network between the plurality of processing elements, the interconnection network being configured to receive an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the interconnection network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are configured to perform operations when an incoming operand set arrives at the plurality of processing elements; as well as a streamer element for prefetching the incoming operand sets from two or more levels of a memory system, Wherein each of the two or more levels can be added a path implemented as a discrete wiring or extension of a network packet format, and wherein the path enables bypassing a level proximate to an originating request.
2. The processor of claim 1, wherein: The streamer element is used to prefetch based on a programmable memory access pattern.
3. The processor of claim 2, wherein: The streamer element includes a plurality of tracking registers for populating in advance of demand streams.
4. The processor of claim 3, wherein: The plurality of tracking registers includes an x-dimensional register for prioritizing fetching in a first dimension of a multi-dimensional streaming fetch mode.
5. The processor of claim 4, wherein: The plurality of tracking registers includes a y-dimension register for prior popping in a second dimension of the multi-dimensional streaming pop mode.
6. A processor for data flow graph processing, comprising: multiple processing elements; an interconnect network between the plurality of processing elements, the interconnect network configured to receive input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the interconnect network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are configured to perform operations upon incoming operand sets arriving at the plurality of processing elements, wherein the incoming operand sets are prefetched from two or more levels of a memory system; as well as a storage management unit for storing output data from the operation in a memory for processing, Wherein each of the two or more levels can be added a path implemented as a discrete wiring or extension of a network packet format, and wherein the path enables bypassing a level proximate to an originating request.
7. The processor of claim 6, wherein: The memory management unit comprises an address register for storing an address of a cache line in the address register, wherein the cache line is for storing a plurality of data values, at least two of the data values being from two different processing elements.
8. The processor of claim 7, wherein: The storage management unit is configured to track storage of the plurality of data values in the cache line.
9. The processor of claim 8, wherein: The memory management unit also includes a plurality of mask bits for performing a masked write in response to determining that less than a complete cache line is to be stored in the memory.
10. A processor for data flow graph processing, comprising: multiple processing elements; an interconnect network between the plurality of processing elements, the interconnect network configured to receive an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is configured to be overlaid onto the interconnect network and the plurality of processing elements, and each node is represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are configured to perform a first operation upon arrival of an incoming operand set at the plurality of processing elements, wherein the incoming operand set is prefetched from two or more levels of a memory system; and a microcontroller configured to perform a second operation, wherein the second operation is an atomic operation, Wherein each of the two or more levels can be added a path implemented as a discrete wiring or extension of a network packet format, and wherein the path enables bypassing a level proximate to an originating request.
11. A method for data flow graph processing, comprising: receiving an input of a data flow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of a processor and an interconnection network between the plurality of processing elements of the processor, with each node being represented as a dataflow operator in the plurality of processing elements; prefetching, by a streamer element, incoming operand sets from two or more levels of a memory system; as well as executing operations of the dataflow graph using the interconnect network and the plurality of processing elements when the incoming operand sets arrive at the plurality of processing elements, Wherein each of the two or more levels can be added a path implemented as a discrete wiring or extension of a network packet format, and wherein the path enables bypassing a level proximate to an originating request.
12. The method of claim 11, wherein: The streamer element is used to prefetch based on a programmable memory access pattern.
13. The method of claim 12, wherein: The streamer element includes a plurality of tracking registers for populating in advance of demand streams.
14. The method of claim 13, wherein: The plurality of tracking registers includes an x-dimensional register for prioritizing fetching in a first dimension of a multi-dimensional streaming fetch mode.
15. The method of claim 14, wherein: The plurality of tracking registers includes a y-dimension register for prior fetching in a second dimension of a multi-dimensional streaming fetch mode.
16. A method for data flow graph processing, comprising: receiving an input of a data flow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of a processor and an interconnection network between the plurality of processing elements of the processor, with each node being represented as a dataflow operator in the plurality of processing elements; executing operations of the dataflow graph using the interconnect network and the plurality of processing elements when incoming operand sets arrive at the plurality of processing elements, wherein the incoming operand sets are prefetched from two or more levels of a memory system; as well as The storage management unit processes the output data from the operation in the memory, Wherein each of the two or more levels can be added a path implemented as a discrete wiring or extension of a network packet format, and wherein the path enables bypassing a level proximate to an originating request.
17. The method of claim 16, wherein: The memory management unit comprises an address register for storing an address of a cache line in the address register, wherein the cache line is for storing a plurality of data values, at least two of the data values being from two different processing elements.
18. The method of claim 17, wherein: Processing by the memory management unit includes tracking storage of the plurality of data values in the cache line.
19. The method of claim 18, comprising: determining that less than a complete cache line is to be stored in the memory; as well as A plurality of mask bits in the memory management unit is used to perform masked writes.
20. A method for data flow graph processing, comprising: receiving an input of a data flow graph comprising a plurality of nodes; overlaying the dataflow graph onto a plurality of processing elements of a processor and an interconnection network between the plurality of processing elements of the processor, with each node being represented as a dataflow operator in the plurality of processing elements; executing a first operation of the dataflow graph using the interconnect network and the plurality of processing elements when an incoming operand set arrives at the plurality of processing elements, wherein the incoming operand set is prefetched from two or more levels of a memory system; as well as performing, by the microcontroller, a second operation, wherein the second operation is an atomic operation, Wherein each of the two or more levels can be added a path implemented as a discrete wiring or extension of a network packet format, and wherein the path enables bypassing a level proximate to an originating request.
Citation Information
Patent Citations
Method and system for encoding instructions for a VLIW that reduces instruction memory requirements
US20030023830A1
Apparatus and method for improving data prefetching efficiency using history based prefetching
US20120166733A1