Electronic device and implementation method thereof
By introducing variable-length pipelines into the processor, the inefficiency problem caused by fixed pipeline length in light loading bucket-type multi-threaded processors is solved, and more efficient instruction execution and thread scheduling is achieved.
Patent Information
- Application Number
- CN202510129984.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-20
- Filing Date
- 2021-10-20
- Publication Date
- 2025-05-30
AI Technical Summary
Lightly loaded bucket multithreaded processors have inefficiency problems when executing instructions, especially at fixed pipeline lengths, resulting in waste of clock cycles.
A variable-length pipeline is introduced, which allows instructions to complete early and reschedule threads by tracking the thread identifier and completion time, solving the problem of register write back conflict.
Through variable length pipelines, the processor efficiency is improved under light loading conditions, reduce clock cycle waste, and enable faster thread rescheduling.
Smart Images

Figure CN120066583A_ABST
Abstract
Description
[0001] Relevant information of divisional application
[0002] This application is a divisional application, and its parent application is a patent application for an invention named "Electronic Device and Method Implemented Thereby" with an application date of October 20, 2021, an application number of 202111222968.5. Technical Field
[0003] Generally speaking, the present application relates to a small chip system. Specifically, the present application relates to variable pipeline lengths in barrel multithreaded processors. Background Art
[0004] A small chip is an emerging technology for integrating various processing functions. Generally speaking, a small chip system is composed of discrete modules (each called a "small chip"), which are integrated on an interposer, and in many instances are interconnected through one or more established networks as needed to provide the desired functions to the system. The interposer and the included small chips can be packaged together to facilitate interconnection with other components of a larger system. Each small chip may include one or more individual integrated circuits (ICs) or "chips", which may be combined with discrete circuit components and commonly coupled to a corresponding substrate to facilitate attachment to the interposer. Most or all of the small chips in the system will be individually configured for communication through one or more established networks.
[0005] The configuration of small chips as individual modules of a system is different from that of this system implemented on a single chip, which contains different device blocks (such as intellectual property (IP) blocks) on one substrate (such as a single die), such as a system on a chip (SoC), or multiple discrete packaged devices integrated on a printed circuit board (PCB). Generally speaking, small chips provide better performance (such as lower power consumption, reduced latency, etc.) than discrete packaged devices, and small chips provide greater production efficiency than single-die chips. These production efficiencies may include higher yield rates or reduced development costs and time.
[0006] A small chip system may include, for example, one or more application (or processor) small chips and one or more support small chips. Here, the distinction between the application small chip and the support small chip is only a reference to possible design scenarios for the small chip system. Thus, for example, a synthetic vision small chip system may include (by way of example only) an application small chip for generating a synthetic vision output, and support small chips such as a memory controller small chip, a sensor interface small chip, or a communication small chip. In a typical use case, a synthetic vision designer may design the application small chip and obtain the support small chips from other parties. Thus, by avoiding the functions included in designing and producing the support small chips, design expenditures (e.g., in terms of time or complexity) are reduced. Small chips also support tight integration of IP blocks that may otherwise be difficult, such as IP blocks fabricated using different processing technologies or using different feature sizes (or utilizing different contact technologies or pitches). Thus, multiple ICs or IC assemblies having different physical, electrical, or communication characteristics can be assembled in a modular manner to provide an assembly that implements the desired functions. The small chip system can also facilitate adaptation to the requirements of different larger systems into which the small chip system will be incorporated. In an example, an IC or other assembly can be optimized for the power, speed, or heat of a particular function, as may occur with a sensor, and can be more easily integrated with other devices compared to attempting to integrate with other devices on a single die. Additionally, by reducing the overall size of the die, the yield of small chips is often higher than the yield of more complex single-die devices. Summary of the Invention
[0007] In one aspect, the present application provides an apparatus including: pipeline circuitry configured to implement a pipeline; instruction write-back circuitry configured to: determine a completion time of an instruction before insertion into the pipeline; detect a conflict between the instruction and a different instruction based on the completion time, the different instruction being in the pipeline, the conflict being detected when the completion time is equal to a second completion time of the different instruction; and calculate a difference between the completion time and a non-conflicting completion time; and write delay circuitry configured to delay completion of the instruction by the difference.
[0008] In another aspect, the present application provides a method including: determining a completion time of an instruction before insertion into a pipeline of a processor; detecting a conflict between the instruction and a different instruction based on the completion time, the different instruction being in the pipeline, the conflict being detected when the completion time is equal to a second completion time of the different instruction; calculating a difference between the completion time and a non-conflicting completion time; and delaying completion of the instruction by the difference. Brief Description of the Drawings
[0009] The present disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments of the present disclosure. However, the drawings should not be regarded as limiting the present disclosure to a particular embodiment, but are for explanation and understanding only.
[0010] Figure 1A and 1B Illustrate an example of a chiplet system according to an embodiment.
[0011] Figure 2 Illustrate components of an example of a memory controller chiplet according to an embodiment.
[0012] Figure 3 Illustrate components in an example of a programmable atomic unit (PAU) according to an embodiment.
[0013] Figure 4 Illustrate a processing flow of a processor component with variable pipeline length for a barrel multithreaded processor according to an embodiment.
[0014] Figure 5 Is a flowchart of an example of a method with variable pipeline length for a barrel multithreaded processor according to an embodiment.
[0015] Figure 6 Is a block diagram of an example of a machine in which, or by which, embodiments of the present disclosure can be operated. Detailed Description
[0016] FIG. 1 described below provides an example of a chiplet system and components operating therein. The illustrated chiplet system includes a memory controller. This memory controller includes a programmable atomic unit (PAU) that is used to execute a custom program, a programmable atomic operator, in response to a memory request for a programmable atomic operator. Additional details regarding the Figure 2 and 3 will be described. The processor of the PAU can be barrel multithreaded and pipelined. The barrel multithreaded processor provides several benefits, such as tolerating the latency of external memory requests while executing many simultaneous threads in a single core, while maintaining high instruction execution throughput. However, problems can occur with a lightly loaded barrel multithreaded processor. Here, lightly loaded means executing few (e.g., one or two) of the available (e.g., eight) hardware threads. In such cases, many clock cycles may be wasted on idle hardware threads.
[0017] Typical barrel multithreaded processors have a fixed pipeline length based on the number of hardware threads. Generally, at each clock, the register file switches to a new thread for a given pipeline stage. This means that the minimum time between thread instruction executions (e.g., between adjacent instructions in a thread) is equal to the number of supported hardware threads. Thus, the latency before a thread can execute the next instruction is fixed and independent of the instruction operation. This approach is used in many barrel multithreaded processors to keep the processor core control hardware simple. However, as noted above, this is inefficient for lightly loaded processors.
[0018] To address the problem of lightly loaded barrel multithreaded processors, a variable length pipeline can be used in the processor. Here, the instructions of a thread can complete early, perform a write-back to the register as soon as possible, and without waiting for the full processor pipeline to complete. This is achieved by tracking the thread identifier (ID) of the thread while the instruction is in the pipeline and reinserting the thread ID into the thread ready-to-run queue when the register write-back is complete. Once the thread ID returns to the thread ready-to-run queue, the thread can be rescheduled, potentially achieving a faster rescheduling of the thread compared to what is possible in traditional barrel multithreaded processor implementations.
[0019] A problem that may arise with a variable pipeline length is a register write-back conflict between threads within a single clock cycle. In traditional barrel multithreaded implementations, this conflict does not occur because the write-back cycle is at the same stage of the pipeline across threads. Since threads start the pipeline at different clock cycles, only one thread attempts to perform a write-back in any given cycle. However, in the case of a variable length pipeline, a conflict can occur.
[0020] To address the register write-back conflict, the instruction register write-back is checked to determine if a conflict exists. If no conflict exists, then once the instruction is able to write-back to the register, the instruction can write-back to the register. Thus, for a move instruction completed at the second stage of the pipeline, when the second stage is complete, the instruction can write-back to the register at the clock cycle. However, if a conflict exists, then the instruction is tracked with a delay (e.g., in clock cycles). Thus, a move instruction may conflict with an earlier instruction with a longer completion time (e.g., later). Here, the move instruction can be given a delay, such as a single clock cycle. When the move instruction is complete, the delay is queried and the move instruction is prevented from performing the write-back. Here, the move instruction remains in the pipeline until the delay expires and the write-back can be completed. At this point, the thread can be rescheduled.
[0021] By providing a variable-length pipeline, a significant drawback of barrel multithreaded processors is addressed. Although barrel multithreaded processors are typically used when many threads are expected to be executed, many programs start by executing a single thread and later switch to using more threads. Thus, for many programs, when the pipeline has a fixed length, barrel multithreaded processors can be inefficient until there is sufficient parallelism (e.g., hardware threads). By introducing a variable-length pipeline, this problem can be improved. Additional details and examples are provided below.
[0022] Figure 1A and 1B Illustrates an example of a die system 110 according to an embodiment. Figure 1A Is a representation of a die system 110 mounted on a peripheral board 105, which may be connected to a wider computer system, for example, via a Peripheral Component Interconnect Express (PCIe). The die system 110 includes a package substrate 115, an interposer 120, and four dies, namely an application die 125, a host interface die 135, a memory controller die 140, and a memory device die 150. Other systems may include many additional dies to provide additional functionality, as will be apparent from the following discussion. The package of the die system 110 is illustrated as a capping or cover plate 165, but other packaging techniques and structures for die systems may be used. Figure 1B Is a block diagram marking the components in the die system for clarity.
[0023] The application die 125 is illustrated as including a Network-on-Chip (NOC) 130 to support a die network 155 for die-to-die communication. In an example embodiment, the NOC 130 may be included on the application die 125. In an example, the NOC 130 may be defined in response to the selected supporting dies (e.g., dies 135, 140, and 150), enabling the designer to select an appropriate number of die network connections or switches for the NOC 130. In an example, the NOC 130 may be located on a separate die or even within the interposer 120. In the examples discussed herein, the NOC 130 implements a Chip Protocol Interface (CPI) network.
[0024] CPI is a packet-based network that supports virtual channels to enable flexible and high-speed interaction between dielets. CPI implements the bridging from the on-die network to the dielet network 155. For example, the Advanced eXtensible Interface (AXI) is a widely used specification for designing in-die communication. However, the AXI specification covers a large number of physical design options, such as the number of physical channels, signal timing, power, etc. Within a single die, these options are typically selected to meet design goals, such as power consumption, speed, etc. However, to achieve the flexibility of the dielet system, an adapter such as CPI is used to mediate between various AXI design options that can be implemented in various dielets. By implementing the mapping of physical channels to virtual channels and encapsulating time-based signaling using a packetization protocol, CPI bridges the on-die network across the dielet network 155.
[0025] CPI can use a variety of different physical layers to transmit packets. The physical layer can include simple conductive connections or can include drivers to increase voltage or otherwise facilitate signal transmission over longer distances. An example of such a physical layer can include the Advanced Interface Bus (AIB), which can be implemented in the interposer layer 120 in various instances. AIB uses source-synchronous data transfer with a forward clock to transmit and receive data. Packets are transmitted across the AIB at single data rate (SDR) or double data rate (DDR) relative to the transmitted clock. AIB supports various channel widths. When operating in SDR mode, the AIB channel width is a multiple of 20 bits (20, 40, 60...), and for DDR mode, the AIB channel width is a multiple of 40 bits: (40, 80, 120...). The AIB channel width includes both transmitted and received signals. The channel can be configured to have a symmetric number of transmit (TX) and receive (RX) input / output (I / O), or an asymmetric number of transmitters and receivers (e.g., all transmitters or all receivers). The channel can act as the AIB master or slave depending on which dielet provides the master clock. The AIB I / O unit supports three timing modes: asynchronous (i.e., non-timed), SDR, and DDR. In various instances, the non-timed mode is used for clocks and some control signals. The SDR mode can use dedicated SDR-only I / O units or double-use SDR / DDR I / O units.
[0026] In an example, a CPI packet protocol (e.g., point-to-point or routable) may use symmetric receive and transmit I / O units within an AIB channel. The CPI streaming protocol allows for more flexible use of the AIB I / O units. In an example, an AIB channel in a streaming mode may configure the I / O units to be all TX, all RX, or half TX and half RX. The CPI packet protocol may use the AIB channel in SDR or DDR operation modes. In an example, the AIB channel is configured in increments of 80 I / O units (i.e., 40 TX and 40 RX) for the SDR mode and in increments of 40 I / O units for the DDR mode. The CPI streaming protocol may use the AIB channel in SDR or DDR operation modes. Here, in an example, the AIB channel is in increments of 40 I / O units for both SDR and DDR modes. In an example, a unique interface identifier is assigned to each AIB channel. The identifier is used during CPI reset and initialization to determine paired AIB channels across adjacent dies. In an example, the interface identifier is a 20-bit value that includes a seven-bit die identifier, a seven-bit column identifier, and a six-bit link identifier. The AIB physical layer uses an AIB out-of-band shift register to transmit the interface identifier. Bits 32 to 51 of the shift register are used to convey the 20-bit interface identifier in both directions across the AIB interface.
[0027] The AIB defines a stacked group of AIB channels as an AIB channel column. An AIB channel column has a certain number of AIB channels, plus an auxiliary channel. The auxiliary channel contains signals for AIB initialization. All AIB channels (except the auxiliary channel) within a column have the same configuration (e.g., all TX, all RX, or half TX and half RX, and have the same number of data I / O signals). In an example, the AIB channels are numbered in consecutive increasing order starting with the AIB channel adjacent to the AUX channel. The AIB channel adjacent to the AUX is defined as AIB channel zero.
[0028] Generally, the CPI interface of an individual die may include serialization - deserialization (SERDES) hardware. SERDES interconnections are well-suited for scenarios that require high-speed signaling and low signal counts. However, for multiplexing and demultiplexing, error detection or correction (e.g., using block-level cyclic redundancy check (CRC)), link-level retry, or forward error correction, SERDEs may cause additional power consumption and longer latency. However, when low latency or power consumption is the main concern for ultra-short distance die-to-die interconnections, a parallel interface with a clock rate that allows for data transfer with minimal latency may be utilized. CPI includes components for minimizing both the latency and power consumption of these ultra-short distance die-to-die interconnections.
[0029] For flow control, CPI employs credit-based techniques. For example, the receiving side of application die 125 provides credit representing available buffers to the sending side of, for example, memory controller die 140. In an instance, the CPI receiver includes buffers for each virtual channel for a given transmission time unit. Thus, if the CPI receiver supports five messages and a single virtual channel in time, then the receiver has five buffers arranged in five rows (e.g., one row per unit time). If four virtual channels are supported, then the receiver has twenty buffers arranged in five rows. Each buffer holds the payload of one CPI packet.
[0030] When the sending side transmits to the receiving side, the sending side decrements the available credit based on the transmission. Once all of the receiver's credit is used up, the sending side stops sending packets to the receiver. This ensures that the receiver always has available buffers to store transmissions.
[0031] When the receiver processes the received packets and frees up buffers, the receiver conveys the available buffer space back to the sending side. Then, the sending side can use this credit return to allow the transmission of additional information.
[0032] A die mesh network 160 is also illustrated, which uses direct die-to-die techniques and does not require a NOC 130. The die mesh network 160 can be implemented in CPI or another die-to-die protocol. The die mesh network 160 typically implements a die pipeline, where one die acts as an interface to the pipeline and the other dies in the pipeline interface only interface with themselves.
[0033] Additionally, dedicated device interfaces can be used to interconnect dies, such as one or more industry standard memory interfaces 145 (e.g., synchronous memory interfaces such as DDR5, DDR6). The connection of a die system or individual dies to external devices (e.g., a larger system) can be through the desired interface (e.g., a PCIE interface). In an instance, for example, the external interface can be implemented through a host interface die 135, which provides a PCIE interface external to the die system 110 in the depicted instance. Such interfaces are typically employed when industry conventions or standards have converged on such dedicated interfaces 145. The illustrated instance of a double data rate (DDR) interface 145 connecting the memory controller die 140 to a dynamic random access memory (DRAM) memory device 150 is such an industry convention.
[0034] Among the various possible support die, the memory controller die 140 is likely to be present in the die system 110 because of the almost ubiquitous use of memory for computer processing and the use of advanced technologies for memory devices. Thus, using the memory device die 150 and the memory controller die 140 produced by other technologies enables die system designers to obtain robust products produced by mature manufacturers. Generally, the memory controller die 140 provides a memory device specific interface to read, write, or erase data. Typically, the memory controller die 140 may provide additional features such as error detection, error correction, maintenance operations, or atomic operator execution. For some types of memory, maintenance operations are often specific to the memory device 150, such as garbage collection in NAND flash or storage class memory, temperature adjustment in NAND flash memory (e.g., cross temperature management). In an example, the maintenance operation may include logical to physical (L2P) mapping or management to provide an intermediate level between the physical and logical representations of data. In other types of memory such as DRAM, some memory operations such as refresh may be controlled by the host processor or the memory controller at certain times and by the DRAM memory device or logic associated with one or more DRAM devices at other times, the logic such as an interface chip (in an example, a buffer).
[0035] An atomic operator is a data manipulation that can be performed by the memory controller die 140. In other die systems, atomic operators can be performed by other die. For example, an atomic operator of "increment" can be specified in a command by the application die 125, the command including a memory address and possibly an increment value. After receiving the command, the memory controller die 140 retrieves the number from the specified memory address, increments the number by the amount specified in the command, and stores the result. After successful completion, the memory controller die 140 provides an indication of command success to the application die 125. Atomic operators avoid transmitting data across the die network 160, thereby reducing the latency in executing such commands.
[0036] Atomic operators can be classified as built-in atoms or programmable (e.g., custom) atoms. Built-in atoms are a limited set of operations invariantly implemented in hardware. Programmable atoms are small programs that can be executed on the programmable atom unit (PAU) (e.g., custom atom unit (CAU)) of the memory controller die 140. FIG. 1 illustrates an example of a memory controller die that discusses the PAU.
[0037] The memory device die 150 can be or include any combination of volatile memory devices or non-volatile memory. Examples of volatile memory devices include, but are not limited to, random access memory (RAM), such as DRAM, synchronous DRAM (SDRAM), graphics double data rate type 6 SDRAM (GDDR6 SDRAM), and so on. Examples of non-volatile memory devices include, but are not limited to, NAND-type flash memory, storage class memory (e.g., phase change memory or memristor-based technology), ferroelectric RAM (FeRAM), and so on. The illustrated example includes the memory device 150 as a die, however, the memory device 150 can reside elsewhere, such as in a different package on the peripheral board 105. For many applications, multiple memory device dies can be provided. In an example, these memory device dies can each implement one or more storage technologies. In an example, the memory die can include multiple stacked memory dies of different technologies, such as one or more static random access memory (SRAM) devices stacked with or otherwise communicating with one or more dynamic random access memory (DRAM) devices. The memory controller 140 can also be used to coordinate operations between multiple memory dies in the die system 110; for example, utilizing one or more memory dies in one or more levels of cache storage devices and using one or more additional memory dies as main memory. The die system 110 can also include multiple memory controllers 140, which can be used to provide memory control functionality for separate processors, sensors, networks, and so on. For example, the die architecture of the die system 110 provides the advantage of allowing adaptation to different memory storage technologies and different memory interfaces through updated die configurations without the need to re-design the rest of the system architecture.
[0038] Figure 2Describe the components of an example of the memory controller die 205 according to an embodiment. The memory controller die 205 includes a cache 210, a cache controller 215, an off-die memory controller 220 (e.g., for communicating with off-die memory 275), a network communication interface 225 (e.g., for interfacing with the die network 285 and communicating with other dies), and a set of atomic and coalescing units 250. This set of components may include, for example, a write coalescing unit 255, a memory hazard unit 260, a built-in atomic unit 265, or a PAU 270. The various components are described logically, and they may not necessarily be implemented. For example, the built-in atomic unit 265 may potentially include different devices along the path to the off-die memory. For example, the built-in atomic unit 265 may be in an interface device / buffer on the memory die, as discussed above. In contrast, the programmable atomic unit 270 may be implemented in a separate processor on the memory controller die 205 (but in various instances, may be implemented in other locations, such as on the memory die).
[0039] The off-die memory controller 220 is directly coupled to the off-die memory 275 (e.g., via a bus or other communication connection) to provide write operations and read operations to and from one or more off-die memories such as the off-die memory 275 and the off-die memory 280. In the depicted example, the off-die memory controller 220 is also coupled to the atomic and coalescing units 250 for output and to the cache controller 215 (e.g., a memory-side cache controller) for input.
[0040] In an example configuration, the cache controller 215 is directly coupled to the cache 210 and may be coupled to the network communication interface 225 for input (e.g., incoming read or write requests) and is coupled to the off-die memory controller 220 for output.
[0041] The network communication interface 225 includes a packet decoder 230, a network input queue 235, a packet encoder 240, and a network output queue 245 to support a packet-based die network 285, such as CPI. The die network 285 may provide packet routing between and among processors, memory controllers, hybrid-thread processors, configurable processing circuits, or communication interfaces. In such a packet-based communication system, each packet typically includes destination and source addressing, as well as any data payload or instructions. In an example, depending on the configuration, the die network 285 may be implemented as a set of crossbar switches with a folded Clos configuration, or a mesh network providing additional connectivity.
[0042] In various examples, the dielet network 285 can be part of an asynchronous switching fabric. Here, data packets can be routed along any one of various paths such that any selected data packet can reach the addressed destination at any one of multiple different times, depending on the routing. Additionally, the dielet network 285 can be at least partially implemented as a synchronous communication network, such as a synchronous mesh communication network. Both configurations of the communication network are contemplated for use in examples in accordance with the present disclosure.
[0043] The memory controller dielet 205 can receive a packet having, for example, a source address, a read request, and a physical address. In response, the off-die memory controller 220 or the cache controller 215 will read data from the specified physical address (which can be in the off-die memory 275 or the cache 210) and assemble a response packet with the source address containing the requested data. Similarly, the memory controller dielet 205 can receive a packet having a source address, a write request, and a physical address. In response, the memory controller dielet 205 will write the data to the specified physical address (which can be in the cache 210 or the off-die memory 275 or 280) and assemble a response packet with the source address containing an acknowledgement that the data was stored in memory.
[0044] Thus, where possible, the memory controller dielet 205 can receive read and write requests via the dielet network 285 and use the cache controller 215 interfaced with the cache 210 to process the requests. If the cache controller 215 is unable to handle the requests, then the off-die memory controller 220 handles the requests by communicating with the off-die memory 275 or 280, the atomic and merge unit 250, or both. As described above, one or more levels of cache can also be implemented in the off-die memory 275 or 280; and in some such examples can be directly accessed by the cache controller 215. Data read by the off-die memory controller 220 can be cached by the cache controller 215 in the cache 210 for subsequent use.
[0045] The atomic and merge unit 250 is coupled to receive (as input) the output of the off-die memory controller 220 and provide the output to the cache 210, the network communication interface 225, or directly to the dielet network 285. The memory hazard unit 260, the write merge unit 255, and the built-in (e.g., predefined) atomic unit 265 may each be implemented as a state machine with other combinational logic circuitry (such as adders, shifters, comparators, AND gates, OR gates, XOR gates, or any suitable combination thereof) or other logic circuitry. These components may also include one or more registers or buffers to store operands or other data. The PAU 270 may be implemented as one or more processor cores or control circuitry, and various state machines with other combinational logic circuitry or other logic circuitry, and may also include one or more registers, buffers, or memories to store addresses, executable instructions, operands, and other data, or may be implemented as a processor.
[0046] The write merge unit 255 receives read data and requested data and merges the requested data and the read data to create a single unit with the read data and the source address to be used in the response or return packet. The write merge unit 255 provides the merged data to the write port of the cache 210 (or equivalently, to the cache controller 215 for writing to the cache 210). Optionally, the write merge unit 255 provides the merged data to the network communication interface 225 to encode and prepare the response or return packet for transmission on the dielet network 285.
[0047] When the requested data is for a built-in atomic operator, the built-in atomic unit 265 receives the request and the read data from the write merge unit 255 or directly from the off-die memory controller 220. The atomic operator is executed, and the resulting data is written to the cache 210 using the write merge unit 255, or provided to the network communication interface 225 to encode and prepare the response or return packet for transmission on the dielet network 285.
[0048] The built-in atomic unit 265 disposes of predefined atomic operators, such as fetch-and-increment or compare-and-swap. In an example, these operations perform simple read-modify-write operations on a single memory location that is 32 bytes or smaller in size. The atomic memory operation is initiated from a request packet transmitted via the die network 285. The request packet has a physical address, an atomic operator type, an operand size, and optionally up to 32 bytes of data. The atomic operator performs a read-modify-write on a cache memory line of the cache 210, thus filling the cache memory when necessary. The atomic operator response can be a simple completion response or a response with up to 32 bytes of data. Example atomic memory operators include fetch-and-and, fetch-and-or, fetch-and-xor, fetch-and-add, fetch-and-subtract, fetch-and-increment, fetch-and-decrement, fetch-and-min, fetch-and-max, fetch-and-swap, and compare-and-swap. In various example embodiments, 32-bit and 64-bit operations are supported as well as operations on 16 or 32 bytes of data. The methods disclosed herein are also compatible with hardware that supports larger or smaller operations and more or less data.
[0049] The built-in atomic operator may also involve a request for a "standard" atomic operator on the requested data, such as relatively simple single-cycle integer atoms, such as fetch-and-increment or compare-and-swap, whose throughput will be the same as that of a conventional memory read or write operation that does not involve an atomic operator. For these operations, the cache controller 215 can typically reserve the cache line in the cache 210 by setting a hazard bit (in hardware) such that the cache line cannot be read by another process during a transition. The data is obtained from the off-die memory 275 or the cache 210 and provided to the built-in atomic unit 265 to perform the requested atomic operator. After the atomic operator, in addition to providing the resulting data to the packet encoder 240 to encode the output packet for transmission on the die network 285, the built-in atomic unit 265 provides the resulting data to the write-combining unit 255, which also writes the resulting data to the cache 210. After writing the resulting data to the cache 210, the memory hazard unit 260 clears any corresponding hazard bits that were set.
[0050] The PAU 270 implements high performance (high throughput and low latency) of programmable atomic operators (also known as "custom atomic transactions" or "custom atomic operators"), which is comparable to the performance of built-in atomic operators. Instead of performing multiple memory accesses, in response to an atomic operator request specifying a programmable atomic operator and a memory address, the circuitry in the memory controller die 205 transmits the atomic operator request to the PAU 270 and sets the hazard bit corresponding to the memory address of the memory row used in the atomic operator stored in the memory hazard register to ensure that no other operations (reads, writes, or atoms) are performed on the memory row, and then clears the hazard bit after the atomic operator is completed. The additional, direct data path provided to the PAU 270 for executing programmable atomic operators allows additional write operations without being subject to any limitations imposed by the bandwidth of the communication network and without adding any congestion to the communication network.
[0051] The PAU 270 includes a multi-threaded processor having one or more processor cores, such as a multi-threaded processor based on the RISC-V ISA, and further has an extended instruction set for executing programmable atomic operators. When having an extended instruction set for executing programmable atomic operators, the PAU 270 can be embodied as one or more hybrid-thread processors. In some example embodiments, the PAU 270 provides barrel-round-robin instantaneous thread switching to maintain a high instructions-per-clock rate.
[0052] The programmable atomic operator can be executed by the PAU 270, and the programmable atomic operator involves a request for a programmable atomic operator regarding the requested data. The user can prepare programming code to provide such programmable atomic operators. For example, the programmable atomic operator can be a relatively simple multi-cycle operation, such as floating-point addition, or can be a relatively complex multi-instruction operation, such as Bloom filter insert. The programmable atomic operator can be the same as or different from the predetermined atomic operator, as long as they are defined by the user rather than the system vendor. For these operations, the cache controller 215 can reserve the cache line in the cache 210 by setting a hazard bit (in hardware) so that the cache line cannot be read by another process during the transition. Data is obtained from the cache 210 or the off-die memory 275 or 280, and the data is provided to the PAU 270 to execute the requested programmable atomic operator. After the atomic operator, the PAU 270 provides the resulting data to the network communication interface 225 to directly encode the output packet with the resulting data for transmission on the dielet network 285. Additionally, the PAU 270 provides the resulting data to the cache controller 215, and the cache controller 215 also writes the resulting data to the cache 210. After writing the resulting data to the cache 210, the cache control circuit 215 clears any corresponding hazard bits that were set.
[0053] In the selected example, the approach taken for the programmable atomic operator is to provide multiple general-purpose custom atomic request types, which can be sent from a starting source such as a processor or other system component to the memory controller dielet 205 via the dielet network 285. The cache controller 215 or the off-die memory controller 220 identifies the request as a custom atom and forwards the request to the PAU 270. In a representative embodiment, the PAU 270: (1) is a programmable processing element capable of effectively executing user-defined atomic operators, (2) can perform load and store on memory, arithmetic and logical operations, and control flow decisions; and (3) utilizes the RISC-V ISA with a new set of dedicated instructions to facilitate interaction with such controllers 215, 220, thereby performing user-defined operations atomically. In a desirable example, the RISC-V ISA contains a complete instruction set that supports high-level language operators and data types. The PAU 270 can utilize the RISC-V ISA, but typically supports a more limited instruction set and a limited register file size to reduce the die size of the unit when included within the memory controller dielet 205.
[0054] As mentioned above, before writing the read data into the cache 210, the memory hazard clearing unit 260 will clear the set hazard bits of the reserved cache lines. Therefore, when the write combining unit 255 receives a request and the read data, the memory hazard clearing unit 260 can transmit a reset or clear signal to the cache 210 to reset the set memory hazard bits of the reserved cache lines. Also, resetting this hazard bit will also release the pending read or write requests involving the specified (or reserved) cache line, thereby providing the pending read or write requests to the inbound request multiplexer for selection and processing.
[0055] Figure 3 Describe the components in an example of the programmable atomic unit 300 (PAU) according to an embodiment, such as the components mentioned above with respect to FIG. 1 (e.g., in the memory controller 140) and Figure 2 (e.g., PAU 270). As illustrated, the PAU 300 includes a processor 305, a local memory 310 (e.g., SRAM), and a controller 315 for the local memory 310.
[0056] In the example, the processor 305 is pipelined such that multiple stages of different instructions are executed together in each clock cycle. The processor 305 is also a barrel multi-threaded processor with circuitry for switching between different register files (e.g., a register bank containing the current processing state) after each clock cycle of the processor 305. This enables efficient context switching between the currently executing threads. In the example, the processor 305 supports eight threads, resulting in eight register files. In the example, some or all of the register files are not integrated into the processor 305 but reside in the local memory 310 (registers 320). This reduces the circuit complexity in the processor 305 by eliminating the conventional flip-flops for these registers 320.
[0057] The local memory 310 can also accommodate a cache 330 and instructions 325 for atomic operators. The atomic instructions 325 include an instruction set that supports atomic operators for various application loads. When an atomic operator is requested, for example, by the application die 125, the instruction set corresponding to the atomic operator is executed by the processor 305. In the example, the atomic instructions 325 are partitioned to form an instruction set. In this example, a particular programmable atomic operator requested by a requesting process can be identified by a partition number for the programmable atomic operator. The partition number can be established when a programmable atomic operator is hosted (e.g., loaded onto) the PAU 300. Additional metadata for the programmable atomic instructions 325 can also be stored in the local memory 310, such as a partition table.
[0058] The atomic operator manipulates cache 330, which is typically synchronized (e.g., flushed) when the thread of the atomic operator completes. Thus, latency is reduced for most memory operations during the execution of the programmable atomic operator thread, in addition to the initial load from an external memory such as off-die memory 275 or 280.
[0059] Processor 305 is configured to have a variable length pipeline to address inefficiencies caused by situations when the processor 305 is lightly loaded (e.g., when using fewer than all available threads). To achieve this, processor 305 is configured to enable instruction completion at one stage in the pipeline before the end of the pipeline. Thus, assuming a pipeline length of eight stages, but an instruction can complete in two stages (e.g., a register move instruction) or in six stages (e.g., a double-precision multiply), then the instruction can complete at the second stage and the sixth stage, respectively. Completion in this context indicates that the instruction is ready to be written back to the register. As described above, and in the discussion below regarding Figure 4 When multiple instructions attempt to complete at the same time (e.g., within the same clock cycle), a variable length pipeline can cause complexity. To address this complexity, processor 305 is configured to determine the completion time of an instruction before it is inserted into the pipeline. The determination is a lookup for a particular instruction or a lookup for the type of instruction. When the instruction will complete, the lookup thus provides a number of clock cycles or one processor stage. This is the completion time relative to the current clock cycle of processor 305. Thus, if an instruction takes five clock cycles to complete, the completion time is five clock cycles from the current clock cycle of the processor. However, other timing elements other than clock cycles can be used, and as described below, the point of the completion time will determine register write-back conflicts with other instructions. Thus, any metric indicating these measures can be used.
[0060] In an example, to determine the completion time of an instruction, processor 305 is configured to determine that the completion time is less than all possible pipeline stages in the pipeline. This condition is consistent with the fixed completion time instructions discussed below regarding Figure 4 Generally, conflict avoidance latency techniques apply to fixed completion time instructions and not to variable completion time instructions.
[0061] Processor 305 is configured to detect a conflict between the instruction and a different instruction based on the completion time. Here, the different instruction is in the pipeline, and a conflict occurs when the completion time is equal to the second completion time of the different instruction. Thus, the detection confirms that at the bottleneck of completing the instruction, e.g., a single cycle of register write-back, both the current instruction and the different instruction will reach the bottleneck at the same moment.
[0062] To track conflicts, processor 305 may include a scoreboard (e.g., a register write-back scoreboard). A scoreboard is a structure in which instructions can be marked as completed. Thus, for example, when a clock cycle is used to determine the completion time, the scoreboard has a field for a given future clock cycle that is marked (e.g., in the case of a logic one) when an instruction (e.g., a different instruction) in the pipeline is expected to complete after that cycle. Other structures can be used to represent the time that will be used, such as a shift structure (e.g., a register) encoded such that each bit represents the number of future cycles, and the structure is shifted down by one at each clock cycle.
[0063] Processor 305 is configured to detect a conflict by looking in the scoreboard for the clock cycle in which the completion time falls to obtain a result. This result is then examined by processor 305 to detect that different instructions will perform a register write-back at that clock cycle. Thus, there is a conflict between the instruction and the different instruction.
[0064] Processor 305 is configured to calculate the difference between the completion time and the conflict-free completion time. In an example, when there is no conflict (e.g., the conflict-free completion time), processor 305 uses the scoreboard to look ahead to future cycles. This difference is the completion latency, which if enforced will ensure that the instruction can be written back to the register. If there is no conflict, then the difference will be zero. In an example, processor 305 is configured to write an indication of the instruction that will complete at the second clock cycle in the scoreboard as part of inserting the instruction into the pipeline. Here, the second clock cycle corresponds to the conflict-free completion time. Thus, processor 305 updates the scoreboard with the completion time plus the difference to enable conflict detection for future instructions.
[0065] Processor 305 is configured to make the completion latency of the instruction the difference. Thus, if an instruction can complete at the second stage of the pipeline but conflicts with a different instruction and a difference of one is given to resolve the conflict, then processor 305 will perform the register write-back of the instruction in the clock cycle after the second stage in the pipeline completes.
[0066] In an example, to support latency, processor 305 may include a write latency structure (e.g., Figure 4write latency circuitry or write-back latency component 445). Here, the processor 305 is configured to detect that an instruction is ready to perform a register write-back at a stage in the pipeline. The processor 305 is then configured to perform a read in the write latency data structure corresponding to the stage in the pipeline. If it is determined that the recorded value is not zero, then the instruction is prevented from performing the register write-back. Here, when the record is implemented as a shift register, the value in the register can be decremented every clock cycle. Thus, if a second-stage instruction is ready to write back with a one-cycle delay, the value in the record is one rather than zero, and thus the write-back of the instruction will be blocked and another instruction will perform the write-back. However, after the next clock cycle, the value in the second-stage record will decrement to zero, and the instruction will be able to perform the write-back. In an example, the processor 305 is configured to write the difference as a value into the record as part of inserting the instruction into the pipeline.
[0067] Since thread instructions can complete at different times (e.g., clock cycles, pipeline stages, etc.), flexible techniques for rescheduling them can include a buffer of the thread IDs currently in the pipeline (e.g., a pipelined thread ID buffer). The processor 305 can be configured to place the thread ID of an instruction into the pipelined thread ID buffer as part of inserting the instruction into the pipeline. Then, when the instruction completes, the processor 305 is configured to move the thread ID from the pipelined thread ID buffer to the thread ready-to-run queue. Here, each thread ID can be retrieved in order to schedule the next instruction.
[0068] Figure 4 Describe the processing flow through processor components with variable pipeline lengths in a bucket multithreaded processor according to an embodiment. As illustrated, a thread ID is obtained from the ready-to-run queue. When processing the next instruction, a thread ID is obtained from the thread ready-to-run queue 405, and the state of the thread (thread state component) 410 is obtained. The next instruction of the thread is cached (instruction cache component 415). Next, the instruction is decoded (e.g., cracked) into components such as operands, instruction type, etc. at the instruction cracking component 420.
[0069] The instruction write-back slot component 425 (e.g., the instruction write-back circuitry) examines the instruction and determines at which stage in the pipeline 150 the instruction will complete. In this context, instruction completion means the point at which the instruction will attempt to write back to the register file 435. For example, a move instruction (e.g., moving data from one register to another register) may be extremely fast and complete at the second stage of the pipeline 450, or complete within two clock cycles of entering the pipeline. In contrast, a double-precision addition instruction may complete at the fourth stage of the pipeline 450 or take four clock cycles. These types of instructions have a known or fixed completion within one stage of the pipeline 450 and are thus referred to herein as fixed-completion-time instructions. Other instructions (e.g., requests to off-chip memory or to components external to the processor (e.g., a coprocessor, an encryption component, or other IP blocks)) may have a completion time that extends beyond the pipeline stage, which may also be variable, and are thus referred to as variable-completion-time instructions.
[0070] If the instruction write-back slot 425 determines that the instruction is a variable-completion-time instruction, then a form of thread self-scheduling may be used. Here, the processor hands over the timing of rescheduling the thread to an external or long-running component to some extent. To achieve this, the thread ID travels with the instruction. Thus, for example, when a memory request instruction enters the pipeline 450 to process the memory request, the thread ID of the instruction accompanies the memory request. At this time, the thread ID is no longer in the thread ready-to-run queue 405 or in any of the other illustrated structures. Thus, the thread will not be rescheduled. However, when the memory request completes, the response from the memory is received by the processor and stored in the buffer 465. A separate buffer 470 may be used for responses from another IP block (e.g., a coprocessor), and buffer 465 may be used for all variable-completion-time instruction responses. The multiplexer 460 determines when the thread corresponding to the response in buffer 465 or 470 will be put back into the thread ready-to-run queue for rescheduling.
[0071] If the instruction is a fixed-completion-time instruction, then the instruction write-back slot component 425 examines the register write-back scoreboard 430 to determine whether any other register write-back can be attempted at the same clock cycle as the current instruction. For example, if two clock cycles ago, an instruction could write back in four clock cycles and the current instruction can write back in two clock cycles, then these two instructions will be able to write back in the same clock cycle. This potential conflict is tracked in the register write-back scoreboard 430. The register write-back scoreboard 430 tracks which clock cycles do not have potential write-backs. Thus, when the 4-cycle instruction is processed by the instruction write-back slot component 425, the fourth clock cycle is marked to indicate the write-back of the 4-cycle component. Then, two cycles later, when the 2-cycle instruction is processed, the clock cycle at which the 2-cycle instruction may complete is marked by the 4-cycle instruction, indicating a register write-back conflict.
[0072] When the instruction write-back slot component 425 detects a register write-back conflict, the instruction write-back slot component 425 resolves the conflict by calculating the latency of the conflicting instruction based on the non-conflicting cycles from the register write-back scoreboard 430. Thus, if the cycle after the conflicting cycle is idle, a 1-cycle write-back latency can be provided for a 2-cycle instruction. If there is no conflict, a 0-cycle write-back latency is provided for the instruction.
[0073] Once the instruction write-back slot component 425 determines the register write-back latency, the instruction write-back slot component 425 writes the register write-back latency to the write-back latency component 445. An example implementation of the write-back latency component 445 can be a set of shift registers, one register for each stage in the pipeline 450, or a subset of the stages in the pipeline 450 that can trigger a register write-back. As illustrated, the arrows leading from the pipeline 450 and the write-back latency component correspond such that the topmost arrow of each arrow corresponds to the write-back of the second stage of the pipeline 450, the second topmost arrow corresponds to the fourth stage, and the last arrow corresponds to the sixth stage.
[0074] After each clock cycle, the shift register in the write-back latency component 445 decrements the value stored therein. Thus, if a 2-cycle instruction has a three-cycle latency, the numerical value three is stored in the shift register corresponding to the second stage of the pipeline 450. After the first cycle, the value in the register will be two, then one after another cycle, and finally zero after the third cycle.
[0075] When an instruction is ready to write back from the pipeline 450, the stage in the write-back latency component 445 where the register is located is queried to determine if the write-back will occur. For example, when the fourth stage of the pipeline 450 is complete, a 4-cycle instruction has zero in the write-back latency component 445 and can thus be immediately written back to the register file 435. However, the register in the write-back latency component 445 corresponding to the 2-cycle instruction does not have zero, and thus no write-back will occur for this instruction. Rather, the thread of the 2-cycle instruction will remain at the second stage of the pipeline 450. In the next cycle, the register in the write-back latency component 445 for the second stage will decrement to zero, and the write-back to the register file 435 for the 2-cycle instruction can proceed.
[0076] Instructions with a later completion time (e.g., a later pipeline stage) have priority. Thus, within a given clock cycle, if instructions at the sixth stage and the fourth stage are ready to be written back to the register file 435, then the sixth-stage instruction will perform the write-back. This ensures that the maximum write-back latency does not cause the write-back to exceed the pipeline length, which is a typical latency imposed on other barrel multithreaded processors. When the latency of an instruction is shorter than the pipeline length, then fewer clock cycles are wasted in a lightly loaded processor because the shorter-cycle instruction can complete in fewer cycles compared to the cycles taken for each stage of the pipeline 450 to complete.
[0077] Once the write-back to the register file 435 is complete, the threads can be rescheduled. This is accomplished by the pipelined thread ID component 440. The pipelined thread ID component 440 contains the thread IDs of the instructions moving through the pipeline 450. When an instruction is able to perform a write-back, the thread ID of the instruction is removed from the pipelined thread ID component 440 and the thread ID of the instruction is delivered to the multiplexer 460 for reinsertion into the thread ready-to-run queue 405. In an example, the multiplexer selects an available thread ID from the pipelined thread ID component 440 before any other source (e.g., a memory response). Thus, the buffers 465 or 470 maintain the thread IDs from the variable completion time instructions until there is no input from the pipelined thread ID component 440.
[0078] Figure 5 A flowchart of an example of a method 500 for variable pipeline length in a barrel multithreaded processor according to an embodiment. The operations of the method 500 are performed by computer hardware (e.g., processing circuitry).
[0079] In an example, the method is performed by a processor (e.g., processor 305) in a PAU (e.g., PAU 300 or PAU 270) in a memory controller (e.g., memory controller 140 or memory controller 205). In an example, the memory controller is a die (e.g., memory controller 140). In an example, the memory controller die is integrated into a die system (e.g., die system 110).
[0080] At operation 505, the completion time of an instruction is determined before it is inserted into the pipeline of the barrel multithreaded processor. In an example, the completion time is the number of clock cycles. In an example, determining the completion time of an instruction includes determining that the completion time is less than all possible pipeline stages in the pipeline.
[0081] At operation 510, a conflict between an instruction and a different instruction is detected based on the completion time. Here, the different instruction is in the pipeline, and a conflict occurs when the completion time is equal to the second completion time of the different instruction.
[0082] In an example, detecting a conflict includes looking up the clock cycle in which the completion time falls in a register write-back scoreboard to obtain a result. Next, the result is examined to detect that the result indicates that different instructions will perform register write-back at the clock cycle. In an example, the operation of method 500 is extended to include writing an indication that the instruction will complete at a second clock cycle in the register write-back scoreboard as part of inserting the instruction into the pipeline. Here, the second clock cycle corresponds to the conflict-free completion time.
[0083] At operation 515, the difference between the completion time and the conflict-free completion time is calculated. In an example, when the completion time is measured within a clock cycle, the difference is a second number of clock cycles. In an example, the maximum latency (e.g., the completion time plus the difference) is eight clock cycles.
[0084] At operation 520, the completion of the instruction is delayed by the difference. In an example, delaying the completion of the instruction by the difference includes detecting that the instruction is ready to perform register write-back at a stage in the pipeline, and performing a read on a record corresponding to the stage in the pipeline in a write-delay data structure. If it is determined that the value of the record is not zero, then the instruction is prevented from performing register write-back. In an example, the operation of method 500 can be extended to include writing the difference as a value into the record as part of inserting the instruction into the pipeline. In an example, the operation of method 500 can be extended to include decrementing the value with each clock cycle.
[0085] In an example, the operation of method 500 can be extended to include moving the thread ID of the instruction from a pipelined thread ID buffer to a thread ready-to-run queue in response to the completion of the instruction. In an example, the thread ID is inserted into the pipelined thread ID buffer as part of inserting the instruction into the pipeline.
[0086] Figure 6A block diagram of an example machine 600 is shown, which can be used to implement, within the machine, or by the machine, any one or more of the techniques (e.g., methods) discussed herein. As described herein, an example can include logic or multiple components or mechanisms within machine 600, or can be operated by it. A circuit system (e.g., a processing circuit system) is a collection of circuits implemented in a tangible entity of machine 600, the entity including hardware (e.g., simple circuits, gates, logic, etc.). Circuit system membership can become flexible over time. A circuit system includes components that can perform specified operations individually or in combination when operating. In an example, the hardware of a circuit system can be invariantly designed to perform a particular operation (e.g., hardwired). In an example, the hardware of a circuit system can include physically reconfigurable components (e.g., execution units, transistors, simple circuits, etc.), including a machine-readable medium that is physically modified (e.g., magnetically, electrically, movable placement of immutable mass particles, etc.) to encode instructions for a particular operation. When connecting physical components, the underlying electrical properties of the hardware configuration change, for example, from an insulator to a conductor, or vice versa. The instructions enable the embedded hardware (e.g., an execution unit or a loading mechanism) to create components of the circuit system in the hardware through physical reconfiguration to perform parts of a particular operation when operating. Thus, in an example, a machine-readable medium element is part of a circuit system, or is communicatively coupled to other components of the circuit system when the device is operating. In an example, any one of the physical components can be used in more than one component in more than one circuit system. For example, in operation, an execution unit can be used for a first circuit of a first circuit system at one point in time, and reused by a second circuit in the first circuit system, or reused by a third circuit in a second circuit system at a different time. The following are additional examples of these components of machine 600.
[0087] In an alternative embodiment, machine 600 can operate as a stand-alone device, or can be connected (e.g., network-connected) to other machines. In a network-connected deployment, machine 600 can operate in a server-client network environment as a server machine, a client machine, or both. In an example, machine 600 can act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Machine 600 can be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a network appliance, a network router, a switch, or a bridge, or any machine capable of executing (sequentially or otherwise) instructions specifying actions to be taken by the machine. Further, although only a single machine is illustrated, the term "machine" should also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein (e.g., cloud computing, software as a service (SaaS), other computer cluster configurations).
[0088] A machine (e.g., a computer system) 600 may include a hardware processor 602 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 604, a static memory (e.g., a memory or storage device for firmware, microcode, basic input / output (BIOS), unified extensible firmware interface (UEFI), etc.) 606, and a mass storage device 608 (e.g., a hard disk drive, a tape drive, a flash storage device, or other block device), some or all of which may communicate with each other via an interconnection (e.g., a bus) 630. The machine 600 may further include a display unit 610, an alphanumeric input device 612 (e.g., a keyboard), and a user interface (UI) navigation device 614 (e.g., a mouse). In an example, the display unit 610, the input device 612, and the UI navigation device 614 may be a touchscreen display. The machine 600 may additionally include a storage device (e.g., a drive unit) 608, a signal generation device 618 (e.g., a speaker), a network interface device 620, and one or more sensors 616, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. The machine 600 may include an output controller 628, such as a serial (e.g., universal serial bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, to communicate with or control one or more peripheral devices (e.g., a printer, a card reader, etc.).
[0089] The registers of the processor 602, the main memory 604, the static memory 606, or the mass storage device 608 may be, or may include, a machine-readable medium 622 on which is stored one or more sets of data structures or instructions 624 (e.g., software) that embody any one or more of the techniques or functions described herein, or are utilized by the techniques or functions. The instructions 624 may also reside, completely or at least partially, within any one of the processor 602, the main memory 604, the static memory 606, or the registers of the mass storage device 608 during execution by the machine 600. In an example, one or any combination of the hardware processor 602, the main memory 604, the static memory 606, or the mass storage device 608 may constitute the machine-readable medium 622. Although the machine-readable medium 622 is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, or associated caches and servers) configured to store one or more instructions 624.
[0090] The term "machine-readable medium" can include any medium that can store, encode, or carry instructions for execution by a machine 600 and cause the machine 600 to perform any one or more of the techniques of the present disclosure, or any medium that can store, encode, or carry data structures used by or associated with such instructions. Non-limiting examples of machine-readable media can include solid-state memory, optical media, magnetic media, and signals (e.g., radio frequency signals, other photon-based signals, sound signals, etc.). In an example, a non-transitory machine-readable medium includes a machine-readable medium having a plurality of particles that have an invariant (e.g., rest) mass and are thus composed of matter. Thus, a non-transitory machine-readable medium is a machine-readable medium that does not include transitory propagated signals. Specific examples of non-transitory machine-readable media can include: non-volatile memory, such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0091] In an example, the information stored or otherwise provided on the machine-readable medium 622 can represent instructions 624, such as the instructions 624 themselves or a format from which the instructions 624 can be derived. This format from which the instructions 624 can be derived can include source code, encoded instructions (e.g., in compressed or encrypted form), encapsulated instructions (e.g., split into multiple encapsulations), etc. The information representing the instructions 624 in the machine-readable medium 622 can be processed by processing circuitry into the instructions to implement any of the operations discussed herein. For example, deriving the instructions 624 from the information (e.g., processed by processing circuitry) can include: compiling (e.g., from source code, object code, etc.), interpreting, loading, organizing (e.g., dynamically or statically linking), encoding, decoding, encrypting, decrypting, encapsulating, de-encapsulating, or otherwise manipulating the information into the instructions 624.
[0092] In an example, the derivation of the instructions 624 can include (e.g., by processing circuitry) the assembly, compilation, or interpretation of the information to create the instructions 624 from some intermediate or pre-processed format provided by the machine-readable medium 622. When the information is provided in multiple parts, the information can be combined, de-encapsulated, and modified to create the instructions 624. For example, the information can be in multiple compressed source code encapsulations (or object code, or binary executable code, etc.) on one or several remote servers. The source code encapsulations can be encrypted when transmitted over a network and, when necessary, decrypted, decompressed, assembled (e.g., linked), and compiled or interpreted (e.g., into independently executable libraries, etc.) at the local machine and executed by the local machine.
[0093] Instruction 624 can be further transmitted or received over communication network 626 using a transmission medium via network interface device 620 using any one of a plurality of transport protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Example communication networks can include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), mobile telephone networks (e.g., cellular networks), plain old telephone (POTS) networks, and wireless data networks (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 802.11 standard series known as , the IEEE 802.16 standard series known as , the IEEE 802.15.4 standard series, peer-to-peer (P2P) networks, etc.). In an example, network interface device 620 can include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas to connect to communication network 626. In an example, network interface device 620 can include multiple antennas to communicate wirelessly using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) techniques. The term "transmission medium" should be considered to include any non-transitory medium that is capable of storing, encoding, or carrying instructions for execution by machine 600, and includes digital or analog communication signals or other non-transitory media to facilitate such communication of software. The transmission medium is a machine-readable medium. To better illustrate the methods and devices described herein, a set of non-limiting example embodiments are set forth below as numbered examples.
[0094] Example 1 is a processor that includes: a pipeline; instruction write-back circuitry configured to: determine a completion time of an instruction before insertion into the pipeline; detect a conflict between the instruction and a different instruction based on the completion time, the different instruction being in the pipeline, the conflict being detected when the completion time equals a second completion time of the different instruction; and calculate a difference between the completion time and an unconflicted completion time; and write-delay circuitry configured to delay the completion of the instruction by the difference.
[0095] In Example 2, the subject matter of Example 1, wherein the completion time is a number of clock cycles.
[0096] In Example 3, the subject matter of Example 2, wherein the difference is a second number of clock cycles.
[0097] In Example 4, the subject matter of Example 3, wherein, to determine the completion time of the instruction, the instruction write-back circuitry is configured to determine that the completion time is less than all possible pipeline stages in the pipeline.
[0098] In Example 5, the subject matter according to Example 4, wherein the maximum delay is eight clock cycles.
[0099] In Example 6, the subject matter according to any one of Examples 1 to 5, wherein, to detect a conflict, the instruction write-back circuitry is configured to: look up the clock cycle in which the completion time falls in a register write-back scoreboard to obtain a result; and detect that the result indicates that different instructions will perform register write-back at the clock cycle.
[0100] In Example 7, the subject matter according to Example 6, wherein the instruction write-back circuitry is configured to: write an indication that the instruction will complete at a second clock cycle as part of inserting the instruction into the pipeline, the second clock cycle corresponding to an unconflicted completion time, in the register write-back scoreboard.
[0101] In Example 8, the subject matter according to any one of Examples 1 to 7, wherein, to delay the completion of an instruction by the difference, the pipeline is configured to: detect that the instruction is ready to perform register write-back at a stage in the pipeline; perform a read on a record corresponding to the stage in the pipeline in a write-delay data structure of a write-delay circuitry; determine that the value of the record is not zero; and prevent the instruction from performing register write-back.
[0102] In Example 9, the subject matter according to Example 8, wherein the instruction write-back circuitry is configured to: write the difference as a value into the record as part of inserting the instruction into the pipeline.
[0103] In Example 10, the subject matter according to any one of Examples 8 to 9, wherein the write-delay circuitry is configured to: decrement the value with each clock cycle.
[0104] In Example 11, the subject matter according to any one of Examples 1 to 10, wherein the pipeline is configured to: in response to the completion of an instruction, move the thread identifier (ID) of the instruction from a pipelined thread ID buffer to a thread ready-to-run queue, the thread ID being inserted into the pipelined thread ID buffer as part of inserting the instruction into the pipeline.
[0105] In Example 12, the subject matter according to any one of Examples 1 to 11, wherein the processor is included in a programmable atomic unit in a memory controller.
[0106] In Example 13, the subject matter according to Example 12, wherein the memory controller is a die in a multi-die system.
[0107] Example 14 is a method that includes: determining the completion time of an instruction before it is inserted into the pipeline of a processor; detecting a conflict between the instruction and different instructions based on the completion time, where the different instructions are in the pipeline, and detecting the conflict when the completion time is equal to the second completion time of the different instructions; calculating the difference between the completion time and the non-conflicting completion time; and delaying the completion of the instruction by the difference.
[0108] In Example 15, for the subject matter according to Example 14, wherein the completion time is the number of clock cycles.
[0109] In Example 16, for the subject matter according to Example 15, wherein the difference is the second number of clock cycles.
[0110] In Example 17, for the subject matter according to Example 16, wherein determining the completion time of the instruction includes determining that the completion time is less than all possible pipeline stages in the pipeline.
[0111] In Example 18, for the subject matter according to Example 17, wherein the maximum delay is eight clock cycles.
[0112] In Example 19, for the subject matter according to any one of Examples 14 to 18, wherein detecting the conflict includes: looking up the clock cycle in which the completion time falls in a register writeback scoreboard to obtain a result; and detecting that the result indicates that different instructions will perform register writeback at the clock cycle.
[0113] In Example 20, for the subject matter according to Example 19, it includes: writing an indication that the instruction will complete at a second clock cycle into the register writeback scoreboard as part of inserting the instruction into the pipeline, where the second clock cycle corresponds to the non-conflicting completion time.
[0114] In Example 21, for the subject matter according to any one of Examples 14 to 20, wherein delaying the completion of the instruction by the difference includes: detecting that the instruction is ready to perform register writeback at a stage in the pipeline; performing a read on a record corresponding to the stage in the pipeline in a write delay data structure; determining that the value of the record is not zero; and preventing the instruction from performing register writeback.
[0115] In Example 22, for the subject matter according to Example 21, it includes: writing the difference as a value into the record as part of inserting the instruction into the pipeline.
[0116] In Example 23, for the subject matter according to any one of Examples 21 to 22, it includes: decrementing the value with each clock cycle.
[0117] In Example 24, a subject matter according to any one of Examples 14 to 23 includes: moving a thread identifier (ID) of an instruction from a pipelined thread ID buffer to a thread ready-to-run queue in response to completion of the instruction, where the thread ID is inserted into the pipelined thread ID buffer as part of inserting the instruction into the pipeline.
[0118] In Example 25, a subject matter according to any one of Examples 14 to 24, where a processor is included in a programmable atomic unit in a memory controller.
[0119] In Example 26, a subject matter according to Example 25, where the memory controller is a die in a chiplet system.
[0120] Example 27 is a machine-readable medium including instructions that, when executed by a circuit system of a processor, cause the processor to perform operations including: determining a completion time of an instruction before inserting it into a pipeline of the processor; detecting a conflict between the instruction and a different instruction based on the completion time, where the different instruction is in the pipeline and the conflict is detected when the completion time is equal to a second completion time of the different instruction; calculating a difference between the completion time and a non-conflicting completion time; and delaying the completion of the instruction by the difference.
[0121] In Example 28, a subject matter according to Example 27, where the completion time is a number of clock cycles.
[0122] In Example 29, a subject matter according to Example 28, where the difference is a second number of clock cycles.
[0123] In Example 30, a subject matter according to Example 29, where determining the completion time of the instruction includes determining that the completion time is less than all possible pipeline stages in the pipeline.
[0124] In Example 31, a subject matter according to Example 30, where the maximum delay is eight clock cycles.
[0125] In Example 32, a subject matter according to any one of Examples 27 to 31, where detecting the conflict includes: looking up a clock cycle in a register writeback scoreboard in which the completion time falls to obtain a result; and detecting that the result indicates that a different instruction will perform a register writeback at the clock cycle.
[0126] In Example 33, a subject matter according to Example 32, where the operation includes: writing an indication that the instruction will complete at a second clock cycle as part of inserting the instruction into the pipeline, where the second clock cycle corresponds to the non-conflicting completion time, in the register writeback scoreboard.
[0127] In example 34, the subject matter according to any one of examples 27 to 33, wherein delaying the completion of the instruction by the difference includes: detecting that the instruction is ready to perform a register write-back at a stage in the pipeline; performing a read on a record corresponding to the stage in the pipeline in a write delay data structure; determining that the value of the record is not zero; and preventing the instruction from performing the register write-back.
[0128] In example 35, the subject matter according to example 34, wherein the operation includes: writing the difference as a value into a record as part of inserting the instruction into the pipeline.
[0129] In example 36, the subject matter according to any one of examples 34 to 35, wherein the operation includes: decrementing the value with each clock cycle.
[0130] In example 37, the subject matter according to any one of examples 27 to 36, wherein the operation includes: in response to the completion of the instruction, moving the thread identifier (ID) of the instruction from a pipelined thread ID buffer to a thread ready-to-run queue, the thread ID being inserted into the pipelined thread ID buffer as part of inserting the instruction into the pipeline.
[0131] In example 38, the subject matter according to any one of examples 27 to 37, wherein the processor is included in a programmable atomic unit in a memory controller.
[0132] In example 39, the subject matter according to example 38, wherein the memory controller is a die in a multi-die system.
[0133] Example 40 is a system that includes: means for determining a completion time of an instruction before inserting it into a pipeline of a processor; means for detecting a conflict between the instruction and a different instruction based on the completion time, the different instruction being in the pipeline, the conflict being detected when the completion time is equal to a second completion time of the different instruction; means for calculating a difference between the completion time and a non-conflicting completion time; and means for delaying the completion of the instruction by the difference.
[0134] In example 41, the subject matter according to example 40, wherein the completion time is the number of clock cycles.
[0135] In example 42, the subject matter according to example 41, wherein the difference is a second number of clock cycles.
[0136] In example 43, the subject matter according to example 42, wherein the means for determining the completion time of the instruction includes means for determining that the completion time is less than all possible pipeline stages in the pipeline.
[0137] In instance 44, the subject matter according to instance 43, wherein the maximum delay is eight clock cycles.
[0138] In instance 45, the subject matter according to any one of instances 40 to 44, wherein the component for detecting a conflict includes: a component for looking up the clock cycle in which the completion time falls in a register write-back scoreboard to obtain a result; and a component for detecting that the result indicates that different instructions will perform register write-back at the clock cycle.
[0139] In instance 46, the subject matter according to instance 45, which includes: a component for writing an indication that an instruction will complete at a second clock cycle as part of inserting the instruction into the pipeline, the second clock cycle corresponding to an unconflicted completion time.
[0140] In instance 47, the subject matter according to any one of instances 40 to 46, wherein the component for delaying the completion of an instruction by the difference includes: a component for detecting that an instruction is ready to perform register write-back at a stage in the pipeline; a component for performing a read on a record corresponding to the stage in the pipeline in a write delay data structure; a component for determining that the value of the record is not zero; and a component for preventing the instruction from performing register write-back.
[0141] In instance 48, the subject matter according to instance 47, which includes: a component for writing the difference as a value into a record as part of inserting the instruction into the pipeline.
[0142] In instance 49, the subject matter according to any one of instances 47 to 48, which includes: a component for decrementing the value with each clock cycle.
[0143] In instance 50, the subject matter according to any one of instances 40 to 49, which includes: a component for moving the thread identifier (ID) of an instruction from a pipelined thread ID buffer to a thread ready-to-run queue in response to the completion of the instruction, the thread ID being inserted into the pipelined thread ID buffer as part of inserting the instruction into the pipeline.
[0144] In instance 51, the subject matter according to any one of instances 40 to 50, wherein the processor is included in a programmable atomic unit in a memory controller.
[0145] In instance 52, the subject matter according to instance 51, wherein the memory controller is a die in a chiplet system.
[0146] Example 53 is at least one machine-readable medium comprising instructions that, when executed by processing circuitry, cause the processing circuitry to perform operations to implement any one of Examples 1 to 52.
[0147] Example 54 is a device that includes components to implement any one of Examples 1 to 52.
[0148] Example 55 is a system that is configured to implement any one of Examples 1 to 52.
[0149] Example 56 is a method that is configured to implement any one of Examples 1 to 52.
[0150] The foregoing detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings illustrate, by way of example, specific embodiments in which the invention may be practiced. These embodiments are also referred to herein as "examples." Such examples may include elements in addition to those shown or described. However, the inventors also contemplate examples in which only those elements shown or described are provided. In addition, the inventors also contemplate examples using any combination or permutation of those elements shown or described with respect to a particular example (or one or more aspects thereof) or with respect to other examples shown or described herein (or one or more aspects thereof).
[0151] In this document, as is common in patent documents, the term "a" is used to include one or more than one, independent of any other instance or use of "at least one" or "one or more." In this document, the term "or" is used to refer to a non-exclusive or, such that "A or B" may include "A but not B," "B but not A," as well as "A and B," unless otherwise indicated. In the appended claims, the terms "comprising" and "in which" are used as the plain-English equivalents of the respective terms "including" and "wherein." Also, in the appended claims, the terms "comprising" and "including" are open-ended, meaning that a system, device, article, or process that includes elements in addition to those listed after such a term in a claim is still considered to be within the scope of that claim. Further, in the following claims, the terms "first," "second," and "third," etc. are used merely as labels, and are not intended to impose numerical requirements on their objects.
[0152] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more aspects thereof) can be used in combination with each other. For instance, other embodiments can be used by those of ordinary skill in the art after reviewing the above description. It should be understood that the embodiments will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the above detailed description, various features can be grouped together to simplify the present disclosure. This should not be construed as intending that the disclosed features not claimed are essential for any claim. In fact, the subject matter of the present invention can lie in less than all of the features of a particular disclosed embodiment. Accordingly, the appended claims are hereby incorporated into the detailed description, where each claim stands on its own as a separate embodiment, and such embodiments can be combined with each other in various combinations or permutations. The scope of the present invention should be determined by reference to the appended claims and the full scope of equivalents to such claims.
Claims
1. A programmable atomic processor for a small chip memory controller, the programmable atomic processor comprises: a memory configured to maintain a register write-back scoreboard; and a write-back slot circuit configured to: determine a register write-back slot for the instruction based on a completion cycle of the instruction; retrieve a value from the register write-back scoreboard at the register write-back slot to determine scheduling of another instruction for the register write-back slot; and establish a delay for execution of the instruction, the delay resulting in a later and uncontended write-back slot for the instruction.
2. The programmable atomic processor according to claim 1, wherein the small chip memory controller is one of a plurality of small chips in a small chip system.
3. The programmable atomic processor according to claim 2, wherein the instruction is part of a programmable atomic operation executed by the programmable atomic processor based on a request for the small chip memory controller from a second small chip in the small chip system.
4. The programmable atomic processor according to claim 3, wherein the small chip memory controller is configured to receive the request via an inter-small chip network interface.
5. The programmable atomic processor according to claim 1, wherein the completion cycle is a number of clock cycles until the instruction is completed.
6. The programmable atomic processor according to claim 5, wherein, to establish the delay, the write-back slot circuit is configured to insert no-operation instructions equal in number to the number of clock cycles to move the completion cycle to the clock cycle represented by the uncontended write-back slot.
7. The programmable atomic processor according to claim 5, wherein, to establish the delay, the write-back slot circuit is configured to reschedule the instruction by some number of clock cycles to move the completion cycle to the clock cycle represented by the uncontended write-back slot.
8. The programmable atomic processor according to claim 1, wherein, to establish the delay, the write-back slot circuit is configured to write an indication to the write-back scoreboard that the later and uncontended write-back slot is being used for the instruction.
9. The programmable atomic processor according to claim 1, comprising processing circuitry configured to execute the instruction after the delay.
10. The programmable atomic processor according to claim 9, wherein the write-back slot circuit is placed between a scheduler and an execution unit in a pipeline of the programmable atomic processor.
11. A method, which comprises: maintaining a register write-back scoreboard in a memory of a programmable atomic processor of a small chip memory controller; determining a register write-back slot for the instruction based on a completion cycle of the instruction; retrieving a value from the register write-back scoreboard at the register write-back slot to determine scheduling of another instruction for the register write-back slot; and establishing a delay for execution of the instruction, the delay resulting in a later and uncontended write-back slot for the instruction.
12. The method according to claim 11, wherein the die memory controller is one of a plurality of dies in a die system.
13. The method according to claim 12, wherein the instruction is part of a programmable atomic operation executed by the programmable atomic processor based on a request for the die memory controller from a second die in the die system.
14. The method according to claim 13, comprising receiving the request via an inter-die network interface of the die memory controller.
15. The method according to claim 11, wherein the completion cycle is the number of clock cycles until the instruction is completed.
16. The method according to claim 15, wherein establishing the delay comprises inserting no-operation instructions equal in number to the number of clock cycles to move the completion cycle to the clock cycle represented by the uncontested write-back slot.
17. The method according to claim 15, wherein establishing the delay comprises rescheduling the instruction by a number of clock cycles to move the completion cycle to the clock cycle represented by the uncontested write-back slot.
18. The method according to claim 11, wherein establishing the delay comprises indicating a write to the write-back scoreboard that the later and uncontested write-back slot is being used for the instruction.
19. The method according to claim 11, comprising executing the instruction after the delay.
20. The method according to claim 19, wherein establishing the delay is performed by a write-back slot circuit of the programmable atomic processor, the write-back slot circuit being located between a scheduler and an execution unit in a pipeline of the programmable atomic processor.
21. A system comprising means for performing the method according to any one of claims 11 to 20.
22. A machine-readable medium comprising instructions that, when executed by a processing circuit, cause the processing circuit to perform the method according to any one of claims 11 to 20.