Apparatus and method for efficiently packing data for transmission over interconnect fabric
Through interconnect structure and transparent queue technology, multiple physically separate dies are connected into a monolithic cache consistency domain, solving the failure risk and efficient connection problems in the processor die manufacturing process, and achieving a high bandwidth and low latency processor system.
Patent Information
- Application Number
- CN202510205350.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-23
- Filing Date
- 2025-02-24
- Publication Date
- 2025-08-26
AI Technical Summary
In the prior art, there is a risk of failure caused by process defects in the manufacturing process of processor dies, and it is difficult to efficiently connect multiple physically separate dies to form a high-performance processor system.
Through an interconnect structure, multiple physically separated dies are connected together to form a monolithic cache consistency domain, and transparent queue and grid interconnection technology are used to achieve efficient data transmission and improved processor performance.
It reduces the risk of single large die manufacturing, improves the bandwidth and performance of processor systems, supports flexible combination and expansion of different dies, and realizes efficient data transmission and sharing of processor resources.
Smart Images

Figure CN120541015A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to computer systems and, more particularly, to apparatus and methods for efficiently packaging data for transmission over an interconnect fabric. Background Art
[0002] A processor or a collection of processors executes instructions from an instruction set (e.g., an instruction set architecture (ISA)). An instruction set is the programming-related part of a computer's architecture and generally includes native data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I / O). BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The present disclosure is illustrated by way of example and not limitation in the accompanying figures in which like references indicate similar elements, and in which:
[0004] Figure 1 The diagram illustrates a hardware processor according to an embodiment of the present disclosure.
[0005] Figure 2A The diagram illustrates a hardware processor according to an embodiment of the present disclosure.
[0006] Figure 2B The diagram illustrates a hardware processor according to an embodiment of the present disclosure.
[0007] Figure 3 The diagram illustrates a hardware processor according to an embodiment of the present disclosure.
[0008] Figure 4 The diagram illustrates transmitter circuitry of a first die coupled to receiver circuitry of a second die via interconnects according to an embodiment of the disclosure.
[0009] Figure 5 The figure shows a data timing diagram and a clock timing diagram for a first clock rate according to an embodiment of the present disclosure.
[0010] Figure 6 The figure shows a data timing diagram and a clock timing diagram for a second clock rate according to an embodiment of the present disclosure.
[0011] Figure 7 The diagram illustrates transmitter circuitry of a first die coupled to receiver circuitry of a second die via interconnects according to an embodiment of the disclosure.
[0012] Figure 8 The figure shows a data timing diagram and a clock timing diagram for a first clock rate according to an embodiment of the present disclosure.
[0013] Figure 9 The figure shows a data timing diagram and a clock timing diagram for a second clock rate according to an embodiment of the present disclosure.
[0014] Figure 10 The diagram illustrates a hardware processor according to an embodiment of the present disclosure.
[0015] Figure 11 The diagram illustrates a hardware processor according to an embodiment of the present disclosure.
[0016] Figure 12 The diagram illustrates a hardware processor according to an embodiment of the present disclosure.
[0017] Figures 13A-13B The figures illustrate example embodiments with interconnections between different dies / structures.
[0018] Figure 14 The diagram illustrates a method for compressing a message according to some embodiments of the present invention.
[0019] Figure 15 The diagram illustrates a method for decompressing a message according to some embodiments of the present invention.
[0020] Figure 16 The diagram illustrates a transaction sequence according to one embodiment.
[0021] Figure 17 The diagram illustrates compression circuitry according to some embodiments.
[0022] Figure 18 The diagram illustrates decompression circuitry in accordance with some embodiments.
[0023] Figure 19 The figure shows an example of a data transmission unit including a plurality of slots.
[0024] Figure 20A-Figure 20B The figure shows an embodiment of an apparatus for packetizing messages using mini-slots.
[0025] Figure 21 The figure shows an example of a collection of messages having a size of one or more minislots.
[0026] Figure 22 The figure shows an example in which messages with different minislot sizes are packed into slots.
[0027] Figure 23 The diagram illustrates one embodiment in which multiple slots are chained to provide more efficient message packing.
[0028] Figures 24A-24BThe figures illustrate different examples in which messages are packed into slots according to their minislot size.
[0029] Figure 25 The figure illustrates a method according to one embodiment.
[0030] Figure 26A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to embodiments of the present disclosure.
[0031] Figure 26B is a block diagram illustrating both an exemplary embodiment of an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor according to embodiments of the present disclosure.
[0032] Figure 27A is a block diagram of a single processor core and its connection to the on-die interconnect network and its local subset of Level 2 (L2) cache according to an embodiment of the present disclosure.
[0033] Figure 27B According to an embodiment of the present disclosure Figure 20A An expanded view of part of the processor core.
[0034] Figure 28 is a block diagram of a processor that may have more than one core, may have an integrated memory controller, and may have integrated graphics according to an embodiment of the present disclosure.
[0035] Figure 29 is a block diagram of a system according to one embodiment of the present disclosure.
[0036] Figure 30 is a block diagram of a more specific exemplary system according to an embodiment of the present disclosure.
[0037] Figure 31 Shown is a block diagram of a system on a chip (SoC) according to an embodiment of the present disclosure.
[0038] Figure 32 is a block diagram illustrating converting binary instructions in a source instruction set into binary instructions in a target instruction set using a software instruction converter according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] In the following description, numerous specific details are set forth. However, it should be understood that embodiments of the present disclosure may be practiced without these specific details. In other instances, well-known circuits, structures, and techniques are not shown in detail to avoid obscuring the understanding of this description.
[0040] References in the specification to "one embodiment," "an embodiment," "an example embodiment," etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment will necessarily include that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an embodiment, it is understood that it is within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in conjunction with other embodiments, whether or not explicitly described.
[0041] A (e.g., hardware) processor or set of processors executes instructions from an instruction set (e.g., an instruction set architecture (ISA)). An instruction set is the programming-related portion of a computer architecture and generally includes native data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I / O). It should be noted that the term instruction herein may refer to macroinstructions, e.g., instructions provided to a processor for execution, or microinstructions, e.g., instructions obtained by decoding a macroinstruction by a decoding unit (decoder) of the processor. A processor (e.g., having one or more cores for decoding and / or executing instructions) may operate on data, e.g., when performing arithmetic, logical, or other functions.
[0042] A processor may be formed on a single die (e.g., a single (semiconductor) block of an integrated circuit). In one embodiment, a single die may have (e.g., manufacturing) errors or defects that impede or remove certain functionality of the die. Such liability for process defects may increase with increasing die area, as may the manufacturing investment at risk of loss in the construction of a (e.g., large) processor. A processor may be formed on a single die (e.g., manufacturing) that has all hardware functionality at one design version, for example, and no hardware-supported functionality has been added, enhanced, or optimized, where these new capabilities were not in the original design version.
[0043] Certain embodiments herein provide for multiple physically separate (e.g., discrete) dies to be connected together (e.g., electrically) via an interconnect to form a processor. Certain embodiments herein provide for a single (e.g., monolithic) cache coherence domain on the interconnect. Certain embodiments herein include not packetizing and / or serializing data (e.g., transmitted and / or received) via the interconnect (e.g., between dies). Certain embodiments herein reduce the risks associated with a single (e.g., large) die size. Certain embodiments herein allow processors to be formed from multiple copies of the same die (and / or mirrored versions of the die) to create (e.g., larger) monolithic domains. Certain embodiments herein allow for redundancy for yield recovery and / or die testability. For example, different dies and / or different groupings of dies can allow for a variety of unique processors (e.g., SKUs) with minimal or no redesign effort. Certain embodiments herein allow for a decision late in the design cycle to decide whether to manufacture a monolithic design of a die or multiple dies (e.g., a 2-way or 4-way split of a single die). Certain interconnects herein include transparent queues across clock and / or power domains, e.g., which can be tuned post-silicon. In certain embodiments, an interconnect (e.g., with transparent queues) can have no latency impact, e.g., if two domains operate at the same frequency but on different power sources. In certain embodiments, a transceiver circuit (e.g., a transmitter circuit and a receiver circuit) includes transparent queues on both the transmitter and receiver circuits, e.g., where data is crossing a physical die boundary, e.g., across a power domain with different power sources on each die.
[0044] Certain embodiments herein provide a monolithic cache domain that spans multiple dies (e.g., allowing very large cross-bandwidth, but also with minimal latency and power impact). Certain embodiments herein allow scaling in two dimensions (e.g., XY) and / or three dimensions (e.g., XYZ). Certain embodiments herein provide a larger die connected to a smaller die (e.g., multiple dies with different numbers of physical connections on their dies). Certain embodiments herein allow transfers according to multiple (e.g., any) protocols between dies (e.g., not limited to a single protocol). Certain embodiments herein provide a mesh loop (e.g., micro) architecture, e.g., to tolerate die-to-die variations. Certain embodiments herein add entries to a lookup table (LUT) to indicate whether data (e.g., a cache line) is to cross a physical die boundary, e.g., through an interconnect between two dies. Certain embodiments herein allow for independent (e.g., power and / or cache) domains as needed, e.g., by disabling rows and / or columns of (e.g., mesh) interconnects to aid yield recovery. Certain embodiments herein allow one die to run at a different frequency than another die of the hardware processor. Some of the transmission protocols herein enable high-speed interconnection between multiple dies and / or seamless crossing of die boundaries. As an alternative to using those protocols to serve as die-to-die connections, some embodiments herein may use other solutions, such as utilizing an interposer.
[0045] Certain embodiments of interconnects between multiple dies provide one or more of: (e.g., very high) increased bandwidth, reduced pin count but allowing full cross-sectional bandwidth, 1 / 4 pins used with 4x the frequency of the die, 1 / 2 pins used with dynamic 1x / 2x modes, for example, 1x: half bandwidth (e.g., operating frequency matches the die because of 1 / 2 pins, 1 / 2 bandwidth) with low power and / or latency impact, no packetization (e.g., for any die-to-die connection) for minimal latency impact, lower frequency and / or lower error rate (e.g., error rate similar to or less than on-silicon error rate) (e.g., to allow not utilizing error protection on the inter-die interconnect link, or utilizing error protection of the on-die interconnect on the inter-die interconnect link), and for example 2x: full bandwidth full performance with increased power and / or latency, doubling the operating frequency relative to the die frequency, and algorithm(s) for switching between the two modes. Certain embodiments of interconnects between multiple dies herein provide reduced latency and / or increased bandwidth of the interconnect, eg, much less than current die-to-die interconnect technology and / or equal or substantially equal to on-die interconnect.
[0046] Certain embodiments herein provide for sharing processor master resources on high bandwidth and low latency electrical interconnects, such that the performance of accessing remote die resources is substantially similar to or very close to the performance of an integrated die manufactured on a single chip. Certain embodiments herein provide for sharing processor infrastructure resources so that power, heat, clock, reset, configuration, error handling, etc. can be tightly managed using electrical interconnects, such that the performance of accessing remote die resources is substantially similar to or very close to the performance of an integrated die manufactured on a single chip. Certain embodiments herein reduce the manufacturing yield risk associated with a single large die size. Certain embodiments herein allow scaling to a certain (e.g., larger) number of functional logic circuit components to provide redundancy for yield recovery and / or special purposes such as die testability. Certain embodiments herein allow for deciding at a later stage of the design cycle (e.g., or at any time) whether to manufacture a monolithic design of a die or multiple dies (e.g., a 2-way or 4-way split of a single die).
[0047] Certain embodiments herein allow for the combination of dissimilar dies to enable phased design completion of some dies over time or to manufacture some dies in more mature or specialized manufacturing processes, as well as to better monetize some older dies from previous products. Certain embodiments herein allow for the combination of dissimilar dies and / or different numbers of dies to enable a variety of unique processor products (e.g., SKUs) with minimal or no redesign effort.
[0048] Certain embodiments herein provide for a larger die connected to a smaller die and / or multiple dies having different numbers of physical connections on their dies. Certain embodiments herein allow forming a processor from multiple copies of the same die and / or mirrored versions of a die to create a larger monolithic domain. Certain embodiments herein allow scaling in two dimensions (e.g., X and Y axes in Cartesian coordinates) and / or three dimensions (e.g., X, Y, and Z axes in Cartesian coordinates).
[0049] Certain embodiments herein provide circuit systems (e.g., PHYs) to deliver low latency high bandwidth die-to-die consistent connectivity, e.g., substantially similar to a monolithic experience. Certain embodiments herein provide performance neutrality and power savings equivalent to the monolithic case. Certain embodiments herein provide a coherent flow from individual dies in a wafer to a packaged modular die product. Certain embodiments herein provide modularity and scalability for splicing together several modular dies (e.g., heterogeneous modular dies). Certain embodiments herein allow dies to affect each other seamlessly and unimpeded, even with security protections exposed to dies with private sideband messaging between them.
[0050] Figure 1The figure shows a hardware processor 100 according to an embodiment of the present disclosure. Although not depicted, certain circuitry (e.g., decode unit(s), execution unit(s), core(s), cache coherence circuitry, cache(s), or other components) may be utilized, for example, as described below. In one embodiment, the processor components on a single die 102 may be interconnected via interconnects such as Figure 1 104). For example, die 102 may include component 108 and component 110 that communicate with each other via the mesh interconnect. In one embodiment, physically separate die 102 communicates with physically separate die 104 via interconnect 106. The die and / or the interconnect may include transceivers for transmitting data between die 102 and die 104. Note that a unidirectional arrow herein may not require unidirectional communication; for example, it may indicate bidirectional communication (e.g., to or from that component). Any or all combinations of communication paths may be utilized in certain embodiments herein.
[0051] In one embodiment, each of die 102 and die 104 are identical. In another embodiment, die 104 is a mirror image of die 102. In one embodiment, die 102 and die 104 are distinct, e.g., each die represents part of a single die design that has been diced into multiple physical dies that are then connected together (e.g., electrically coupled) via interconnects.
[0052] In one embodiment, the mesh interconnect of a die does not rely on a connection to another die to function, e.g., data signals (e.g., requests and / or responses) can be looped back to the die if interconnect 106 is not functioning or present. In one embodiment, such data signals are not blocking signals (e.g., not fences).
[0053] The cache coherence circuitry in each of the plurality of physically separate dies may be switched between a master mode and a slave mode. In one embodiment, management circuitry (e.g., a controller) is configured to set one of the cache coherence circuits in each of the plurality of physically separate dies as, for example, a master device and to set the remaining cache coherence circuits as slaves of the master device. The cache coherence circuitry may be located in a controller (e.g., Figure 25-28 within the (one or more) controllers).
[0054] Figure 2AThe figure illustrates a hardware processor 200A according to an embodiment of the present disclosure. In the depicted embodiment, die 202 and die 204 are smaller than die 206, die 208, die 210, and die 212. Each of the depicted dies is coupled to an adjacent die via an interconnect (INT). Die 202 is depicted as having two connections (e.g., discrete interconnects) with die 206. Die 204 is depicted as having a different number (e.g., three) of connections (e.g., discrete interconnects) with die 208. Die 206 is depicted as having four connections (e.g., discrete interconnects) with die 208. Die 210 is depicted as having a different number (e.g., three) of connections (e.g., discrete interconnects) with die 212.
[0055] The intersections of the mesh interconnects of the dies (e.g., intersection 214 or intersection 216 of die 206) can be access points into the mesh interconnects, for example, by circuit components. In one embodiment, multiple (e.g., any) mesh configurations having different sizes on their respective dies are coupled together by certain embodiments herein. In one embodiment, a die with a mesh interconnect is coupled to a die without a mesh interconnect, for example, die 218 at Figure 2A 2 is depicted as a mesh interconnect coupled to die 206 via a single interconnect (INT).
[0056] Figure 2B The diagram illustrates a hardware processor 200B according to an embodiment of the present disclosure. In the depicted embodiment, dies 202 and 204 are smaller than dies 206, 220, 222, and 212. Die 220 is depicted as including a different mesh interconnect than die 222 (e.g., having a different number of crosspoints). Figure 2B The figures illustrate that in some embodiments some of the plurality of dies may be different (eg, in one embodiment, they are not symmetrical). Figure 2B The figures illustrate that in certain embodiments, a mesh interconnect on a die may be different from another mesh interconnect on a different die (eg, in one embodiment, they are asymmetric).
[0057] Figure 3 The figure shows a hardware processor 300 according to an embodiment of the present disclosure. For clarity, the mesh interconnect is not shown in each tube core, but the mesh interconnect can be used, for example, as shown in FIG. Figure 1 Or as shown in Figure 2. Figure 3The diagram illustrates a three-dimensional stacked architecture. Multiple dies can extend in any single direction (e.g., with interconnect(s) between each die). In the depicted embodiment, die 302 and die 304 extend in a first single plane, and die 306 and die 308 extend in a second, different single plane that is laterally spaced apart from the first single plane. The dies can be attached to another substrate, such as a mounting substrate (not depicted).
[0058] In certain embodiments, a first die communicates with one or more other dies (e.g., to and / or from one or more other dies), for example, via electrical connections therebetween. A transceiver (e.g., comprising transmitter circuitry and / or receiver circuitry) may be used in one or more of the dies in the die and / or in the interconnects between the dies. The transceiver (e.g., transceiver circuitry) may include a physical transport layer (e.g., PHY) circuit (e.g., input / output PHY, i.e., I / O PHY). The transceiver may be used for communication between multiple dies (e.g., multiple dies including a split-die processor arrangement). In one embodiment, one or more of the I / O ports (e.g., grid lines) of one or more of the multiple dies is electrically coupled to the I / O ports (e.g., grid lines) of one or more of the other dies. In one embodiment, one or more of the multiple dies includes a grid interconnect within the die, and one or more of the I / O ports (e.g., grid lines) of each grid interconnect may be electrically coupled to the I / O ports (e.g., grid lines) of the grid interconnect of another die, for example, at a die boundary intersection. The electrical coupling of the die can be customized for optimized power and latency performance. The coupling (e.g., line) can be bidirectional, unidirectional, or a combination of both. The physical medium that connects and allows signaling between multiple die transceivers (e.g., I / O PHY) can be an interconnect or other electrical connection.
[0059] Transceiver (e.g., I / O PHY) lanes and / or interconnect lanes (e.g., communication lanes) can be programmed to operate at multiples of a processor (e.g., mesh interconnect) (e.g., on-die) line data transmission rate (e.g., data rate). For example, one times (1X) the clock rate (e.g., PHY) of data (e.g., clock rate) is a 1:1 ratio between the interconnect and / or transceiver (e.g., PHY I / O) (e.g., lane) data transmission rate (e.g., data rate) and the die (e.g., mesh interconnect or mesh line) data transmission rate (e.g., data rate). For example, two times (2X) the clock rate (e.g., PHY) of data (e.g., clock rate) is a 2:1 ratio between the interconnect and / or transceiver (e.g., PHY I / O) (e.g., lane) data transmission rate (e.g., data rate) and the die (e.g., mesh interconnect or mesh line) data transmission rate (e.g., data rate). In one embodiment, the interconnect and the portion of the transceiver directly coupled to the interconnect have the same data rate, e.g., different from the internal (e.g., within-mesh) interconnect data rate of the die. As another example, other ratios are possible, e.g., 3 times (3X), 4 times (4X), 5 times (5X), 6 times (6X), 7 times (7X), 8 times (8X), 9 times (9X), 10 times (10X), etc. The clocking scheme for the transceiver (e.g., PHY I / O) can be source synchronous (e.g., for higher per-line bandwidth performance) or common clock (e.g., for lower bandwidth targets).
[0060] Figure 4 The diagram shows transmitter circuitry 402 of a first die coupled to receiver circuitry 404 of a second die via interconnect 406 according to an embodiment of the disclosure. Figure 4 A high-level (e.g., source-synchronous clock) circuit diagram of a transceiver (e.g., PHY I / O) for connecting two dies together (e.g., for data transmission therebetween) is shown. Transmitter circuit 402 includes multiple transmitters (412A, 412B, 412C, 412D) that generate (e.g., amplify) signals. Receiver circuit 404 includes multiple receivers (414A, 414B, 414C, 414D) (e.g., samplers) that receive transmitted signals. Interconnect 406 includes multiple paths (416A, 416B, 416C, 416D). In some embodiments, the interconnect can have any one or more of these paths. In some embodiments, the interconnect can include each of these paths. In one embodiment, each of these paths is a discrete wire of the interconnect. Although a single data path 416 is depicted, multiple data paths (e.g., including one or more respective instances of one or more of the components of transmitter circuitry 402 and / or receiver circuitry 404) may be utilized, for example, along with a single clock path associated with those multiple data paths.
[0061] In some embodiments, transmitter circuitry 402, interconnect 406, and / or receiver circuitry 404 (e.g., any one or any combination thereof) includes circuitry (e.g., clock circuitry) for varying the operating frequency and / or the clock rate for the operating frequency. In some embodiments, a clock phase placement is determined (e.g., predetermined) for one or more operating frequencies and / or one or more clock rates for those one or more operating frequencies (e.g., as discussed herein). As an example, data to be transmitted from a first die to a second die may be received by transmitter circuitry 402 of the first die and then sent to the second die via interconnect 406 via receiver circuitry 404. The first die may operate at the operating frequency and the second die may operate at (e.g., the same) operating frequency, but the clock circuitry (e.g., clock circuitry 408) may adjust the clock phase placement for the operating frequency (e.g., and the clock rate for the operating frequency) from a plurality of clock phase placements (e.g., for the same clock cycle). For example, the clock phase placement for the operating frequency may be selected so that no data is lost or a minimal amount of data is lost during transmission. In one embodiment, the intra-die interconnect operates at multiple clock rates relative to the operating frequency of different (eg, inter-die) interconnects of one or more dies coupled to the intra-die interconnect.
[0062] As an example, the transmitter circuit 402 may receive data to be transmitted to the receiver circuit 404 (e.g., a second die including the receiver circuit 404) from the data generator 421 of the first die. The data generator 421 of the first die may be a processor of the first die (e.g., a processor including a decoder for decoding instructions and an execution unit for executing the decoded instructions to generate data). The data to be transmitted may include first data (e.g., a data stream) (e.g., data D0) and (e.g., a separate) second data (e.g., a data stream) (e.g., data D1).
[0063] A clock signal from transmitter circuit 402 (e.g., transmitter side) (e.g., from or based on a clock signal in the first die) can be sent (e.g., forwarded) along with (e.g., simultaneously with) data (e.g., payload data) being sent to receiver circuit 404. Clock circuit 420 can be an internal (e.g., master) clock of the first die (e.g., a grid in the first die). Clock circuit 410 can be a separate clock generator (e.g., separate from the internal (e.g., master) clock of the first die) and / or a dedicated clock circuit for transmitter circuit 402. A multiplexer can select and output one of multiple inputs based on a control signal. Multiplexer (mux) 428 can be configured to provide a clock signal from clock circuit 410 or clock circuit 420, for example, based on a control signal. Multiplexer 428 can be controlled by power management circuit 432, for example, based on a control signal received from a power management circuit (e.g., a power management controller). The power management circuitry may control switching of operating frequencies and / or clock rates (e.g., operating frequencies and / or clock rates in the first die and / or the second die (e.g., connected to the first die via an interconnect). Local and / or dedicated clock circuitry (e.g., clock circuitry 410) (e.g., phase-locked loop (PLL) circuitry) (e.g., in the I / O PHY) may be employed to achieve higher I / O bandwidth by filtering (e.g., meshing) barrier clock jitter components.
[0064] In the depicted embodiment, the multiplexer 428 multiplexes the received clock signal (e.g., Figure 5 and Figure 6 ) as a control signal to multiplexer 424. Multiplexer 424 may also take a second input from valid signal circuit 418, for example, so that when valid signal circuit 418 indicates inactive (e.g., logic zero), multiplexer 424 does not provide an output. Multiplexer 424 may then output data (e.g., payload data) from its output to data path 416B, for example, via transmitter 412B.
[0065] A multiplexer 430 may be included so that the clock signal output from multiplexer 428 passes through both multiplexer 424 and multiplexer 430, for example, to replicate the delay through multiplexer 424. Multiplexer 430 may have a first input that is a ground and a second input that is a power source. In the depicted embodiment, multiplexer 430 outputs its signal to clock path 416C (e.g., via transmitter 412C) and clock return path 416D (e.g., via transmitter 412D).
[0066] Although two data sources (e.g., D0 and D1) (e.g., two lines or two signals that will cross the die boundary to another die) are depicted in some figures herein as sharing a single data path, it should be understood that a single data source (e.g., line or signal) may utilize a single data path, such as data path 412.
[0067] For example, for each different operating frequency, one or more components of circuit 400 may switch from a first clock rate to a different second clock rate.
[0068] Clock gating can be employed to conserve power by enabling a (e.g., data) valid signal (e.g., active only when data on a connection (e.g., a data link, e.g., one or more lanes of the link) is active (e.g., intended for data transmission)). Valid signal controller 418 can generate a valid signal, for example, when a first die is about to transmit data to a second die. In some embodiments, the data signal (e.g., data payload) is separate from the control signal. Valid signal circuit 418 (e.g., valid signal controller) can be part of a power management circuit (e.g., a power management controller). The power management circuit can be a component of the die. Each die can have its own power management controller. Valid signal circuit 418 can assert a valid signal or an invalid signal, for example, to start or stop receiving and / or transferring data from a first die (e.g., from transmitter circuit 402) to a second die (e.g., to receiver circuit 404) and / or away from the second die (e.g., away from receiver circuit 404), for example, by shutting down receivers 414B and / or 414C (respectively). Retimer circuit 425 may retime the data valid signal (eg, exiting receiver 414A) based on the clock phase placement.
[0069] Receiver circuit 404 may receive a valid signal on valid path 416A of interconnect 406, a data signal on data path 416B of interconnect 406, and / or a clock signal (or an inverse signal, or a combination of those signals as a strobe signal) on clock path 416C and / or clock path 416D of interconnect 406. Retimer circuit 425 may retime the valid signal so that it is synchronized with the data and / or clock signal(s) transmitted with it. For example, a valid data signal may be transmitted for one or more data streams and may be output to AND gate 422. AND gate 422 may receive a clock signal from clock circuit 408 of receiver circuit 404, for example, such that the output of AND gate 422 is used to turn on one of multiple receivers 414B and 414C (e.g., where a NOT gate (inverter) is included before the control signal is input to receiver 414B).
[0070] like Figure 5 As shown in , this allows data to be transmitted serially from source D0, then source D1, then source D0 again, and so on, so that the data signal alternates between D0 and D1 (e.g., subject to any data signal being output, such as a logic high (e.g., 1) or a logic low (e.g., 0)). Thus, multiplexer 426 can alternate between outputting data from receiver 414B and from receiver 414C. A control signal (e.g., the output of AND gate 422) is used to switch the multiplexer 426 input between obtaining output from receiver 414B and from receiver 414C.
[0071] The depicted clock circuit 408 receives one or more input clock signals from the transmitter circuit 402 and is used to align one or more of the clock edges with the received data signal (e.g., payload data on the data path 416B (which may be more than one data path)) so that the received data is correctly received (e.g., so that the data sent from the transmitter circuit 402 matches the data received at the receiver circuit 404). In one embodiment, the clock circuit 408 is used to shift the phase (rather than the frequency) of the received clock signal to align it with the received data signal (e.g., payload data on the data path 416B) as needed.
[0072] In one embodiment, the clock circuit 408 of the receiver circuit 404 includes circuitry for aligning (e.g., shifting) the (e.g., source-synchronous) clock edges of a received clock signal (e.g., waveform) from the transmitter circuit 402 with a corresponding received data signal (e.g., different from the clock signal) for high-performance timing, e.g., so that data in the data signal is not altered, lost, corrupted, or any combination thereof. The clock circuit 408 may include a clock phase delay generator 408A (e.g., a DLL circuit) and / or a phase interpolator circuit 408B. In one embodiment, the clock phase alignment is performed by a phase interpolator (e.g., phase interpolator circuit 408B). In one embodiment, the phase interpolator is a circuit that adjusts (e.g., shifts) the phase of the clock signal. In one embodiment, the phase interpolator has a granularity level of, for example, equally spaced steps for each clock phase (e.g., 2, 4, 6, 8, 10, 12, etc.), and the phase interpolator can set the rising clock edge and / or falling clock edge at any of these steps, for example, as further discussed below with reference to FIG. 13 .
[0073] Clock circuitry 408 (e.g., including a delay-locked loop (DLL) circuit) can be employed at the receiver circuitry 404 of the receiver die to properly align source-synchronous clock edges for high-performance timing (e.g., to enable efficient high-speed signal transmission). A DLL circuit can be a negative-delay gate placed in the clock path of a digital circuit. In one embodiment, clock circuitry 408 is a component of receiver circuitry 404. Local and / or dedicated clock circuitry (e.g., clock circuitry 410) (e.g., a phase-locked loop (PLL) circuit) (e.g., within an I / O PHY) can be used to achieve higher I / O bandwidth by filtering (e.g., grid) barrier clock jitter components. A PLL circuit can be a control circuit that generates an output signal whose phase is related to the phase of an input signal. While there are different types of PLL circuits, one example is a circuit having a variable frequency oscillator and a phase detector in a feedback loop, e.g., where the oscillator generates a periodic signal, the phase detector compares the phase of the signal to the phase of an input periodic signal, and adjusts the oscillator to maintain phase matching. The PLL may be an all-digital PLL (ADPLL). In one embodiment, the DLL circuit uses a variable phase (e.g., delay) block and the PLL circuit uses a variable frequency block. Clock circuit 408 may include a control register 409, for example, for storing clock phase placement settings, for example, so that clock circuit 408 applies those settings.
[0074] To maintain high power efficiency of the transmitter circuitry and / or receiver circuitry (e.g., I / O PHY), techniques such as low swing signaling, clock gating, and aggregating source synchronous clock power across multiple (e.g., a large number) of serviced datapaths may be employed. For example, one forwarded source synchronous clock may be used for each of 2, 3, 4, 5, 6, 7, 8, 16, 32, 64, 128, 256, etc. datapaths, or any subset thereof. Datapath 416B is merely an example, and multiple pathways may be utilized. In some embodiments, clock phase delay generator 408A (e.g., DLL circuitry) generates locked (e.g., non-clocked) timing for a clock rate of an operating frequency (e.g., such as a clock source). Figure 16 ) (e.g., 90 degrees or 180 degrees of clock phase lock, e.g., as shown in Figure 6 and Figure 7). In some embodiments, phase interpolator circuit 408B subdivides those clock signals into a finer granularity. In some embodiments, clock circuit 408 utilizes predetermined clock phase placement data (e.g., prior to current data transmission), for example, clock phase delay generator 408A (e.g., DLL circuit) and / or phase interpolator circuit 408B both utilize predetermined clock phase placement data. In one embodiment, clock phase delay generator 408A is a clock phase controller or clock phase adjuster. In one embodiment, clock phase delay generator 408A maintains a certain phase relationship of the clock arriving at a receiver (e.g., sampler) (e.g., of the second die) relative to one or more input clocks coming in from a transmitter (e.g., of the first die). In some embodiments, clock phase delay generator 408A generates the clock phase delay, and phase interpolator circuit 408B is used to further subdivide those clock signals into a finer granularity. In one embodiment, the clock phase delay generator 408A looks up and utilizes a lock code for a particular clock rate and / or operating frequency, and / or the phase interpolator circuit 408B looks up and utilizes a buffer setting for the phase interpolator for a particular clock rate and / or operating frequency. For example, the lock code (e.g., of the DLL) can be changed for each frequency and / or each process, voltage, and / or temperature point (e.g., multiple points), and the phase interpolator circuit can perform (e.g., finer granularity) clock (e.g., edge) placement within the (e.g., DLL) lock code. Once the (e.g., predetermined) clock phase placement for the operating frequency and clock rate is looked up and updated into the circuit system (e.g., clock circuit 408), data can be received by the receiver circuit, for example, output to the data buffer 434 (e.g., as Figure 21 ).
[0075] Figure 5 The diagram illustrates a data timing diagram 501 and a clock timing diagram 502 for a first clock rate according to an embodiment of the present disclosure. In the depicted embodiment, the clock timing diagram 501 illustrates a clock signal (e.g., Figure 16 The data timing diagram 501 illustrates that data at the 1x clock rate can be read in at each falling edge of the clock (e.g., using Figure 4 As discussed herein, clock edges may be placed using a predetermined clock phase placement (eg, relative to data timing).
[0076] Figure 6The diagram illustrates a data timing diagram 601 and a clock timing diagram 602 for a second clock rate according to an embodiment of the present disclosure. In the depicted embodiment, the clock timing diagram 601 illustrates a clock signal (e.g., Figure 16 The data timing diagram 601 illustrates that data at the 2x clock rate can be read in at each of the rising and falling edges of the clock (e.g., using Figure 4 As discussed herein, predetermined clock phase placement (eg, relative to data timing) may be used to place clock edges.
[0077] Figure 7 The diagram shows transmitter circuitry 702 of a first die coupled to receiver circuitry 704 of a second die via interconnect 706 according to an embodiment of the disclosure. Figure 7 A high-level (e.g., source-synchronous clock) circuit diagram of a transceiver (e.g., PHY I / O) for connecting two dies together (e.g., for data transmission therebetween) is shown. Transmitter circuit 702 includes multiple transmitters (712A, 712B, 712C, 712D) that generate (e.g., amplify) signals. Receiver circuit 704 includes multiple receivers (714A, 714B, 714C, 714D, 714E, 714F) that receive the transmitted signals. Interconnect 706 includes multiple paths (716A, 716B, 716C, 716D). In some embodiments, the interconnect can have any one or more of these paths. In some embodiments, the interconnect can include each of these paths. In one embodiment, each of these paths is a discrete line of the interconnect. Although two data paths are depicted (i.e., data paths 716B and 716D), a single data path or three or more data paths (e.g., including one or more respective instances of one or more of the components of transmitter circuitry 702 and / or receiver circuitry 704) may be utilized, for example, along with a single clock path associated with those multiple data paths. For example, a single data source (e.g., D0) may be utilized, for example, by removing the control signal line from clock circuitry 710 to multiplexer 724 (and / or removing multiplexer 724 and / or outputting data from data path 716B directly to a single receiver (e.g., receiver 714E) without using multiplexer 726).
[0078] In some embodiments, transmitter circuitry 702, interconnect 706, and / or receiver circuitry 704 (e.g., any one or any combination thereof) includes circuitry (e.g., clock circuitry) for varying the operating frequency and / or the clock rate for the operating frequency. In some embodiments, a clock phase placement is determined (e.g., predetermined) for one or more operating frequencies and / or the clock rate for those one or more operating frequencies (e.g., as discussed herein). As an example, data (e.g., payload data) to be transmitted from a first die to a second die may be received by transmitter circuitry 702 and then sent to the second die via receiver circuitry 704 via interconnect 706. The first die may operate at the operating frequency and the second die may operate at (e.g., switch to) the (e.g., same) operating frequency, but the clock circuitry (e.g., clock circuitry 708) may adjust the clock phase placement for the operating frequency (e.g., and the clock rate for the operating frequency) from a plurality of clock phase placements (e.g., for the same clock cycle). For example, the clock phase placement for the operating frequency may be selected such that no data or a minimal amount of data is lost during transmission.
[0079] As an example, the transmitter circuit 702 may receive data to be transmitted to the receiver circuit 704 (e.g., a second die including the receiver circuit 704) from the data generator 720 and / or the data generator 730 of the first die (e.g., which may be combined into a single data generator). The data generator 720 and / or the data generator 730 of the first die may be one or more processors of the first die (e.g., each processor including a decoder for decoding instructions and an execution unit for executing the decoded instructions to generate data). The data to be transmitted may include any one of the first data (e.g., a data stream) (e.g., data D0), the (e.g., separate) second data (e.g., a data stream) (e.g., data D1), the (e.g., separate) third data (e.g., a data stream) (e.g., data D2), the (e.g., separate) fourth data (e.g., a data stream) (e.g., data D3), or any combination thereof.
[0080] A clock signal from transmitter circuit 702 (e.g., transmitter side) (e.g., from or based on a clock signal in the first die) can be sent (e.g., forwarded) along with (e.g., simultaneously with) data (e.g., payload data) sent to receiver circuit 704. Clock circuit 710 can be an internal (e.g., master) clock of the first die (e.g., of a grid in the first die), a separate clock generator separate from the internal (e.g., master) clock of the first die, and / or a dedicated clock circuit for transmitter circuit 702.
[0081] As a component of interconnect 706 or separate from interconnect 706, circuit 700 (or other circuits herein) may include a control path for sending control signals from a first die (e.g., via transmitter circuit 702) to a second die (e.g., via receiver circuit 704). The control signal may be sent by power management circuit 740 (e.g., a power management controller), for example, to receiver circuit 704 (e.g., clock circuit 708 of receiver circuit 704 and / or the second die). The control signal may switch a circuit (e.g., a clock circuit) between closed-loop mode and open-loop mode. The power management circuit may control switching of operating frequency and / or clock rate (e.g., operating frequency and / or clock rate in the first die and / or the second die (e.g., connected to the first die via the interconnect)). Local and / or dedicated clock circuits (e.g., clock circuit 710) (e.g., phase-locked loop (PLL) circuitry) (e.g., in an I / O PHY) may be used to achieve higher I / O bandwidth by filtering (e.g., grid) barrier clock jitter components. In one embodiment, a first die is used to request a second die (e.g., two dies) to operate at different frequencies and / or clock rates based on usage, for example, operating at a (e.g., single) frequency and increasing the clock rate when data is being backed up (e.g., in a buffer in the first die), and / or operating at a (e.g., single) frequency and decreasing the clock rate when data is not being backed up (e.g., an empty or unfilled buffer in the first die).
[0082] In the depicted embodiment, clock circuit 710 converts a clock signal (e.g., Figure 8 and Figure 9 ) as a control signal to multiplexer 724 and / or multiplexer 734. Multiplexer 724 can then output data (e.g., payload data) from its output to datapath 716B, for example, via transmitter 712B, and / or multiplexer 734 can then output data (e.g., payload data) from its output to datapath 716D, for example, via transmitter 712D. A clock signal can be transmitted from transmitter circuit 702 to transmitter 712C, transmitted through clock (e.g., strobe) path 716C (e.g., of interconnect 706) to receiver 714C of receiver circuit 704, and then transmitted to clock circuit 708, for example.
[0083] Although two pairs of data sources (e.g., D0 / D1 and D2 / D3) (e.g., four lines or four signals that will cross the die boundary to another die) are depicted in some figures herein as sharing a single data path, it should be understood that a single data source (e.g., line or signal) may utilize a single data path, e.g., data path 716B or data path 716D.
[0084] For example, for each different operating frequency, one or more components of circuit 700 may switch from a first clock rate to a different second clock rate.
[0085] Clock gating can be employed to conserve power by enabling (e.g., data) control signals (e.g., active only when data on a connection (e.g., a data link, such as one or more lanes of the link) is active (e.g., to be used for data transmission)). For example, when a first die is about to transmit data to a second die, power management circuitry 740 (e.g., a power management controller) can generate valid data and / or frequency change and / or clock rate change signals. In some embodiments, the data signal (e.g., data payload) is separate from the control signals. The power management circuitry can be a component of the die. Each die can have its own power management controller. The power management circuitry can assert valid or invalid signals, for example, to start or stop receiving and / or transferring data from a first die (e.g., from transmitter circuitry 702) to a second die (e.g., to receiver circuitry 704) and / or away from the second die (e.g., away from receiver circuitry 704) by shutting down transmitter(s) and / or receiver(s), respectively.
[0086] The receiver circuit 704 may receive a control signal (e.g., to change the frequency and / or clock rate) on a control path 716A of the interconnect 706, a data signal on a data path 716B of the interconnect 706, a data signal on a data path 716D of the interconnect 706, and / or a clock signal (or an inverse signal, or a combination of those signals as a gating signal) on a clock path 716C of the interconnect 706. For example, the power management circuit 740 may send a signal to the receiver circuit 704 (e.g., its clock circuit 708) to enable a particular frequency and / or clock rate for the receiver circuit 704 (e.g., its clock circuit 708), e.g., the same frequency and / or clock rate as the transmitter circuit 702.
[0087] The receiver 722 may receive a clock signal from the clock circuit 708 of the receiver circuit 704, for example, so that the output of the receiver 722 is used to turn on one of the plurality of receivers 714B and 714E (for example, wherein a NOT gate (inverter) is included before the control signal is input into the receiver 714B) (for example, and turn off the other receiver in the pair) and / or turn on one of the plurality of receivers 714D and 714F (for example, wherein a NOT gate (inverter) is included before the control signal is input into the receiver 714D) (for example, and turn off the other receiver in the pair). Figure 8As shown in , this allows data to be serially transmitted from source D0, then source D1, then source D0 again, and repeatedly so that the data signal alternates between D0 and D1 (e.g., subject to any data signal being output, such as a logic high (e.g., 1) or a logic low (e.g., 0)) and / or (e.g., in parallel with the serial transmission of D0 and D1) data to be serially transmitted from source D2, then source D3, then source D2 again, and repeatedly so that the data signal alternates between D2 and D3 (e.g., subject to any data signal being output, such as a logic high (e.g., 1) or a logic low (e.g., 0)). Thus, multiplexer 726 can alternate between outputting data from receiver 714B and from receiver 714E. A control signal (e.g., the output of receiver 722) (e.g., a received source synchronous clock after it has passed through the DLL / PI / clock distribution circuitry) is used to switch the multiplexer 726 input between obtaining output from receiver 714B and from receiver 714E. Thus, multiplexer 728 can alternate between outputting data from receiver 714D and from receiver 714F. A control signal (e.g., the output of receiver 722) (e.g., a received source synchronous clock after it has passed through the DLL / PI / clock distribution circuitry) is used to switch the multiplexer 728 input between obtaining output from receiver 714D and from receiver 714F.
[0088] The depicted clock circuit 708 receives one or more input clock signals from the transmitter circuit 702 and is used to align one or more of the clock edges with the received data signal (e.g., payload data on data path 716B and / or data path 716D (which may be more than two data paths)) so that the received data is correctly received (e.g., so that the data sent from the transmitter circuit 702 matches the data received at the receiver circuit 704). In one embodiment, the clock circuit 708 is used to shift the phase (rather than the frequency) of the received clock signal so that it is aligned with the received data signal (e.g., payload data on data path 716B and / or data path 716D) as needed.
[0089] In one embodiment, the clock circuit 708 of the receiver circuit 704 includes circuitry for aligning (e.g., shifting) the (e.g., source-synchronous) clock edges of a received clock signal (e.g., waveform) from the transmitter circuit 702 with a corresponding received data signal (e.g., different from the clock signal) for high-performance timing, e.g., so that data in the data signal is not altered, lost, corrupted, or any combination thereof. The clock circuit 708 may include a clock phase delay generator 708A (e.g., a DLL circuit) and / or a phase interpolator circuit 708B. In one embodiment, the clock phase alignment is performed by a phase interpolator (e.g., phase interpolator circuit 708B). In one embodiment, the phase interpolator is a circuit that adjusts (e.g., shifts) the phase of the clock signal. In one embodiment, the phase interpolator has a granularity level of, for example, equally spaced steps for each clock phase (e.g., 2, 4, 6, 8, 10, 12, etc.), and the phase interpolator can set the rising clock edge and / or falling clock edge at any of these steps, for example, as further discussed below with reference to FIG. 13 .
[0090] Clock circuitry 708 (e.g., including a delay-locked loop (DLL) circuit) can be employed at the receiver circuitry 704 of the receiver die to properly align source-synchronous clock edges for high-performance timing (e.g., to enable efficient high-speed signal transmission). A DLL circuit can be a negative-delay gate placed in the clock path of a digital circuit. In one embodiment, clock circuitry 708 is a component of receiver circuitry 704. Local and / or dedicated clock circuitry (e.g., clock circuitry 710) (e.g., a phase-locked loop (PLL) circuit) (e.g., within an I / O PHY) can be used to achieve higher I / O bandwidth by filtering (e.g., grid) barrier clock jitter components. A PLL circuit can be a control circuit that generates an output signal whose phase is related to the phase of an input signal. While different types of PLL circuits exist, one example is a circuit with a variable frequency oscillator and a phase detector in a feedback loop, e.g., where the oscillator generates a periodic signal, the phase detector compares the phase of the signal to the phase of an input periodic signal, and adjusts the oscillator to maintain phase matching. The PLL can be an all-digital PLL (ADPLL). In one embodiment, the DLL circuit uses a variable phase (eg, delay) block and the PLL circuit uses a variable frequency block. The clock circuit 708 may include a control register 709, eg, for storing clock phase placement settings, eg, to cause the clock circuit 708 to apply those settings.
[0091] To maintain high power efficiency of transmitter circuitry and / or receiver circuitry (e.g., I / O PHY), techniques such as low-swing signaling, clock gating, and aggregating source-synchronous clock power across multiple (e.g., a large number) of serviced data lanes may be employed. For example, one forwarded source-synchronous clock may be used for each of 2, 3, 4, 5, 6, 7, 8, 16, 32, 64, 128, 256, etc. data lanes, or any subset thereof. Data lane 716B is merely an example, and multiple lanes may be utilized. In some embodiments, the clock phase delay generator 708A (eg, a DLL circuit) generates a locked (eg, non-clocked) timing for a clock rate of an operating frequency (eg, as Figure 16 ) (e.g., 90 degrees or 180 degrees of clock phase lock, e.g., as shown in Figure 8 and Figure 9 ). In some embodiments, the phase interpolator circuit 708B subdivides those clock signals into a finer granularity. In some embodiments, the clock circuit 708 utilizes predetermined clock phase placement data (e.g., before the current data transmission), for example, the clock phase delay generator 708A (e.g., a DLL circuit) and / or the phase interpolator circuit 708B both utilize predetermined clock phase placement data. In one embodiment, the clock phase delay generator 708A is a clock phase controller or a clock phase adjuster. In one embodiment, the clock phase delay generator 708A maintains a certain phase relationship of the clock arriving at the receiver (e.g., a sampler) (e.g., of the second die) relative to one or more input clocks coming in from the transmitter (e.g., of the first die). In some embodiments, the clock phase delay generator 708A generates the clock phase delay, and the phase interpolator circuit 708B is used to further subdivide those clock signals into a finer granularity. In one embodiment, the clock phase delay generator 708A looks up and utilizes a lock code for a particular clock rate and / or operating frequency, and / or the phase interpolator circuit 708B looks up and utilizes a buffer setting of the phase interpolator for a particular clock rate and / or operating frequency. For example, the lock code (e.g., of the DLL) can be changed for each frequency and / or each process, voltage, and / or temperature point (e.g., multiple points), and the phase interpolator circuit can perform (e.g., finer granularity) clock (e.g., edge) placement within the (e.g., DLL) lock code. Once the (e.g., predetermined) clock phase placement for the operating frequency and clock rate is looked up and updated into the circuit system (e.g., clock circuit 708), data can be received by the receiver circuit, for example, output to data buffer 735 and / or data buffer 736 (e.g., as Figure 21 In one embodiment, the first die includes one or more transmitter circuits (e.g., Figure 4 The transmitter circuit 402 or Figure 7 702 ), and the second die includes one or more receiver circuits (e.g., Figure 4 The receiver circuit 404 or Figure 7 Additionally or alternatively, the second die may include one or more transmitter circuits (e.g., Figure 4 The transmitter circuit 402 or Figure 7 702), and the first die may include one or more receiver circuits (e.g., Figure 4 The receiver circuit 404 or Figure 7 Receiver circuit 704), for example, to allow bidirectional communication between the dies.
[0092] Figure 8 The diagram illustrates a data timing diagram 801 and a clock timing diagram 802 for a first clock rate according to an embodiment of the present disclosure. In the depicted embodiment, the clock timing diagram 801 illustrates a clock signal (e.g., Figure 16 The data timing diagram 801 illustrates that data at the 1x clock rate can be read in at each falling edge of the clock (e.g., using Figure 7 As discussed herein, clock edges may be placed using a predetermined clock phase placement (eg, relative to data timing).
[0093] Figure 9 The diagram illustrates a data timing diagram 901 and a clock timing diagram 902 for a second clock rate according to an embodiment of the present disclosure. In the depicted embodiment, the clock timing diagram 901 illustrates a clock signal (e.g., Figure 16 The data timing diagram 901 illustrates that data at the 2x clock rate can be read in at each of the rising and falling edges of the clock (e.g., using Figure 7 As discussed herein, clock edges may be placed using a predetermined clock phase placement (eg, relative to data timing).
[0094] In one embodiment, I / O PHY circuitry (e.g., transmitter circuitry on one die and receiver circuitry on one or more other dies) is capable of (e.g., rapidly) changing between different clock rates (e.g., data rates) (e.g., 1x, 2x, 4x, etc.) and / or clock frequency rate changes, e.g., to support interconnects employed in a grid of dies. In certain embodiments, e.g., at initial boot time, one or more clock circuits (e.g., delay-locked loops (DLLs) and phase interpolators (PIs)) used for (e.g., receiver) clock edge alignment are calibrated for multiple (e.g., all) possible clock rates (e.g., data rates) and / or frequencies. In embodiments employing digitally controlled DLLs+PIs, calibration information for each of the clock rate (e.g., data rate) and operating frequency configurations is stored (e.g., in a memory array, e.g., in the clock circuitry) and is invoked when a circuit (e.g., a die) initiates a clock rate (e.g., data rate) and / or frequency change (e.g., of an interconnect connecting two or more dies). This can also be implemented for analog-controlled DLL+PI circuits, for example, by using an analog-to-digital (A / D) converter to convert the analog bias points into digital information for storage in a memory array, and then using a digital-to-analog (D / A) converter to convert them back to analog bias points when updating the operating point. These recalled clock (e.g., DLL+PI) calibration settings can be used to override the current clock (e.g., DLL+PI) calibration settings to allow a fast clock (e.g., DLL+PI) to lock and / or calibrate to the new settings and / or operating point. Thus, certain embodiments herein allow for fast transitions between different clock rates (e.g., data rates) and / or frequencies.
[0095] Certain embodiments herein provide novel circuit systems and algorithms to allow for fast and dynamic I / O clock rate (e.g., data rate) and / or frequency changes on the fly. In one embodiment, I / O timing (e.g., clock rate and / or operating frequency) between dies is facilitated by tuned clock phases (e.g., by a combination of DLL auto-tracking circuitry and training PI scans). In one embodiment, training occurs all at one time (e.g., one training session) (e.g., at manufacturing time, before the processor is utilized by the end user). The I / O clock architecture can be source synchronous, e.g., a forwarded clock that is tuned to a specific phase relationship relative to one or more data paths to maximize I / O timing margins. Figure 4 and Figure 7 The figure shows an example of an advanced clock architecture. Figure 5 、 Figure 6 、 Figure 8 and Figure 9 The diagram depicts the relationship between the data eye (e.g., Figure 5 、 Figure 6 、 Figure 8 and Figure 9 1 and 2. In some embodiments, fine-grained control of clock gating placement allows for maximum performance. Some embodiments achieve this through a combination of DLL+PI for small phase step granularity (e.g., 1 or about 1 picosecond (ps) increments). Figure 13 (discussed further below) shows example circuit architecture details of the digital delay line within the DLL and the digital PI. The output of the DLL+PI can be one clock (e.g., clocked using two clock edges), or two outputs (e.g., clocked using one clock edge of each), or four outputs (e.g., in the case of 4x clock rate) (e.g., using one clock edge of each clock or, alternatively, sending out 2 clocks and using two clock edges of each clock to clock all 4 data bits per cycle). Note that Figure 5 、 Figure 6 、 Figure 8 and Figure 9 A single clock output is shown (e.g., clocked using one clock edge for 1x clock rate or two edges for 2x clock rate), but FIG13 shows two outputs to illustrate that the circuit and method can also be used for 2x (2X) clocking, e.g., by clocking using only one clock edge per clock cycle. In some embodiments, the tuned clock phase will be unique for each frequency bin and clock rate at that frequency bin (e.g., and unique for each hardware instantiation within a die and / or from die to die).
[0096] Figure 10 The figure shows a hardware processor 1000 according to an embodiment of the present disclosure. For clarity, the mesh interconnect is not shown in each tube core, but the mesh interconnect can be used, for example, as shown in FIG. Figure 1 、 Figure 2A or Figure 2B As shown in . Figure 10 The diagram illustrates a three-dimensional stacked architecture. Multiple dies can extend in any single direction, with one or more electrical interconnects between each die. In the depicted embodiment, die 1002 and die 1004 extend in a first single plane, and die 1006 and die 1008 extend in a second, different single plane that is laterally spaced apart from the first single plane. The dies can be attached to another substrate, such as a mounting substrate (not depicted).
[0097] In one embodiment, a multi-die architecture is implemented using a silicon interposer as a physical manufacturing technology. In this implementation, the metal lines used to implement the bridge between two or more dies can be implemented in a different die (e.g., silicon), which forms the basis of all other dies. The base die can have through-silicon vias (TSVs) for delivering power to the die and / or routing I / O signals out to the board / external connector. Alternatively, the base die can have no TSVs, and power delivery and I / O breakout can be provided by some form of peripheral wire bonding.
[0098] Certain embodiments herein provide for multiple physically separate discrete dies to be electrically connected together via an electrical interconnect to form a larger and more capable processor. Certain embodiments herein provide for a single shared cache coherence domain on the interconnect to form a monolithic cache domain across the entire processor. Certain embodiments herein include communication with native protocols for data transfer within each die and do not require the overhead of packetizing or serializing data transmitted or received over the electrical interconnect between the dies. Certain embodiments herein allow for transfers between dies according to single or multiple simultaneous transaction protocols.
[0099] Certain embodiments herein allow for multiple dies to have relative clock alignment uncertainties, different power sources, different die manufacturing process deviations, and different die temperatures. Certain embodiments herein allow for one die to run at a different frequency than one or more other dies of the hardware processor. Certain embodiments herein allow for interconnects to have separable independent power, clock, and / or reset domains to aid in yield recovery, for example, by disabling rows and / or columns of a mesh interconnect. In certain embodiments, the electrical interconnects allow for (e.g., very large) cross-bandwidth, but also have minimal latency and power impact. Certain embodiments herein provide for mesh loop designs, for example, to tolerate die-to-die variations.
[0100] Certain embodiments herein add entries to a lookup table (LUT) (e.g., within a transceiver) to indicate whether data (e.g., a cache line) is to cross a physical die boundary in order to traverse the interconnect between two dies. Certain transmission protocols herein enable seamless (e.g., high-speed) interconnects between multiple dies and / or across die boundaries. As an alternative to using those protocols for die-to-die connections, certain embodiments herein may utilize other solutions, such as utilizing an interposer. Certain interconnects herein include fabric arbitration block circuitry (e.g., within a transceiver) to accommodate uncertainty in the state of transaction destination resources without forcing the source to delay potential indications, and to accommodate transaction merging into open transaction routing slots in the remote die fabric. In certain embodiments, the electrical interconnect fabric arbitration block circuitry (e.g., a controller) is located only in one of the receiver circuitry or the transmitter circuitry. Certain interconnects herein include post-silicon tunable buffers (e.g., transparent queues (TQs)), for example, to support high bandwidth and low latency to facilitate die crossing under clock alignment uncertainty, varying power sources, varying die manufacturing process variations, and / or varying die temperatures. In some embodiments, the electrical interconnect buffer can have no latency impact if both domains operate at the same frequency and manage clock uncertainty, despite the dies being at different power sources, different die manufacturing process variations, and different die temperatures. In some embodiments, the electrical interconnect buffer is located only at one of the receiver circuitry or the transmitter circuitry. In some embodiments, the interconnect buffer is located at both the transmitter and receiver circuitry.
[0101] Figure 11The figure illustrates a hardware processor 1100 according to an embodiment of the present disclosure. In the depicted embodiment, die 1102 and die 1104 are smaller than die 1106, die 1108, die 1110, and die 1112. Each of the depicted dies is coupled to an adjacent die via (e.g., inter-die) interconnects (INT). Die 1102 is depicted as having two discrete interconnects with die 1106, e.g., interconnects including one or more instances of the receiver circuit(s) and / or one or more instances of the transmitter circuit(s) disclosed herein. Die 1104 is depicted as having a different number (e.g., three) of discrete interconnects with die 1108. Die 1106 is depicted as having four discrete interconnects with die 1108. Die 1110 is depicted as having a different number (e.g., three) of discrete interconnects with die 1112. The intersections of the mesh interconnects of the dies (e.g., intersection 1114 or intersection 1116 of die 1106) can be access points into the mesh interconnects by circuit components. In one embodiment, multiple (e.g., any) mesh configurations having different sizes on their respective dies are coupled together by certain embodiments herein. In one embodiment, a die with a mesh interconnect is coupled to a die without a mesh interconnect, e.g., die 1118 at Figure 11 1106. Although a mesh interconnect is discussed in certain embodiments, other interconnect topologies (e.g., ring, star, tree, fully connected mesh, partially connected mesh, etc.) may be utilized.
[0102] Figure 12 The diagram illustrates a hardware processor 1200 in accordance with an embodiment of the present disclosure. In the depicted embodiment, die 1202 and die 1204 (e.g., having the same size) are smaller than die 1206, die 1208, die 1210, and die 1212. Die 1206 is depicted as including a different mesh interconnect than die 1208 (e.g., having a different number of crosspoints (e.g., crosspoint 1214) and / or transceivers (e.g., transceiver 1216)). Figure 12 The figures illustrate that in some embodiments some of the plurality of dies may be different (eg, in one embodiment, they are not symmetrical). Figure 12 The figures illustrate that in certain embodiments, a mesh interconnect on a die may be different from another mesh interconnect on a different die (eg, in one embodiment, they are asymmetric).
[0103] Certain embodiments herein provide consistent resources and grid transactions. Certain embodiments herein provide a master die controller to discover resource conditions across all die in order to build resource capabilities, resource address tables, and / or routing performance deviation tables. Certain embodiments of the master controller walk through the expected possible resources and perform subtractions, for example, by reading remote fuses or registers and based on a successful handshake. Certain embodiments of the master controller have a pre-programmed mapping set to configure resource tables (e.g., credits), a grid lookup table (LUT), address translation services (e.g., system address mapping), etc., to allow grid traversal across die. The selected pre-programmed mapping can be based on the identified resources.
[0104] Certain embodiments of electrical interconnects (e.g., and / or (one or more) transceiver circuits) between multiple die provide very high bandwidth that matches the bandwidth of an integrated (e.g., mesh) interconnect on a die. Certain embodiments of electrical interconnects (e.g., and / or (one or more) transceiver circuits) between multiple die provide (e.g., very) low latency, e.g., latency that matches or substantially matches the latency of an integrated interconnect on a die. Certain interconnects (e.g., and / or (one or more) transceiver circuits) herein include communications with native protocols for data transmission within each die and / or do not require overhead for packetizing or serializing data transmitted or received via the electrical interconnects between die (e.g., thereby minimizing the latency impact of the interconnect). Certain interconnects (e.g., and / or (one or more) transceiver circuits) herein include bandwidth reduction for communications without error protection as a means to increase data transmission efficiency and reduce latency. Certain interconnects herein (e.g., and / or transceiver circuit(s)) include on-the-fly transfer rate transitions (e.g., to match on-die communication bus frequency changes) with minimal (e.g., single-digit) clock cycles to update and transition timing synchronization of the electrical interconnects.
[0105] Certain interconnects (e.g., and / or transceiver circuit(s) herein) provide reduced pin counts while allowing full cross-sectional bandwidth (e.g., clock rate), such as using ¼ pins at 4x the data rate compared to the data frequency within the die, or using ½ pins at 2x the data rate compared to the data frequency within the die. Certain interconnects (e.g., and / or transceiver circuit(s) herein) provide reduced pin counts while allowing selectable bandwidth, such as 2x bandwidth at 4x the data rate compared to the data frequency within the die, or 1x bandwidth at 2x the data rate compared to the data frequency within the die. Certain interconnects (e.g., and / or transceiver circuit(s) herein) include dynamic and rapid transitions between a first (e.g., 1x) bandwidth and a second, different (e.g., 2x) bandwidth as two modes that conditionally provide an optimal choice of bandwidth performance benefits versus power savings benefits, a reduced penalty in delay caused by additional clock crossings into a low-jitter clock domain, and / or a reduction in the error rate that high-performance transmissions may have. Certain interconnects (e.g., and / or transceiver circuit(s) herein) provide for dynamic and rapid transitions between a first (e.g., 1x) bandwidth and a second, different (e.g., higher or lower) (e.g., 2x) bandwidth mode. Certain interconnects (e.g., and / or transceiver circuit(s) herein) include traffic flow control circuitry for temporarily pausing traffic during transitions (e.g., when transitioning between clock rates (e.g., 1x, 2x, 4x, etc.) and / or when transitioning between different operating frequencies (e.g., frequency rates).
[0106] Certain interconnects herein (e.g., and / or transceiver circuit(s)) provide for individual and independent tuning of receiver, transmitter, and / or clock circuitry for each bandwidth (e.g., clock rate) and frequency mode, e.g., at each instance and on each die, to compensate for intra-die and die-to-die process variations, as well as temporal temperature and voltage supply variations. Certain interconnects herein (e.g., and / or transceiver circuit(s)) include communication error detection mechanisms (e.g., parity checking, etc.) that allow for appropriate handling (e.g., rebooting, etc.) at the processor level.
[0107] Certain embodiments herein provide an electrical interconnect (e.g., and / or transceiver circuit(s)) with facilities for boot-time multi-point characterization sweeps of transmitter and receiver circuit parameters across multiple variables, and with storage for fast parameter lookups during runtime changes (e.g., changes in clock frequency, voltage level, or clock rate (e.g., 1x, 2x, 4x, etc.)). Certain embodiments herein provide an electrical interconnect (e.g., and / or transceiver circuit(s)) that provides periodic refresh of stored transmitter and receiver circuit parameter re-characterizations to recapture changing environmental and circuit conditions. Certain embodiments herein provide fast processor clock, power, and / or data rate transitions during critical runtime operations, and apply low-run multi-point characterization and parameter logging, for example, only at boot time or during runtime periods that are not sensitive to processor performance. Certain embodiments of the electrical interconnects (e.g., and / or transceiver circuit(s)) herein provide for optimizing explicit state updates (e.g., Rx DLL locked, Tx PLL locked, Tx duty cycle corrector (DCC) locked, etc.) and / or reducing latency from assumption timers for die-to-die swaps. Certain embodiments of the electrical interconnects (e.g., and / or transceiver circuit(s)) herein provide for autonomous management within the interconnect circuitry after multi-point scan characterization, e.g., without requiring management from firmware, BIOS, and / or drivers. Apparatus and method for scene-based compression
[0108] In current SoC implementations, various IP blocks communicate via an interconnect fabric, which can include on-die fabric links and, in the case of multi-chip packages, die-to-die (D2D) links. A major challenge with D2D implementations is the limited on-die area, or coastline, required for protocol routing. Consequently, the total link bandwidth can be lower than desired.
[0109] One way to address the problem of insufficient on-die area or coastline is to use a relatively narrow, high-speed interface that transmits inter-chip messages in a packetized format. For example, such packetization is defined in protocols such as PCIe, Compute Express Link (CXL), and Ultrapath Interconnect (UXI). The same set of physical wires of a D2D link can be used for all protocol messages, including requests, responses, and data transfers. The packetization protocol defines how to arrange the original protocol messages into raw data path blocks called "flits."
[0110] Figure 13A The figure shows an example of a first die / fabric 1395 including a first plurality of fabric endpoints or routers 1391A-1391E, and a second die / fabric 1396 including a second plurality of fabric endpoints or routers 1392A-1392D. Fabric endpoints / routers 1391A-1391E are coupled via intra-fabric / intra-die links 1395, and fabric endpoints / routers 1392A-1392D are coupled via intra-fabric / intra-die links 1396. Inter-fabric / die-to-die (D2D) links 1372 couple fabric endpoint / router 1391A of the first die / fabric 1395 to fabric endpoint / router 1392A of the second die / fabric 1396. However, due to the issues described above, inter-fabric D2D links 1372 may not be able to support the desired bandwidth between die / fabric 1395 and die / fabric 1396.
[0111] To address the bandwidth limitations of these inter-fabric / D2D links 1372, embodiments of the present invention include message compression and decompression logic 1378 to compress messages prior to transmission and decompress messages upon reception, thereby allowing a greater number of messages to be communicated for a given physical bandwidth.
[0112] Figure 13B The diagram shows a specific example of a processor die 1301 including a D2D interface 1370 coupled to a D2D interface 1371 of a peripheral control device 1302 via a D2D link 1372. In this example, the processor die 1301 includes multiple cores 1310-1311 and a home agent 1352 that couples the cores 1310-1311 to an on-die fabric 1309 for accessing various on-die components, including a cache subsystem 1315, a memory controller 1350 (to access system memory 1355), a display engine 1342, graphics processing circuitry 1330, tensor / AI circuitry 1334 (e.g., for performing matrix operations), and the D2D interface 1370 (e.g., for conducting inter-die transactions over the D2D link 1372).
[0113] The peripheral control die 1302 includes a PCIe interface 1360 , audio circuitry 1362 , a boot / security IP block 1364 , a memory controller 1366 (for coupling to a storage device 1370 ), a network interface 1368 , and a USB interface 1369 .
[0114] While the specific arrangement of processor die 1301 and peripheral control die 1302 is illustrated as an example, the underlying principles of the invention are not limited to this specific configuration. Rather, embodiments of the invention may be implemented in any multi-die or multi-fabric data processing device in which inter-die or inter-fabric links are used.
[0115] Each of the D2D interfaces 1370-1371 includes compression / decompression circuitry 1378A-1378B for implementing message compression as described herein. In one embodiment, each compression / decompression circuitry 1378A-1378B includes at least one message control register, referred to as a message match control register (MMCR), for storing constant message fields. In some embodiments, each MMCR is statically programmed with constant message fields, for example, based on a pre-runtime analysis of expected traffic. Alternatively or additionally, in some embodiments, message fields transmitted / received multiple times during one or more transactions and / or over a period of time can be identified at runtime and the MMCR updated accordingly.
[0116] Regardless of how the MMCR is programmed, a compressed message format is defined in which variable message fields are combined with indications of constant message fields stored in the MMCR. For example, a constant message field within a message can be replaced with a bit field that identifies the MMCR in which the corresponding constant message field is stored.
[0117] In operation, when a transmitter such as D2D interface 1370 detects a match between a message field of a message to be transmitted and a message field in the MMCR, the corresponding compression / decompression circuitry 1378A compresses the message by replacing the message field with an indication of the corresponding location in the MMCR (i.e., the location where a copy of the message field is stored). The compression / decompression circuitry 1378B of the receiving D2D interface 1371 then uses the indication to identify the message field in its corresponding MMCR to reconstruct the packet. The end result is that each link can support the provision of increased bandwidth that can be provided over a given link, which provides value to customers because it enables a wider range of uses, for example, higher I / O bandwidth makes disk scenarios (such as file copies) faster.
[0118] exist Figure 13BIn the specific example shown in , the D2D interfaces 1370-1371 are used for different types of transactions, including IO read requests (memory read requests from various IO devices connected through the peripheral control die 1302) and responses (e.g., data from memory 1355 returned from the processor die 1301 to the peripheral control die 1302). In addition, IO write requests including data to be written to memory 1355 can originate from the peripheral control die 1302 and require a completion response from the processor die 1301 to the peripheral control die 1302. Address translation requests can also be generated to service IO read and write requests. As another example, memory-mapped IO (MMIO) read and write requests generated by the processor die 1301 are directed to devices connected through the peripheral control die 1302.
[0119] Figure 14 The middle figure illustrates one embodiment of a method for performing compression (e.g., via a transmitting D2D interface). At 1401, a new message arrives for transmission, and at 1402, message fields are compared with fields stored in one or more MMCRs. In this embodiment, the transmitting D2D interface and the receiving D2D interface maintain a synchronized set of MMCR registers by updating their respective MMCR registers in response to the same detected messaging event.
[0120] At 1403, if the transmitter D2D interface detects a match between the message field and any MMCR, the message field is replaced with the corresponding MMCR identifier (ID), which identifies the specific MMCR and / or a specific position within the MMCR. The transmitter D2D interface may also set a bit to indicate that at least one message field is compressed. If there is no match at 1403, the message is transmitted uncompressed.
[0121] Figure 15 One embodiment of a method for performing decompression (e.g., by a receiving D2D interface) is shown in FIG. At 1501, a message arrives at the receiving D2D interface, which determines at 1502 (e.g., by reading compression bits) whether the message is compressed. If so, at 1504, the receiving D2D interface uses fields of the compressed message and fields from the MMCR(s) identified by the MMCR(s) ID(s) to reconstruct the original message. If not, at 1503, the uncompressed message is delivered to its destination.
[0122] Therefore, scenario-based compression reduces the amount of data transmitted by sending only parts of the message in combination with the MMCR ID(s) to identify certain fields stored in the MMCR register of the receiving D2D interface.
[0123] Figure 16 The figure illustrates a set of D2D operations for one embodiment of read and write transactions. A read operation initiated by peripheral control die (PCD) 1302 from IO system fabric (IOSF) 1601 is routed across D2D link 1372 via D2D interfaces 1371-1372 to home agent 1352. In response, home agent 1352 generates a forward read command to memory controller 1350, specifying the memory address from which data is to be read. Memory controller 1350 performs a read from system memory (not shown) and responds to PCD 1302 with two messages (MemData), each of which includes a header and a block of 32B of data read from memory.
[0124] The write initiated by PCD 1302 from IOSF 1602 travels across D2D link 1372 via D2D interfaces 1371-1372 to reach home agent 1352, which responds with an acknowledgment (GO-E). PCD 1302 then generates two messages (MemWr), each including a header and a 32-byte block of data to be written to memory. In response, home agent 1352 generates a write operation (MemWr) to memory controller 1350. Memory controller 1350 performs the write to system memory and responds with a completion message, which home agent 1352 forwards to PCD 1302.
[0125] In some embodiments, each D2D interface subdivides the data to be transmitted into data transmission units called "slots." For example, a 16-byte slot can be used to transmit a quarter of a cache line (assuming a 64-byte cache line) or as a "header slot," which includes a combination of request, response, and data header information used in a transaction's messages. Some embodiments of compression / decompression circuitry 1378A-1378B implement scenario-based compression as described herein to reduce the number of required header slots, which are typically duplicated across messages in a given transaction.
[0126] Without any compression or optimization, Figure 16The read and write transactions shown in the figure will require 7 slots (4 data slots in each direction + 3 header slots in the upward direction from the UP direction (i.e., from the PCD to the processor die). The effective link bandwidth (the portion of the bandwidth used for data transfer) is 4 / 7 of the total link bandwidth. Therefore, if the total link bandwidth is 16GB / s / direction, the effective bandwidth is 9.14GB / s / direction.
[0127] In contrast, using an embodiment of context-based compression as described herein, 6 slots are required (4 data slots in each direction + 2 header slots in the UP direction (i.e., from the PCD to the processor die). The effective link bandwidth is 4 / 6 of the total link bandwidth, or (4 / 6) * (16 GB / s / direction) = 10.67 GB / s / direction, a 17% overall bandwidth improvement.
[0128] In some embodiments, each type of transaction can be optimized using the corresponding 32b MMCR register in each die or structure. By way of example and not limitation, to achieve Figure 16 For a 17% bandwidth increase, six 32-bit MMCR registers can be used, one for each D2D message: RdCurr (request), MemData (data header), SpecFlOwn (request), GO-E (response), MemWr (data header), and Cmp (response).
[0129] In some embodiments, MMCR register lookups are performed in parallel with other bandwidth optimization and compression techniques. Depending on the implementation, scenario-based compression adds no latency or may add up to two cycles (one for compression in the transmitter D2D interface and one for decompression in the receiver D2D interface).
[0130] These embodiments can also provide significant power reductions. For example, in a D2D link that can transition between a higher power active data transmission state and a lower power active idle state, the number of active cycles drops from 13 to 11, which is a 15% reduction in active power.
[0131] Figure 17 Compression circuitry 1778 is illustrated for performing message compression in accordance with an embodiment of the present invention. The illustrated compression circuitry 1778 may be included in compression / decompression circuitry 1378A-1378B to perform message compression as described above.
[0132] The illustrated embodiment includes four MMCR registers 1700-1703, each of which is uniquely identified with a different MMCR identifier (ID). The constant field 1717 extracted from the original message 1720 is compared with the field in the MMCR registers 1700-1703. If a match is detected, the corresponding MMCR ID 1710 and variable field 1715 are combined and inserted into the message, replacing the corresponding constant field 1717. In some embodiments, the data field position can also be added to the message to identify the position within the MMCR register storing the constant field. As described above, the "constant" message field is a message field (e.g., such as header information associated with a transaction) used in multiple messages. If the constant field is not identified in the MMCR, the original message 1720 is transmitted without compression. The MMCR 1700-1703 can be updated with one or more constant fields 1717 so that it can be used to compress subsequent messages.
[0133] Figure 18 Illustrated is decompression circuitry 1878 for performing message decompression in accordance with an embodiment of the present invention. The illustrated decompression circuitry 1878 may be included in compression / decompression circuitry 1378A-1378B to perform message decompression as described above.
[0134] Decompression circuitry 1878 includes four MMCR registers 1800-1803 that correspond to the four MMCR registers 1700-1703 of compression circuitry 1778 and are associated with the same MMCR ID as MMCR registers 1700-1703. Decompression circuitry uses MMCR ID 1701 inserted by compression circuitry 1778 to identify corresponding constant fields 1717 in MMCRs 1800-1803 (e.g., via select MUX 1821 or similar logic). It reconstructs original message 1850 by inserting constant fields 1717 in place of corresponding MMCR ID 1701. As indicated, variable fields are copied to original message 1850 without modification.
[0135] To further illustrate the operation of certain embodiments of the present invention, examples are provided below using CXL.mem messages. However, it should be noted that the underlying principles of the present invention are not limited to any particular message type.
[0136] The left-hand side of the table below includes the original fields and their sizes as defined in the CXL specification. The right-hand side classifies the fields (or parts of the fields) into bits that must be sent over the link (and therefore included in the compressed message format) and bits that can be stored in the MMCR register (and therefore are not included in the compressed message). Table A
[0137] In one example, certain MMCRs may only be used by certain types of opcodes and attributes (e.g., Snp types, meta fields, and meta values in the CXL.mem example). When a specific opcode type is identified, the opcode and attributes are stored in the corresponding MMCR and not sent with the compressed message. The receiver then uses its corresponding MMCR to determine the opcode and attributes.
[0138] In some embodiments, a field indicating a rare condition (which may or may not occur) may be stored in the MMCR, such as a poison bit. In the example above, the poison value stored in the MMCR indicates that the poison condition has not occurred. In the relatively rare case that a poison condition occurs, the message will not be compressed (and will be sent in its original format).
[0139] Some fields may be partially stored in the MMCR, and only the unstored portion of the field is included in the compressed message. For example, as indicated in Table A, the source port ID (SPID) and destination port ID (DPID) (source / destination port ID) are 12 bits in the original message. If it is known that the message to be compressed is always sent from or to a small group of agents that share the same PortID prefix, the prefix value may be stored in the MMCR register. As an example and not limitation, memory access requests may always be sent to the home agent, and address translation requests may always be sent to the IOMMU agent.
[0140] Continuing with the same example, in some embodiments, a new message format is defined. This may include all variable fields of the original message, as well as a new field, MMCR ID, that identifies the MMCR used to compress the message. Table B provides an example of a CXL.mem message. Table B
[0141] And the supplementary MMCR including the "Constant" message field is defined as follows: Table C Apparatus and method for efficiently packing data for transmission over an interconnect structure
[0142] There are challenges associated with transmitting messages over the fabric used to connect IP blocks within a die and to connect two dies over a D2D interface. For example, there may not be enough area or coastline for all required protocol lines, and the total link bandwidth may be lower than required.
[0143] The solution to the area problem is to use a high-speed, narrow interface that transmits messages in a packetized format. This type of packetization is defined in protocols such as PCIe, Compute Express Link (CXL), and Ultrapath Interconnect (UXI). The same physical wires are used for all protocol messages (such as requests, responses, and data transfers). The packetization protocol defines how the original protocol messages are arranged in the original datapath blocks (also known as "flites").
[0144] Embodiments of the present invention improve bandwidth by efficiently arranging various messages within a flit. The size of each message can vary, especially when using different compression techniques. The number of pending messages varies over time, and the embodiments described herein provide an efficient way to arbitrate between the various messages and pack them within the flit with minimal area and latency.
[0145] In some embodiments, the packetization protocol defines a "slot" size. For example, in UXI, the slot size is 128 bits, which can be used for data (e.g., 1 / 4 of a cache line) or a combination of other message types. Some embodiments of the present invention define a constant "minislot" size, such that each slot includes N minislots. For example, with a 25-bit minislot size, a 128-bit slot includes 5 minislots (with the remaining 3 bits used as a "slot header").
[0146] In some implementations, each protocol message (including message variants) is assigned a predefined number of minislots into which the message can be efficiently adapted. By way of example and not limitation, a 70-bit request that can be compressed to 40 bits or 20 bits is defined as a 3-minislot message for an uncompressed format, a 2-minislot message for a first compressed format, and a single-minislot message for a second compressed format.
[0147] In some embodiments, the arbitrator selects among pending messages based on the number of minislots required to be packed into those messages and the number of minislots available for packing. The number of minislots required to pack into a message is called the required minislots (RMS), and the number of minislots available for packing is called the available minislots (AMS). In at least one implementation, the condition for packing messages is "RMS ≤ AMS." If this condition does not apply to all pending messages, the arbitrator selects a subset of pending messages that meet the condition. Because the packetization process is independent of slot size, it can be applied when the slots are small (but have a smaller "AMS"). Additionally, in some embodiments, packetization links more than one slot together for more efficient packing (e.g., based on the packetization process but with a larger "AMS").
[0148] As a brief overview, in some protocols such as CXL and UXI, a flits is the link layer transmission unit. For example, CXL uses 68 or 256-byte flits. A flits includes fields reserved for the physical and link layers of the protocol, while the rest of the flits are divided into slots used by the protocol layers. Figure 19 The figure shows one of the CXL flit formats. The fields: HDR, CRC, CRD, FEC are used by the link layer and the physical layer. The protocol layer has 15 slots, 13 of which are G slots (global, 16 bytes), one H slot (header, 14 bytes), and one HS slot (header small, 12 bytes). The G slot can be used for a combination of data (one-quarter of a 64B cache line) or "header" (non-data) messages. The H* slot can be used only for header information.
[0149] The CXL protocol defines a large number of predefined packetization arrangements for each slot type. These static, predefined packetization requirements limit the flexibility of slot formation and add a layer of complexity to the arbiter, which must consider various formats when determining which message to pack (among the currently pending messages). Furthermore, the hardware must implement several shifters to place each message in all possible slot formats. The arbiter is limited to the predefined packetization formats, even though slightly different optimizations would be beneficial.
[0150] Embodiments of the present invention provide a generic definition of a packetizer format that is independent of the length of different messages, can be easily extended to use new message types of different lengths, can be optimized for different use cases, different priorities between messages, temporary overload of specific message types, etc., effectively increases link bandwidth by reducing the number of unused bits, and operates independently of slot size (for example, in CXL-like packetization, the same algorithm can be used for G, H or HS slots), and allows chaining of multiple slots for better utilization.
[0151] Figure 20A The diagram illustrates an example implementation of an interface 2001 for generating a flit or other form of data transmission structure based on messages received from a fabric or IP block 2000, where the transmission structure is configured for transmission over a specific fabric / interconnect 2090 (e.g., an on-die fabric, a die-to-die (D2D) fabric / interconnect, etc.). Incoming messages from the fabric / IP block 2000 are stored in one or more pending message queues 2050. A packetizer 2060 efficiently packs messages or portions thereof into arbitration slots 2070 based on values indicating required minislots (RMS) 2061 and available minislots (AM) 2062. In one embodiment, minislot tracking circuitry 2065 continuously updates these values by tracking the number of minislots required to encode the messages currently stored in the pending message queue 2050 and the number of minislots available in the current arbitration slot 2070. Based on these values, the packetizer 2060 selects groups of messages and / or portions thereof to be packed into the current set of arbitration slots 2070 (i.e., where the RMS of the selected messages is less than or equal to the AM). In some embodiments, the packetizer 2060 is programmed with a specific minislot configuration 2063 (eg, defining minislot sizes), which may be statically or dynamically configurable based on the implementation (eg, depending on the characteristics of the data transmission unit slots).
[0152] Figure 20B The diagram illustrates an embodiment in which compression circuitry 1778 is applied to compress messages prior to storage in pending message queue 2050. Compression circuitry 1778 may perform various types of compression, including but not limited to those described above with respect to Figure 14-17 As further described below with respect to some embodiments, performing message compression can reduce the number of minislots consumed by messages, thereby reducing the RMS value 2061 and allowing the packetizer 2060 to insert a greater number of messages in each slot. Minislot and message definitions
[0153] Therefore, the basic building block of these embodiments is a minislot, and a minislot is a bit stream of constant length and defines the granularity of a protocol message (for example, the protocol message length can be defined as an integer number of minislots). In certain embodiments, the minislot length is selected so that the set of N minislots will be adapted to the higher-level transmission granularity with high efficiency / minimum waste, and so that an integer number of minislots with high efficiency / minimum waste can be used to define various protocol messages.
[0154] For example, in many protocols such as CXL or Universal Chiplet Interconnect Express (UCIe), the transmission granularity is a 128-bit "slot." For a 25-bit minislot length, a slot consists of 5 minislots (a total of 125 bits) + 3 bits for the slot header. Another option is to choose a minislot length of 32 bits and allow 4 minislots without a slot header.
[0155] Therefore, in these embodiments, each protocol message can be specified based on a set of minislots. Using the example of 25-bit minislots, a protocol message can include 1, 2, 3, 4, and 5 minislots, with message lengths of 25 bits, 50 bits, 75 bits, 100 bits, and 125 bits, respectively.
[0156] In some embodiments, the message definition also includes an "overhead" field, including a field defining the message length in minislots (to be used to determine where a message ends and a new message begins) and a field defining the message type or format (if more than one message type uses the same message length). Alternatively, the message length and format can be decoded in the same field. Alternatively or additionally, the protocol can be defined so that the length and format are not included in the minislot chain, but are included elsewhere (e.g., in the slot header). Message placement in the "head slot"
[0157] Figure 21 The diagram illustrates an example of the above message definitions using a 1x minislot message 2101, a 2x minislot message 2102, a 3x minislot message 2103, a 4x minislot message 2104, and a 5x minislot message 2105 that utilize 1, 2, 3, 4, or 5 minislots, respectively. In this particular example, each message includes a 3-bit LEN (length) field that indicates the number of minislots in the corresponding message and a 2-bit FMT (format) field that defines what protocol message is being transmitted. Thus, the "overhead" is 5 bits, so the actual message size is 20 bits, 45 bits, 70 bits, 95 bits, or 120 bits.
[0158] As another example, a particular protocol has a 110-bit uncompressed request message, which in some cases can be compressed to 90 bits, and in other cases can be compressed to 70 bits (e.g., via compression circuitry 1778). In this example, the uncompressed format may use a message of 5 minislots, the corresponding first compressed message will use a message size of 4 minislots, and the second (more highly compressed) message will use a message size of 3 minislots. This example illustrates how different compression techniques can directly translate into a reduction in the number of minislots required to transmit a message, and thus effectively increase link bandwidth.
[0159] Figure 22 The diagram illustrates how protocol messages constructed with 1-5 minislots are placed in the higher granularity transmission unit of a slot. A 128-bit slot with 3 HDR (header) bits and 5 25-bit minislots is illustrated as template 2200. Using this template provides a variety of ways to pack messages into slots, including (but not limited to): a 3-minislot message and a 2-minislot message 2201; a 2-minislot message, a 1-minislot message, and a 2-minislot message 2202; five 1-minislot messages 2203; a 1-minislot message combined with a 4-minislot message 2204; and a single 5-minislot message 2205, or any other combination of messages totaling up to 5 minislots in any order (e.g., {3,2}, {2,1,2}, {1,4}, etc.). Header slot link
[0160] Some embodiments of the present invention increase the effective link bandwidth by linking multiple slots, which does not require changing the message definitions. Figure 23 In the example, three messages need to be packaged, two of which are sized at 3 minislots and one sized at 4 minislots. Without chaining, since slots hold 5 minislots, three slots 2301-2303 are required to transmit the three messages. In contrast, with slot chaining, a dual-headed slot with a capacity of 10 minislots is formed, as indicated by template data structure 2300. Using this arrangement, the three messages can be packaged using two chained slots 2304, resulting in a 33% savings. Slot size independence
[0161] Another benefit of the embodiments described herein is that there is no need to define different formats for G slots and H slots. Therefore, there is no need to modify the operation of the packetizer 2060 based on the slot type.
[0162] As background, in protocols like CXL, not all slots are of equal length. The basic slot has 128 bits, but some slots include protocol overhead, such as link layer and data integrity information (such as CRC or parity bits). The protocol defines a G slot as a slot with no overhead (so it can use the entire 128 bits) or an H* slot that includes an overhead field (so fewer bits can be used for protocol messages). As an example, a 16x 128b slot flit includes 13 G slots (1-7, 9-14) with 16 bytes, 1 H slot (slot 0) with 14 bytes, and 1 HS slot (slot 8) with only 12 bytes.
[0163] Embodiments of the present invention are to G groove, H groove and HS groove and realize identical packing operation.Unique difference is the quantity that can be used for the microgrooves of packing, and grouper is configured to be used for operating with this quantity as parameter.When packing single header groove, the G groove has 5 microgrooves, and only the H groove has 4 microgrooves, and the HS groove has 3 microgrooves.Similarly, when linking two grooves, if these two grooves are G type grooves, then use 10 microgrooves.If the groove of two links is H* groove or G groove, then only have 9 or 8 microgrooves respectively to be used for packing.The quantity of available microgrooves (AM) is offered to grouper 2060 as input, but in addition, packing operation is not subject to the influence of groove type. Grouper process
[0164] 20, the packetizer 2060 selects among the pending messages in the pending message queue 2050 according to the minislot structure to add to the current arbitration slot 2070. One embodiment of the packetizer 2060 performs the following sequence of operations:
[0165] Operation 1: The number of required minislots (RMS) 2061 is determined, which includes a summary of the lengths of all pending messages, represented by the minislot tracking circuitry 2065 as a number of minislots.
[0166] Operation 2: The number of available minislots (AMS) 2062 is determined, which is represented by the minislot tracking circuitry 2065 as the number of minislots available for packing in the current arbitration slot and the size of each slot (e.g., in the above example, the slot has 5, 4, or 3 minislots). If the packetizer 2060 determines that the message is to be packed into two linked slots, there will be 6-10 minislots.
[0167] Operation 3: If all pending messages can fit into the available slots (ie, RMS ≤ AMS), then the grouper 2060 adds all messages to the slots.
[0168] Operation 4: Otherwise, a subset of pending messages that can be adapted is selected.For example, the grouper 2060 implements arbitration based on a set of priority rules, uses round-robin arbitration, and / or may consider other variables based on the implementation.
[0169] Figures 24A-24B The diagram illustrates example operations of different implementations of the packetizer 2060, which adds messages from the pending message queue 2050 to a 125-bit arbitration slot ( Figure 24A ) and two linked 125-bit arbitration slots ( Figure 24B). In these examples, pending message queue 2050 includes request queue 2450A, data header queue 2450B, and response queue 2450. Each queue can send two messages per packet. Request queue 2450A stores a 4-minislot message 2401 and a 5-minislot message 2402; data header queue 2450B stores a 3-minislot message 2403; and response queue 2450C stores a 1-minislot message 2404 and a 2-minislot message 2405. These metrics are tracked by minislot tracking circuitry 2065, which determines a required minislot (RMS) value of 15 (i.e., 4+5+3+1+2).
[0170] exist Figure 24A In , a single slot is packed, so the minislot tracking circuitry 2065 detects that available minislots = 5, and the packetizer 2060 selects a request message 2401 and a response message 2404 that effectively fit within the 5 minislots. Figure 24B In the packet, the second slot is available for packetization (e.g., via slot chaining), and the minislot tracking logic determines a new available minislot value of 10. Packetizer 2060 additionally selects a 3-minislot data header message 2403 and another response message that is 2 minislots, for a total of 10 minislots.
[0171] Figure 25 A method according to one embodiment is illustrated in The method can be implemented on various architectures described herein, but is not limited to any particular processor or system architecture.
[0172] At 2400, a message received from a first fabric or IP block is optionally compressed, and the (compressed) message is queued in a corresponding message queue at 2501. In some embodiments described herein, different message queues are configured to store different types of messages and / or different message components (e.g., a request message queue, a response message queue, a data header message queue, etc.).
[0173] The size of the message in minislots is determined at 2502, and the number of available minislots in one or more arbitration slots is determined at 2503. For example, with a minislot size of 25 bits, a 125-bit slot can be used to transmit a single message of size 5 minislots, a message of size 5 minislots, or any combination therebetween (see, e.g., Figure 22 and associated text).
[0174] At 2504, potential message packing options (e.g., different combinations of messages to be packed into the arbitration slot) are evaluated, and one or more messages to be packed into the transport fabric slot(s) are selected to minimize unused bits. As previously mentioned, some embodiments allow multiple slots to be chained to provide more efficient packing (see, e.g., Figure 23 and associated text).
[0175] At 2505, after the message has been packetized, the transport structure is transmitted over the second fabric / interconnect. Performance improvements
[0176] As mentioned above, one of the benefits of the embodiments described herein is to increase the effective link bandwidth by packing more messages in each slot.By combining the embodiments described herein with other compression techniques such as zero removal and message combining, bandwidth savings are further improved. Exemplary Core and System Architectures In-order and out-of-order core block diagram
[0177] Figure 26A is a block diagram illustrating both an exemplary in-order pipeline and an exemplary register renaming, out-of-order issue / execution pipeline according to embodiments of the present disclosure. Figure 26B is a block diagram illustrating both an exemplary embodiment of an in-order architecture core and an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor according to embodiments of the present disclosure. Figures 26A-26B The solid-line boxes in illustrate the in-order pipeline and in-order core, while the optional addition of dashed-line boxes illustrates the out-of-order issue / execution pipeline and core with register renaming. Considering the in-order aspect to be a subset of the out-of-order aspect, the out-of-order aspect will be described.
[0178] exist Figure 26A , the processor pipeline 2600 includes a fetch stage 2602, a length decode stage 2604, a decode stage 2606, an allocate stage 2608, a rename stage 2610, a dispatch (also known as dispatch or issue) stage 2612, a register read / memory read stage 2614, an execute stage 2616, a write back / memory write stage 2618, an exception handling stage 2622, and a commit stage 2624.
[0179] Figure 26BA processor core 2690 is shown, comprising a front end unit 2630 coupled to an execution engine unit 2650, and both the front end unit 2630 and the execution engine unit 2650 are coupled to a memory unit 2670. The core 2690 may be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As a further option, the core 2690 may be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, or the like.
[0180] The front end unit 2630 includes a branch prediction unit 2632 coupled to an instruction cache unit 2634, which is coupled to an instruction translation lookaside buffer (TLB) 2636, which is coupled to an instruction fetch unit 2638, which is coupled to a decode unit 2640. The decode unit 2640 (or decoder or decoder unit) can decode an instruction (e.g., a macroinstruction) and generate as output one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals decoded from, or otherwise reflecting, or derived from, the original instruction. The decode unit 2640 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLA), microcode read-only memories (ROM), and the like. In one embodiment, core 2690 includes a microcode ROM or other medium for storing microcode for certain macroinstructions (e.g., in decode unit 2640 or otherwise within front end unit 2630). Decode unit 2640 is coupled to rename / allocator unit 2652 in execution engine unit 2650.
[0181] The execution engine unit 2650 includes a rename / allocator unit 2652 coupled to a retirement unit 2654 and a set of one or more scheduler units 2656. The scheduler unit(s) 2656 represent any number of different schedulers, including reservation stations, central instruction windows, and the like. The scheduler unit(s) 2656 are coupled to the physical register file(s) 2658. Each of the physical register file(s) 2658 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating point, packed integers, packed floating point, vector integers, vector floating point, status (e.g., an instruction pointer, which is the address of the next instruction to be executed), and the like. In one embodiment, the physical register file(s) 2658 include a vector register unit, a write mask register unit, and a scalar register unit. These register units may provide architectural vector registers, vector mask registers, and general purpose registers. The physical register file(s) unit(s) 2658 are overlapped by the retirement unit 2654 to illustrate various ways in which register renaming and out-of-order execution can be implemented (e.g., using reorder buffer(s) and retirement register file(s); using future file(s), history buffer(s), and retirement register file(s); using register maps and register pools, etc.). The retirement unit 2654 and the physical register file(s) unit(s) 2658 are coupled to the execution cluster(s) 2660. The execution cluster(s) 2660 include a set of one or more execution units 2662 and a set of one or more memory access units 2664. The execution units 2662 can perform various operations (e.g., shifts, additions, subtractions, multiplications) and can operate on various data types (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some embodiments may include multiple execution units dedicated to a particular function or set of functions, other embodiments may include only one execution unit or multiple execution units that all perform all functions. The scheduler unit(s) 2656, physical register file(s) units 2658, and execution cluster(s) 2660 are shown as possibly multiple because certain embodiments create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline each with its own scheduler unit, physical register file(s) units, and / or execution cluster—and in the case of separate memory access pipelines, implementing certain embodiments in which only the execution cluster of that pipeline has memory access unit(s) 2664).It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution and the remaining pipelines may be in-order issue / execution.
[0182] A set of memory access units 2664 is coupled to a memory unit 2670, which includes a data TLB unit 2672, which is coupled to a data cache unit 2674, which is coupled to a level 2 (L2) cache unit 2676. In one exemplary embodiment, the memory access unit 2664 may include a load unit, a store address unit, and a store data unit, each of which is coupled to the data TLB unit 2672 in the memory unit 2670. The instruction cache unit 2634 is further coupled to a level 2 (L2) cache unit 2676 in the memory unit 2670. The L2 cache unit 2676 is coupled to one or more other levels of cache and ultimately to main memory.
[0183] As an example, the exemplary register renaming, out-of-order issue / execution core architecture may implement the pipeline 2600 as follows: 1) instruction fetch 2638 executes the fetch stage 2602 and the length decode stage 2604; 2) the decode unit 2640 executes the decode stage 2606; 3) the rename / allocator unit 2652 executes the allocate stage 2608 and the rename stage 2610; 4) (one or more) scheduler units 2656 execute the schedule stage 2612; 5) (one or more) physical register file units 2658 and memory units 2670 execute the register read / memory read stage 2614; the execution cluster 2660 executes the execute stage 2616; 6) the memory unit 2670 and (one or more) physical register file units 2658 execute the write back / memory write stage 2618; 7) each unit may be involved in the exception handling stage 2622; and 8) the retirement unit 2654 and (one or more) physical register file units 2658 execute the commit stage 2624.
[0184] Core 2690 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions that have been added with newer versions); the MIPS instruction set from MIPS Technologies, Inc. of Sunnyvale, California; the ARM instruction set from ARM Holdings, Inc. of Sunnyvale, California (with optional additional extensions such as NEON)), including the instruction(s) described herein. In one embodiment, core 2690 includes logic to support packed data instruction set extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.
[0185] It should be understood that a core may support multithreading (executing two or more sets of operations or threads in parallel) and that this multithreading may be accomplished in a variety of ways, including time-shared multithreading, simultaneous multithreading (where a single physical core provides a logical core for each of the threads that the physical core is simultaneously multithreading), or a combination thereof (e.g., time-shared fetching and decoding and thereafter, such as with Hyper-threading technology to synchronize multithreading).
[0186] Although register renaming is described in the context of out-of-order execution, it should be understood that register renaming can be used in an in-order architecture. Although the illustrated embodiment of the processor also includes separate instruction and data cache units 2634 / 2674 and a shared L2 cache unit 2676, alternative embodiments may have a single internal cache for both instructions and data, such as, for example, a first level (Level 1, L1) internal cache or multiple levels of internal cache. In some embodiments, the system may include a combination of internal caches and external caches external to the core and / or processor. Alternatively, all caches may be external to the core and / or processor. Specific exemplary in-order core architecture
[0187] Figures 27A-27B A block diagram illustrating a more specific exemplary in-order core architecture, which would be one logic block among several logic blocks in a chip (including other cores of the same and / or different types). Depending on the application, the logic block communicates with some fixed function logic, memory I / O interfaces, and other necessary I / O logic via a high-bandwidth interconnect network (e.g., a ring network).
[0188] Figure 27A 2704 and 2716. The block diagram of a single processor core and its connection to an on-die interconnect network 2702 and its local subset 2704 of a level 2 (L2) cache according to an embodiment of the present disclosure. In one embodiment, the instruction decode unit 2700 supports the x86 instruction set with the packed data instruction set extension. The L1 cache 2706 allows low-latency access to cache memory in the scalar and vector units. Although in one embodiment (to simplify the design), the scalar unit 2708 and the vector unit 2710 use separate register sets (scalar registers 2712 and vector registers 2714, respectively), and data transferred between these registers is written to memory and then read back from the level 1 (L1) cache 2706, alternative embodiments of the present disclosure may use a different approach (e.g., using a single register set or including a communication path that allows data to be transferred between the two register files without being written and read back).
[0189] The local subset 2704 of the L2 cache is part of the global L2 cache, which is divided into separate local subsets, one for each processor core. Each processor core has a direct access path to its own local subset 2704 of the L2 cache. Data read by a processor core is stored in its L2 cache subset 2074 and can be quickly accessed in parallel with other processor cores accessing their own local L2 cache subsets. Data written by a processor core is stored in its own L2 cache subset 2074 and flushed from other subsets when necessary. The ring network ensures the consistency of shared data. The ring network is bidirectional to allow agents such as processor cores, L2 caches, and other logic blocks to communicate with each other within the chip. Each ring data path is 1012 bits wide in each direction.
[0190] Figure 27B According to an embodiment of the present disclosure Figure 27A An expanded view of part of the processor core. Figure 27B Includes the L1 data cache 2706A portion of the L1 cache 2704, as well as more details about the vector unit 2710 and vector registers 2714. Specifically, the vector unit 2710 is a 16-wide vector processing unit (VPU) (see 16-wide ALU 2728) that executes one or more of integer, single-precision floating-point, and double-precision floating-point instructions. The VPU supports blending of register inputs via blend unit 2720, numerical conversion using value conversion units 2722A-2722B, and copying of memory inputs using copy unit 2724. Write mask register 2726 allows vector writes resulting from predicates.
[0191] Figure 28 is a block diagram of a processor 2800 that may have more than one core, may have an integrated memory controller, and may have integrated graphics, according to an embodiment of the present disclosure. Figure 28 The solid line box diagram in the figure shows a processor 2800 having a single core 2802A, a system agent 2810, a set 2816 of one or more bus controller units, while the optional addition of the dashed line box illustrates an alternative processor 2800 having multiple cores 2802A-2802N, a set 2814 of one or more integrated memory controller units in the system agent unit 2810, and dedicated logic 2808.
[0192] Thus, different implementations of processor 2800 may include: 1) a CPU where the specialized logic 2808 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores) and cores 2802A-2802N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of the two); 2) a coprocessor where cores 2802A-2802N are a large number of specialized cores intended primarily for graphics and / or scientific (throughput); and 3) a coprocessor where cores 2802A-2802N are a large number of general-purpose in-order cores. Thus, processor 2800 may be a general-purpose processor, a coprocessor, or a specialized processor, such as, for example, a network or communications processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput many-integrated-core (MIC) coprocessor (including 30 or more cores), an embedded processor, or the like. The processor may be implemented on one or more chips. Processor 2800 may be part of and / or implemented on one or more substrates using any of a variety of process technologies, such as, for example, BiCMOS, CMOS, or NMOS.
[0193] The memory hierarchy includes one or more levels of cache within the core, a set or one or more shared cache units 2806, and external memory (not shown) coupled to a set of integrated memory controller units 2814. The set of shared cache units 2806 may include one or more mid-level caches, such as a level 2 (L2), level 3 (L3), level 4 (L4), or other level of cache, a last level cache (LLC), and / or combinations thereof. While in one embodiment, a ring-based interconnect unit 2812 interconnects the integrated graphics logic 2808, the set of shared cache units 2806, and the system agent unit 2810 / integrated memory controller unit(s) 2814, alternative embodiments may use any number of well-known techniques to interconnect such units. In one embodiment, coherency is maintained between the one or more cache units 2806 and the cores 2802A-2802N.
[0194] In some embodiments, one or more cores 2802A-2802N may be multi-threaded. System agent 2810 includes components that coordinate and operate cores 2802A-2802N. System agent unit 2810 may include, for example, a power control unit (PCU) and a display unit. The PCU may include or may include the logic and components required to regulate the power state of cores 2802A-2802N and integrated graphics logic 2808. The display unit is used to drive one or more externally connected displays.
[0195] The cores 2802A-2802N may be homogeneous or heterogeneous in terms of architectural instruction sets; that is, two or more of the cores 2802A-2802N may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of the instruction set or a different instruction set. Exemplary Computer Architecture
[0196] Figure 29-Figure 31 is a block diagram of an exemplary computer architecture. Other system designs and configurations known in the art for laptops, desktop computers, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, a wide variety of systems or electronic devices that can include a processor and / or other execution logic as disclosed herein are generally suitable.
[0197] Now see Figure 29, shown is a block diagram of a system 2900 according to one embodiment of the present invention. System 2900 may include one or more processors 2910, 2915 coupled to a controller hub 2920. In one embodiment, controller hub 2920 includes a graphics memory controller hub (GMCH) 2990 and an input / output hub (IOH) 2950 (which may be on separate chips); GMCH 2990 includes memory and a graphics controller, to which memory 2940 and a coprocessor 2945 are coupled; and IOH 2950 couples input / output (I / O) devices 2960 to GMCH 2990. Alternatively, one or both of the memory and graphics controller are integrated within the processor (as described herein), and the memory 2940 and coprocessor 2945 are directly coupled to the processor 2910 and the controller hub 2920, which is in a single chip with the IOH 2950. The memory 2940 may include, for example, a cache coherence and / or interconnect management module 2940A for storing code that, when executed, causes the processor to perform any of the methods of the present disclosure.
[0198] The optional additional processor 2915 is Figure 29 Each processor 2910 , 2915 may include one or more of the processing cores described herein and may be a version of processor 4000 .
[0199] The memory 2940 may be, for example, dynamic random access memory (DRAM), phase change memory (PCM), or a combination of the two. For at least one embodiment, the controller hub 2920 communicates with the processor(s) 2910, 2915 via a multi-drop bus such as a frontside bus (FSB), a point-to-point interface such as a Quick Path Interconnect (QPI), or a similar connection 2995.
[0200] In one embodiment, the coprocessor 2945 is a special purpose processor such as, for example, a high throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, etc. In one embodiment, the controller hub 2920 may include an integrated graphics accelerator.
[0201] There may be various differences between the physical resources 2910 , 2915 in terms of a range of quality metrics including architectural, microarchitectural, thermal, power consumption characteristics, and the like.
[0202] In one embodiment, the processor 2910 executes instructions that control general types of data processing operations. Embedded within these instructions may be coprocessor instructions. The processor 2910 recognizes these coprocessor instructions as being of a type that should be executed by the attached coprocessor 2945. Accordingly, the processor 2910 issues these coprocessor instructions (or control signals representing coprocessor instructions) to the coprocessor 2945 over a coprocessor bus or other interconnect. The coprocessor(s) 2945 accept and execute the received coprocessor instructions.
[0203] Now see Figure 30 , shown is a block diagram of a first more specific exemplary system 3000 according to an embodiment of the present disclosure. Figure 30 As shown in FIG, multiprocessor system 3000 is a point-to-point interconnect system and includes a first processor 3070 and a second processor 3080 coupled via a point-to-point interconnect 3050. Each of processors 3070 and 3080 may be a version of processor 2800. In one embodiment of the present disclosure, processors 3070 and 3080 are processors 2910 and 2915, respectively, and coprocessor 3038 is coprocessor 2945. In another embodiment, processors 3070 and 3080 are processor 2910 and coprocessor 2945, respectively.
[0204] Processors 3070 and 3080 are shown as including integrated memory controller (IMC) units 3072 and 3082, respectively. Processor 3070 also includes point-to-point (PP) interfaces 3076 and 3078 as part of its bus controller unit; similarly, second processor 3080 includes PP interfaces 3086 and 3088. Processors 3070, 3080 can exchange information via PP interface 3050 using point-to-point (PP) interface circuits 3078, 3088. Figure 30 As shown, IMCs 3072 and 3082 couple the processors to respective memories, namely, memory 3032 and memory 3034, which may be portions of main memory locally attached to the respective processors.
[0205] The processors 3070, 3080 may each exchange information with a chipset 3090 via respective PP interfaces 3052, 3054 using point-to-point interface circuits 3076, 3094, 3086, 3098. The chipset 3090 may optionally exchange information with a coprocessor 3038 via a high-performance interface 3039. In one embodiment, the coprocessor 3038 is a special-purpose processor such as, for example, a high-throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, or the like.
[0206] A shared cache (not shown) may be included in either processor, or external to both processors but connected to the processors via the PP interconnect, such that if the processors are placed in a low power mode, local cache information for either or both processors may be stored in the shared cache.
[0207] Chipset 3090 may be coupled to first bus 3016 via interface 3096. In one embodiment, first bus 3016 may be a Peripheral Component Interconnect (PCI) bus or a bus such as PCI Express or another third generation I / O interconnect bus, but the scope of the present disclosure is not limited in this regard.
[0208] like Figure 30 As shown in , various I / O devices 3014 may be coupled to the first bus 3016, along with a bus bridge 3018 that couples the first bus 3016 to a second bus 3020. In one embodiment, one or more additional processors 3015, such as a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (such as, for example, a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array, or any other processor, are coupled to the first bus 3016. In one embodiment, the second bus 3020 may be a low pin count (LPC) bus. Various devices may be coupled to the second bus 3020, including, in one embodiment, a keyboard and / or mouse 3022, communication devices 3027, and a storage unit 3028, such as a disk drive or other mass storage device, which may include instructions / code and data 3030. Further, an audio I / O 3024 may be coupled to the second bus 3020. Note that other architectures are possible. For example, instead of Figure 30 Instead of a point-to-point architecture, the system can implement a multi-drop bus or other such architectures.
[0209] Now see Figure 31 , shown is a block diagram of a SoC 3100 according to an embodiment of the present disclosure. Figure 31Similar elements in the FIGURE 1 use similar reference numerals. In addition, the dashed boxes are optional features on more advanced SoCs. Figure 31 In the embodiment, interconnect unit(s) 3102 are coupled to: application processor 3110, which includes a set of one or more cores 3102A-3102N and shared cache unit(s) 3106; system agent unit 3110; bus controller unit(s) 3116; integrated memory controller unit(s) 3114; one or more coprocessor(s) 3120, which may include integrated graphics logic, image processor, audio processor, and video processor; static random access memory (SRAM) unit 3130; direct memory access (DMA) unit 3132; and display unit 3140 for coupling to one or more external displays. In one embodiment, coprocessor(s) 3120 include specialized processors such as, for example, network or communication processors, compression engines, GPGPUs, high-throughput MIC processors, embedded processors, and the like.
[0210] The embodiments (e.g., mechanisms) disclosed herein may be implemented in hardware, software, firmware, or a combination of such implementations. The embodiments of the present disclosure may be implemented as a computer program or program code executed on a programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0211] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0212] The program code can be implemented with a high-level process-oriented programming language or an object-oriented programming language to communicate with the processing system. If necessary, the program code can also be implemented with assembly language or machine language. In fact, the mechanism described herein is not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0213] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium that represent various logic within a processor, which, when read by a machine, causes the machine to fabricate logic for performing the techniques described herein. Such representations, known as "IP cores," may be stored on a tangible, machine-readable medium and supplied to various customers or manufacturing facilities to be loaded into fabrication machines that actually fabricate the logic or processor.
[0214] Such machine-readable storage media may include, but are not limited to, non-transitory tangible arrangements of articles of manufacture manufactured or formed by a machine or apparatus, including storage media such as: a hard disk; any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks); semiconductor devices (such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM); phase change memory (PCM); magnetic or optical cards; or any other type of medium suitable for storing electronic instructions.
[0215] Therefore, embodiments of the present disclosure also include non-transitory, tangible, machine-readable media containing instructions or design data, such as a hardware description language (HDL), that defines the structures, circuits, devices, processors, and / or system features described herein. Such embodiments may also be referred to as program products. Simulation (including binary translation, code deformation, etc.)
[0216] In some cases, an instruction converter can be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter can translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), deform, simulate, or otherwise convert an instruction into one or more other instructions to be processed by the core. The instruction converter can be implemented with software, hardware, firmware, or a combination thereof. The instruction converter can be on the processor, off the processor, or partially on the processor and partially off the processor.
[0217] Figure 32 1 is a block diagram illustrating a method for converting binary instructions in a source instruction set into binary instructions in a target instruction set using a software instruction converter according to an embodiment of the present disclosure. In the illustrated embodiment, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 32 3202 to generate x86 binary code 3206 that can be natively executed by a processor 3216 having at least one x86 instruction set core. The processor 3216 having at least one x86 instruction set core represents a processor that can execute programs that are natively executed by a processor 3216 having at least one x86 instruction set core by compatibly executing or otherwise processing the following: Any processor with substantially the same functionality as the processor: (1) a substantial portion of the instruction set of an x86 instruction set core, or (2) is targeted at a computer having at least one x86 instruction set core. processors with at least one x86 instruction set core. x86 compiler 3204 represents a compiler operable to generate x86 binary code 3206 (e.g., object code) that can be executed on a processor 3216 having at least one x86 instruction set core with or without additional linking processing. Similarly, Figure 32It is shown that an alternative instruction set compiler 3208 can be used to compile a program in a high-level language 3202 to generate alternative instruction set binary code 3210 that can be natively executed by a processor 3214 that does not have at least one x86 instruction set core (e.g., a processor having a core that executes the MIPS instruction set of MIPS Technologies, Inc. of Sunnyvale, California, and / or the ARM instruction set of ARM Holdings, Inc. of Sunnyvale, California). An instruction converter 3212 is used to convert the x86 binary code 3206 into code that can be natively executed by the processor 3214 that does not have an x86 instruction set core. This converted code is unlikely to be identical to the alternative instruction set binary code 3210 because an instruction converter that can do so is difficult to manufacture; however, the converted code will perform general operations and be composed of instructions from the alternative instruction set. Thus, the instruction converter 3212 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have an x86 instruction set processor or core to execute x86 binary code 3206 through simulation, emulation, or any other process.
[0218] Certain embodiments provide a coherent flow from individual dies in a wafer to a packaged modular die product. Additionally, these embodiments provide modularity and scalability for stitching together several modular dies (e.g., heterogeneous modular dies), and further provide interconnection between dies on an interconnection grid or fabric.
[0219] Different embodiments implement connections between these dies in different ways. For example, in a 2.5D packaging solution, a silicon interposer and a through-substrate via (TSV) connect the dies at silicon interconnect speeds with minimal footprint. In another example, a bridge die can be used. For example, an embedded multi-die interconnect bridge (EMIB) is a silicon bridge embedded below the edges of two interconnected dies, which helps to electrically couple them. In a three-dimensional (3D) architecture, one die is stacked on top of another, resulting in a smaller footprint overall. Typically, TSVs and high-pitch solder-based bumps (e.g., C4 interconnects) are used to implement electrical connections and mechanical coupling in such 3D architectures. EMIB and 3D stacked architectures can also be combined using omni-directional interconnects (ODIs), which allow the top-packaged chip to communicate with other chips horizontally using EMIBs and vertically using molded through-holes (TMVs) that are typically larger than TSVs.
[0220] However, as the number of individual IC dies integrated into a single microprocessor or other such system-in-package increases, the available floor space on a fixed-size package substrate for interconnecting these IC dies becomes challenging. To help alleviate the floor space challenge, the sizes of the IC dies can be designed to be uniform and arranged in a grid pattern in a tiled computing architecture. This tiled architecture allows for the addition of more core complex IC dies or the replacement of input / output (IO) dies to accommodate different products. As used herein, the terms "core complex" and "core" are used interchangeably to refer to a circuit comprising reusable logic cells, cells, or IC layout designs with specific functions and defined interfaces that serve as building blocks in IC chip design. For example, a core may include a collection of memory registers, arithmetic logic units (ALUs), power converters, high-speed I / O interfaces, peripherals, a programmable microprocessor, a microcontroller, a digital signal processor, analog-digital mixed signal processing blocks, a configurable computing architecture, and the like. Smaller cores (e.g., computing cores) can be combined with other smaller cores (e.g., memory) to form larger cores. For example, a core may include a computing core coupled to IO circuits that bring data into and out of the computing core, power delivery circuits that deliver power to the computing core, and aggregated or disaggregated memory blocks that serve as cache for the computing core. Multiple such cores may be referred to as a core complex, although they may also be simply referred to as cores. Because computing cores typically require additional components to create a fully functional chip or SOC, it is assumed that these supplementary components are inherent in the microelectronic assemblies of the various embodiments disclosed herein, either directly coupled to the core in question or coupled to the core in question through other cores or circuit blocks (e.g., portions, i.e., "blocks" of circuitry).
[0221] On the electrical and logic protocol side, this is accommodated by standardizing the die-to-die interface to accommodate connecting different dies together. In some scenarios, for example, when moving a core or IO die to a different process node or to a different manufacturer, the die size ends up being a bit larger or smaller. This can be accommodated in standard organic packages where the die-to-die wiring density is not very high and having matching die edge sizes is not a primary requirement. However, this is very challenging in 3D ICs with fixed (e.g., silicon) interposer or EMIB size / width and tight channel specifications.
[0222] Example
[0223] The following are example implementations of different embodiments of the invention.
[0224] Example 1. A processor comprising: message queue circuitry for implementing one or more pending message queues to store a plurality of messages received from a first interconnect fabric or an IP block; and a packetizer for determining a size in units of minislots for each of the plurality of messages, each minislot comprising a defined portion of a slot of a data transfer unit, the packetizer for further determining a number of available minislots in one or more current slots of the data transfer unit, and for packing all or a selected subset of the plurality of messages into the one or more slots based on the minislot size of each of the plurality of messages and the number of available minislots to minimize a number of unused bits in the one or more slots, wherein the data transfer unit is transmitted over a second interconnect fabric after the selected subset or all of the plurality of messages have been packed.
[0225] Example 2. The processor of Example 1, further comprising: minislot tracking circuitry integrated into or coupled to the packetizer to track a size in minislots of each of the plurality of messages and a number of available minislots in the current one or more slots of the data transmission unit.
[0226] Example 3. The processor of example 1 or 2, wherein the packetizer is configured to cause at least two slots of the data transmission unit to be linked, and to pack a selected subset or all of the plurality of messages into the at least two slots.
[0227] Example 4. The processor of any of Examples 1-3, further comprising compression circuitry to compress one or more of the plurality of messages prior to storage in the one or more pending message queues.
[0228] Example 5. The processor of any of Examples 1-4, wherein the one or more pending message queues include a request message queue, a response message queue, and a data header queue, wherein each message of the plurality of messages or a portion thereof is to be stored in one of the request message queue, the response message queue, and the data header queue.
[0229] Example 6. The processor of any of Examples 1-5, further comprising: configuration circuitry integrated with or coupled to the grouper, the configuration circuitry configured to configure the grouper based on minislot characteristics including minislot size, the grouper operable according to the minislot characteristics.
[0230] Example 7. The processor of any of Examples 1-6, wherein the data transfer unit comprises a flit, and the slot comprises one of a plurality of slots of the flit.
[0231] Example 8. The processor of any of Examples 1-7, wherein the flits comprise 68-byte or 256-byte Compute Express Link (CXL) flits.
[0232] Example 9. A method includes: storing a plurality of messages received from a first interconnect fabric or an IP block in one or more pending message queues; determining a size of each of the plurality of messages in units of minislots, each minislot comprising a defined portion of a slot of a data transfer unit; determining a number of available minislots in one or more current slots of the data transfer unit; and packing all or a selected subset of the plurality of messages into the one or more slots based on the minislot size of each of the plurality of messages and the number of available minislots to minimize a number of unused bits in the one or more slots, wherein the data transfer unit is transmitted over a second interconnect fabric after the selected subset or all of the plurality of messages have been packed.
[0233] Example 10. The method of Example 9, further comprising: tracking a size of each of the plurality of messages in units of minislots; and tracking a number of available minislots in the current one or more slots of the data transmission unit.
[0234] Example 11. The method of Example 9 or 10, further comprising: linking at least two slots of the data transmission unit; and packing a selected subset or all of the plurality of messages into the at least two slots to minimize the number of unused bits in the at least two slots.
[0235] Example 12. The method of any of Examples 9-11, further comprising compressing one or more of the plurality of messages before storing in the one or more pending message queues.
[0236] Example 13. The method of any of Examples 9-12, wherein the one or more pending message queues include a request message queue, a response message queue, and a data header queue, wherein each message in the plurality of messages or a portion thereof is to be stored in one of the request message queue, the response message queue, and the data header queue.
[0237] Example 14. The method of any of Examples 9-13, further comprising: configuring minislot characteristics including a minislot size, wherein a selected subset or all of the plurality of messages are to be packed according to the minislot characteristics.
[0238] Example 15. The method of any of Examples 9-14, wherein the data transfer unit comprises a flitter, and the slot comprises one of a plurality of slots of the flitter.
[0239] Example 16. The method of any of Examples 9-15, wherein the flits comprise 68-byte or 256-byte Compute Express Link (CXL) flits.
[0240] Example 17. A machine-readable medium having program code stored thereon, the program code, when executed by a machine, causes the machine to perform operations comprising: storing a plurality of messages received from a first interconnect fabric or IP block in one or more pending message queues; determining a size of each of the plurality of messages in units of minislots, each minislot comprising a defined portion of a slot of a data transfer unit; determining a number of available minislots in one or more current slots of the data transfer unit; and packing all or a selected subset of the plurality of messages into the one or more slots based on the minislot size of each of the plurality of messages and the number of available minislots to minimize a number of unused bits in the one or more slots, wherein after the selected subset or all of the plurality of messages have been packed, transmitting the data transfer unit over a second interconnect fabric.
[0241] Example 18. The machine-readable medium of Example 17, further comprising program code for causing the machine to: track a size in minislots of each of the plurality of messages; and track a number of available minislots in the current one or more slots of the data transmission unit.
[0242] Example 19. The machine-readable medium of Example 17 or 18, further comprising program code for causing the machine to: link at least two slots of the data transmission unit; and pack a selected subset or all of the plurality of messages into the at least two slots to minimize a number of unused bits in the at least two slots.
[0243] Example 20. The machine-readable medium of any of Examples 17-19, further comprising program code that causes the machine to compress one or more of the plurality of messages before storing in the one or more pending message queues.
[0244] Example 21. The machine-readable medium of any of Examples 17-20, wherein the one or more pending message queues include a request message queue, a response message queue, and a data header queue, wherein each message of the plurality of messages, or a portion thereof, is to be stored in one of the request message queue, the response message queue, and the data header queue.
[0245] Example 22. The machine-readable medium of any of Examples 17-21, further comprising program code that causes the machine to configure minislot characteristics including minislot sizes, wherein a selected subset or all of the plurality of messages are to be packed according to the minislot characteristics.
[0246] Example 23. The machine-readable medium of any of Examples 17-22, wherein the data transfer unit comprises a flit, and the slot comprises one of a plurality of slots of the flit.
[0247] Example 24. The machine-readable medium of any of Examples 17-23, wherein the flits comprise 68-byte or 256-byte Compute Express Link (CXL) flits.
Claims
1. A processor, comprising: message queue circuitry for implementing one or more pending message queues to store a plurality of messages received from the first interconnect fabric or IP blocks; as well as a packetizer for determining a size of each of the plurality of messages in units of minislots, each minislot comprising a defined portion of a slot of a data transmission unit, the packetizer for further determining a number of available minislots in a current one or more slots of the data transmission unit, and for packing all or a selected subset of the plurality of messages into the one or more slots based on the minislot size of each of the plurality of messages and the number of available minislots to minimize a number of unused bits in the one or more slots, wherein the data transmission unit is transmitted over a second interconnect structure after the selected subset or all of the plurality of messages have been packaged.
2. The processor of claim 1 , further comprising: Minislot tracking circuitry is integrated into or coupled to the packetizer to track the size of each of the plurality of messages in units of minislots and to track the number of available minislots in the current one or more slots of the data transmission unit.
3. The processor according to claim 1 or 2, wherein: The packetizer is configured to cause at least two slots of the data transmission unit to be linked, and to pack the selected subset or all of the plurality of messages into the at least two slots.
4. The processor according to any one of claims 1 to 3, further comprising: Compression circuitry is provided for compressing one or more of the plurality of messages prior to storage in the one or more pending message queues.
5. The processor according to any one of claims 1 to 4, wherein: The one or more pending message queues include a request message queue, a response message queue, and a data header queue, wherein each message of the plurality of messages or a portion thereof is to be stored in one of the request message queue, the response message queue, and the data header queue.
6. The processor according to any one of claims 1 to 5, further comprising: Configuration circuitry is integrated with or coupled to the grouper, the configuration circuitry being operable to configure the grouper based on minislot characteristics including minislot size.
7. The processor according to any one of claims 1 to 6, wherein: The data transfer unit includes a flitter, and the slot of the data transfer unit includes one of a plurality of slots of the flitter.
8. The processor according to claim 7, wherein: The microchip includes a 68-byte or 256-byte Compute Express Link CXL microchip.
9. A method comprising: storing a plurality of messages received from the first interconnect fabric or IP block in one or more pending message queues; determining a size of each of the plurality of messages in units of minislots, each minislot comprising a defined portion of a slot of a data transmission unit; determining a number of available minislots in one or more current slots of the data transfer unit; as well as packing all or a selected subset of the plurality of messages into the one or more slots based on a minislot size of each of the plurality of messages and the number of available minislots to minimize a number of unused bits in the one or more slots, wherein the data transmission unit is transmitted over a second interconnect structure after the selected subset or all of the plurality of messages have been packaged.
10. The method of claim 9, further comprising: tracking a size in minislots of each of the plurality of messages; as well as The number of available minislots in the current one or more slots of the data transfer unit is tracked.
11. The method according to claim 9 or 10, further comprising: linking at least two slots of the data transmission unit; as well as The selected subset or all of the plurality of messages are packed into the at least two slots to minimize a number of unused bits in the at least two slots.
12. The method of any one of claims 9 to 11, further comprising: One or more of the plurality of messages are compressed prior to storage in the one or more pending message queues.
13. The method according to any one of claims 9 to 12, characterized in that The one or more pending message queues include a request message queue, a response message queue, and a data header queue, wherein each message of the plurality of messages or a portion thereof is to be stored in one of the request message queue, the response message queue, and the data header queue.
14. The method of any one of claims 9 to 13, further comprising: Minislot characteristics including minislot sizes are configured, wherein the selected subset or all of the plurality of messages are to be packetized according to the minislot characteristics.
15. The method according to any one of claims 9 to 14, characterized in that The data transfer unit includes a flitter, and the slot of the data transfer unit includes one of a plurality of slots of the flitter.
16. The method according to claim 15, wherein The microchip includes a 68-byte or 256-byte Compute Express Link CXL microchip.
17. A machine-readable medium having program code stored thereon, the program code, when executed by a machine, causing the machine to perform operations comprising: storing a plurality of messages received from the first interconnect fabric or IP block in one or more pending message queues; determining a size of each of the plurality of messages in units of minislots, each minislot comprising a defined portion of a slot of a data transmission unit; determining a number of available minislots in one or more current slots of the data transfer unit; as well as packing all or a selected subset of the plurality of messages into the one or more slots based on a minislot size of each of the plurality of messages and the number of available minislots to minimize a number of unused bits in the one or more slots, wherein the data transmission unit is transmitted over a second interconnect structure after the selected subset or all of the plurality of messages have been packaged.
18. The machine-readable medium of claim 17, further comprising program code that causes the machine to: tracking a size in minislots of each of the plurality of messages; and The number of available minislots in the current one or more slots of the data transfer unit is tracked.
19. The machine-readable medium of claim 17 or 18, further comprising program code for causing the machine to: linking at least two slots of the data transmission unit; and The selected subset or all of the plurality of messages are packed into the at least two slots to minimize a number of unused bits in the at least two slots.
20. The machine-readable medium of any one of claims 17 to 19, further comprising program code for causing a machine to: One or more of the plurality of messages are compressed prior to storage in the one or more pending message queues.
Citation Information
Cited By
Template-based method for reducing analysis complexity of receiving side
CN121166201A