Latency optimization in partial-width link conditions
The implementation of layered protocol stacks and point-to-point interconnects with optimized flit and packet structures addresses latency and bandwidth challenges in computing systems, enhancing performance and power efficiency in high-performance computing environments.
Patent Information
- Application Number
- JP2023569594
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-20
- Filing Date
- 2022-11-21
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2042-11-21
AI Technical Summary
As computing systems evolve with increased processing power and complexity, existing interconnect architectures face challenges in managing communication latency and bandwidth requirements across multiple processors and devices, particularly in high-performance computing environments where power efficiency and performance are critical.
Implementing a layered protocol stack and point-to-point interconnects, such as PCIe and UPI, with optimized flit and packet structures to enhance communication efficiency, including credit-based flow control and flexible routing, to address latency and bandwidth demands.
Enhances communication latency optimization and power efficiency in computing systems, supporting high-performance computing environments by optimizing interconnect architectures for improved performance and power management.
Smart Images

Figure 0007791907000001 
Figure 0007791907000002 
Figure 0007791907000003
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims priority to U.S. Patent Application No. 17 / 556,685, entitled "Latency Optimization in Partial-Width Link Conditions," filed December 20, 2021. This priority application is incorporated herein by reference.
[0002] The present disclosure relates to computing systems, and particularly (but not exclusively) to physical interconnects and associated link protocols. [Background technology]
[0003] Advances in semiconductor processing and logic design have enabled an increase in the amount of logic that can reside in an integrated circuit device. As a natural consequence, computer system configurations have evolved from single or multiple integrated circuits within a system to multiple cores, multiple hardware threads, and multiple logical processors residing on individual integrated circuits, as well as other interfaces integrated with such processors. A processor or integrated circuit typically comprises a single physical processor die, which may include any number of cores, hardware threads, logical processors, interfaces, memory, controller hubs, etc.
[0004] Smaller computing devices have grown in popularity as a result of their greater capabilities, allowing them to fit greater processing power in smaller packages. Smartphones, tablets, ultra-thin notebooks, and other user devices have grown exponentially. However, these smaller devices rely on servers for both data storage and complex processing that exceeds their form factor. As a result, demand in the high-performance computing market (i.e., the server space) has also increased. For example, in today's servers, increasing computing power typically involves not only a single processor with multiple cores, but also multiple physical processors (also referred to as multiple sockets). However, as processing power increases along with the number of devices in a computing system, communication between sockets and other devices becomes more important.
[0005] In fact, interconnects have grown from more traditional multi-drop buses that primarily handled telecommunications to full-fledged interconnect architectures that facilitate high-speed communication. Unfortunately, with the demands of future processors consuming at much higher rates, there is a corresponding demand for the capabilities of existing interconnect architectures. [Brief explanation of the drawings]
[0006] [Figure 1] 1 illustrates an embodiment of a computing system that includes an interconnect architecture.
[0007] [Figure 2] 1 illustrates an embodiment of an interconnect architecture that includes a hierarchical stack.
[0008] [Figure 3] 1 illustrates an embodiment of a request or packet that may be generated or received within an interconnect architecture.
[0009] [Figure 4]1 illustrates an embodiment of a transmitter and receiver pair for an interconnect architecture.
[0010] [Figure 5] 1 illustrates an embodiment of a potential high performance processor-to-processor interconnect configuration.
[0011] [Figure 6] 1 illustrates an embodiment of a layered protocol stack associated with an interconnect.
[0012] [Figure 7] 1 shows a simplified block diagram of an exemplary computing system utilizing a link conforming to the Compute Express Link (CXL) based protocol.
[0013] [Figure 8] 1 shows a simplified block diagram of an exemplary system including multiple integrated circuit blocks.
[0014] [Figure 9] 1 shows a simplified block diagram of an exemplary protocol circuit.
[0015] [Figure 10] 1 shows a simplified block diagram of an exemplary protocol circuit of a port.
[0016] [Figure 11A] 1 shows a representation of exemplary standard flit formats. [Figure 11B] 1 shows a representation of exemplary standard flit formats. [Figure 11C] 1 shows a representation of exemplary standard flit formats. [Figure 11D]1 shows a representation of exemplary standard flit formats.
[0017] [Figure 12A] 10 illustrates a representation of an exemplary modified flit format used during partial width link conditions. [Figure 12B] 10 illustrates a representation of an exemplary modified flit format used during partial width link conditions. [Figure 12C] 10 illustrates a representation of an exemplary modified flit format used during partial width link conditions.
[0018] [Figure 13A] 1 is a simplified block diagram illustrating an exemplary modified flit format during partial-width link conditions.
[0019] [Figure 13B] 1 is a simplified block diagram illustrating throttling to address latency with selective retry events.
[0020] [Figure 14] 1 illustrates an embodiment of a block diagram for a computing system including a multi-core processor.
[0021] [Figure 15] 1 illustrates an embodiment of a block for a computing system including multiple processors. DETAILED DESCRIPTION OF THE INVENTION
[0022] The following description includes numerous specific details, such as examples of specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific dimensions / heights, specific processor pipeline stages and operations, etc., to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that these specific details are not required to implement the solutions provided in the present disclosure. In other examples, well-known components or methods, such as specific and alternative processor architectures, specific logic circuits / code for the described algorithms, specific firmware code, specific interconnect operations, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, specific representations of algorithms in code, specific power-down and gating techniques / logic, and other specific operational details of computer systems, are not described in detail to avoid unnecessarily obscuring the present disclosure.
[0023] Although the following embodiments may be described with reference to energy conservation and energy efficiency in specific integrated circuits, such as in computing platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein can be applied to other types of circuits or semiconductor devices that can also benefit from better energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to desktop computer systems or Ultrabooks™. They may also be used in other devices, such as handheld devices, tablets, other thin notebooks, system-on-a-chip (SOC) devices, and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications generally include microcontrollers, digital signal processors (DSPs), system-on-a-chips, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other systems capable of performing the functions and operations taught below. Furthermore, the apparatus, methods, and systems described herein are not limited to physical computing devices, but may also relate to software optimization for energy conservation and efficiency. As will become readily apparent from the description below, the method, apparatus, and system embodiments described herein (whether relating to hardware, firmware, software, or a combination thereof) are critical to the future of "green technology" balanced with performance considerations.
[0024] As computing systems advance, the components therein become more complex. As a result, the interconnect architectures used to link and communicate between components also become more complex to ensure that bandwidth requirements for optimal component operation are met. Furthermore, different market segments require different aspects of interconnect architectures to meet their market needs. For example, servers require higher performance, while the mobile ecosystem may in some cases sacrifice overall performance for power savings. However, providing the highest possible performance with maximum power savings is the sole purpose of most fabrics. Many interconnects, as described below, could potentially benefit from aspects of the solution described herein.
[0025] One interconnect fabric architecture includes the Peripheral Component Interconnect (PCI) Express (PCIe) architecture. The primary goal of PCIe is to enable components and devices from different vendors to interoperate in an open architecture across multiple market segments: client (desktop and mobile), server (standard and enterprise), and embedded communications devices. PCI Express is a high-performance, general-purpose I / O interconnect defined for a variety of future computing and communications platforms. Some PCI attributes, such as its usage model, load-store architecture, and software interface, have been maintained throughout its revisions, whereas the previous parallel bus implementation has been replaced by a highly scalable, fully serial interface. More recent versions of PCI Express leverage advances in point-to-point interconnects, switch-based technology, and packetized protocols to deliver new levels of performance and features. Some of the advanced features supported by PCI Express include power management, quality of service (QoS), hot-plug / hot-swap support, data integrity, and error handling.
[0026] Referring to Figure 1, one embodiment of a fabric composed of point-to-point links interconnecting a set of components is shown. System 100 includes processor 105 and system memory 110 coupled to controller hub 115. Processor 105 may include any processing element, such as a microprocessor, host processor, embedded processor, coprocessor, or other processor. Processor 105 is coupled to controller hub 115 through front side bus (FSB) 106. In one embodiment, FSB 106 is a serial point-to-point interconnect, as described below. In another embodiment, link 106 includes a serial differential interconnect architecture that conforms to a different interconnect standard.
[0027] System memory 110 includes any memory device, such as random access memory (RAM), non-volatile (NV) memory, or other memory accessible by multiple devices in system 100. System memory 110 is coupled to controller hub 115 through memory interface 116. Examples of memory interfaces include a double data rate (DDR) memory interface, a dual channel DDR memory interface, and a dynamic RAM (DRAM) memory interface.
[0028] In one embodiment, controller hub 115 is a root hub, root complex, or root controller in a Peripheral Component Interconnect Express (PCIe or PCIE) interconnect hierarchy. Examples of controller hub 115 include a chipset, a memory controller hub (MCH), a northbridge, an interconnect controller hub (ICH), a southbridge, and a root controller / hub. The term chipset often refers to two physically separate controller hubs: a memory controller hub (MCH) coupled to an interconnect controller hub (ICH). Note that current systems often include an MCH integrated into processor 105, but controller 115 communicates with I / O devices in a manner similar to that described below. In some embodiments, peer-to-peer routing is optionally supported through root complex 115.
[0029] Here, controller hub 115 is coupled to switch / bridge 120 via serial link 119. Input / output modules 117 and 121, which may also be referred to as interfaces / ports 117 and 121, contain / implement layered protocol stacks to provide communication between controller hub 115 and switch 120. In one embodiment, multiple devices may be coupled to switch 120.
[0030] Switch / bridge 120 routes packets / messages upstream from devices 125, i.e., up the hierarchy toward the root complex to controller hub 115, and downstream, i.e., down the hierarchy away from the root controller, from processor 105 or system memory 110 to devices 125. In one embodiment, switch 120 is referred to as a logical assembly of multiple virtual PCI-to-PCI bridge devices. Devices 125 include any internal or external device or component coupled to an electronic system, such as I / O devices, network interface controllers (NICs), add-in cards, audio processors, network processors, hard drives, storage devices, CD / DVD ROMs, monitors, printers, mice, keyboards, routers, portable storage devices, FireWire devices, Universal Serial Bus (USB) devices, scanners, and other input / output devices. In PCIe terminology, such devices are often referred to as endpoints. Although not specifically shown, devices 125 may include PCIe-to-PCI / PCI-X bridges to support legacy or other versions of PCI devices. Endpoint devices in PCIe are often categorized as legacy, PCIe, or root complex integrated endpoints.
[0031] Graphics accelerator 130 is also coupled to controller hub 115 via serial link 132. In one embodiment, graphics accelerator 130 is coupled to an MCH which is coupled to an ICH. Switch 120, and therefore I / O devices 125, are then coupled to the ICH. I / O modules 131 and 118 also implement a layered protocol stack for communicating between graphics accelerator 130 and controller hub 115. Similar to the MCH discussion above, the graphics controller or graphics accelerator 130 itself may be integrated into processor 105. Additionally, one or more links (e.g., 123) of the system may include one or more expansion devices (e.g., 150), such as retimers, repeaters, etc.
[0032] Turning to FIG. 2, one embodiment of a layered protocol stack is shown. Layered protocol stack 200 may include any form of layered communication stack, such as a Quick Path Interconnect (QPI) stack, a PCIe stack, a next-generation high-performance computing interconnect stack, or other layered stack. While the description immediately below regarding FIGS. 1-4 is directed to a PCIe stack, the same concepts may apply to multiple other interconnect stacks. In one embodiment, protocol stack 200 is a PCIe protocol stack including transaction layer 205, link layer 210, and physical layer 220. Interfaces such as interfaces 117, 118, 121, 122, 126, and 131 in FIG. 1 may be represented as communication protocol stack 200. The representation as a communication protocol stack may also refer to modules or interfaces that implement / contain the protocol stack.
[0033] PCI Express communicates information between components using packets. Packets are formed in the transaction layer 205 and data link layer 210 and carry information from the transmitting component to the receiving component. As transmitted packets flow through other layers, they are extended with additional information required for those layers to process the packet. On the receiving side, reverse processing occurs, converting packets from their physical layer 220 representation to their data link layer 210 representation and finally (for transaction layer packets) to a form that can be processed by the transaction layer 205 of the receiving device.
[0034] [Transaction Layer]
[0035] In one embodiment, transaction layer 205 provides an interface between the processing cores of a device and the interconnect architecture, such as data link layer 210 and physical layer 220. In this regard, the primary role of transaction layer 205 is the assembly and disassembly of packets (i.e., transaction layer packets or TLPs). Conversion layer 205 typically manages credit-based flow control for TLPs. PCIe implements split transactions, i.e., multiple transactions with requests and responses separated in time, allowing the link to carry other traffic while the target device collects data for the response.
[0036] PCIe also utilizes credit-based flow control. In this scheme, a device advertises an initial amount of credits for each of multiple receive buffers in transaction layer 205. An external device at the opposite end of the link, such as controller hub 115 in FIG. 1, counts the number of credits consumed by each TLP. If the transaction does not exceed the credit limit, the transaction can be sent. Upon receiving a response, the credit amount is restored. The advantage of the credit scheme is that the latency of credit return does not affect performance unless the credit limit is encountered.
[0037] In one embodiment, the four transaction address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more read and write requests that transfer data to or from memory-mapped locations. In one embodiment, memory space transactions can use two different address formats, for example, a short address format such as a 32-bit address or a long address format such as a 64-bit address. Configuration space transactions are used to access the configuration space of a PCIe device. Transactions to the configuration space include read and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.
[0038] Thus, in one embodiment, transaction layer 205 assembles packet header / payload 206. The current packet header / payload format can be found in the PCIe specification at the PCIe specification website.
[0039] Referring briefly to Figure 3, one embodiment of a PCIe transaction descriptor is shown. In one embodiment, transaction descriptor 300 is a mechanism for conveying transaction information. In this regard, transaction descriptor 300 supports the identification of multiple transactions within a system. Other potential uses include changing the default transaction ordering and tracking the association of transactions with multiple channels.
[0040] Transaction descriptor 300 includes a global identifier field 302, an attribute field 304, and a channel identifier field 306. In the illustrated example, global identifier field 302 is shown to include a local transaction identifier field 308 and a source identifier field 310. In one embodiment, global transaction identifier 302 is unique for all outstanding requests.
[0041] According to one implementation, the local transaction identifier field 308 is a field generated by the requesting agent that is unique for all outstanding requests requiring completion for that requesting agent. Additionally, in this example, the source identifier 310 uniquely identifies the requesting agent within the PCIe hierarchy. Thus, the local transaction identifier 308 field, together with the source ID 310, provides global identification of the transaction within the hierarchy domain.
[0042] The attribute field 304 specifies characteristics and relationships of a transaction. In this regard, the attribute field 304 is potentially used to provide additional information that allows for modification of the default handling of transactions. In one embodiment, the attribute field 304 includes a priority field 312, a reserved field 314, an ordering field 316, and a no-snoop field 318, where the priority subfield 312 may be modified by the initiator to assign a priority to a transaction. The reserved attribute field 314 is left reserved for future or vendor-defined use. Possible usage models that use priority or security attributes may be implemented using the reserved attribute field.
[0043] In this example, the ordering attribute field 316 is used to provide optional information conveying the type of ordering that may change the default ordering rules. According to one exemplary implementation, an ordering attribute of "0" indicates that the default ordering rules apply, and an ordering attribute of "1" indicates relaxed ordering, where multiple writes can pass multiple writes in the same direction, and multiple read completions can pass multiple writes in the same direction. The snoop attribute field 318 is utilized to determine whether multiple transactions are snooped. As shown, the channel ID field 306 identifies the channel with which the transaction is associated.
[0044] [Link Layer]
[0045] Link layer 210, also referred to as data link layer 210, acts as an intermediary between transaction layer 205 and physical layer 220. In one embodiment, the role of data link layer 210 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components on a link. One side of data link layer 210 accepts TLPs assembled by transaction layer 205, applies a packet sequence identifier 211, i.e., an identification number or packet number, calculates and applies an error detection code, i.e., CRC 212, and sends the modified TLPs to physical layer 220 for transmission from the physical device to an external device.
[0046] [Physical layer]
[0047] In one embodiment, physical layer 220 includes logical sub-block 221 and electrical sub-block 222 for physically transmitting packets to an external device, where logical sub-block 221 is responsible for the "digital" functions of physical layer 220. In this regard, the logical sub-blocks include a transmit section for preparing outgoing information for transmission by physical sub-block 222 and a receive section for identifying and preparing received information before passing it to link layer 210.
[0048] Physical block 222 includes a transmitter and a receiver. The transmitter is provided with symbols by logical sub-block 221, which the transmitter serializes and transmits to an external device. The receiver is provided with serialized symbols from the external device and converts the received signal into a bit stream. The bit stream is deserialized and provided to logical sub-block 221. In one embodiment, an 8b / 10b transmission code is used, where 10-bit symbols are transmitted / received. Special symbols are used here to frame packets using frame 223. Additionally, in one example, the receiver also provides a symbol clock recovered from the incoming serial stream.
[0049] As described above, transaction layer 205, link layer 210, and physical layer 220 are described with respect to a specific embodiment of a PCIe protocol stack, but the layered protocol stack is not limited thereto. In fact, any layered protocol may be included / implemented. As an example, a port / interface represented as a layered protocol includes: (1) a first layer, i.e., transaction layer, that assembles packets; a second layer, i.e., link layer, that sequences the packets; and a third layer, i.e., physical layer, that transmits the packets. As a specific example, the Common Standard Interface (CSI) layered protocol is utilized.
[0050] Referring now to Figure 4, one embodiment of a PCIe serial point-to-point fabric is shown. While one embodiment of a PCIe serial point-to-point link is shown, a serial point-to-point link includes, but is not limited to, any transmit path for transmitting serial data. In the embodiment shown, the basic PCIe link includes two low-voltage, differentially driven signal pairs: transmit pair 406 / 411 and receive pair 412 / 407. Thus, device 405 includes transmit logic 406 for transmitting data to device 410 and receive logic 407 for receiving data from device 410. In other words, the PCIe link includes two transmit paths, paths 416 and 417, and two receive paths, paths 418 and 419.
[0051] A transmission path refers to any path for transmitting data, such as a transmission line, copper wire, optical line, wireless communication channel, infrared communication link, or other communication path. For example, a connection between two devices, such as device 405 and device 410, is called a link, such as link 415. A link may support one lane, where each lane represents a set of differential signal pairs (one pair for transmit and one pair for receive). To scale bandwidth, a link may aggregate multiple lanes, denoted xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or more.
[0052] A differential pair refers to two transmission paths, such as lines 416 and 417, for transmitting differential signals. As an example, when line 416 switches from a low voltage level to a high voltage level, i.e., a rising edge, line 417 drives from a high logic level to a low logic level, i.e., a falling edge. Differential signals potentially exhibit better electrical properties, such as better signal integrity, i.e., cross-coupling, voltage overshoot / undershoot, ringing, etc., which allows for better timing windows and, therefore, faster transmission frequencies.
[0053] In one embodiment, Ultra Path Interconnect™ (UPI™) may be utilized to interconnect two or more devices. UPI is capable of implementing a next-generation cache-coherent, link-based interconnect. As one example, UPI may be used in high-performance computing platforms such as workstations or servers, including in systems where PCIe or another interconnect protocol is typically used to connect multiple processors, accelerators, I / O devices, etc. However, UPI is not so limited. Instead, UPI may be used in any system or platform described herein. Furthermore, the individual concepts developed may be applied to other interconnects and platforms, such as PCIe, MIPI, QPI, etc.
[0054] To support multiple devices, in one exemplary implementation, the UPI can include instruction set architecture (ISA) independence (i.e., the UPI can be implemented in multiple different devices). In another scenario, the UPI may also be used to connect to high-performance I / O devices, not just processors or accelerators. For example, a high-performance PCIe device may be coupled to the UPI through an appropriate conversion bridge (i.e., UPI to PCIe). Furthermore, multiple UPI links may be used by multiple UPI-based devices, such as multiple processors, in various ways (e.g., star, ring, mesh, etc.). FIG. 5 shows exemplary implementations of several potential multi-socket configurations. As shown, a two-socket configuration 505 can include two UPI links; however, in other implementations, one UPI link may be used. For larger topologies, any configuration may be used as long as identifiers (IDs) can be assigned and some form of virtual path exists, among other additional or alternative features. As shown, in one example, a four-socket configuration 510 has a UPI link from each processor to another processor. However, in the eight-socket implementation shown in configuration 515, not all sockets are directly connected to each other through UPI links. However, such configurations are supported if virtual paths or channels exist between multiple processors. The range of supported processors includes 2 to 32 in the native domain. Among other examples, larger numbers of processors may be achieved through the use of multiple domains or other interconnects between node controllers.
[0055] The UPI architecture, in some examples, includes the definition of a layered protocol architecture, including a protocol layer (coherent, non-coherent, and optionally other memory-based protocols), a routing layer, a link layer, and a physical layer. Additionally, UPI may further include improvements related to a power manager (such as a power control unit (PCU)), design for test and debug (DFT), fault handling, registers, and security, among other examples. FIG. 5 illustrates an embodiment of an exemplary UPI layered protocol stack. In some implementations, at least some of the layers shown in FIG. 5 may be optional. Each layer handles its own level of granularity or quantity of information (protocol layers 620a,b handle packets 630; link layers 610a,b handle flits 635; and physical layers 605a,b handle phits 640). Note that in some embodiments, a packet may include multiple partial flits, a single flit, or multiple flits, depending on the implementation.
[0056] As a first example, the width of phit 640 includes a one-to-one mapping of link width to number of bits (e.g., a 20-bit link width includes a 20-bit phit, etc.). Flits may have larger sizes, such as 184, 192, or 200 bits. Note that if phit 640 is 20 bits wide and the size of flit 635 is 184 bits, then a fractional number of phits 640 are required to transmit one flit 635 (e.g., 9.2 phits at 20 bits to transmit an 184-bit flit 635, or 9.6 phits at 20 bits to transmit a 192-bit flit, among other examples). Note that the widths of the underlying link at the physical layer may vary. For example, the number of lanes per direction may include 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, etc. In one embodiment, the link layer 610a,b can embed pieces of different transactions within a single flit, and one or more headers (e.g., 1, 2, 3, 4) may be embedded within the flit. In one example, the UPI divides the headers into corresponding slots, allowing multiple messages in the flit to be destined for different nodes.
[0057] In one embodiment, the physical layer 605a, b may be responsible for high-speed transfer of information over a physical medium (such as electrical or optical). A physical link may be point-to-point between two link layer entities, such as layers 605a and 605b. The link layer 610a, b may abstract the physical layer 605a, b from multiple upper layers, providing the ability to reliably transfer data (as well as requests) and manage flow control between two directly connected entities. The link layer may also be responsible for virtualizing physical channels into multiple virtual channels and message classes. The protocol layer 620a, b relies on the link layer 610a, b to map protocol messages to appropriate message classes and virtual channels before passing them to the physical layer 605a, b for transfer across multiple physical links. The link layer 610a, b may support messages such as request, snoop, response, writeback, and noncoherent data, among other examples.
[0058] The UPI physical layer 605a,b (or PHY) may be implemented above the electrical layer (i.e., the electrical conductors connecting two components) and below the link layer 610a,b, as shown in FIG. 6. The physical layer and corresponding logic may reside on each agent, connecting link layers on two agents (A and B) that are separate from each other (e.g., in devices on either side of a link). The local and remote electrical layers are connected by a physical medium (e.g., wires, conductors, light, etc.). In one embodiment, the physical layer 605a,b has two main phases: initialization and operation. During initialization, the connection is opaque to the link layer, and signaling may include a combination of timed states and handshake events. During operation, the connection is transparent to the link layer, and signaling is at a specific speed when all lanes operate together as a single link. During the operational phase, the physical layer transports flits from agent A to agent B and from agent B to agent A. A connection, also called a link, abstracts some physical aspects including medium, width, and speed from the link layer, exchanging flits and control / status of the current configuration (e.g., width) with the link layer. The initialization phase includes minor phases, e.g., polling, configuration. The operational phase also includes minor phases (e.g., link power management states).
[0059] In one embodiment, the link layer 610a,b may be implemented to provide reliable data transfer between two protocol or routing entities. The link layer may abstract the physical layer 605a,b from the protocol layer 620a,b, handle flow control between two protocol agents (A, B), and provide multiple virtual channel services for the protocol layer (multiple message classes) and the routing layer (multiple virtual networks). The interface between the protocol layer 620a,b and the link layer 610a,b may typically be at the packet level. In one embodiment, the smallest transmission unit at the link layer is called a flit, which is a specified number of bits, e.g., 192 bits or some other unit name. The link layer 610a,b relies on the physical layer 605a,b to frame the transmission unit (flit) of the physical layer 605a,b into the transmission unit (flit) of the link layer 610a,b. Additionally, the link layer 610a,b may be logically divided into two parts: a transmitter and a receiver. A transmitter / receiver pair on one entity may be connected to a transmitter / receiver pair on another entity. Flow control is often performed on both a flit and packet basis. Error detection and correction is also potentially performed on a flit-level basis.
[0060] In one embodiment, the routing layer 615a,b can provide a flexible and distributed method for routing UPI transactions from source to destination. This scheme is flexible because multiple routing algorithms for multiple topologies can be specified via programmable routing tables in each router (in one embodiment, the programming is performed by firmware, software, or a combination thereof). The routing function can be distributed; routing can be done via a series of routing steps, with each routing step defined via a table lookup in either the source, intermediate, or target router. A lookup at the source can be used to insert a UPI packet into the UPI fabric. A lookup at an intermediate router can be used to route a UPI packet from an input port to an output port. A lookup at the destination port can be used to target a destination UPI protocol agent. Note that because multiple routing tables, and therefore multiple routing algorithms, are not specifically defined by the specification, the routing layer can be thin in some implementations. This allows for flexibility and various usage models, including flexible platform architecture topologies, to be defined by the system implementation. The routing layer 615a,b relies on the link layer 610a,b to provide for the use of up to three (or more) virtual networks (VNs), in one example, two deadlock-free VNs, VN0 and VN1, with multiple message classes defined in each virtual network. A shared adaptive virtual network (VNA) may be defined at the link layer, but this adaptive network may not be directly exposed to multiple routing concepts, as each message class and virtual network may have multiple dedicated resources and guaranteed forwarding progress, among other features and examples.
[0061] In one embodiment, the UPI may include a coherency protocol layer 620a,b to support agents that cache lines of data from memory. An agent that wants to cache memory data may use the coherency protocol to read the line of data to load it into its cache. An agent that wants to modify a line of data in its cache may use the coherency protocol to obtain ownership of the line before modifying the data. After modifying the line, the agent may follow protocol requirements to keep it in its cache until either writing the line back to memory or including it in a response to an external request. Finally, the agent may respond to external requests to invalidate the line in its cache. The protocol ensures data coherency by prescribing rules that all caching agents can follow. It also provides a means for agents without caches to coherently read and write memory data.
[0062] The optimized implementations discussed herein may be applied to the flit and packet structures of various protocols, including PCIe and UPI. Such features and improvements may also be applied to other protocols. For example, the features discussed herein may be applied to the Compute Express Link (CXL) interconnect protocol, which is designed to provide improved high-speed CPU-to-device and CPU-to-memory interconnects designed to accelerate next-generation datacenter performance, among other applications. CXL maintains memory coherency between CPU memory space and memory on attached devices, which enables resource sharing for higher performance, reduced software stack complexity, and lower overall system costs, among other exemplary benefits. CXL enables communication between a host processor (e.g., a CPU) and a set of workload accelerators (e.g., graphics processing units (GPUs), field-programmable gate array (FPGA) devices, tensor and vector processor units, machine learning accelerators, and dedicated accelerator solutions, among other examples). Indeed, CXL is designed to provide a standard interface for high-speed communications as accelerators are increasingly used to complement CPUs to support emerging computing applications such as artificial intelligence, machine learning, and other applications.
[0063] A CXL link can be a low-latency, high-bandwidth discrete or on-package link that supports dynamic protocol multiplexing of coherency, memory access, and input / output (I / O) protocols. Among other uses, a CXL link can enable accelerators to access system memory as a caching agent and / or host system memory, among other examples. CXL is a dynamic multi-protocol technology designed to support a vast spectrum of accelerators. Over a discrete or on-package link, CXL provides a rich set of protocols, including PCIe (CXL.io)-like I / O semantics, cache protocol semantics (CXL.cache), and memory access semantics (CXL.mem). Depending on the particular accelerator usage model, all of the CXL protocols or only a subset of these protocols can be enabled. In some implementations, CXL may be built on the well-established and widely adopted PCIe infrastructure (e.g., PCIe 5.0), leveraging the PCIe physical and electrical interface to provide advanced protocols for areas including I / O, memory protocols (e.g., allowing a host processor to share memory with an accelerator device), and coherency interfaces.
[0064] Referring to FIG. 7, a simplified block diagram 700 illustrating an exemplary system utilizing a CXL link 750 is shown. For example, the link 750 may interconnect a host processor 705 (e.g., a CPU) to an accelerator device 710. In this example, the host processor 705 includes one or more processor cores (e.g., 715a-b) and one or more I / O devices (e.g., 718). A host memory (e.g., 760) may be provided with the host processor (e.g., on the same package or die). The accelerator device 710 may include accelerator logic 720 and, in some implementations, may include its own memory (e.g., accelerator memory 765). In this example, the host processor 705 may include circuitry for implementing coherence / cache logic 725 and interconnect logic (e.g., PCIe logic 730). CXL multiplexing logic (e.g., 755a-b) may also be provided to enable multiplexing of CXL protocols (e.g., I / O protocols 735a-b (e.g., CXL.io), cache protocols 740a-b (e.g., CXL.cache), and memory access protocols 745a-b (CXL.mem)), thereby allowing data of any one of the supported protocols (e.g., 735a-b, 740a-b, 745a-b) to be sent in a multiplexed manner over link 750 between host processor 705 and accelerator device 710.
[0065] In some implementations, Flex Bus™ ports may be utilized in concert with CXL-compliant links to flexibly adapt devices to interconnect with a wide variety of other devices (e.g., other processor devices, accelerators, switches, memory devices, etc.). Flex Bus ports are flexible, high-speed ports statically configured to support either PCIe or CXL links (and potentially links of other protocols and architectures). Flex Bus ports allow designs to choose between providing native PCIe protocol or CXL over a high-bandwidth off-package link. The selection of the protocol applied at the port may occur during boot time via auto-negotiation and may be based on the device plugged into the slot. The Flex Bus uses PCIe circuitry to be compatible with PCIe retimers and adheres to the PCIe form factor, a standard for add-in cards.
[0066] 8, an example system is shown (in simplified block diagram 800) that utilizes Flex Bus ports (e.g., 835-840) to implement CXL (e.g., 815a-b, 850a-b) and PCIe links (e.g., 830a-b) to couple various devices (e.g., 710, 810, 820, 825, 845, etc.) to a host processor (e.g., CPU 705, 805). In this example, the system may include two CPU host processor devices (e.g., 705, 805) interconnected by an inter-processor link 870 (e.g., utilizing UltraPath Interconnect (UPI), Infinity Fabric™, or other interconnect protocol). Each host processor device 705, 805 may be coupled to a local system memory block 760, 860 (e.g., a double data rate (DDR) memory device), which is coupled to the respective host processor 705, 805 via a memory interface (e.g., a memory bus or other interconnect).
[0067] As discussed above, CXL links (e.g., 815a, 815b) may be utilized to interconnect various accelerator devices (e.g., 710, 810). Accordingly, corresponding ports (e.g., Flex Bus ports 835, 840) may be configured (e.g., a selected CXL mode) to establish a CXL link to enable interconnection of a corresponding host processor device (e.g., 705, 805) to the accelerator device (e.g., 710, 810). As shown in this example, Flex Bus ports (e.g., 836, 839), or other similarly configurable ports, may be configured to implement generic I / O links (e.g., PCIe links) 830a-b instead of CXL links to interconnect the host processor (e.g., 705, 805) to I / O devices (e.g., smart I / O devices 820, 825, etc.). In some implementations, the memory of the host processor 705 can be expanded, for example, through memory (e.g., 765, 865) of an attached accelerator device (e.g., 710, 810) or a memory expander device (e.g., 845 connected to the host processor 705, 805 via a corresponding CXL link (e.g., 850a-b) implemented on a Flex Bus port (837, 838), among other example implementations and architectures.
[0068] FIG. 9 is a simplified block diagram illustrating an example port architecture 900 (e.g., Flex Bus) utilized to implement a CXL link. For example, the Flex Bus architecture may be organized as multiple layers for implementing multiple protocols supported by the port. For example, a port may include transaction layer logic (e.g., 905), link layer logic (e.g., 910), and physical layer logic (e.g., 915) (e.g., implemented in whole or in part in circuitry). For example, the transaction (or protocol) layer (e.g., 905) may be subdivided into transaction layer logic 925 that implements PCIe transaction layer 955 and CXL transaction layer enhancements 960 (for CXL.io) of the base PCIe transaction layer 955, and logic 930 for implementing cache (e.g., CXL.cache) and memory (e.g., CXL.mem) protocols for the CXL link. Similarly, link layer logic 935 may be provided to implement a base PCIe data link layer 965 and a CXL link layer (for CXL.io) that represents an enhanced version of the PCIe data link layer 965. The CXL link layer 910 may also include cache and memory link layer enhancement logic 940 (e.g., for CXL.cache and CXL.mem).
[0069] Continuing with the example of FIG. 9 , the CXL link layer logic 910 may interface with CXL arbitration / multiplexing (ARB / MUX) logic 920, which, among other example implementations, interleaves traffic from two logic streams (e.g., PCIe / CXL.io and CXL.cache / CXL.mem). During link training, the transaction layer and link layer are configured to operate in either PCIe mode or CXL mode. In some examples, a host CPU, among other examples, may support implementation of either PCIe mode or CXL mode, while other devices, such as accelerators, may only support CXL mode. In some implementations, a port (e.g., a Flex Bus port) may utilize a physical layer 915 based on a PCIe physical layer (e.g., PCIe electrical circuit PHY 950). For example, the Flex Bus physical layer may be implemented as a converged logical physical layer 945 that can operate in either PCIe mode or CXL mode based on the results of alternate mode negotiation during the link training process. In some implementations, the physical layer may support multiple signaling rates (e.g., 8 GT / s, 16 GT / s, 32 GT / s, etc.) and multiple link widths (e.g., x16, x8, x4, x2, x1, etc.). In PCIe mode, the link implemented by port 900 may be fully compliant with native PCIe features (e.g., as defined in the PCIe specification), while in CXL mode, the link supports all features defined for CXL. Thus, among other examples, the Flex Bus port may transmit native PCIe protocol data or dynamic multi-protocol CXL data to provide a point-to-point interconnect capable of providing I / O protocols, coherency protocols, and memory protocols over PCIe circuitry.
[0070] The CXL I / O protocol, or CXL.io, provides a non-coherent load / store interface for I / O devices. Transaction types, transaction packet formatting, credit-based flow control, virtual channel management, and transaction ordering rules in CXL.io may follow all or part of the PCIe definition. The CXL cache coherency protocol, or CXL.cache, defines interactions between devices and hosts as multiple requests, each with at least one associated response message and possibly a data transfer. The interface consists of three channels in each direction: request, response, and data.
[0071] The CXL memory protocol, or CXL.mem, is a transactional interface between processors and memory that uses the CXL physical and link layers when communicating between dies. Among other examples, CXL.mem can be used for multiple different memory attachment options, including when the memory controller is located in the host CPU, when the memory controller is in an accelerator device, or when the memory controller is moved to a memory buffer chip. Among other exemplary features, CXL.mem can be applied to transactions involving different memory types (e.g., volatile, persistent, etc.) and configurations (e.g., flat, hierarchical, etc.). In some implementations, a host processor's coherency engine can interface with memory using CXL.mem requests and responses. In this configuration, the CPU coherency engine is considered a CXL.mem Master, and the Mem device is considered a CXL.mem Subordinate. The CXL.mem Master is the agent responsible for sourcing CXL.mem requests (e.g., read, write, etc.), and the CXL.mem Subordinate is the agent responsible for responding to CXL.mem requests (e.g., data, completion, etc.). When the Subordinate is an accelerator, the CXL.mem protocol assumes the presence of a Device Coherency Engine (DCOH). This agent is expected to be responsible for implementing coherency-related functions such as snooping the device cache and updating metadata fields based on CXL.mem commands. In implementations where metadata is supported by device-attached memory, the metadata may be used by the host to implement a coarse-grained snoop filter for the CPU socket, among other example uses.
[0072] In some implementations, an interface may be provided to couple circuitry or other logic (e.g., an intellectual property (IP) block or other hardware element) implementing a link layer (e.g., 910) to circuitry or other logic (e.g., an IP block or other hardware element) implementing at least a portion of a protocol's physical layer (e.g., 915). For example, the interface may be based on the Logical PHY Interface (LPIF) specification, which defines a common interface between a link layer controller, module, or other logic and a module implementing a logical physical layer ("logical PHY" or "logPHY") to promote interoperability, design, and validation reuse between one or more link and physical layers for interfaces to physical interconnects, as in the example of FIG. 9. Furthermore, as in the example of FIG. 9, the interface may be implemented with logic (e.g., 935, 940) to implement and support multiple protocols simultaneously. Additionally, in such implementations, an arbitration and multiplexer layer (e.g., 920) may be provided between the link layer (e.g., 910) and the physical layer (e.g., 915). In some implementations, each block (e.g., 915, 920, 935, 940) in a multi-protocol implementation may interface with other blocks via an independent LPIF interface (e.g., 980, 985, 990). Among other examples, if bifurcations are supported, each bifurcated port may have its own independent LPIF interface as well.
[0073] While examples discussed herein may reference the use of an LPIF-based link layer-to-logical PHY interface, it should be understood that the details and principles discussed herein may equally apply to non-LPIF interfaces. Similarly, while some examples may reference the use of a common link layer-to-logical PHY interface to couple a PHY to a controller for implementing CXL or PCIe, other link layer protocols may also use such an interface. Similarly, while some reference may be made to a Flex Bus physical layer, in some implementations other physical layer logic may similarly be employed and use the common link layer-to-logical PHY interface as discussed herein, among other example variations within the scope of this disclosure.
[0074] 10 is a simplified block diagram 1000 illustrating an example implementation of a device 1005 having a port 1025 that includes circuitry for establishing a physical communication link 1075 (e.g., according to a link training state machine defined in an interconnect protocol) and communicating over the link 1075 with a link partner device 1050 at the other end of the link 1075 according to one or more interconnect protocols supported by the link 1075. In one example, the device 1005 may include computer processing circuitry (e.g., processor core 1010), computer memory blocks (e.g., 1015), and corresponding memory controller 1020, among other or alternative components such as a graphics processing component, a tensor processing component, accelerator hardware, a caching agent, etc. The port 1025 may implement a transmitter (TX) for transmitting data over the link 1075 and a corresponding receiver (RX) for receiving data over the link 1075. Similarly, link partner device 1050 may have a corresponding port 1055 (e.g., having similar logic and characteristics as port 1025 and similarly supporting and capable of communicating via link 1075). Port 1025 may include protocol logic 1030 implemented at least in part in hardware circuitry for implementing logic supporting one additional layered interconnect protocol (e.g., PCIe, CXL, UPI, USB, Gen-Z, etc.) to support training, equalizing, and otherwise establishing link 1075 and then communicating via link 1075 and managing link state transitions on link 1075.
[0075] According to some of the example features and solutions discussed herein, some implementations of a port (e.g., 1025) may utilize enhanced protocol logic (e.g., 1030) to implement improved features of the interconnect protocol. For example, in some examples, the protocol logic 1030 may include a flit mode controller circuit 1035 to implement multiple alternative flit-based transport modes. For example, the flit-based transport modes may include a standard flit-based transport mode, a latency-optimized flit-based transport mode, and a partial link width-optimized flit-based transport mode, among other examples. In each of the alternative modes, a different respective flit format may be utilized, and the flit mode controller 1035 may negotiate with the link partner device 1050 to determine which of these flit modes are mutually supported and how and when the link partners (e.g., 1005 and 1050) transition between flit modes during operation. For example, the flit mode controller 1035 may, among other examples, identify a transition from a first active link state operating at a first link width to a second active link state operating at a different, second, narrower link width and trigger a corresponding transition (e.g., from a standard flit-based transport mode or a latency-optimized flit-based transport mode) to a fractional link width-optimized flit-based transport mode. Such a feature may help mitigate the latency impact of the narrower link width (e.g., fractional link width), among other example benefits.
[0076] Additionally, some implementations of protocol logic may include throttle controller logic 1040. In some implementations, selective retry is supported by the protocol (e.g., PCIe6, CXL, etc.) and may be implemented using receive and / or transmit buffers (e.g., 1042, 1044). If a receiving device on a link (e.g., 1075) detects that a received flit (in a sequence of flits) contains an error, the receiving device may send a request over the link to replay the particular erroneous flit (rather than, e.g., requesting a replay of not only the erroneous flit but also flits received following the erroneous flit (e.g., as in conventional flit retry implementations)). In an exemplary implementation of selective retry, a receiving device may request a replay of a particular erroneous flit, buffer (e.g., in a replay buffer) any subsequent flits that are not idle (idle flits can be dropped because they do not carry any meaningful information), and consume these buffered flits after a corrected retry version of the particular erroneous flit is received. Because these buffered flits will be consumed sequentially after the erroneous flit is replayed, selective retry can introduce significant latency not only to the buffered flits but also to any non-idle flits that follow the replayed flit in sequence. In some implementations, periodic idle flits and / or SKP ordered sets (OS) may be transmitted (e.g., opportunistically or according to a defined frequency), and these moments of idle data may be utilized to “recover” some of this lost latency (e.g., by instead transmitting these delayed flits in place of scheduled idle flits or SKP symbols). While providing idle periods in the data flow can largely prevent the accumulation of latency from selective retries, if the error rate increases, the rate of selective retries will increase as well, and as a result, the latency introduced by selective retry buffering may accumulate unbearably quickly.Accordingly, among other example solutions, in some implementations, the throttle controller 1040 may detect an increase in flit errors (e.g., with the help of a software controller (e.g., 1090) or by tracking replay requests in hardware) and throttle or temporarily increase the frequency of idle windows in the data stream (e.g., by intentionally increasing the number or frequency of idle flits on the link).
[0077] As introduced above, interconnects may support flit-mode transport using defined flow control units, or flits, to encapsulate data (e.g., packet data) for transmission over the physical layer of a link. An entire packet or a segment of a packet may be transmitted within a single flit. A flit may be defined by the respective protocol to have a fixed length and format (e.g., with defined fields). One or more of these fields may be defined to carry error detection or error correction codes, such as a cyclic redundancy check (CRC) code, a forward error correction (FEC) code, or other error correction code (ECC). For example, to handle high error rates in 64.0 GT / s PAM-4 signaling, PCIe 6.0 introduces flit mode to PCIe and defines a 256B flit size with a 3-way interleaved forward error correction (FEC) mechanism (e.g., 2 bytes each, for a total of 6 bytes) and a powerful 8-byte Reed-Solomon code-based CRC mechanism. Other protocols, such as CXL and UPI, may also utilize flits (albeit with different defined formats), which include error detection and correction codes. Indeed, versions of CXL (e.g., CXL3) and versions of UPI (e.g., UPI3) are defined to utilize PCIe PHYs (e.g., PCIe 6.0 PHYs to operate at 64.0 GT / s), but have 128B levels of sub-flits with 6B CRCs to maintain zero latency impact.
[0078] 11A-11D are diagrams 1100a-d illustrating flit configurations for various protocol implementations. For example, FIG. 11A shows a representation of an exemplary PCIe flit mode flit. A flit may have a defined length of 256 B with a single error detection code provided for the entire flit. For example, flit 1105 may include 236 B of transaction layer packet (TLP) data 1110 and 6 B of data link layer packet (DLLP) data 1115, with a single error detection code (e.g., an 8 B CRC 1120) and a forward error correction code (e.g., a 6 B FEC 1125) calculated for the entire flit. Similarly, as shown in the example of FIG. 11B, a CXL flit format may have a 256 B flit length, including a 2 B flit header 1130 and 240 B of data 1135, covered by a single error detection code (e.g., a CRC 1140 and an FEC 1145). In some implementations, forward error correction may be performed (using an FEC code (e.g., 1145)) only if a bit error is determined from the CRC. This may be advantageous in that the latency introduced through forward error correction may be avoided in all cases except when it is determined to be necessary (e.g., when an error is detected in a flit based on the CRC code). Among other example implementations, forward error correction may be performed for every flit (e.g., with a CRC applied after the FEC) in other cases (e.g., in more latency-tolerant applications).
[0079] One drawback of a single error detection code covering the entire flit is that a receiving device (e.g., using corresponding protocol circuitry (e.g., logical PHY circuitry) provided at a port of the device) must wait to receive all bytes of the flit in order to perform error checking on the corresponding error detection code (e.g., CRC) in the flit. This cumulative latency may be unacceptable in some applications. Therefore, some implementations may employ an alternative latency-optimized flit format, such as that shown in FIGS. 11C-11D, in which multiple error detection codes are provided within a flit. For example, as shown in the example CXL latency-optimized flit format shown in FIG. 11C, a flit may be logically subdivided into sub-flits (e.g., a first sub-flit 1150 and a second sub-flit 1155), and a respective error detection code may be provided for each sub-flit 1150, 1155. For example, the error detection code provided for the first sub-flit 1150 may be CRC 1160, and the error detection code provided for the second sub-flit 1155 may be CRC 1165 and FEC 1170. A latency-optimized flit may reduce accumulation latency because a CRC check can be performed after only the corresponding sub-flit is received, allowing for immediate consumption of the preceding data (e.g., of the first sub-flit 1150) if the CRC check does not indicate an error on this data. Figure 11D shows a similar latency-optimized flit structure (e.g., for UPI3) in which a 512B flit is subdivided into four sub-flits, each with a unique respective CRC code. In the examples of Figures 11C and 11D, in some implementations, forward error correction for the entire flit may be skipped if each of the CRC checks for each of the sub-flits does not indicate an error condition. However, among other example implementations, if one or more of the CRC checks indicate an error in the associated sub-flit, forward error correction (e.g., using FEC code 1170) may be used to attempt to correct the error.
[0080] In some implementations, the interconnect link protocol state machine may define multiple active link states. When a link is operating in a full-width active link state (e.g., L0), it utilizes all of the physical lanes available to the link. In some implementations, a partial width link state may provide a power-saving active link state that will utilize (e.g., temporarily) less than all of the available lanes or link width. In a partial width link state (e.g., partial L0 or L0p state), power savings may be achieved by entering a partial width state in which one or more available lanes of the link are idled (e.g., from a full-width link state or even an idle link state). Asymmetric partial width refers to each direction of a bidirectional link having different widths, which may be supported in some designs. During a partial width state, the transmission rate may remain the same, but the link width is reduced, thereby reducing the total throughput rate of the link. Furthermore, in partial width link states, flits are transmitted at a narrower link width, resulting in longer transmission times than if the flits were transmitted in a wider link state (e.g., L0 or wider partial width link state).
[0081] Therefore, the latency accumulation problem discussed above (arising from the time it takes to accumulate data corresponding to CRC codes) is exacerbated in partial-width link states, as the time to accumulate data increases as the link width is reduced. In fact, in L0p, power savings are achieved at the expense of latency. For example, a x16 link with native 256B flits has an accumulation latency of 2 ns at 64.0 GT / s, and a sub-flit accumulation latency of 1 ns for 128B sub-flits. The same link in L0p has accumulation latencies of 4 / 8 / 16 / 32 ns for x8 / x4 / x2 / x1 widths on 256B native flits, and 2 / 4 / 8 / 16 ns for x8 / x4 / x2 / x1 widths on 128B sub-flits. Thus, for example, for a x16 link at 64.0 GT / s, a 256B flit would have an added additional latency of 31 ns when going from a x16 link width to a x1 width. This latency impact is intolerable for coherency or memory protocols where most transactions are contained within 14B or 16B slots. For example, a cache line data transfer consists of four back-to-back 16B slots. Therefore, among other exemplary advantages, an improved implementation may support a special flit mode that utilizes a flit format adapted to address the latency challenges introduced through transmission on fewer lanes during partial link conditions.
[0082] As introduced above, in some implementations, when a link transitions to a partial-width link state utilizing fewer lanes than a threshold (and fewer lanes than would be utilized in a full-width active link state), the port's protocol logic may transition from a first flit format to a different second flit format during the partial-width state. More specifically, if the first flit format provides a first number of error-detecting codes (e.g., CRC codes) to protect the flit's data, the second flit format provides a larger second number of error-detecting codes to protect a larger number of relatively small segments (e.g., sub-flits) of the flit's data. As with latency-optimized flits, increasing the prevalence and frequency of error-detecting codes specifically during a partial-width link state can be utilized to alleviate accumulated latency pressure exacerbated by reducing the link's active lanes.
[0083] As an illustrative example, FIGS. 12A-12C show an exemplary implementation of a partial-width link state flit format. During certain partial-width link conditions, such a specialized flit format may be utilized, and upon exiting the corresponding partial-width link state, the link partners transition back to a different standard or latency-optimized flit format. For example, FIG. 12A shows an exemplary partial-width state flit format. In the example of FIG. 12A, a latency-optimized CXL flit format may be utilized when the link is active at full width (or at a relatively high partial link width in a partial-width state). The exemplary flit format depicted in FIG. 12A may represent a partial-width state flit format corresponding to a latency-optimized CXL flit format (e.g., similar to that shown in the example of FIG. 11C). FIG. 12A depicts the partial-width state flit format in tabular form 1200a, with each cell in table 1200a representing two bytes of data for the flit. The cells labeled S0 through S14 represent the bytes of numbered slots 0 through 14 of the CXL flit slot. For example, slot 0 (S0) may be configured with 14 bytes, slot 1 (S1) with 16 bytes, and so on. The flit may include a CXL flit header (e.g., 1205). Additionally, seven instances of a 6B CRC error detection code (e.g., crc0 through crc6) are provided for every 34B of slot data. For example, to cover two slots, a 6B CRC is applied every 40 bytes (34 bytes of information + 6B CRC). The same H matrix as in (6) is used, but with 96 columns from the right to form a 3-way interleaved FEC. The payload size of this exemplary partial-width status flit may be identical to its full-width counterpart (e.g., the exemplary latency-optimized flit in FIG. 11C) to ensure that the retry mechanism also operates seamlessly as the link width changes dynamically. In the example of Figure 12A, slots 0 and 8 are 14B each; the remaining 13 slots are 16B each. Slot 8 is a high latency slot because it takes a full flit to aggregate.The remaining slots have accumulated latencies of 14B to 40B, which is <2ns to 5ns even for a single lane (x1) link. Note that even in full-width mode with 128B sub-flits, the accumulated latency can be up to 2ns, 4ns, and 8ns for x16, x8, and x4 link widths, respectively.
[0084] In the example of FIG. 12A , the increased (e.g., 7) CRC instances represent a roughly four-fold increase in the occurrence of CRC codes within the partial-width flit format compared to a corresponding CXL latency-optimized flit format (e.g., shown in FIG. 11C ) that includes only two CRC code instances, one protecting 40 B of data and the other protecting 128 B of data. However, as a result of providing more frequent error detection codes, the length of the partial-width flit format is longer than the corresponding flit format transmitted in a wider link width state (e.g., 288 B vs. 256 B). During operation, a port may utilize a first flit format (e.g., of FIG. 11C ) and, upon detecting a transition to a particular partial-width state, transition to instead using the corresponding partial-width flit format (e.g., of FIG. 12A ) while in the particular partial-width state. When the particular partial-width state ends, the port and link may revert to utilizing a more “standard” flit format (e.g., of FIG. 11C ).
[0085] FIG. 12B is a tabular representation 1200b of a partial-width state flit format that represents a partial-width-optimized translation of a more standard UPI3 latency-optimized flit format (e.g., of FIG. 11D). For example, for a 512-B standard flit format, the corresponding partial-width state (L0p) flit format could be 576 B, using a layout similar to that of the 288-B L0p flit format shown in the example of FIG. 12A. For example, the first 13 sub-flits are 40 B each (34 B information, 6 B CRC), followed by 50 B sub-flits each (40 B information, 4 B reserved, 6 B CRC), accommodating a 480 B payload (+ 2 B flit header) and 6 B interleaved FEC. The partial-width state flit format may be defined to terminate on a clean flit or symbol boundary. Among other example features, a reserved field may be provided to define the partial-width state flit format at a clean flit or symbol boundary.
[0086] In yet another example of a partial-width state flit format, Figure 12C shows a representation 1200c of a partial-width-optimized version of a non-latency-optimized 256B flit for PCIe 6.0 or CXL (see Figures 11A-11B). The partial-width state flit format in this example may be composed of four sub-flits (e.g., 1250, 1255, 1260, 1270), with an 8B CRC provided for each of the 64B TLP sub-flits (e.g., 1250, 1255, 1260) and the final sub-flit 1270 having a 44B TLP, a 6B DLP, an 8B CRC, and a 6B FEC (expanding the flit format to 280B). Thus, the first three sub-flits are 72B each, and the final sub-flit is 64B. In a CXL implementation, among other example implementations, the flit header may occupy the first 2 B of the first TLP sub-flit 1250.
[0087] Support for a special partial-width-optimized flit format may be an optional feature, and link partners may first determine whether such flit format is mutually supported by the link partner port and under what conditions before utilizing such flit format. For example, a configuration register may identify whether a particular device supports the partial-width-optimized flit format (and other optional flit formats, such as the latency-optimized flit format). The configuration information in the register may further identify the conditions under which the partial-width-optimized flit format will or may be used. In some cases, a BIOS, operating system, or other software-based controller may define values in configuration registers to control whether and under what conditions the partial-width-optimized flit format will be deployed. In some implementations, among other features, link partner devices may advertise their inherent capabilities and / or discover the respective capabilities of their link partners during link training (e.g., during alternate protocol negotiation), including whether the partial-width-optimized flit format is supported and what threshold partial link width should trigger the use of the partial-width-optimized flit format (e.g., x4, x2, x1 link width). Negotiation of terms regarding the use of partial width optimized flit formats on a link may be performed through the exchange of training sequences (e.g., TS1, TS2, or other training sequences) or ordered sets (e.g., SKP Ordered Sets (SKP OS)). For example, specific fields or bits within the training sequences or ordered sets may be defined to hold information identifying support for partial width optimized flit formats, partial width threshold information, and other information usable to implement such features during operation of the link.
[0088] A mode or link state may be enabled using various techniques in which, in response to entering a partial-width state (or even a partial-width state utilizing a link width lower than a set threshold link width), the link partner automatically identifies the link width change and adjusts the flit format used from the standard or default flit format to a corresponding partial-width-optimized flit format while the link is in this partial-width operating mode. After this partial-width operating mode ends (and a link width above a defined threshold is used), the link partner may automatically revert to using the standard flit format. For example, FIG. 13A is a flow diagram 1300a illustrating an exemplary technique for transitioning between the use of the standard flit format and the partial-width-optimized flit format. For example, upon negotiating support and use of the partial-width-optimized flit format on a link, the link may be trained for an active or operational link state. An active link state may support operation up to a specific number of lanes (supported by the physical layer of the link). A partial-width link state may utilize fewer lanes than this maximum number for transmitting and / or receiving data. For example, a link may be configured to operate at a width of 16 lanes (x16), but in a partial-width link state utilizes only eight (x8), four (x4), two (x2), or one (x1) lane. A partial-width link state may be entered 1302, and it may be determined 1304 whether the link should operate at one of its supported partial link widths. Ordered sets or other data may be communicated between the link partners to negotiate the transition to the partial-width link state and the link width to be used.
[0089] If the identified partial link width is less than a threshold (or less than or equal to, depending on the definition of the threshold) 1305, the link partners may each identify that a modified flit format adapted for partial-width operation should be used and may transition from using a standard flit format (e.g., a default flit format or a latency-optimized flit format adapted for use during full-link-width or higher link-width operation) to the modified flit format 1306. If the partial-width link state utilizes a link width that exceeds the threshold partial-width threshold, the link partners may continue (at 1310) to send and receive data over the link using the standard flit format (e.g., negotiated by the link partners during training). When in modified flit mode, flits may be generated and sent over the link (at 1308) according to the modified flit format, and the link partners may receive the modified flit format and accurately analyze and process the flits based on the link partners' mutual transition to using the modified flit format during the partial-width link state.
[0090] The link partners may ultimately determine whether the partial-width link state should be changed by either exiting the partial-width link state for a full-width link state (e.g., L0) or increasing the link width used in the partial-width link state, and the link partners may negotiate this transition (e.g., 1312). Among other examples, the flit format used on the link following transition 1312 may change in some cases, such as when the conditions for using the changed flit format are no longer met, causing the link partner to transition back to using the standard flit format. Following this transition, flits in this format may be transmitted and received (e.g., 1314). In some implementations, the latency for exiting or transitioning out of a particular partial-width link state (e.g., L0p) may be reduced by, for example, beginning the transition to the higher link width as soon as the idle lanes are trained to the next flit boundary (e.g., aligned to a 16B boundary) rather than waiting for the next scheduled SKP OS (e.g., within about 1.5 us). In one example, for x1 and x2 link widths, 288B flits may always be aligned to 16B boundaries. For x4 widths, boundaries may exist on a per-flit basis. Using the starting data stream ordered set, lane-to-lane deskew can be obtained across those lanes that will be activated, aligning them to flit boundaries on the active lanes. This results in an average width-increasing latency reduction of 750 ns in this example, among other exemplary implementations.
[0091] In addition to supporting modified flit formats to help reduce link latency experienced in partial-width link states, to mitigate selective replay overhead on coherency and memory interconnects, devices (e.g., in software or hardware) may implement opportunistic self-throttling of the link to mitigate the presence of selectively replayed flits. When throttling, the protocol layer sends idle flits even if it may have several transactions to be sent in that flit's slot. Throttling allows the receiver to catch up from its receiver replay buffer. For example, some systems may support or implement a selective replay mechanism to help improve link bandwidth utilization. Traditionally, when an erroneous flit or packet is received, a replay request is sent by the receiver to trigger replay of not only the erroneous flit but also any other flits sent by the link partner device after the erroneous flit. With selective replay, the receiver may instead request replay of only the erroneous flit and buffer flits received after this flit, thus avoiding "wasting" link bandwidth by otherwise retransmitting error-free flits (i.e., flits following the erroneous flit that would have been replayed under traditional replay). However, this optimization of link bandwidth does not come without a trade-off. Link latency may be adversely affected during selective replay, since the receiving device may not be able to consume the buffered, correct flit until the preceding replayed flit is received. The added latency due to this processing delay may last until placeholder data (e.g., a SKP OS symbol, an idle flit, or other non-substantive data) is received on the link, allowing the receiver to "catch up" in its processing of the buffered flits.
[0092] In clearing the retry buffer that accumulates flits during selective retries, idle flits, NOP flits, periodic SKP OS, etc. may present the receiver with an opportunity to catch up, allowing the receiver replay buffer to decay over time, but there is a risk that another error resulting from another selective replay will occur while the next idle flit or SKP OS is awaited and before the retry buffer clears. Accordingly, this additional selective retries can compound the latency experienced at the receiver. Accordingly, depending on the error rate and traffic pattern on the link (e.g., the link is not fully utilized but has very few idle flits), the delay through the replay buffer may be very high.
[0093] In some implementations, a system may attempt to reduce link latency introduced through selective replay if the transmitter throttles its flit transmission rate after transmitting a selective replay flit. For example, when a receiver issues a selective replay request (e.g., a selective Nak) and is waiting for a selective replay, it stores valid payload flits in an RX retry buffer and processes these flits after the selective replay is received. Payload flits that arrive after a selective replay must also be stored in the RX retry buffer (e.g., in a system where all payload flits are to be processed in sequence number order). Without transmitter throttling, the RX retry buffer must consume all entries in order, adding latency to the receive path. Throttling may involve the transmitter selectively and strategically withholding transmission of substantial payload flits in favor of placeholder data (e.g., idle flits, SKP OS, etc.), which accelerates the receiver's ability to catch up with received flits and drain its RX retry buffer. Latency can be reduced and this performance impact mitigated if the transmitter throttles flit transmission earlier to drain the RX retry buffer. This early transmitter throttling does not affect overall performance because flits (e.g., carrying TLPs) delayed at the transmitter will be delayed at the receiver anyway due to the processing backlog of flits stored in the RX retry buffer. Therefore, performing early throttling at the transmitter drains the RX retry buffer, which reduces latency for subsequent TLPs.
[0094] Because of selective retry, the receiver can only catch up on idle flits, both for the latency-optimized path and for storage and forwarding in the receiver retry buffer. If the link has low to medium utilization, the receiver retry buffer may have, for example, 100-200 ns worth of traffic (replay round-trip delay), which translates into additional latency for all subsequent traffic from an erroneous flit until 100-200 ns of idle flits or SKP ordered sets are communicated on the link (assuming no new errors occur). The transmitter may track the number of non-idle flits in the receiver replay buffer. Throttling may be tailored to the receiver's drain rate. For example, if the link is operating at a reduced bandwidth (e.g., due to L0p), the RX retry buffer drain rate may be faster than the link rate, so throttling may not be necessary. In some implementations, a software controller may determine or predict traffic patterns and / or error rates affecting the link and trigger throttling in response to selective replay based on these detected patterns. For example, upon receiving a selective replay request, a certain bandwidth efficiency may be enforced on each new non-idle flit sent by the transmitter, so that the receiver retry buffer has a chance to process the buffered flit and allow the receiver to bypass it.
[0095] In some implementations, a software-based controller and / or port protocol hardware may implement logic to calculate a beneficial throttling rate and / or the timing of such throttling. For example, the controller or transmitter logic may assume that the initial link width is the receiver's drain rate. Upon receiving a selective NAK, the transmitter, knowing how many payload flits are currently in its TX retry buffer, can calculate the number of payload flits in (or en route to) the receiver's RX retry buffer. The transmitter can use this information to throttle or enforce certain bandwidth efficiencies (e.g., limit the number of no-op (NOP) TLPs in a payload flit) on each new payload flit it sends from its transaction layer, so as to drain the receiver's RX retry buffer.
[0096] As an illustrative example, two devices may be connected by a link (e.g., an x8 link), the first device may receive a selective NAK for a particular flit (e.g., flit number X), and the transmitter of the first device may transmit an additional four payload flits after initially transmitting flit number X. In this case, the transmitter of the first device or other slotting logic may detect or estimate the latency added by RX replay buffering in the second device and cause its upper protocol layer (e.g., the transaction or protocol layer) to reorder the data transmitted on the link to create a "bubble" (e.g., a 4-flit bubble) of limited activity on the link corresponding to the predicted latency experienced on the link. For example, with respect to an x8 link, in one exemplary implementation, a flit may be 16 symbols, and a SKP ordered set may be 32 symbols, and thus a 4-flit "bubble" (or potentially a bubble of any size resulting from a selective replay event) may be constructed using placeholder data such as any of 4 NOP or idle flits, 2 NOP or idle flits, and 1 SKP ordered set, or 2 SKP ordered sets. Among other exemplary implementations, in one example, the slotting logic attempts to perform slotting in its transmitter such that it enforces <Y% NOP TLP (e.g., Y = 25) in its payload flits until it transmits a "bubble", thereby allowing the flits in its replay buffer to be processed and cleared by the other device before additional payload flits are transmitted (e.g., this may take 2 SKP ordered sets if TLP is being transmitted constantly).
[0097] FIG. 13B is a simplified flow diagram 1300b illustrating a technique utilized to throttle a link in response to a selective replay event. A selective flit replay request may be identified 1324. In response, logic may attempt to identify the status of the receiver's replay buffer at the other end of the link (e.g., 1330). Additionally, conditions under which throttling is enabled may be defined (e.g., by software that calculates error rates, bandwidth utilization, etc. and compares this to thresholds). The logic may track transmissions of idle flits (e.g., at 1332), SKP OS (e.g., 1336), and other placeholder data that do not add to the backlog of payload data or other substantive data that will in turn be consumed by the receiver, and that represent opportunities for the receiver to consume flits stored in its retry buffer and gradually clear the buffer. Even while operating to clear the retry buffer (e.g., at 1326), additional selective flit replay requests may be detected. Receipt of such data may cause the logic to adjust its understanding of retry buffer latency (e.g., at 1338, 1340). Among other example implementations, if throttling is determined to have cleared the retry buffer (e.g., to terminate selective replay), throttling may be turned off (e.g., at 1328) until another eligible selective retry event is received.
[0098] It should be noted that the above apparatus, methods, and systems may be implemented in any electronic device or system, as described above. By way of specific illustration, the following figures provide exemplary systems for utilizing the solutions described herein. As the systems below are described in more detail, numerous different interconnects are disclosed, described, and reviewed from the above description. While some of the above examples are based on a UPI system, it should be understood that the solutions and features discussed above can be just as easily applied to other cache-coherent interconnects used to link sockets, packages, boards, etc. within various computing platforms. As will be readily apparent, the advances described above can be applied to any of the interconnects, fabrics, or architectures discussed herein, as well as other equivalent interconnects, fabrics, or architectures not explicitly or specifically described herein.
[0099] Referring to Figure 14, an embodiment of a block diagram for a computing system including a multi-core processor is depicted. Processor 1400 may include any processor or processing device, such as a microprocessor, embedded processor, digital signal processor (DSP), network processor, handheld processor, application processor, co-processor, system-on-chip (SOC), or other device for executing code. In one embodiment, processor 1400 includes at least two cores, namely, cores 1401 and 1402, which may include asymmetric cores or symmetric cores (in the illustrated embodiment). However, processor 1400 may include any number of processing elements, which may be symmetric or asymmetric.
[0100] In one embodiment, a processing element refers to hardware or logic that supports a software thread. Examples of hardware processing elements include a thread unit, thread slot, thread, processing unit, context, context unit, logical processor, hardware thread, core, and / or any other element capable of maintaining processor state, such as execution state or architectural state. In other words, in one embodiment, a processing element refers to any hardware that can be associated with code, such as a software thread, an operating system, an application, or independently other code. A physical processor (or processor socket) typically refers to an integrated circuit that potentially contains any number of other processing elements, such as cores or hardware threads.
[0101] A core often refers to logic located on an integrated circuit capable of maintaining independent architectural states, where each independently maintained architectural state is associated with at least some dedicated execution resources. In contrast to multiple cores, a hardware thread typically refers to any logic located on an integrated circuit capable of maintaining independent architectural states, where multiple such independently maintained architectural states share access to multiple execution resources. As will be appreciated, the nomenclature boundaries between hardware threads and cores overlap when certain resources are shared and other resources are dedicated to architectural state. However, cores and hardware threads are often viewed by the operating system as individual logical processors, and the operating system can schedule operations on each logical processor independently.
[0102] As shown in FIG. 14 , physical processor 1400 includes two cores, namely, core 1401 and 1402. Here, cores 1401 and 1402 are considered to be symmetric cores, i.e., cores with the same configuration, functional units, and / or logic. In another embodiment, core 1401 includes an out-of-order processor core, and core 1402 includes an in-order processor core. However, cores 1401 and 1402 may be individually selected from any type of core, such as a native core, a software-managed core, a core adapted to execute a native instruction set architecture (ISA), a core adapted to execute a translated instruction set architecture (ISA), a co-designed core, or other known core. In a heterogeneous core environment (i.e., an asymmetric core), some form of translation, such as binary translation, may be utilized to schedule or execute code on one or both cores. However, in the illustrated embodiment, the functional units shown in core 1401 are described in more detail below for further discussion, as the units in core 1402 operate similarly.
[0103] As shown, core 1401 includes two hardware threads 1401a and 1401b, which may also be referred to as hardware thread slots 1401a and 1401b. Thus, in one embodiment, software entities, such as an operating system, potentially view processor 1400 as four separate processors, i.e., four logical processors or processing elements, capable of simultaneously executing four software threads. As noted above, a first thread may be associated with multiple architecture state registers 1401a, a second thread may be associated with multiple architecture state registers 1401b, a third thread may be associated with multiple architecture state registers 1402a, and a fourth thread may be associated with multiple architecture state registers 1402b. Here, each of the architecture state registers (1401a, 1401b, 1402a, and 1402b) may also be referred to as a processing element, thread slot, or thread unit, as described above. As shown, architecture state registers 1401a are replicated in architecture state registers 1401b so that individual architecture states / contexts can be stored for logical processor 1401a and logical processor 1401b. In core 1401, other smaller resources, such as instruction pointers and renaming logic in allocator and renamer block 1430, may also be replicated for threads 1401a and 1401b. Some resources, such as the reorder buffer, ILTB 1420, load / store buffers, and queues in reorder / retirement unit 1435, can be shared through partitioning. Other resources, such as general-purpose internal registers, page table base registers, low-level data cache and data TLB 1415, execution unit 1440, and portions of out-of-order unit 1435, are potentially fully shared.
[0104] Processor 1400 often includes other resources, which may be fully shared, shared through partitioning, or dedicated by / to processing elements. In FIG. 14, one embodiment of a purely exemplary processor is shown, with exemplary logical units / resources associated with the processor. Note that the processor may include or omit any of these functional units, as well as any other known functional units, logic, or firmware not shown. As shown, core 1401 includes a simple, representative out-of-order (OOO) processor core. However, different embodiments may utilize an in-order processor. The OOO core includes a branch target buffer 1420 for predicting branches to be executed / taken, and an instruction translation buffer (I-TLB) 1420 for storing address translation entries for instructions.
[0105] The core 1401 further includes a decode module 1425 coupled to the fetch unit 1420 for decoding fetched elements. In one embodiment, the fetch logic includes separate sequencers associated with each of the thread slots 1401a, 1401b. Typically, the core 1401 is associated with a first ISA that defines / specifies instructions executable on the processor 1400. Often, the machine code instructions that are part of the first ISA include a portion of the instruction (called an opcode) that references / specifies the instruction or operation to be performed. The decode logic 1425 includes circuitry that recognizes these instructions from their opcodes and passes the decoded instructions down the pipeline for processing defined by the first ISA. For example, as described in more detail below, in one embodiment, the decoders 1425 include logic designed or adapted to recognize specific instructions, such as transactional instructions. As a result of recognition by decoder 1425, architecture or core 1401 performs specific, predefined operations to perform tasks associated with the appropriate instruction. It is important to note that any of the tasks, blocks, operations, and methods described herein may be performed in response to a single or multiple instructions; some of these may be new or old instructions. Note that in one embodiment, multiple decoders 1426 recognize the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, decoder 1426 recognizes a second ISA (either a subset of the first ISA or a separate ISA).
[0106] In one example, allocator and renamer block 1430 includes an allocator for reserving resources such as a register file for storing instruction processing results. However, threads 1401a and 1401b are potentially capable of out-of-order execution, and allocator and renamer block 1430 also reserves other resources such as a reorder buffer for tracking instruction results. Unit 1430 may also include a register renamer for renaming program reference registers / instruction reference registers to other registers within processor 1400. Reorder / retirement unit 1435 includes components such as the reorder buffer and load buffer described above to support out-of-order execution and subsequent storage buffers to support in-order retirement of instructions that were executed out-of-order.
[0107] In one embodiment, scheduler and execution units block 1440 includes a scheduler unit for scheduling instructions / operations on the execution units. For example, floating-point instructions are scheduled on ports of execution units that have available floating-point execution units. Register files associated with these execution units are also included for storing information instruction processing results. Exemplary execution units include floating-point execution units, integer execution units, jump execution units, load execution units, store execution units, and other known execution units.
[0108] A lower-level data cache and data translation buffer (D-TLB) 1450 is coupled to the execution unit 1440. The data cache stores recently used / operated items, such as data operands potentially held in multiple memory coherency states, on multiple elements. The D-TLB stores recent virtual / linear address to physical address translations. As a specific example, a processor may include a page table structure that divides physical memory into multiple virtual pages.
[0109] Here, cores 1401 and 1402 share access to a higher-level or further cache, such as a second-level cache associated with on-chip interface 1410. Note that higher-level or further-out refers to cache levels increasing or further away from the execution units. In one embodiment, the higher-level cache is a last-level data cache, i.e., the last cache, such as a second- or third-level data cache, in the memory hierarchy of processor 1400. However, a higher-level cache is not so limited, as it may be associated with or include an instruction cache. Rather, a trace cache, which is a type of instruction cache, may be coupled after decoder 1425 to store multiple recently decoded traces. Here, instruction potentially refers to a macroinstruction (i.e., a general instruction recognized by a decoder), which can be decoded into several microinstructions (micro-operations).
[0110] In the illustrated configuration, processor 1400 also includes an on-chip interface module 1410. Historically, a memory controller, described in more detail below, has been included within a computing system external to processor 1400. In this scenario, on-chip interface 1410 communicates with multiple devices external to processor 1400, such as system memory 1475, a chipset (which often includes a memory controller hub that connects to memory 1475 and an I / O controller hub that connects to multiple peripheral devices), a memory controller hub, a northbridge, or other integrated circuit. Also in this scenario, bus 1405 can include any known interconnect, such as a multi-drop bus, a point-to-point interconnect, a serial interconnect, a parallel bus, a coherent (e.g., cache coherent) bus, a layered protocol architecture, a heterogeneous bus, and a GTL bus.
[0111] Memory 1475 may be dedicated to processor 1400 or shared with other devices in the system. Common examples of types of memory 1475 include DRAM, SRAM, non-volatile memory (NV memory), and other known storage devices. Note that devices 1480 may include a graphics accelerator, a card coupled to a processor or memory controller hub, data storage coupled to an I / O controller hub, a wireless transceiver, a flash device, an audio controller, a network controller, or other known devices.
[0112] Recently, however, as more logic and devices are integrated on a single die, such as in a SOC, each of these devices may be integrated onto the processor 1400. For example, in one embodiment, a memory controller hub resides on the same package and / or die as the processor 1400. Here, part of the core (on-core portion) 1410 includes one or more controllers for interfacing with other devices, such as memory 1475 or graphics device 1480. Configurations including interconnects and controllers for interfacing with such devices are often referred to as on-core (or un-core configurations). As an example, the on-chip interface 1410 includes a ring interconnect for on-chip communication and a high-speed serial point-to-point link 1405 for off-chip communication. However, in a SOC environment, much more devices, such as a network interface, coprocessor, memory 1475, graphics processor 1480, and any other known computing device / interface, may be integrated onto a single die or integrated circuit to provide a small form factor with high functionality and low power consumption.
[0113] In one embodiment, processor 1400 is capable of executing compiler, optimization, and / or transformation code 1477 to compile, transform, and / or optimize application code 1476 to support or interface with the apparatus and methods described herein. A compiler often includes a program or set of programs that transform source text / code into target text / code. Compilation of program / application code using a compiler typically occurs in multiple phases and passes to convert high-level programming language code into low-level machine or assembly language code. However, single-pass compilers can still be used for simple compilation. The compiler can use any known compilation techniques and perform any known compiler operations, such as lexical analysis, preprocessing, parsing, semantic analysis, code generation, code transformation, and code optimization.
[0114] Larger compilers often include multiple phases, but in most cases these phases are contained within two general phases: (1) the front-end, i.e., generally where syntactic processing, semantic processing, and some transformations / optimizations may occur, and (2) the back-end, i.e., generally where analysis, transformations, optimizations, and code generation occur. Some compilers are referred to as "middle," indicating that the distinction between the front and back ends of a compiler is blurred. As a result, references to insertion, association, generation, or other compiler operations may occur in any of the above phases or passes, as well as in any other known phases or passes of a compiler. As an illustrative example, a compiler may potentially insert multiple operations, calls, functions, etc., within one or more phases of compilation, such as inserting multiple calls / operations within the front-end phase of compilation and then translating those calls / operations into lower-level code during the transformation phase. Note that during dynamic compilation, compiler code or dynamic optimization code may insert such operations / calls and optimize the code for execution during runtime. As a specific illustrative example, binary code (already compiled code) may be dynamically optimized during runtime, where program code may include dynamically optimized code, binary code, or a combination thereof.
[0115] Similar to compilers, translators, such as binary translators, translate code either statically or dynamically to optimize and / or transform the code. Thus, references to the execution of code, application code, program code, or other software environments may refer to (1) the dynamic or static execution of a compiler program, code optimizer, or translator to compile program code, maintain software structures, perform other operations, optimize code, or transform code; (2) the execution of main program code that includes operations / calls, such as optimized / compiled application code; (3) the execution of other program code, such as libraries associated with the main program code, to maintain software structures, perform other software-related operations, or optimize code; or (4) combinations thereof.
[0116] Referring now to Figure 15, a block diagram of a second system 1500 is shown in accordance with an embodiment of the present disclosure. As shown in Figure 15, the multiprocessor system 1500 is a point-to-point interconnect system and includes a first processor 1570 and a second processor 1580 coupled via a point-to-point interconnect 1550. Each of the processors 1570 and 1580 may be some version of a processor. In one embodiment, 1552 and 1554 are part of a serial point-to-point coherent interconnect fabric, such as a high-performance architecture. As a result, the solutions described herein may be implemented within a UPI or other architecture.
[0117] Although shown with only two processors 1570, 1580, it should be understood that the scope of the present disclosure is not so limited. In other embodiments, there may be one or more additional processors within a given processor.
[0118] Processors 1570 and 1580 are shown to include integrated memory controller units 1572 and 1582, respectively. Processor 1570 also includes point-to-point (PP) interfaces 1576 and 1578 as part of its bus controller unit, and similarly, second processor 1580 includes PP interfaces 1586 and 1588. Processors 1570, 1580 can exchange information over point-to-point (PP) interface 1550 using PP interface circuits 1578, 1588. As shown in FIG. 15, IMCs 1572 and 1582 couple the processors to their respective memories, i.e., memory 1532 and memory 1534, which may be portions of main memory attached locally to the respective processors.
[0119] Processors 1570, 1580 exchange information with chipset 1590 via individual PP interfaces 1552, 1554 using point-to-point interface circuits 1576, 1594, 1586, 1598, respectively. Chipset 1590 also exchanges information with high-performance graphics circuit 1538 via interface circuit 1592 along high-performance graphics interconnect 1539.
[0120] A shared cache (not shown) may be included within either processor or external to both processors; however, it is connected to the processors via the PP interconnect so that local cache information of either or both processors can be stored in the shared cache when the processors are placed in a low power mode.
[0121] Chipset 1590 may be coupled to a first bus 1516 via an interface 1596. In one embodiment, first bus 1516 may be a bus such as a Peripheral Component Interconnect (PCI) bus, or a PCI Express bus or another third generation I / O interconnect bus, although the scope of the disclosure is not so limited.
[0122] As shown in FIG. 15 , various I / O devices 1514 are coupled to a first bus 1516 along with a bus bridge 1518, which couples the first bus 1516 to a second bus 1520. In one embodiment, the second bus 1520 comprises a low pin count (LPC) bus. In one embodiment, various devices are coupled to the second bus 1520, including, for example, a keyboard and / or mouse 1522, a communication device 1527, and a storage unit 1528, such as a disk drive or other mass storage device, which often contains instructions / code and data 1530. Additionally, an audio I / O 1524 is shown coupled to the second bus 1520. Note that other architectures are possible, where the included components and interconnect architectures vary. For example, instead of the point-to-point architecture of FIG. 15 , a system could implement a multi-drop bus or other such architecture.
[0123] While the solutions discussed herein have been described with respect to a limited number of embodiments, those skilled in the art will recognize numerous modifications and variations therefrom, and it is intended that the appended claims cover all such modifications and variations that fall within the true spirit and scope of the present disclosure.
[0124] A design may go through various stages, from creation to simulation to manufacturing. Data representing the design may represent the design in a number of ways. First, hardware may be represented using a hardware description language or another functional description language, as is useful in simulation. In addition, a circuit-level model using logic and / or transistor gates may be generated at some stages of the design process. Furthermore, most designs, at some stage, reach a level of data representing the physical placement of various devices in the hardware model. When conventional semiconductor manufacturing techniques are used, the data representing the hardware model may be data specifying the presence or absence of various features on different mask layers of a mask used to fabricate an integrated circuit. In any representation of the design, the data may be stored in the form of any machine-readable medium. Memory, or magnetic or optical storage such as a disk, may be a machine-readable medium that stores such information transmitted via optical or radio waves that are modulated or otherwise generated to transmit the information. When an electrical carrier wave indicating or carrying the code or design is transmitted, a new copy is created to the extent that copying, buffering, or transmission of the electrical signal is performed. Thus, a communications provider or network provider may store, at least temporarily, an article such as information encoded in a carrier wave on a tangible, machine-readable medium, embodying techniques according to embodiments of the present disclosure.
[0125] As used herein, a module refers to any combination of hardware, software, and / or firmware. As an example, a module includes hardware such as a microcontroller associated with a non-transitory medium that stores code adapted to be executed by the microcontroller. Thus, in one embodiment, reference to a module refers to hardware specifically configured to recognize and / or execute code held on the non-transitory medium. Furthermore, in another embodiment, the use of a module refers to a non-transitory medium containing code specifically adapted to be executed by a microcontroller to perform a plurality of predetermined operations. As can be inferred, in yet another embodiment, the term module (in this example) may refer to a combination of a microcontroller and a non-transitory medium. In many cases, the boundaries of multiple modules shown as separate typically vary and potentially overlap. For example, a first and second module may share hardware, software, firmware, or a combination thereof, but potentially maintain some independent hardware, software, or firmware. In one embodiment, use of the term logic includes hardware such as transistors, registers, or other hardware such as programmable logic devices.
[0126] In one embodiment, use of the phrase "configured to" refers to arranging, assembling, manufacturing, offering for sale, importing, and / or designing a device, hardware, logic, or element to perform a specified or determined task. In this example, a device or element thereof that is not in operation is still "configured to" perform said specified task if it is designed, coupled, and / or interconnected to perform said specified task. As a purely illustrative example, a logic gate may provide a 0 or a 1 during operation. However, a logic gate that is "configured" to provide an enable signal to a clock does not include all potential logic gates that can provide a 1 or a 0. Instead, the logic gate is coupled in some manner such that a 1 or 0 output enables said clock during operation. Note that use of the term "configured to" does not require operation, but instead focuses on the potential state of the device, hardware, and / or element, in which the device, hardware, and / or element is designed to perform a particular task when the device, hardware, and / or element is in operation.
[0127] Additionally, in one embodiment, the use of the phrases "to," "capable of / to," and / or "operable to" refers to a device, logic, hardware, and / or element that is designed to enable use of that device, logic, hardware, and / or element in a specified manner. It should also be noted above that in one embodiment, the use of the phrases "capable of" or "operable to" refers to an underlying state of a device, logic, hardware, and / or element, where the device, logic, hardware, and / or element is not operating but is designed to enable use of the device in a specified manner.
[0128] As used herein, a value includes any known representation of a number, state, logical state, or binary logical state. Often, the use of logic levels, logic values, or logical values, also referred to as 1 and 0, simply represents a binary logic state. For example, 1 refers to a high logic level and 0 refers to a low logic level. In one embodiment, a storage cell, such as a transistor or flash cell, may be capable of holding a single logical value or multiple logical values. However, other representations of values in computer systems are used. For example, the decimal number 10 may be represented as the binary value 1010 and as the letter A in hexadecimal. Thus, a value includes any representation of information that can be held in a computer system.
[0129] Furthermore, a state may be represented by a value or portion of a value. As an example, a first value, such as a logical one, may represent a default or initial state, while a second value, such as a logical zero, may represent a non-default state. Additionally, in one embodiment, the terms reset and set refer to default and updated values or states, respectively. For example, a default value potentially includes a high logical value, i.e., reset, while an updated value potentially includes a low logical value, i.e., set. Note that any combination of multiple values may be utilized to represent any number of states.
[0130] The above-described method, hardware, software, firmware, or code embodiments may be implemented via instructions or code stored on a machine-accessible, machine-readable, computer-accessible, or computer-readable medium that is executable by a processing element. A non-transitory machine-accessible / readable medium includes any mechanism that provides (i.e., stores and / or transmits) information in a form readable by a machine, such as a computer or electronic system. For example, non-transitory machine-accessible media include random access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; magnetic or optical storage media; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; other types of storage devices for holding information received from a transitory (propagated) signal (e.g., carrier wave, infrared signal, digital signal); etc., which should be distinguished from non-transitory media that can receive information therefrom.
[0131] The instructions used to program the logic to implement the example embodiments herein may be stored in a system's memory, such as DRAM, cache, flash memory, or other storage. Additionally, the instructions may be distributed over a network or using other computer-readable media. Thus, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including, but not limited to, a floppy disk, an optical disk, a compact disk, a read-only memory (CD-ROM), and a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, a flash memory, or any tangible machine-readable storage used to transmit information over the Internet via electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, computer-readable media includes any type of tangible, machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0132] The following examples relate to embodiments according to the present specification. Example 1 is an apparatus comprising a port, the port comprising: a transmitter; and protocol circuitry configured to generate a first flit according to a first flit format, the first flit format defining that a first number of error detection codes are to be provided for an amount of data to be transmitted in the first flit, the first flit to be transmitted by the transmitter over a link while the link is operating at a first link width, the link coupling the port to another computing device; and generating a second flit according to a second flit format based on the transition to the second link width, wherein the second flit is to be transmitted while the link is operating at the second link width, the second flit format defining that a second number of error detection codes are to be provided for a same amount of data to be transmitted in the second flit, wherein the second number is greater than the first number.
[0133] Example 2 includes the subject matter of Example 1, where the first link width utilizes a first number of active physical lanes and the second link width utilizes a second, fewer number of active physical lanes.
[0134] Example 3 includes the subject matter of any one of Examples 1-2, wherein the first flit format has a defined first length and the second flit format has a defined second length that is longer than the first length.
[0135] Example 4 includes the subject matter of any one of Examples 1-3, wherein the amount of data in the first flit holds first transaction layer data and the amount of data in the second flit holds second transaction layer packet data.
[0136] Example 5 includes the subject matter of any one of Examples 1-4, wherein the first number of error detection codes include respective cyclic redundancy check (CRC) codes for respective portions of the first flit, and each of the second number of error detection codes includes a respective CRC code for respective portions of the second flit.
[0137] Example 6 includes the subject matter of example 5, wherein the first flit format includes a first error correction code calculated based on the amount of data of the first flit, and the second flit format includes a second error correction code calculated based on the amount of data of the second flit.
[0138] Example 7 includes the subject matter of example 6, wherein the first error correcting code includes a first forward error correcting (FEC) code and the second error correcting code includes a second FEC code.
[0139] Example 8 includes the subject matter of any one of Examples 1-7, wherein the protocol circuitry is further adapted to transition from a first active link state to a partial-width active link state, the transition from the first link width to the second link width accompanying the transition from the first active link state to the partial-width active link state.
[0140] Example 9 includes the subject matter of any one of Examples 1-8, wherein the first flit format and the second flit format comply with an interconnect protocol.
[0141] Example 10 includes the subject matter of Example 9, wherein the interconnect protocol includes one of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or UltraPath Interconnect (UPI).
[0142] Example 11 is a method comprising: transmitting flits according to a first flit format when a link is operating at a first link width, wherein the first flit format includes a first number of cyclic redundancy check (CRC) codes to protect a first amount of data; transitioning the link from the first link width to a second link width, wherein the second link width comprises a fractional link width and the second link width is narrower than the first link width; and transmitting flits according to a second flit format when the link is operating at the second link width, wherein the second flit format includes a second number of CRC codes to protect a second amount of data, wherein the first amount of data is equal to the second amount of data and the second number of CRC codes is greater than the first number of CRC codes.
[0143] Example 12 includes the subject matter of Example 11, further comprising: determining that the second link width is narrower than a threshold link width; and determining, based on the second link width being narrower than the threshold link width, that the second flit format will be used instead of the first flit format.
[0144] Example 13 includes the subject matter of any one of Examples 11-12, further comprising: participating in training of the link, wherein the link couples a first device to a second device; and determining during training of the link that both the first device and the second device support use of the second flit format, wherein the second flit format is used at the second link width based on a determination of mutual support of the second flit format by the first device and the second device.
[0145] Example 14 includes the subject matter of Example 13, and further includes determining, during training of the link, a threshold link width corresponding to use of the second flit format, wherein the second flit format will be used if the link width of the link is narrower than the threshold link width.
[0146] Example 15 includes the subject matter of any one of Examples 11-14, wherein in a partial width state, both the first link width and the second link width are used.
[0147] Example 16 includes the subject matter of any one of Examples 11-15, wherein the first flit format includes a first error correction code calculated based on the amount of data of the first flit, and the second flit format includes a second error correction code calculated based on the amount of data of the second flit.
[0148] Example 17 includes the subject matter of Example 16, wherein the first error correcting code comprises a first forward error correcting (FEC) code and the second error correcting code comprises a second FEC code.
[0149] Example 18 includes the subject matter of any one of Examples 11-17, wherein the first flit format and the second flit format comply with an interconnect protocol.
[0150] Example 19 includes the subject matter of Example 18, wherein the interconnect protocol includes one of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or UltraPath Interconnect (UPI).
[0151] Example 20 is a system including means for carrying out the method according to any one of Examples 11-19.
[0152] Example 21 includes the subject matter of Example 20, wherein the means includes a non-transitory machine-readable medium having instructions stored thereon, the instructions being executable by the machine to cause the machine to perform at least a portion of the method of any one of Examples 11-19.
[0153] Example 22 illustrates a method for transmitting a first data unit to a first device over a link while the link is operating at a first link width, the first data units being formatted according to a first format, the first format defining that each of the first data units holds a particular amount of data and includes a first number of error detection codes for the particular amount of data; a method for transmitting a first data unit over a link while the link is operating at a first link width, the first data units being formatted according to a ... second data unit over a link while the link is operating at a first link width, the first data units being formatted according to a first format defining that each of the first data units holds a particular amount of data and includes a first number of error detection codes for the particular amount of data; a second device including circuitry such that: a larger number of lanes are used than in the second link width; and, while the link is operating at the second link width, transmitting second data units over the link to the first device, the second data units being formatted according to a second format, the second format defining that each of the second data units holds the specified amount of data and includes a second number of error detection codes for the specified amount of data, the second number being greater than the first number.
[0154] Example 23 includes the subject matter of Example 22, wherein the circuitry is further configured to determine that the second link width is narrower than a threshold link width defined for the link, and based on determining that the second link width is narrower than the threshold link width, the second format is used for data units transmitted over the link.
[0155] Example 24 includes the subject matter of any one of Examples 22-23, wherein the data units include flits.
[0156] Example 25 includes the subject matter of any one of Examples 22-24, wherein one of the first device or the second device includes a processor device.
[0157] Example 26 includes the subject matter of any one of Examples 22-24, wherein one of the first device or the second device includes an accelerator device.
[0158] Example 27 includes the subject matter of any one of Examples 22-26, wherein the first link width utilizes a first number of active physical lanes and the second link width utilizes a second, fewer number of active physical lanes.
[0159] Example 28 includes the subject matter of any one of Examples 22-27, wherein the amount of data in the first unit holds first transaction layer data and the amount of data in the second unit holds second transaction layer packet data.
[0160] Example 29 includes the subject matter of any one of Examples 22-28, wherein the first format has a defined first length and the second format has a defined second length that is longer than the first length.
[0161] Example 30 includes the subject matter of any one of Examples 22-29, wherein the first number of error detection codes include respective cyclic redundancy check (CRC) codes for respective portions of the first units, and each of the second number of error detection codes includes a respective CRC code for respective portions of the second units.
[0162] Example 31 includes the subject matter of Example 30, wherein the first format includes a first error correction code calculated based on the amount of data of the first unit, and the second format includes a second error correction code calculated based on the amount of data of the second unit.
[0163] Example 32 includes the subject matter of Example 31, wherein the first error correcting code includes a first forward error correcting (FEC) code and the second error correcting code includes a second FEC code.
[0164] Example 33 includes the subject matter of any one of Examples 22-32, where the protocol circuitry is further configured to transition from a first active link state to a partial-width active link state, and the transition from the first link width to the second link width accompanies the transition from the first active link state to the partial-width active link state.
[0165] Example 34 includes the subject matter of any one of Examples 22-33, wherein the first format and the second format comply with an interconnect protocol.
[0166] Example 35 includes the subject matter of Example 34, wherein the interconnect protocol includes one of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or UltraPath Interconnect (UPI).
[0167] Example 36 is a method comprising: determining characteristics of a link, where the link couples a first device to a second device; identifying a selective retry request to retry transmission of a particular flit in a stream of flits based on an error detected in the particular flit; determining latency accumulated through buffering of flits in a retry buffer based on the selective retry request; and throttling traffic on the link to assist in clearing the retry buffer.
[0168] Example 37 includes the subject matter of Example 36, wherein the link complies with a PCIe-based protocol.
[0169] Example 38 includes the subject matter of any one of Examples 36-37, wherein the characteristic includes a predicted bandwidth availability for the link.
[0170] Example 39 includes the subject matter of any one of Examples 36-38, wherein the characteristic includes a predicted error rate for the link.
[0171] Example 40 includes the subject matter of any one of Examples 36-39, wherein the properties are determined by software.
[0172] Example 41 is a system including means for carrying out the method according to any one of Examples 36-40.
[0173] Example 42 includes the subject matter of Example 41, wherein the means includes a non-transitory machine-readable medium having instructions stored thereon, the instructions being executable by the machine to cause the machine to perform at least a portion of the method of any one of Examples 36-40.
[0174] Throughout this specification, a reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0175] In the foregoing specification, the detailed description has been given with reference to certain exemplary embodiments. However, it will be apparent that various modifications or changes may be made to those embodiments without departing from the broader spirit and scope of the invention as set forth in the appended claims. Accordingly, the specification and drawings should be considered in an illustrative rather than a restrictive sense. Moreover, the above use of embodiment and other exemplary language does not necessarily refer to the same embodiment or the same example, but may refer to different individual embodiments and potentially the same embodiment.
Claims
1. 1. A device comprising a port, The port is a transmitter; and A protocol circuit comprising: generating a first flit according to a first flit format, wherein the first flit format defines that a first number of error detection codes are to be provided for an amount of data to be transmitted in the first flit, the first flit to be transmitted by the transmitter over a link while the link is operating at a first link width, wherein the link couples the port to another computing device; Identifying a transition of the link from the first link width to a second link width, where the second link width is narrower than the first link width; and generating a second flit according to a second flit format based on the transition to the second link width, wherein the second flit is to be transmitted while the link is operating at the second link width, and wherein the second flit format defines that a second number of error detection codes is to be provided for an amount of data that is the same as the amount of data of the first flit to be transmitted in the second flit, and wherein the second number is greater than the first number; Protocol Circuit and the first flit format has a defined first length, and the second flit format has a defined second length that is longer than the first length; The protocol circuitry negotiates with the other computing device the use of the second flit format on the link prior to the transition. Device.
2. 2. The apparatus of claim 1, wherein the first link width results in a first number of active physical lanes being utilized and the second link width results in a second, fewer number of active physical lanes being utilized.
3. 2. The apparatus of claim 1, wherein the amount of data in the first flit holds first transaction layer packet data and the amount of data in the second flit holds second transaction layer packet data.
4. 4. The apparatus of claim 1, wherein the first number of error detection codes comprises a cyclic redundancy check (CRC) code for a portion of the first flit, and the second number of error detection codes comprises multiple CRC codes for multiple portions of the second flit.
5. 5. The apparatus of claim 4, wherein the first flit format includes a first error correction code calculated based on the amount of data of the first flit, and the second flit format includes a second error correction code calculated based on the amount of data of the second flit.
6. The apparatus of claim 5 , wherein the first error correcting code comprises a first forward error correcting (FEC) code and the second error correcting code comprises a second FEC code.
7. 4. The apparatus of claim 1, wherein the protocol circuitry is further adapted to transition from a first active link state to a partial-width active link state, the transition from the first link width to the second link width accompanying the transition from the first active link state to the partial-width active link state.
8. The apparatus of any one of claims 1 to 3, wherein the first flit format and the second flit format comply with an interconnect protocol.
9. 10. The apparatus of claim 8, wherein the interconnect protocol comprises one of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or UltraPath Interconnect (UPI).
10. 1. A method performed by a first device, comprising: transmitting flits according to a first flit format when a link coupling the first device to a second device is operating at a first link width, wherein the first flit format includes a first number of cyclic redundancy check (CRC) codes to protect a first amount of data; transitioning the link from the first link width to a second link width, wherein the second link width comprises a partial link width, the second link width being narrower than the first link width; and transmitting flits according to a second flit format when the link is operating at the second link width, wherein the second flit format includes a second number of CRC codes for protecting a second amount of data, the first amount of data equals the second amount of data, and the second number of CRC codes is greater than the first number of CRC codes. Equipped with the first flit format has a defined first length, and the second flit format has a defined second length that is longer than the first length; The method further comprising, prior to the transitioning step, negotiating with the second device the use of the second flit format on the link.
11. determining that the second link width is less than a threshold link width; and determining that the second flit format will be used instead of the first flit format based on the second link width being less than the threshold link width; The method of claim 10 further comprising:
12. participating in training of said link; and determining, during training of the link, that both the first device and the second device support use of the second flit format, wherein, based on a determination of mutual support of the second flit format by the first device and the second device, the second flit format is used at the second link width. The method of claim 10 or 11, further comprising:
13. determining a threshold link width corresponding to use of the second flit format during training of the link, wherein the second flit format will be used if the link width of the link is less than the threshold link width. The method of claim 12 further comprising:
14. 12. The method of claim 10 or 11, wherein while in a partial width state, both the first link width and the second link width are used.
15. 12. The method of claim 10 or 11, wherein the first flit format includes a first error correction code calculated based on the first amount of data, and the second flit format includes a second error correction code calculated based on the second amount of data.
16. 16. The method of claim 15, wherein the first error correcting code comprises a first forward error correcting (FEC) code and the second error correcting code comprises a second FEC code.
17. The method of claim 10 or 11, wherein the first flit format and the second flit format comply with an interconnect protocol.
18. 20. The method of claim 17, wherein the interconnect protocol comprises one of Peripheral Component Interconnect Express (PCIe), Compute Express Link (CXL), or UltraPath Interconnect (UPI).
19. A system comprising the first device for performing the method according to claim 10 or 11.
20. 20. The system of claim 19, wherein the first device comprises a computer program that causes a processor to perform the method of claim 10 or 11.
21. a first device; and a second device coupled to the first device by a link, the second device comprising: transmitting first data units over the link to the first device while the link is operating at a first link width, wherein the first data units are formatted according to a first format, the first format defining that each of the first data units holds a particular amount of data and includes a first number of error detection codes for the particular amount of data; transitioning from the first link width to a second link width, where a greater number of lanes are used in the first link width than in the second link width; and transmitting second data units over the link to the first device while the link is operating at the second link width, wherein the second data units are formatted according to a second format, wherein the second format defines each of the second data units to hold the particular amount of data and to include a second number of error detection codes for the particular amount of data, wherein the second number is greater than the first number; A second device including a circuit such that Equipped with the first format has a defined first length, and the second format has a defined second length that is longer than the first length; The circuitry negotiates with the first device the use of the second format on the link prior to the transition. system.
22. 22. The system of claim 21, wherein the circuitry is further configured to determine that the second link width is narrower than a threshold link width defined for the link, and wherein based on a determination that the second link width is narrower than the threshold link width, the second format is used for data units transmitted on the link.
23. 23. The system of claim 21 or 22, wherein one of the first device or the second device includes a processor device.
24. 23. The system of claim 21 or 22, wherein one of the first device or the second device comprises an accelerator device.
Citation Information
Patent Citations
Method and system for flexible credit exchange in high performance fabric
JP2017506378A
Code Block Segmentation by LDPC Basis Matrix Selection
JP2020511030A
Characterizing and margining multi-voltage signal encoding for interconnects
US20210050941A1