Adjustable re-timer buffer
By using adjustable retimer buffers in computing systems, the problems of latency and power consumption in interconnect architectures are addressed, enabling efficient data transmission and frequency mismatch compensation in high-performance computing environments.
Patent Information
- Application Number
- CN201880014989.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-03-31
- Filing Date
- 2018-02-27
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2038-02-27
AI Technical Summary
As the complexity of components in computing systems increases, the complexity of interconnect architectures also increases, making communication between sockets and other devices more critical. Existing interconnect architectures struggle to effectively manage latency and latency-sensitive channel length issues at high frequencies, especially in high-performance computing environments where the use of retimers increases latency and power consumption.
An adjustable retimer buffer is used, which dynamically adjusts its size to adapt to different communication needs, reducing latency and optimizing power consumption. At the same time, it supports frequency mismatch compensation and elastic buffering to ensure the reliability of data transmission.
It effectively reduces interconnect latency, optimizes power consumption, and improves data transmission reliability and adaptability to frequency mismatch, making it suitable for complex interconnect architectures in high-performance computing environments.
Smart Images

Figure CN110366842B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to computing systems, and in particular, but not exclusively, to a re-timer in an interconnect. BACKGROUND
[0002] Advances in semiconductor processing and logic design have allowed an increase in the amount of logic that can exist on an integrated circuit device. As a corollary, computer system configurations have evolved from a single or multiple integrated circuits in a system to multiple cores, multiple hardware threads, and multiple logical processors existing on a single integrated circuit, as well as other interfaces integrated within such processors. A processor or integrated circuit typically comprises a single physical processor die, where the processor die can include any number of cores, hardware threads, logical processors, interfaces, memory, controller hubs, etc.
[0003] Due to the ability to pack greater processing power in smaller packages, the prevalence of smaller computing devices has increased. Smart phones, tablet computers, ultra-thin notebooks, and other user devices have grown exponentially. However, these smaller devices rely on servers for both data storage as well as complex processing beyond the form factor. Thus, the demand for the high-performance computing market (i.e., the server space) has also increased. For example, in modern servers, not only are there typically single processors with multiple cores, but there are also multiple physical processors (also referred to as multiple sockets) to increase computing power. But as processing power grows along with the number of devices in a computing system, communication between the sockets and other devices becomes more critical.
[0004] Indeed, interconnects have evolved from more traditional multi-drop buses that primarily handled electrical communication to well-developed interconnect architectures that facilitate fast communication. Unfortunately, as the demand for future processors to consume at even faster rates, a corresponding demand is placed on the capabilities of existing interconnect architectures. BRIEF DESCRIPTION OF DRAWINGS
[0005] Figure 1 Embodiments of a computing system including an interconnect architecture are shown.
[0006] Figure 2 Embodiments of an interconnect architecture including a layered stack are shown.
[0007] Figure 3 Embodiments of a request or packet to be generated or received within an interconnect architecture are shown.
[0008] Figure 4 Embodiments of a transmitter and receiver pair for an interconnect architecture are shown.
[0009] Figure 5Embodiments of potentially high performance processor-to-processor interconnect configurations are shown.
[0010] Figure 6 Embodiments of layered protocol stacks associated with an interconnect are shown.
[0011] Figures 7A-7C A simplified block diagram showing an example implementation of a test pattern for determining errors in one or more sublinks of a link is shown.
[0012] Figures 8A-8B A simplified block diagram showing an example link including one or more extension devices is shown.
[0013] Figure 9 A simplified block diagram showing an example re-timer device is shown.
[0014] Figure 10 is a flow diagram showing an example technique for dynamically resizing a re-timer's elastic buffer.
[0015] Figures 11A-11E A simplified block diagram illustrating resizing a re-timer's elastic buffer in one example is shown.
[0016] Figure 12 is a flow diagram showing an example technique in conjunction with a link including a re-timer.
[0017] Figure 13 Embodiments of a block diagram of a computing system including a multi-core processor are shown.
[0018] Figure 14 Embodiments of a block of a computing system including multiple processors are shown. DETAILED DESCRIPTION
[0019] In the following description, numerous specific details are set forth such as examples of specific types of processors and system configurations, specific hardware structures, specific architectural and micro architectural details, specific register configurations, specific instruction types, specific system components, specific measurements / heights, specific processor pipeline stages and operation etc. in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without these specific details. In other instances, well-known components or methods are not described in detail or are presented in a simple manner in order not to unnecessarily obscure the present application. For example, specific functionality, details of certain components, code, code examples, techniques and / or computer models, etc. are set forth in the detailed description section below.
[0020] While the following embodiments can be described with reference to energy conservation and energy efficiency in specific integrated circuits, such as in computing platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein can be applied in other types of circuits or semiconductor devices that can also benefit from better energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to desktop computer systems or ultra-book TM systems, and can also be used in other equipment such as handheld devices, tablets, other thin notebooks, systems on a chip (SOC) devices, and embedded applications. Some examples of handheld devices include cellular phones, Internet protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system on a chip, network computers (NetPC), set-top boxes, network hubs, wide area network (WAN) switches, or any other system that can perform the functions and operations taught below. Moreover, the apparatuses, methods, and systems described herein are not limited to physical computing devices, but can also relate to software optimizations for energy conservation and efficiency. As will become apparent in the following description, embodiments of the methods, apparatuses, and systems described herein (whether referred to as hardware, firmware, software, or a combination thereof) are of enduring importance and are relevant to the "green technology" future balancing performance considerations.
[0021] As computing systems have advanced, the components therein have become increasingly complex. As a result, the interconnect architecture used to couple and communicate between the components has also increased in complexity to ensure bandwidth requirements are met for optimal component operation. Further, different market segments require different aspects of the interconnect architecture to suit the needs of the market. For example, servers require higher performance, while the mobile ecosystem is sometimes able to sacrifice overall performance for power savings. However, most structures have a sole purpose of providing the highest performance possible with the maximum power savings. A variety of interconnects are discussed below that would potentially benefit from the aspects of the application described herein.
[0022] An interconnect fabric architecture includes Peripheral Component Interconnect (PCI) Express (PCIe) architecture. The primary goal of PCIe is to enable components and devices from different vendors to interoperate in an open architecture, spanning multiple market segments: (desktop and mobile) clients, (standard and enterprise) servers, and embedded and communications devices. PCI Express is a high performance, general purpose I / O interconnect defined for a variety of future computing and communications platforms. Some of the PCI attributes (e.g., its usage model, load-store architecture, and software interface) have been maintained through its revisions, while previous parallel bus implementations have been replaced by a highly scalable, fully serial interface. The most recent version of PCI Express leverages improvements in point-to-point interconnect, switch-based technology, and packet protocols to deliver new levels of performance and features. Some of the advanced features supported by PCI Express include power management, Quality of Service (QoS), hot plug / hot swap support, data integrity, and error handling.
[0023] Reference is made to Figure 1 FIG. 1 shows an embodiment of a structure composed of point-to-point links that interconnect a set of components. System 100 includes processor 105 and system memory 110 coupled with controller hub 115. Processor 105 includes any processing element, such as a microprocessor, host processor, embedded processor, co-processor, or other processor. Processor 105 is coupled with controller hub 115 through front side bus (FSB) 106. In one embodiment, FSB 106 is a serial, point-to-point interconnect as described below. In another embodiment, link 106 includes a serial, differential interconnect architecture that conforms to a different interconnect standard.
[0024] System memory 110 includes any memory device, such as random access memory (RAM), non-volatile (NV) memory, or other memory accessible by devices in system 100. System memory 110 is coupled with controller hub 115 through memory interface 116. Examples of memory interfaces include double data rate (DDR) memory interfaces, dual channel DDR memory interfaces, and dynamic RAM (DRAM) memory interfaces.
[0025] In one embodiment, controller hub 115 is a root hub, a root complex, or a root controller in a Peripheral Component Interconnect Express (PCIe or PCIE) interconnect hierarchy. Examples of controller hub 115 include a chipset, a memory controller hub (MCH), a northbridge, an interconnect controller hub (ICH), a southbridge, and a root controller / hub. The term chipset often refers to two physically separate controller hubs, i.e., a memory controller hub (MCH) coupled with an interconnect controller hub (ICH). Note that current systems often include an MCH integrated with processor 105, while controller 115 communicates with I / O devices in a manner similar to that described below. In some embodiments, peer-to-peer routing is optionally supported through root complex 115.
[0026] Here, controller hub 115 is coupled with switch / bridge 120 through serial link 119. Input / output modules 117 and 121 (also referred to as interfaces / ports 117 and 121) include / implement a layered protocol stack to provide communication between controller hub 115 and switch 120. In one embodiment, multiple devices are capable of coupling with switch 120.
[0027] Switch / bridge 120 routes packets / messages from devices 125 upstream (i.e., up the hierarchy toward the root complex) to controller hub 115, and routes packets / messages from processor 105 or system memory 110 downstream (i.e., down the hierarchy away from the root controller) to devices 125. In one embodiment, switch 120 is referred to as a logical assembly of multiple virtual PCI-to-PCI bridge devices. Devices 125 include any internal or external device or component to be coupled with the electronic system, e.g., I / O devices, network interface controllers (NICs), plug-in cards, audio processors, network processors, hard drives, storage devices, CD / DVD ROMs, monitors, printers, mice, keyboards, routers, portable storage devices, Firewire devices, Universal Serial Bus (USB) devices, scanners, and other input / output devices. In PCIe, the common parlance such as devices are referred to as endpoints. Although not specifically shown, devices 125 can include a PCIe-to-PCI / PCI-X bridge to support legacy PCI devices or other versions of PCI devices. Endpoint devices in PCIe are generally classified as legacy integrated endpoints, PCIe integrated endpoints, or root complex integrated endpoints.
[0028] The graphics accelerator 130 is also coupled to the controller center 115 via serial link 132. In one embodiment, the graphics accelerator 130 is coupled to the MCH, which is coupled to the ICH. The switch 120, and therefore the I / O device 125, is then coupled to the ICH. I / O modules 131 and 118 are also used to implement a layered protocol stack for communication between the graphics accelerator 130 and the controller center 115. Similar to the discussion above regarding the MCH, the graphics controller or graphics accelerator 130 itself may be integrated into the processor 105. Furthermore, one or more links in the system (e.g., 123) may include one or more expansion devices (e.g., 150), such as retimers, repeaters, etc.
[0029] Go to Figure 2 This illustrates an embodiment of a layered protocol stack. The layered protocol stack 200 includes any form of layered communication stack, such as a Fast Path Interconnect (QPI) stack, a PCIe stack, a next-generation high-performance computing interconnect stack, or other layered stacks. Although the following references... Figures 1-4 The discussion pertains to the PCIe stack, but the same concepts can be applied to other interconnect stacks. In one embodiment, protocol stack 200 is a PCIe protocol stack, which includes a transaction layer 205, a link layer 210, and a physical layer 220. Interfaces (e.g., Figure 1 Interfaces 117, 118, 121, 122, 126, and 131 in the code can be represented as communication protocol stack 200. A representation of a communication protocol stack can also be called a module or interface that implements / includes the protocol stack.
[0030] PCI uses packets to transmit information between components. Packets are formed in transaction layer 205 and data link layer 210 to carry information from the sending component to the receiving component. As the transmitted packets flow through other layers, they are expanded with additional information necessary for processing at those layers. On the receiving side, the reverse process occurs, and packets are transformed from their physical layer 220 representation to their data link layer 210 representation, and (for transaction layer packets) ultimately transformed into a form that can be processed by the transaction layer 205 of the receiving device.
[0031] Transaction layer
[0032] In one embodiment, transaction layer 205 provides an interface between the device's processing core and the interconnect architecture (e.g., data link layer 210 and physical layer 220). In this respect, the primary responsibility of transaction layer 205 is the assembly and disassembly of packets (i.e., transaction layer packets or TLPs). Transformation layer 205 typically manages credit-based flow control for TLPs. PCIe implements decoupled transactions, i.e., time-separated requests and responses, thereby allowing the link to carry other traffic while the target device collects data for the response.
[0033] Furthermore, PCIe utilizes credit-based flow control. In this scheme, a device advertises an initial credit amount for each of the receive buffers in the transaction layer 205. The number of credits consumed by each TLP is counted by the external device at the opposite end of the link (e.g., the controller hub 115 in Figure 1 If the transaction does not exceed the credit limit, the transaction can be sent. Upon receiving a response, the credit amount will be restored. One advantage of the credit scheme is that the delay in credit return will not impact performance as long as the credit limit is not reached.
[0034] In one embodiment, the four transaction address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more of read requests and write requests to transfer data to or from a memory-mapped location. In one embodiment, memory space transactions can use two different address formats, e.g., a short address format (e.g., 32-bit address) or a long address format (e.g., 64-bit address). Configuration space transactions are used to access the configuration space of a PCIe device. Transactions for the configuration space include read requests and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.
[0035] Thus, in one embodiment, the transaction layer 205 assembles the packet header / payload 206. The format of the current packet header / payload can be found in the PCIe specification at the PCIe specification website.
[0036] Quick Reference Figure 3 FIG. 3, an embodiment of a PCIe transaction descriptor is shown. In one embodiment, the transaction descriptor 300 is a mechanism for carrying transaction information. In this regard, the transaction descriptor 300 supports identification of transactions in a system. Other potential uses include tracking modifications to the ordering of default transactions and association of transactions with channels.
[0037] The transaction descriptor 300 includes a global identifier field 302, an attribute field 304, and a channel identifier field 306. In the illustrated example, the global identifier field 302 is depicted as including a local transaction identifier field 308 and a source identifier field 310. In one embodiment, the global transaction identifier 302 is unique for all outstanding requests.
[0038] According to one implementation, the local transaction identifier field 308 is a field generated by the requester agent and is unique for all outstanding requests requiring completion by that requester agent. Further, in this example, the source identifier 310 uniquely identifies the requester agent within the PCIe hierarchy. Thus, the local transaction identifier 308 field, together with the source ID 310, provides a global identification of a transaction within the hierarchy domain.
[0039] The attributes field 304 specifies the characteristics and relationship of the transaction. In this regard, the attributes field 304 is potentially used to provide additional information that allows modification of the default handling of the transaction. In one embodiment, the attributes field 304 includes a priority field 312, a reserved field 314, an ordering field 316, and a no-snoop field 318. Here, the priority subfield 312 can be modified by the initiator to assign a priority to the transaction. The reserved attributes field 314 is reserved for future use or vendor defined use. The reserved attributes field can be used to implement possible use models using priority or security attributes.
[0040] In this example, the ordering attributes field 316 is used to supply optional information that conveys an ordering type that can modify the default ordering rules. According to one example implementation, an ordering attribute of "0" indicates that the default ordering rules are to be applied, where an ordering attribute of "1" indicates relaxed ordering where writes can pass writes in the same direction and read completions can pass writes in the same direction. The snoop attribute field 318 is used to determine whether the transaction is snooped. As shown, the channel ID field 306 identifies the channel associated with the transaction.
[0041] Link Layer
[0042] The link layer 210 (also referred to as the data link layer 210) serves as an intermediate level between the transaction layer 205 and the physical layer 220. In one embodiment, the responsibility of the data link layer 210 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components of a link. One side of the data link layer 210 accepts TLPs assembled by the transaction layer 205, applies a packet sequence identifier 211 (i.e., tag or packet number), calculates and applies an error detection code (i.e., CRC 212), and submits the modified TLP to the physical layer 220 for transmission across the physical to an external device.
[0043] Physical Layer
[0044] In one embodiment, physical layer 220 includes a logic subblock 221 and an electrical subblock 222 for physically transmitting packets to external devices. Here, logic subblock 221 is responsible for the "digital" functions of physical layer 221. In this respect, the logic subblock includes a transmitting portion for preparing outgoing information to be transmitted by physical subblock 222, and a receiving portion for identifying and preparing received information before passing it to link layer 210.
[0045] Physical block 222 includes a transmitter and a receiver. The transmitter is supplied with symbols by logic subblock 221, serializes these symbols, and transmits them to an external device. The receiver is supplied with serialized symbols from the external device and converts the received signals into a bit stream. The bit stream is deserialized and supplied to logic subblock 221. In one embodiment, an 8b / 10b transmission code is used, where ten-bit symbols are transmitted / received. Here, special symbols are used to frame packets using frame 223. Furthermore, in one example, the receiver also provides a symbol clock recovered from the incoming serial stream.
[0046] As stated above, although the transaction layer 205, link layer 210, and physical layer 220 have been discussed with reference to specific embodiments of the PCIe protocol stack, the layered protocol stack is not limited thereto. In fact, any layered protocol can be included / implemented. As an example, a port / interface represented as a layered protocol includes: (1) a first layer (i.e., the transaction layer) for assembling packets; a second layer (i.e., the link layer) for ordering packets; and a third layer (i.e., the physical layer) for sending packets. As a specific example, the Common Standard Interface (CSI) layered protocol is used.
[0047] Next reference Figure 4 An embodiment of a PCIe serial point-to-point architecture is illustrated. Although an embodiment of a PCIe serial point-to-point link is shown, the serial point-to-point link is not limited to this, as it includes any transmission path for transmitting serial data. In the illustrated embodiment, the basic PCIe link includes two low-voltage differential drive signal pairs: transmit pair 406 / 411 and receive pair 412 / 407. Therefore, device 405 includes transmit logic 406 for transmitting data to device 410 and receive logic 407 for receiving data from device 410. In other words, the PCIe link includes two transmit paths (i.e., paths 416 and 417) and two receive paths (i.e., paths 418 and 419).
[0048] A transmission path refers to any path used to transmit data, e.g., a transmission line, a copper line, an optical line, a wireless communication channel, an infrared communication link, or other communication path. A connection between two devices (e.g., device 405 and device 410) is referred to as a link (e.g., link 415). A link can support one lane - each lane representing a set of differential signal pairs (one pair to send, one pair to receive). To scale bandwidth, a link can aggregate multiple lanes denoted by xN, where N is any supported link width, e.g., 1, 2, 4, 8, 12, 16, 32, 64, or wider.
[0049] A differential pair refers to two transmission paths used to transmit a differential signal, e.g., lines 416 and 417. As an example, when line 416 switches from a low voltage level to a high voltage level (i.e., a rising edge), line 417 drives from a high logic level to a low logic level (i.e., a falling edge). Differential signals potentially exhibit better electrical characteristics, e.g., better signal integrity (i.e., cross-talk, voltage overshoot / undershoot, ringing, etc.). This allows for better timing windows, enabling faster transmission frequencies.
[0050] In one embodiment, a super-path interconnect (UPI) can be utilized to interconnect two or more devices. The UPI can implement a next generation cache coherent link-based interconnect. As one example, the UPI can be used in a high performance computing platform (e.g., a workstation or server), including in systems where PCIe or another interconnect protocol is typically used to connect processors, accelerators, I / O devices, etc. However, the UPI is not so limited. Rather, the UPI can be used in any of the systems or platforms described herein. Moreover, various ideas developed can be applied to other interconnects and platforms, e.g., PCIe, MIPI, QPI, etc.
[0051] To support multiple devices, in one example implementation, the UPI can include an instruction set architecture (ISA) agnostic (i.e., the UPI can be implemented in multiple different devices). In another scenario, the UPI can also be used to connect high performance I / O devices, not just processors or accelerators. For example, a high performance PCIe device can be coupled with the UPI through an appropriate (i.e., UPI to PCIe) translation bridge. Moreover, UPI links can be used by many UPI-based devices (e.g., processors) in various ways (e.g., star, ring, mesh, etc.). Figure 5Example implementations of multiple potential multi-socket configurations are shown. As depicted, a two-socket configuration 505 can include two UPI links; however, in other implementations, one UPI link can be used. For larger topologies, any configuration can be used so long as an identifier (ID) is assignable and there is some form of virtual path and other additional or alternative features. As shown, in one example, a four-socket configuration 510 has UPI links from each processor to another processor. But in an eight-socket implementation shown in configuration 515, not every socket is directly connected to each other via a UPI link. However, this configuration is supported if there is a virtual path or channel between the processors. The range of processors supported in a native domain includes 2-32. Larger numbers of processors can be implemented through the use of multiple domains or other interconnections between node controllers, among other examples.
[0052] In some examples, the UPI architecture includes a definition of a layered protocol architecture including protocol layers (of coherent, incoherent, and optionally other memory-based protocols), a routing layer, a link layer, and a physical layer. In addition, UPI can also include enhancements involving power managers (e.g., power control units (PCUs)), design for test and debug (DFT), fault handling, registers, security, among other examples. Figure 5 An embodiment of an example UPI layered protocol stack is shown. In some implementations, Figure 5 At least some of the layers shown in FIG. 6 can be optional. Each layer handles its own level of granularity or amount of information (protocol layers 620a, 602b have packets 630, link layers 610a, 610b have flits 635, and physical layers 605a, 605b have physical flits (phits) 640). Note that in some embodiments, a packet can include a partial flit, a single flit, or multiple flits based on the implementation.
[0053] As a first example, the width of the physical flit 640 includes a 1-to-1 mapping of link width to bits (e.g., a 20-bit link width includes a 20-bit physical flit, etc.). Flits can have larger sizes, e.g., 184, 192, or 200 bits. Note that if the physical flit 640 is 20 bits wide and the size of the flit 635 is 184 bits, then a fractional number of physical flits 640 are taken to send one flit 635 (e.g., 9.2 physical flits of 20 bits to send a 184-bit flit 635, or 9.6 physical flits of 20 bits to send a 192-bit flit, among other examples). Note that the width of the underlying link at the physical layer can vary. For example, the number of lanes per direction can include 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, etc. In one embodiment, the link layers 610a, 610b are capable of embedding multiple different transactions into a single flit, and one or more headers (e.g., 1, 2, 3, 4) can be embedded within the flit. In one example, the UPI splits the headers into corresponding time slots to enable multiple messages in a flit to go to different nodes.
[0054] In one embodiment, the physical layers 605a, 605b can be responsible for the fast transfer of information over the physical medium (electrical or optical, etc.). The physical link can be a point-to-point between two link layer entities (e.g., layers 605a and 605b). The link layer 610a, 610b can abstract the physical layer 605a, 605b from the upper layers and provide the ability to reliably transfer data (and requests) and manage flow control between two directly connected entities. The link layer can also be responsible for virtualizing the physical channel into multiple virtual channels and message classes. The protocol layers 620a, 620b rely on the link layer 610a, 610b to map protocol messages into the appropriate message class and virtual channel before handing the protocol messages to the physical layer 605a, 605b for transfer across the physical link. The link layer 610a, 610b can support multiple messages, e.g., requests, snoop, responses, writebacks, non-coherent data, among other examples.
[0055] As Figure 6As shown, the physical layer 605a, 605b (or PHY) of UPI can be implemented above the electrical layer (i.e., electrical conductors connecting two components) and below the link layer 610a, 610b. The physical layer and corresponding logic can reside on each agent and connects the link layers on two agents (A and B) that are separate from each other (e.g., on devices on either side of the link). The local and remote electrical layers are connected by a physical medium (e.g., wires, conductors, light, etc.). In one embodiment, the physical layer 605a, 605b has two main phases: initialization and operation. During initialization, the connection is opaque to the link layer and signaling can involve a combination of timing states and handshaking events. During operation, the connection is transparent to the link layer and signaling is at a speed where all lanes operate together as a single link. During the operation phase, the physical layer transmits flits from agent A to agent B and from agent B to agent A. The connection is also referred to as a link and abstracts some physical aspects including medium, width, and speed from the link layer when exchanging flits and control / state of the current configuration (e.g., width). The initialization phase includes secondary phases, e.g., polling, configuration. The operation phase also includes secondary phases (e.g., link power management states).
[0056] In one embodiment, the link layer 610a, 610b can be implemented to provide reliable data transfer between two protocol or routing entities. The link layer can abstract the physical layer 605a, 605b from the protocol layer 620a, 620b and can be responsible for flow control between two protocol agents (A, B) and provide virtual channel services to the protocol layer (message class) and routing layer (virtual network). The interface between the protocol layer 620a, 602b and the link layer 610a, 610b can typically be at the packet level. In one embodiment, the smallest transfer unit at the link layer is referred to as a flit with a specified number of bits, e.g., 192 bits or some other measure. The link layer 610a, 610b relies on the physical layer 605a, 605b to frame the transfer units of the physical layer 605a, 605b (physical flits) into the transfer units of the link layer 610a, 610b (flits). Additionally, the link layer 610a, 610b can be logically split into two parts: the sender and the receiver. The sender / receiver pair on one entity can connect to the receiver / sender pair on the other entity. Flow control is often performed on a flit and packet basis. Error detection and correction can also potentially be performed on a flit level basis.
[0057] In one embodiment, the routing layer 615a, 615b can provide a flexible and distributed approach to routing UPI transactions from source to destination. The scheme is flexible in that routing algorithms for multiple topologies can be specified by programmable routing tables at each router (programmed in one embodiment by firmware, software, or a combination thereof). The routing function can be distributed; routing can be accomplished through a series of routing steps, where each routing step is defined by a lookup table at a source router, an intermediate router, or a destination router. Lookups at the source can be used to inject UPI packets into the UPI fabric. Lookups at intermediate routers can be used to route UPI packets from an input port to an output port. Lookups at the destination port can be used to target a destination UPI protocol agent. Note that in some implementations, the routing layer can be thin in that the routing tables, and thus the routing algorithms, are not explicitly defined by the specification. This allows for flexibility and a variety of usage models, including a flexible platform architecture topology to be defined by a system implementation. The routing layer 615a, 615b relies on the link layer 610a, 610b to provide for the use of up to three (or more) virtual networks (VNs) - in one example, two deadlock-free VNs (VN0 and VN1) with multiple message classes defined in each virtual network. A shared adaptive virtual network (VNA) can be defined in the link layer, but this adaptive network can not be directly exposed in the routing concept, as each message class and virtual network can have dedicated resources and guaranteed forwarding progress, among other features and examples.
[0058] In some implementations, UPI can utilize embedded clocks. Clock signals can be embedded into data sent using the interconnect. In cases where clock signals are embedded in data, a distinct and dedicated clock lane can be omitted. This can be useful, for example, as it can allow more pins of a device to be dedicated to data transfer, particularly in systems where pin space is at a premium.
[0059] In some implementations, a link such as a link conforming to a PCIe, USB, UPI, or other interconnect protocol can include one or more re-timer or other extension devices, e.g., repeaters. A re-timer device (or simply “re-timer”) can include an active electronic device that receives and re-sends (re-times) digital I / O signals. A re-timer can be used to extend the length of a channel that can be used with a digital I / O bus.
[0060] Figures 7A-7Care simplified block diagrams 700a-700c showing example implementations of a link interconnecting two system components or devices (e.g., an upstream component 705 and a downstream component 710). In some instances, the upstream component 705 and the downstream component 710 can be directly connected (e.g., without a retimer, re-driver, or repeater disposed on the link between the two components 705, 710) in instances where the maximum length of the link is less than or equal to the maximum length of the link specified by the protocol or technology used to interconnect the two components 705, 710. Figure 7A In other instances, a retimer (e.g., 715) can be provided to extend the link connecting the upstream component 705 and the downstream component 710 (e.g., as shown in Figure 7B In other implementations, two or more retimers (e.g., 715, 720) can be provided in series to further extend the link connecting the upstream component 705 and the downstream component 710. For example, a particular interconnect technology or protocol can specify a maximum channel length and one or more retimers (e.g., 715, 720) can be provided to extend the physical length of the channel connecting the two devices 705, 710. For example, providing a retimer 715, 720 between the upstream component 705 and the downstream component 710 can allow the link to be three times the maximum length specified for the link without these retimers (e.g., 715, 720), among other example implementations.
[0061] A link containing one or more retimers can form two or more separate electronic links at data rates comparable to the data rates achieved by a link employing a similar protocol but without retimers. For example, a link including a single retimer can form a link having two separate sublinks, each operating at a rate of 8.0 GT / s or higher. Figures 8A-8B Simplified block diagrams 800a-800b showing example links including one or more retimers are shown. For example, in Figure 8A In the example shown in FIG. 8A, a link connecting a first component 805 (e.g., an upstream component) to a second component 810 (e.g., a downstream component) can include a single retimer 815a. A first sublink 820a can connect the first component 805 to the retimer 815a, and a second sublink 820b can connect the retimer 815a to the second component. As shown in Figure 8B As shown in FIG. 8B, multiple retimers 815a, 815b can be used to extend the link. Three sublinks 820a-820c can be defined by the two retimers 815a, 815b, with a first sublink 815a connecting the first component to the first retimer 815a, a second sublink connecting the first retimer 815a to the second retimer 815b, and a third sublink 815c connecting the second retimer 815b to the second component.
[0062] As shown in FIG. 8B, multiple retimers 815a, 815b can be used to extend the link. Three sublinks 820a-820c can be defined by the two retimers 815a, 815b, with a first sublink 815a connecting the first component to the first retimer 815a, a second sublink connecting the first retimer 815a to the second retimer 815b, and a third sublink 815c connecting the second retimer 815b to the second component. Figures 8A-8BIn the example shown in FIG. 8, in some implementations, a retimer can include two pseudo-ports, and the pseudo-ports can dynamically determine their respective downstream / upstream orientation. Each retimer 815a, 815b can have an upstream path and a downstream path. Further, the retimers 815a, 815b can support operational modes including a forwarding mode and a performing mode. In some instances, the retimers 815a, 815b can decode data received on a sublink and re-encode data to be forwarded downstream on another sublink thereof. In certain cases, such as when processing and forwarding in-order set data, the retimer can modify certain values in its received data. Additionally, the retimer can potentially support any width option as its maximum width, such as a set of width options defined by a specification such as PCIe.
[0063] As data rates of serial interconnects (e.g., PCIe, UPI, USB, etc.) increase, retimers are increasingly used to extend channel reach. Retimers can capture a received bitstream before regenerating and retransmitting the bitstream. In some cases, retimers can be protocol-aware and have a full physical layer or even a protocol stack to allow the retimer to participate in link training state machine (LTSSM), including transmitter / receiver equalization and other link training activities. However, in high-speed links that implement retimers with full protocol stack or full physical layer logic or link layer logic, etc., unacceptable latency can be added for links connecting two or more devices through one or more retimers. In fact, as operating frequencies of external interfaces continue to increase while channels improve at a more modest pace, more and more applications can utilize retimers or other channel extension devices. Further, many applications require longer channel lengths, such as data center and server applications where interconnect channels can span several inches, increasing or exceeding a maximum channel length natively supported by an emerging high-speed interconnect. For example, a PCI Express Gen 4 designed to operate at a frequency of 16.0 GT / s can provide a specific limited maximum channel length (e.g., 14” or less). For server applications where channel lengths can typically exceed 20 inches, a retimer, re-driver, or other repeater element can be sought to extend the channel. Similarly, for an Ultra Path Interconnect (UPI) cache coherent interface, an extension device can likewise be utilized to support longer length platforms at 10.4 GT / s, among other examples.
[0064] Developing and implementing a retimer for a high-speed interface can face various problems. For example, in a cache coherency protocol such as UPI, channels can be extremely sensitive to latency, making the additional latency of 30 nsec per retimer hop increasingly intolerable due to the performance penalty introduced by the retimer(s). Latency can also be a problem in examples such as PCIe, for example, in memory applications (e.g., memory drivers and memory service processors), and as next generation non-volatile memory (NVM) technology provides higher bandwidth and lower latency closing the gap with double data rate (DDR) memory (e.g., DDR synchronous dynamic random access memory (SDRAM)), these challenges are expected to only worsen.
[0065] The speed of high-speed differential serial links continues to increase. USB 3.1 has a link speed of 10 GT / s and PCI Express 4.0 has a link speed of 16 GT / s, with future standard and non-standard applications expected to further increase in speed. Despite these advances, the physical size of many systems and devices remains the same - making the design of high-speed differential channels more challenging as the I / O speed increases. Many channel designs now require active extension devices such as retimers, and the percentage of channel designs requiring extension devices is increasing. Extension devices can include examples such as repeaters, re-drivers, and retimers. Retimers provide the most extension (100% per retimer) and guaranteed interoperability in these examples. However, retimers also have some drawbacks compared to simpler analog-only re-driver extension devices, including increased cost, latency, and power. The increased latency of retimers is particularly significant in systems that use a clocking architecture in which the transmitter and receiver use separate / independent reference clocks containing low frequency spread spectrum modulation (SSC) to help pass emission tests. As an example, in PCI Express, a clocking architecture can be employed: spread spectrum with independent SSC (or SRIS). This clocking architecture is also employed in USB 3.0 / 3.1. In some cases, these clocking architectures can be employed so that the link complies with radio communication rules and standards (e.g., Federal Communications Commission standards) and for other reasons. Similar issues can exist for other clocking architectures of other interconnects.
[0066] In many protocols, the task of a retimer is to forward all symbols without dropping or adding any symbols other than the delimiter. In some implementations, a frequency mismatch can occur between endpoints connected by a retimer. This mismatch can be a direct result of the clock at the endpoints connected by the link applying a modulation scheme. In some implementations, a special data symbol can be defined from which bits can be removed or to which bits can be added in order to resolve the frequency mismatch. For example, a SKIP or SKP ordered set (OS) can be defined (e.g., in PCIe and USB) that is designed to be modified to compensate for the frequency difference between the bit rates at the two ends of the link. Additionally, a flexible buffer can be provided to perform this compensation. In some implementations, a flexible buffer can be provided in the physical layer logic of the receiver of the endpoints connected on the link. Likewise, the retimer can also be equipped with a flexible buffer to handle the periodic compensation of the frequency difference between the endpoints.
[0067] In some implementations, the flexible buffer of a retimer can be designed to nominally remain half full and to prevent underflow or overflow in the data stream when the independent transmit and receive clocks have a rate difference (e.g., due to independent reference clocks and spread spectrum modulation). In this example, whenever the buffer becomes more than half full, the retimer removes a SKP symbol from the next SKP ordered set, and whenever the buffer becomes less than half full, the retimer adds a SKP symbol to the next SKP ordered set. For some example protocols (e.g., PCI Express and USB), the size of the flexible buffer of a retimer can be a function of factors such as the rate at which SKP ordered sets are transmitted (e.g., based on whether SRIS (or another clock modulation scheme) is applied to the clock), the link width (the number of lanes forming the link), and the maximum data packet size (e.g., because in some examples these protocols can not allow SKP ordered sets to be transmitted in the middle of a data packet (but these SKP ordered sets will be buffered, resulting in multiple SKP ordered sets being transmitted after a long data packet is completed)), among other examples.
[0068] In an illustrative example, in a system utilizing a clock with applied SRIS timing, worst-case latency can occur when the link width is x1 and the maximum packet payload size is at its maximum value (e.g., 4096 bytes). Considering the possibility of this happening over the link's lifetime, conventional retimer implementations have elastic buffers that are resized to potentially handle the worst-case scenario (e.g., x1 link and 4096-byte maximum packet size) to avoid critical errors due to elastic buffer overflow. However, because resizing the retimer elastic buffer to accommodate worst-case conditions on the link can introduce considerable latency, impacting the performance of the link where the retimer is introduced.
[0069] In some cases, the link width and maximum packet size can be determined dynamically, and even changed dynamically (e.g., according to protocol definitions, such as in PCIe). In some implementations, enhanced functionality can be provided for retimers to detect link width, maximum packet size, and other link attributes, and to dynamically adapt the size of their elastic buffers based on these attributes, thus reducing the latency introduced by the elastic buffers of the retimer. For example, an enhanced retimer can dynamically update its elastic buffer size based on changes to the link width, maximum packet size, clock modulation mode, and other examples. This can lead to a significant reduction in retimer latency, among other benefits.
[0070] Go to Figure 9 A simplified block diagram 900 illustrating an example implementation of an enhanced retimer equipped with logic enabling the retimer 910 to dynamically adjust the size of a buffer (e.g., 930) used by the retimer 910, resulting in dynamic changes to link characteristics. One or more retimer devices 910 can be provided to extend links connecting two devices (e.g., 905, 915), such as a processor device and an endpoint device, or two processor devices. In this example, each of the retimers 910 may include (e.g., implemented in hardware circuitry, hardware logic, firmware, and / or software) retimer logic 920, protocol logic 925, elastic buffer 930, and buffer manager 935 logic to determine changes to link characteristics and automatically adjust the size of the elastic buffer 930 based on these changes to the link. In this example, the retimer logic corresponding to a single channel of a multi-channel link is shown. Therefore, for each of the multiple channels supported by the retimer 910 (and on each side of the retimer), the retimer logic can be replicated. Figure 9The components shown in the middle (as well as in other examples herein) (e.g., at each of its two receivers (e.g., upstream receive port and downstream receive port), each lane has a separate elastic buffer).
[0071] In one example, the re-timer logic 920 of the re-timer 910 can facilitate standard re-timer functionality of the re-timer 910, including receiving data at one of the re-timer's receiver (Rx) ports and regenerating the data for transmission on a corresponding one of the re-timer's transmitter (Tx) ports. The re-timer 910 can additionally possess protocol logic 925 to enable the re-timer 910 to support and in some cases participate in link training, adaptation, and / or other functions defined for one or more interconnect protocols. The protocol logic 925 can contain less than the entirety of a full protocol stack, rather the protocol logic 925 can be the minimum set of functions that allow the re-timer 910 to be inserted into a link without interrupting normal initialization and operation of the link as defined in the interconnect protocol. In some instances, the protocol logic 925 can be limited to supporting a single interconnect protocol (giving the re-timer 910 a protocol-specific design). In other cases, the protocol logic 925 can be provided with protocol logic for multiple different interconnect protocols, and a bit (or bits) can be set (e.g., in a register corresponding to the re-timer 910) to indicate what subset of the protocol logic 925 is to be enabled on a particular link (e.g., such that only a single one of the protocol logics is enabled), among other examples.
[0072] Protocol logic 925 can be provided on the re-timer 910 to augment the basic re-timer logic 920 and allow the re-timer 910 to support activities of a link that are compliant with a particular interconnect specification (e.g., PCIe, USB, SATA, UPI, etc.). For example, the re-timer can include logic to support or be aware of activities such as electrical idle exit / entry, speed change, link width negotiation and change, equalization, and other features defined in the corresponding interconnect specification. For example, in PCIe, a data rate (or “speed”) change can be requested by one of the devices connected on the PCIe link. For example, a downstream port of a device can request a speed change through an equalization (EQ) training sequence 1 (TS1) ordered set (OS) to inform an upstream port of another device. For example, defined ordered set sequences can be sent, each with one or more bits (e.g., speed change bits in a PCIe TS1 or TS2 ordered set) set to indicate a speed change. The receiving device can respond by sending a corresponding ordered set sequence to the requesting device to acknowledge or reject the request to change speed. If the request is acknowledged, the link can be retrained (using sequences of other ordered sequences) to operate at the new speed. Similarly, patterns of ordered sets (e.g., TS1 and / or TS2) can be sent according to parameters defined by the specification and / or with particular OS bits encoded to indicate a link width to be applied within the link, a request to change the link width (and subsequent acknowledgement), a change in clock modulation pattern (e.g., to enable or disable spread spectrum modulation), among other examples. Moreover, rather than equip the re-timer with a full protocol stack or protocol layer, the implementation of the enhanced re-timer can be equipped with only particular portions of the protocol-specific logic required (at the re-timer) to handle certain features such as dynamic speed change, transmitter equalization, electrical idle entry for power state changes, clock modulation change, link width change, and receiver detection (e.g., for hot plug), among other additional or alternative examples.
[0073] A resilience buffer 930 of the retimer can be provided to compensate for frequency differences between bit rates at both ends of the link (e.g., 905, 915). The resilience buffer 930 can be expanded in order to be able to hold enough symbols to handle worst case frequency differences and worst case spacing between symbols that can be used for rate compensation (e.g., SKP OS). In this implementation, the depth of the resilience buffer can be dynamically resized, but the maximum size of the resilience buffer can be according to the worst case identified for the particular protocol supported by the retimer 910. Logic (e.g., physical layer logic of the retimer controlling the resilience buffer 930) can be responsible for inserting or removing symbols from dedicated data sequences (e.g., SKP OS, ALIGN, other OS, etc.) designated or designed to add or subtract at the buffer 930 in order to avoid overflow or underflow of the resilience buffer. For example, the protocol logic 925 can include logic (e.g., logical PHY logic) to monitor the received data stream and determine when these dedicated data sequences are received, providing an opportunity for the resilience buffer to add or remove symbols in an attempt to keep the capacity of the buffer 930 at a particular level (e.g., as close to half full as possible). However, this target capacity can correspond to an average delay that will be introduced due to the retimer resilience buffer 930. Thus, it can be desirable to minimize the (effective) size of the resilience buffer in order to reduce the delay on the link, while still managing frequency mismatches and preventing buffer overflow / underflow.
[0074] The size of the elastic buffer 930 can be managed by a buffer manager 935 component of the re-timer 910 in order to dynamically and autonomously manage the delay introduced by the elastic buffer 930. In one example, the buffer manager 935 can include (e.g., implemented in hardware circuitry, hardware logic, firmware, and / or software) link change detection logic 940, change mapping logic 945, and buffer size change logic 950, among other modules and subcomponents. In one example, the link change detection 940 logic can be equipped with functionality to identify various types of changes to characteristics of the link over which the re-timer 910 is implemented. These characteristics can include those that can affect the maximum size needed for the elastic buffer 930 in order to avoid overflow or underflow, such as the link width of the link, whether and which type of clock modulation is applied at the clock of the link endpoints (e.g., 905, 910), the interval at which a dedicated data sequence (e.g., SKP OS) is sent over the link (thereby providing an opportunity to compensate for capacity in the elastic buffer 930 when it is approaching an underflow condition (by adding bits or symbols to the received dedicated data sequence) or when it is approaching an overflow condition (by removing bits or symbols from the received dedicated data sequence)), the speed of the link, the payload or packet size (e.g., because dedicated data sequences can only be allowed to be sent between packets), among other examples.
[0075] In some cases, the link change detection logic 940 can interoperate with or be combined with the protocol logic 925 and identify requests, sequences, and other information in the data stream sent by the endpoints 905, 915 through the re-timer to facilitate not only detecting that a change in one of these link characteristics is occurring or about to occur, but also the nature or extent of the change (e.g., the amount of lanes added in a link width change, the type of clock modulation applied, the amount of speed change, the amount of payload size change). The messaging of such changes can be protocol defined and the protocol logic 925 can implement detection and proper interpretation of such messaging / signaling between the devices 905, 915. The buffer manager 910 can additionally be equipped with logic 945 to map the type (and extent) of link change to a corresponding change in the size of the elastic buffer. For example, changing the link width can affect a change in the size of the elastic buffer inversely. As an example, configuring the link width of a link from x4 (i.e., 4 lanes) up to x16 (i.e., 16 lanes) can allow the size of the elastic buffer 930 to be reduced by 75%. Additionally, disabling a particular clock modulation scheme (e.g., SRIS) can map (through the change mapping module 945) to a particular reduction in the elastic buffer (e.g., due to an expected ppm reduction by removing the particular clock modulation), as well as a change (increase or decrease) in the maximum payload size employed on the link (which can be mapped by the change mapping module 945 to a corresponding increase or decrease that can be made to the depth of the elastic buffer 930). In essence, the change mapping module 945 can quantify the amount by which the size of the elastic buffer 930 can increase or decrease in response to detecting a particular type (and extent) of change in the link characteristics. The buffer change module 950 can utilize this information to cause the size of the elastic buffer 930 to change by the amount determined by the link change detection 940 and change mapping module 945.
[0076] In some implementations, the changes to the size of the elastic buffer 930 (by the buffer manager logic 935) can be limited to certain states or operational modes of the link. In one example, changing the size of the elastic buffer 930 of the re-timer can potentially cause data sent to the re-timer 910 for re-timing to be lost, changed, or corrupted. This is not acceptable in operational transmission states where critical application or system management data is sent between devices through the re-timer 910 (e.g., in link states that cannot tolerate this degree of potential bit or packet loss (e.g., L0)). Thus, in one example, the changes to the size of the elastic buffer 930 can be limited to link training states that tolerate the symbol discarding that can be required when resizing the elastic buffer.
[0077] In one example, the re-timer monitors the training sets sent over the link (and through the re-timer) to identify when a change to the properties of the link will occur, e.g., a dynamic link width change. The enhanced re-timer can then dynamically adjust the elastic buffer size by dropping or adding training sets at some point in the (re)training. In one implementation, link training can involve a two-step process, where the first step is dedicated to establishing bit-lock, symbol / block alignment, setting up sequence scrambling logic (e.g., a linear feedback shift register (LFSR)), and otherwise bringing the link into an operational state. For example, training in PCIe can involve one endpoint device sending a particular (e.g., TS1) ordered set, and can switch to a TS2 ordered set (e.g., to indicate a transition to a second training step) when it has completed the first step. Further, in some implementations, the link can tolerate any additions / deletions to the (TS1) ordered set that brings the link into an operational state in the training phase. However, once the link is brought to that point, the link can be more sensitive (or even prohibited) to adding or deleting data. For example, in PCIe, once a device starts sending TS2 ordered sets, the link does not feature handling any errors in the TS2 ordered sets in both directions. Thus, in one example, changes to the re-timer buffer size can occur during the training phase that precedes the operational state of the link (e.g., during the TS1 phase of link training). In some instances, the re-timer can even be used to extend the training phase in order to take advantage of that window and "buy time" to properly set the depth of its elastic buffer based on the characteristics of the link. For example, the re-timer can artificially continue to send TS1 on both the upstream and downstream directions to extend the TS1 training phase, and allow its link partner (e.g., 905, 915) to handle any symbol loss / addition that occurs with the change to the elastic buffer.
[0078] In some implementations, the link training phase in which the re-timer elastic buffer size change can be allowed can also conveniently be a link training phase in which one or more of the corresponding link property changes that are the basis for the elastic buffer size readjustment will be detected. For example, in PCIe, a change in link width and a change to SRIS mode can occur during a link training phase by messaging, in which TSl OS is sent and elastic buffer size readjustment can be performed. In other instances, certain types of link property changes can occur outside of these allowed states. For example, a change in maximum payload size can be initiated by a device (e.g., 905 and 915) outside of a link training phase or other link state in which re-timer elastic buffer size readjustment is allowed (due to the potential for discarding or adding symbols during this size readjustment). Thus, in one example, upon identifying certain types of link property changes (e.g., using link change detection module 940) and identifying that the change occurred outside of the operable conditions of the link that allow for size readjustment of the elastic buffer 930, buffer manager 935 can attempt to force the link into a link training phase or other state in which size readjustment is possible (e.g., a recovery or configuration state) in response to such link property changes. For example, in the example of a PCIe compliant link and re-timer, in response to detecting an attempt to change the maximum payload size supported on the link, buffer manager 935 can force entry into a PCIe recovery phase by using mechanisms such as flipping a sync header in 128b / 130b encoding or sending TSl in 8b / 10b encoding, among other techniques, and other examples.
[0079] In some implementations, certain link characteristic changes can not be detectable from the regular in-band data stream of the protocol (e.g., by the protocol logic 925). This can be due to a minimized or otherwise simplified protocol stack logic that resides on the re-timer (e.g., and is used by components such as the link change detection module 940). In some cases, to support detection of certain types of link characteristic changes and thereby facilitate corresponding re-timer elastic buffer size adjustments, a dedicated messaging packet can be defined by which one or both of the endpoints (e.g., 905, 915) can notify the re-timer 910 of an impending or pending change to a relevant link characteristic (e.g., SRIS mode or maximum payload size, etc.). In other examples, one or more sideband ports can be defined on the re-timer 910 to support a sideband channel (e.g., 955) with one or both of its link partners (e.g., 905, 915) by which out-of-band cues, signals, and messages can be sent from one of the devices (e.g., 905) to the re-timer 910 and convey an impending or pending change to a link characteristic that potentially impacts the re-timer's elastic buffer 930 size. In one particular example based on a PCIe-compliant system, a host controller (e.g., on device 905) can inform the re-timer 910 of a maximum payload size (or a change to a maximum payload size) and / or whether SRIS is enabled on the link partner clock through a sideband mechanism (e.g., system management bus (SMBUS), joint test action group (JTAG), etc.) or an alternative in-band mechanism (e.g., by providing a vendor-defined management command in the last three symbols of the enhanced SKP OS), among other example implementations.
[0080] In one example implementation, the enhanced features described above can be provided by incorporating PCI Fast Physical Interface (PIPE) logic implemented on a retimer. The PIPE can provide a standard interface between the MAC / controller and the interconnecting PHY. In one example of a PIPE-based implementation, the PIPE MAC / controller can be enhanced with logic such that when the controller recognizes that a dynamic elastic buffer resizing adjustment is safe based on detected events (e.g., link width change, speed change, ppm difference change (e.g., from a change in clock architecture), or change in maximum payload size), it updates the signals for the elastic buffer size. The PHY updates the elastic buffer and notifies the PHY when the update is complete. In one example, the signals in the PIPE interface used to indicate size changes and when the size change is complete can be updated registers or a set of signals indicating size. Furthermore, the status signals used by the PHY to indicate completion can be signals / lines or registers updated by the PHY. The PHY can be enhanced to implement dynamic elastic buffer changes quickly without causing any data corruption (only lost or duplicated dates). For example, the PHY (or other buffer modification logic (e.g., 930)) can reduce the size of the elastic buffer by decreasing the distance between its read buffer pointer and write buffer pointer and silently discarding data between the old and new read pointer positions, among other potential implementations. In one example, to increase the size of the elastic buffer, the PHY (or other buffer modification logic (e.g., 930)) could increase the distance between the read and write pointers and add repeating entries to fill the increased buffer range between the new read and write pointers, among other example implementations.
[0081] Continuing the example of a PIPE-based implementation, the MAC in the re-timer 910 can be responsible for changing the elastic buffer depth when conditions are appropriate (e.g., maximum payload size, link width, ppm disparity, change in speed change), while TS1 is still in flight. Even if the MAC receives TS2 in its receiver on an upstream port (connected to any of the sublinks of the re-timer 910), the MAC can cause TS1 to be generated and sent out on the re-timer transmitter instead. The MAC (or other example logic) can continue to ensure that TS1 is received by its downstream link partner to preserve the opportunity for dynamic re-timer elastic buffer size readjustment to complete. The MAC can continue to do so until it determines that the elastic buffer depth has been successfully updated and confirms that packets are being received at the output of both elastic buffers (corresponding to each of its receivers). In one example, the MAC can continue to send TS1 until one or more EIEOS (electrical idle exit ordered set) have been sent (e.g., in PCIe) to ensure block alignment is restored before moving to TS2 after the elastic buffer depth change. When the MAC switches to control sending TS1 from the elastic buffer, it can ensure that the same block boundary is maintained. However, when the MAC must switch back to resume sending the contents of the elastic buffer, the boundary can or can not be preserved. The MAC can ensure that it continues to send TS1 with the new block boundary alignment (even though it is receiving TS2) and can continue to do so until it passes at least one EIEOS. This ensures that the link partner (e.g., 905, 915) can again get their block alignment and continue link training.
[0082] As noted above, in some implementations, the event that suggests an elastic buffer size readjustment can occur during the same link training phase that allows such an elastic buffer size readjustment to occur. If the link is in an operational or active state (e.g., L0 state) when the relevant link characteristics are determined to be changing (or have changed), the MAC can be responsible for forcing the link into a recovery state in response. This can operate to ensure that by the time its link partner (e.g., 905, 915) sees TS2, the elastic buffer depth has been successfully changed to reflect the minimum latency required to reliably operate the link.
[0083] Turning to Figure 10Fig. 10 shows a flowchart 1000 illustrating example techniques for dynamically resizing the elastic buffer for enhanced re-timer on a link. Logic can be provided on the enhanced re-timer to determine (at 1005) whether a change in characteristics of the link has been detected, including examples such as link speed, frequency change (e.g., ppm according to changes in the timing architecture and / or mode), link width, and maximum payload size. When a change in link characteristics is detected that can affect the optimal elastic buffer size (or "depth"), the re-timer can reference a buffer depth table that defines a mapping between the type and extent of the detected characteristic change and the potential impact such a change can have on the maximum elastic buffer size that needs to be supported by the re-timer. In some cases, the change can indicate that no change is necessary. In other cases, referencing (1010) the buffer depth table at the re-timer can cause the re-timer logic to determine (at 1015) that the current elastic buffer depth should be modified according to the mapping of the buffer depth table.
[0084] The manner in which the re-timer approaches modification of its elastic buffer depth can be based on the type of link characteristic change. For example, for a link speed change (at 1020), the re-timer can stop forwarding on lanes that are in electrical idle and perform a buffer depth size resizing (at 1025). After determining (at 1050) that any buffer depth change has been successfully completed, the re-timer can then resume forwarding data (1055) on the link on active lanes (which can include fewer or more lanes than before the buffer depth change in the case of a link width change), and monitor the link for any other changes. For a link width change (at 1030) or other changes detected during link training (e.g., ppm changes), the buffer depth can be resized (1035) during a phase of link training that is able to handle the symbols dropped or added by the re-timer in conjunction with the buffer size resizing (according to the discovery in 1010). The extent of the buffer size resizing can depend on the extent of the link width change (e.g., a change in link width from x4 to x8 can cause the elastic buffer size to be halved, while a change in link width from x16 to x4 can cause the elastic buffer size to be quadrupled, among other examples), the amount of ppm increase or decrease in the link characteristic change, and so on.
[0085] In some cases, the detection of link characteristic changes (at 1005) can occur while the link is in an active state (e.g., L0) or in a training phase that is sensitive to symbol loss (or addition) potentially introduced during a flexible buffer size readjustment. For example, for a change to the maximum payload size detected at the re-timer (e.g., via sideband messaging) during L0 (at 1005), the re-timer can first force the link into a recovery state change (1045) (e.g., by intentionally introducing errors into data it re-timed and forwarded downstream) to cause the link to enter a training phase that allows for a size readjustment of the re-timer’s flexible buffer. At this point, the re-timer can make appropriate modifications to the buffer depth (1035) (e.g., commensurate with the length of the payload size increase or decrease) during the link training caused by the forced link recovery (at 1045). The re-timer can then monitor the link (e.g., during the training or active state) for any further link characteristic changes (at 1005), among other examples.
[0086] Figures 11A-11E A series of simplified block diagrams 1100a-1100e illustrating examples of dynamic size readjustment of the depth of example flexible buffers of example re-timers 910 are shown. In each of the block diagrams 1100a-1100e, an example re-timer 910 is shown for extending a link connection between two computing devices (e.g., 905, 915). The re-timer flexible buffer (e.g., 930a-930e) is shown to illustrate the dynamic adjustments in size that the re-timer flexible buffer can undergo as the enhanced re-timer 910 detects various changes to the link characteristics. For example, in Figure 11A the example of FIG. 1100a, the re-timer 910 is shown with its flexible buffer 930a sized according to the worst-case combination of link characteristics. In Figures 11A-11E the representation of FIG. 1100b, the darkened portion of the re-timer flexible buffer represents the “active” portion of the flexible buffer, or the portion that is configured to buffer data. As the re-timer is readjusted in size, the unused portion of the buffer can be represented as the lightened portion of the flexible buffer in Figures 11A-11E the representation of FIG. 1100c. Thus, in Figure 11AIn some examples, the re-timer 910 can initially set the size of its elastic buffer 930a to a default size (e.g., its maximum size (as shown by the fully blacked out re-timer 930a)). The maximum size of the elastic buffer can correspond to the size needed to handle worst case scenarios related to ppm mismatch, SKP intervals, link speed, link width, maximum packet payload size, etc. on the link. The elastic buffer can be designed to keep capacity as close to half full as possible by adding or subtracting a sign from the SKP OS received at the re-timer 910 and buffered into the elastic buffer. In some examples, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples. Figure 11A Continuing with the example above, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples.
[0087] Continuing with the example above, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples. Figure 11A Continuing with the example above, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples. Figure 11B Continuing with the example above, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples.
[0088] Continuing with the example above, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples. Figure 11A Continuing with the example above, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples. 11B Continuing with the example above, the re-timer 910 can be configured to keep the elastic buffer filled to half of its maximum capacity. This initial or default size can be based on initial boot-up or start-up of the system or links of the system, among other examples. Figure 11BIn the example, the change to SRIS mode applied at the corresponding clocks of devices A and B can be detected in data 1110 by the retimer 910, in which case SRIS modulation is disabled. As in Figure 11C The updated elastic buffer shown in 930c reflects that the re-timer 910 can again dynamically resize its elastic buffer, resulting in changes to detected link characteristics, in which case the size of the elastic buffer is further reduced to reflect the reduction in ppm caused by the clock architecture of the devices (e.g., 905, 915).
[0089] While the examples above discussed changes to link characteristics that can be detected by enhanced retimers, leading to a reduction in the resilient buffer size of the retimer, other link characteristic modifications can result in an increase in the resilient buffer size. For example, reducing the resilient buffer size based on data 1110... Figure 11C When the elastic buffer size is small, the link and retimer can restart operation when another data stream (e.g., 1115) is sent, indicating yet another modification to the (different or identical) link characteristics. Retimer 910 can similarly detect and properly interpret the data included in data stream 1115 (e.g., TS or other ordered sets) to determine the link characteristic change and further determine whether the change (e.g., a reduction in link width, re-enabling of SRIS modulation features, an increase in maximum payload size, etc.) is related to an increase in the elastic buffer size (e.g., ...). Figure 11D (As reflected in the updated elastic buffer 930d). Furthermore, although... Figures 11A-11C The example can be interpreted as indicating that device A is the initiator of the link characteristic modification represented by data 1105, 1110, and 1115, but other link partners (device B) of retimer 910 can similarly be able to initiate or signal changes to one or more link characteristics. For example, retimer 910 can parse data 1120 to detect another link characteristic change and, in response, reduce the size of its elastic buffer again (e.g., ...). Figure 11E The representation of the elastic buffer 930e in the text reflects this.
[0090] In some cases (e.g., during link training), the data stream transmitted by the re-timer 910 can indicate multiple related link characteristic modifications. Moreover, detecting a particular link characteristic modification can be based not only on the data stream from one of the link partners (e.g., device A) of the re-timer, but also (based on protocol definition) on data streams received from both link partners (e.g., devices A and B), e.g., by detecting a handshake indicating both a request and an acknowledgement of a suggested link characteristic modification, among other examples. In this sense, the logic of the re-timer, while based at least in part on each lane implementation, can be coupled with consolidated data detected at both its downstream receiver port and upstream receiver port, and make elastic buffer size readjustment decisions based on these findings. Additionally, while some link characteristic modifications can be detected directly from in-band data streams, out-of-band messaging can also be utilized to communicate or indicate link characteristic changes (e.g., a change to maximum payload size) to the re-timer 910, allowing the re-timer to size its elastic buffer accordingly, among other examples.
[0091] Figure 12 is a flowchart 1200 illustrating example techniques involving a re-timer configured to dynamically adjust the size of its elastic buffer based on modifications to characteristics of a link using the re-timer and connecting two devices. A data stream can be received 1205 by the re-timer, such as a data stream generated by one of the two devices. In some cases, the data stream can include data transmitted (e.g., in a downstream direction) from one of the two devices to the other, as well as data transmitted in response (e.g., in an upstream direction). The re-timer can scan its reception and regenerate for forwarding onto the link to detect 1210 (e.g., from an ordered set, training sequence, handshake, encoding, or other signaling in the data stream) a modification to one or more of a set of characteristics of the link. The set of characteristics of interest can include those used as a basis for worst-case sizing of the re-timer’s elastic buffer. Detection of a change (or a specification of an actual value or state) to one of these link characteristics can be processed by the re-timer to determine a size readjustment to the elastic buffer in order to reduce the latency introduced by the elastic buffer on the link. For example, the re-timer can determine a type and extent of the link characteristic modification, and consult a table or other mapping to calculate, identify, or otherwise determine 1215 an amount by which the elastic buffer size can be modified (i.e., increased or decreased) based on the link characteristic modification. Moreover, in response to detecting 1210 a link characteristic modification, the re-timer can automatically and dynamically change 1220 the size of its elastic buffer by the amount determined 1215 based on the link characteristic modification to more accurately reflect the updated “worst case,” among other examples.
[0092] Note that the apparatuses, methods, and systems described above can be implemented in any of the previously described electronic devices or systems. As a particular illustration, the following figures provide exemplary systems for utilizing the present application as described herein. As the following systems are described in more detail, many different interconnections are disclosed, described, and reconsidered in light of the above discussion. And it will be apparent that the improvements described above can be applied to any of those interconnections, structures, or architectures.
[0093] Reference Figure 13 FIG. 1 depicts an embodiment of a block diagram of a computing system including a multi-core processor. Processor 1300 includes any processor or processing device, such as a microprocessor, embedded processor, digital signal processor (DSP), network processor, handheld processor, application processor, co-processor, system on a chip (SOC), or other device for executing code. In one embodiment, processor 1300 includes at least two cores - cores 1301 and 1302, which can include asymmetric cores or symmetric cores (the illustrated embodiment). However, processor 1300 can include any number of processing elements that can be symmetric or asymmetric.
[0094] In one embodiment, a processing element refers to hardware or logic that is used to
[0095] A core typically refers to logic located on an integrated circuit that is capable of maintaining an independent architectural state, where each independently maintained architectural state is associated with at least some dedicated execution resources. In contrast to a core, a hardware thread typically refers to any logic located on an integrated circuit that is capable of maintaining an independent architectural state, where the independently maintained architectural states share access to execution resources. As can be seen, the line between the nomenclature of hardware threads and cores blurs when some resources are shared and others are dedicated to an architectural state. However, typically, cores and hardware threads are treated as separate logical processors by an operating system, where the operating system is capable of scheduling operations on each logical processor independently.
[0096] As Figure 13As shown, physical processor 1300 includes two cores - core 1301 and 1302. Here, cores 1301 and 1302 are considered symmetric cores, i.e., cores that have the same configuration, functional units, and / or logic. In another embodiment, core 1301 includes an out-of-order processor core, while core 1302 includes an in-order processor core. However, cores 1301 and 1302 can be individually selected from any of the types of cores, e.g., native cores, software managed cores, cores adapted to execute a native instruction set architecture (ISA), cores adapted to execute a translated instruction set architecture (ISA), co-designed cores, or other known cores. In a heterogeneous core environment (i.e., asymmetric cores), some form of translation (e.g., binary translation) can be utilized to schedule or execute code on one or both cores. However, for further discussion, the functional units shown in core 1301 are described in further detail below, as the units in core 1302 operate in a similar manner in the depicted embodiment.
[0097] As depicted, core 1301 includes two hardware threads 1301a and 1301b, which can also be referred to as hardware thread slots 1301a and 1301b. Thus, in one embodiment, a software entity such as an operating system potentially views processor 1300 as four separate processors, i.e., four logical processors or processing elements capable of executing four software threads concurrently. As mentioned above, a first thread is associated with architectural state registers 1301a, a second thread is associated with architectural state registers 1301b, a third thread can be associated with architectural state registers 1302a, and a fourth thread can be associated with architectural state registers 1302b. Here, as described above, each of the architectural state registers (1301a, 1301b, 1302a, and 1302b) can be referred to as a processing element, thread slot, or thread unit. As shown, architectural state registers 1301a are replicated in architectural state registers 1301b, thus enabling separate architectural state / context to be stored for logical processor 1301a and logical processor 1301b. In core 1301, other smaller resources can also be replicated for threads 1301a and 1301b, e.g., instruction pointers and renaming logic in allocator and renamer block 1330. Some resources can be shared through partitioning, e.g., reorder buffers in reorder / retirement unit 1335, ILTB 1320, load / store buffers and queues. Other resources (e.g., general-purpose internal registers, page table(s) base register(s), low-level data cache and data TLB 1315, execution unit(s) 1340, and portions of out-of-order unit 1335) are potentially fully shared.
[0098] The processor 1300 typically includes other resources that can be fully shared, shared through partitioning, or dedicated / dedicate-able to a processing element. In Figure 13 In the illustrated embodiment, an exemplary processor is shown having illustrative logical units / resources of a processor. Note that the processor can include or omit any of these functional units, as well as include any other known functional units, logic, or firmware not depicted. As shown, the core 1301 includes a simplified, representative out-of-order (OOO) processor core. However, in-order processors can be utilized in different embodiments. The OOO core includes a branch target buffer 1320 to predict branches to be executed / taken and an instruction translation buffer (I-TLB) 1320 to store address translation entries for instructions.
[0099] The core 1301 also includes a decode module 1325 coupled with the fetch unit 1320 to decode fetched elements. In one embodiment, the fetch logic includes separate sequencers associated with thread slots 1301a, 1301b, respectively. Generally, the core 1301 is associated with a first ISA that defines / specifies instructions executable on the processor 1300. Typically, machine code instructions that are part of the first ISA include a portion of the instruction (referred to as an opcode) that references / specifies an instruction or operation to be performed. The decode logic 1325 includes circuitry that recognizes these instructions from their opcodes and passes the decoded instructions along in the pipeline for processing defined by the first ISA. For example, as discussed in more detail below, in one embodiment the decoder 1325 includes logic designed or adapted to recognize particular instructions (e.g., transactional instructions). As a result of the recognition by the decoder 1325, the architecture or core 1301 takes particular predefined actions to perform tasks associated with the appropriate instructions. It is important to note that any of the tasks, blocks, operations, and methods described herein can be performed in response to a single or multiple instructions; some of which can be new or old instructions. Note that in one embodiment the decoder 1326 recognizes the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, the decoder 1326 recognizes a second ISA (a subset of the first ISA or a different ISA).
[0100] In one example, the allocator and renamer block 1330 includes an allocator for reserving resources (e.g., register files for storing instruction processing results). However, the threads 1301a and 1301b are potentially capable of out-of-order execution, where the allocator and renamer block 1330 also reserves other resources, such as reorder buffers for tracking instruction results. The unit 1330 can also include a register renamer to rename program / instruction reference registers to other registers internal to the processor 1300. The reorder / retirement unit 1335 includes components such as the reorder buffers, load buffers, and store buffers mentioned above to support out-of-order execution and later in-order retirement of out-of-order executed instructions.
[0101] In one embodiment, the scheduler(s) and execution unit block 1340 includes a scheduler unit to schedule instructions / operations on execution units. For example, floating point instructions are scheduled on ports of execution units that have available floating point execution units. Register files associated with the execution units are also included to store information instruction processing results. Example execution units include floating point execution units, integer execution units, jump execution units, load execution units, store execution units, and other known execution units.
[0102] A lower level data cache and data translation buffer (D-TLB) 1350 is coupled with the execution unit(s) 1340. The data cache is used to store recently used / operated on elements, such as data operands, which are potentially held in a memory coherency state. The D-TLB is used to store recent virtual / linear to physical address translations. As a particular example, the processor can include a page table structure to divide physical memory into a plurality of virtual pages.
[0103] Here, the cores 1301 and 1302 share access to higher level or further away caches, such as a second level cache associated with the on-chip interface 1310. Note that higher level or further away refers to an increase in cache level or further away from the execution unit(s). In one embodiment, the higher level cache is a last level data cache - the last cache in the memory hierarchy on the processor 1300 - such as a second or third level data cache. However, the higher level cache is not limited to this as it can be associated with or include an instruction cache. A trace cache - a type of instruction cache - can instead be coupled after the decoders 1325 to store recently decoded traces. Here, instructions potentially refer to macro instructions (i.e., general purpose instructions) that can be decoded into a plurality of micro instructions (micro-ops).
[0104] In the depicted configuration, the processor 1300 also includes an on-chip interface module 1310. Historically, the memory controllers described in greater detail below have been included in computing systems external to the processor 1300. In such a scenario, the on-chip interface 1310 is used to communicate with devices external to the processor 1300 (e.g., system memory 1375, a chipset (typically including a memory controller hub for connecting to memory 1375 and an I / O controller hub for connecting to peripheral devices), a memory controller hub, a northbridge, or other integrated circuit). And in such a scenario, the bus 1305 can include any known interconnect, such as a multi-drop bus, a point-to-point interconnect, a serial interconnect, a parallel bus, a coherent (e.g., cache coherent) bus, a hierarchical protocol architecture, a differential bus, and a GTL bus.
[0105] The memory 1375 can be dedicated to the processor 1300 or shared with other devices in the system. Common examples of memory 1375 types include DRAM, SRAM, non-volatile memory (NV memory), and other known memory devices. Note that the devices 1380 can include a graphics accelerator, a processor or card coupled with a memory controller hub, a data storage device coupled with an I / O controller hub, a wireless transceiver, a flash device, an audio controller, a network controller, or other known devices.
[0106] However, more recently, as more logic and devices are integrated on a single die, such as an SOC, each of these devices can be consolidated on the processor 1300. For example, in one embodiment, the memory controller hub is on the same package and / or die as the processor 1300. Here, a portion of the core (on-core portion) 1310 includes one or more controllers for interfacing with other devices, such as the memory 1375 or the graphics device 1380. The configuration including the interconnects and controllers for interfacing with these devices is often referred to as on-core (or non-core configuration). As an example, the on-chip interface 1310 includes a ring interconnect for on-chip communications and a high-speed serial point-to-point link 1305 for off-chip communications. However, in an SOC environment, even more devices (e.g., network interfaces, coprocessors, memory 1375, graphics processor 1380, and any other known computer device / interface) can be integrated onto a single die to provide a small form factor with high functionality and low power consumption.
[0107] In one embodiment, the processor 1300 is capable of executing a compiler, optimization, and / or translator code 1377 to compile, translate, and / or optimize application code 1376 to support the apparatus and methods described herein or interface therewith. A compiler generally includes a program or set of programs to convert source text / code to object text / code. Generally, the compilation of program / application code using a compiler is done in multiple stages and passes to convert high level programming language code to low level machine or assembly language code. However, a single pass compiler can still be used for simple compilation. The compiler can utilize any known compilation techniques and perform any known compiler operations, such as, for example, lexical analysis, preprocessing, parsing, semantic analysis, code generation, code translation, and code optimization.
[0108] Larger compilers generally include multiple stages, but most of the time these stages are included in two general phases: (1) the front end, i.e., generally where syntax processing, semantic processing, and some transformations / optimizations can occur, and (2) the back end, i.e., generally where analysis, translation, optimization, and code generation occur. Some compilers refer to an intermediate section that accounts for the ambiguity of the delineation between the front end and back end of the compiler. Thus, references to the insertion, association, generation, or other manipulation of a compiler can occur in any of the aforementioned stages or passes as well as any other known stages or passes of a compiler. As an illustrative example, a compiler potentially inserts an operation, call, function, etc. in one or more stages of compilation, such as, for example, inserting a call / operation in the front end stage of compilation and then translating the call / operation to a lower level of code during the translation stage. Note that during dynamic compilation, compiler code or dynamic optimization code can insert such operations / calls as well as optimization code to be executed during runtime. As a specific illustrative example, binary code (compiled code) can be dynamically optimized during runtime. Here, the program code can include dynamic optimization code, binary code, or a combination thereof.
[0109] Similar to a compiler, a translator (e.g., binary translator) statically or dynamically translates code to optimize and / or translate code. Thus, references to the execution of code, application code, program code, or other software environment can refer to: (1) dynamically or statically executing compiler program(s), optimization code optimizer, or translator to compile program code, maintain software structure, perform other operations, optimize code, or translate code; (2) executing host program code that includes operations / calls, such as, for example, application code that has been optimized / compiled; (3) executing other program code (e.g., libraries) associated with the host program code to maintain software structure, perform other software related operations, or optimize code; or (4) a combination thereof.
[0110] Reference is now made to Figure 14FIG. 14 shows a block diagram of a second system 1400 in accordance with an embodiment of the present application. As shown, the multiprocessor system 1400 is a point-to-point interconnect system, and includes a first processor 1470 and a second processor 1480 coupled via a point-to-point interconnect 1450. Each of the processors 1470 and 1480 can be some version of the Intel® Itanium® processor. In one embodiment, 1452 and 1454 are part of a serial point-to-point coherent interconnect fabric, such as Intel® QuickPath Interconnect (QPI). Thus, the present application can be implemented within a QPI fabric. Figure 14
[0111] While only two processors 1470, 1480 are shown, it is to be understood that the scope of the present application is not so limited. In other embodiments, one or more additional processors can be present in a given processor.
[0112] The processors 1470 and 1480 are shown including integrated memory controller units 1472 and 1482, respectively. The processor 1470 also includes point-to-point (P-P) interface circuits 1476 and 1478 as part of a bus controller unit of the processor 1470; similarly, the second processor 1480 includes P-P interface circuits 1486 and 1488. The processors 1470, 1480 can exchange information via a point-to-point (P-P) interface 1450 using P-P interface circuits 1478, 1488. As Figure 14 shown, the IMCs 1472 and 1482 couple the processors to respective memories, namely a memory 1432 and a memory 1434, which can be part of a main memory attached to a respective processor.
[0113] The processors 1470, 1480 each exchange information with a chipset 1490 via individual P-P interfaces 1452, 1454 using point-to-point interface circuits 1476, 1494, 1486, 1498. Chipset 1490 also exchanges information with a high-performance graphics circuit 1438 via an interface circuit 1492 along a high-performance graphics interconnect 1439.
[0114] A shared cache (not shown) can be included in either processor or outside of both processors; connected with them via P-P interconnects, such that each processor's local cache information can be stored in the shared cache if processors are placed into a low power mode.
[0115] The chipset 1490 can be coupled with a first bus 1416 via an interface 1496. In one embodiment, the first bus 1416 can be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, although the scope of the present application is not so limited.
[0116] AsFigure 14 As shown, various I / O devices 1414 are coupled to the first bus 1416 along with a bus bridge 1418 that couples the first bus 1416 to a second bus 1420. In one embodiment, the second bus 1420 includes a low pin count (LPC) bus. In one embodiment, various devices are coupled to the second bus 1420 including, for example, a keyboard and / or mouse 1422, communication device 1427 and storage unit 1428 (e.g., disk drive or other mass storage device that typically includes instructions / code and data 1430). Further, audio I / O 1424 is shown coupled to the second bus 1420. Note that other architectures are possible, wherein the included components and interconnect architectures are different. For example, in addition to or instead of a point-to-point architecture, a system can implement a multi-drop bus or other such architecture. Figure 14
[0117] While the application has been described with respect to a limited number of embodiments, those skilled in the art will appreciate numerous modifications and variations therefrom. It is intended that the appended claims cover all such modifications and variations as fall within the true spirit and scope of this application.
[0118] A design can go through various stages, from creation to simulation to fabrication. Data representing a design can represent the design in a number of manners. First, as is useful in simulations, the hardware can be represented using a hardware description language or another functional description language. Additionally, a circuit level model with logic and / or transistor gates can be produced at some stages of the design process. Furthermore, most designs reach a level at which there are no longer changes made to the datas, but only additions to faster or corrected versions of the datas, so the designers will typically want to the design process to verify operations at some stage and to sign off the resulting design. Designs also typically go through various stages of testing. Running a simulation operation of a version of design's data through a diagnostic program to verify the functionality of the design can be conducted in each of the stages of testing. Following each stage in the design process, designers can use tests based on or including at least a portion of the design's data to make sure that the data for the stages of the designs they are designing are correct and function as the designers intend.
[0119] A module, as used herein, refers to any combination of hardware, software, and / or firmware. As an example, a module includes hardware, such as a micro-controller, associated with a non-transitory medium containing code that is executed by the micro-controller. Therefore, a module can be, but is not limited to, a hardware-only module, a software-only module, or a combination of both hardware and software. Moreover, a module can be implemented as part of a larger system, for example, a micro-controller can be a module that is part of a larger system. In one embodiment, a module is a hardware-only module that includes a micro-controller, and a non-transitory medium containing code that is executed by the micro-controller. In another embodiment, a module is a hardware-only module that includes a micro-controller, and a non-transitory medium containing code that is executed by the micro-controller. In yet another embodiment, the term module (in this example) can refer to the combination of the micro-controller and the non-transitory medium. Generally, the boundaries of modules are therefore arbitrary, and in practice, two adjacent modules can share hardware, software, firmware, or a combination thereof, while potentially retaining some independent hardware, software, firmware or combination thereof. In one embodiment, the use of the term logic includes hardware, for example, transistors, registers or other hardware (such as programmable logic devices).
[0120] In one embodiment, the use of the phrase "configured to" denotes that a device, hardware, logic, or element is designed, placed, arranged, manufactured, provided, introduced, and / or designed to perform a specified or determined task. In this example, a device or element thereof that is not operating is still "configured to" perform a specified task if it is designed, coupled, and / or interconnected to perform the specified task. As a purely illustrative example, a logic gate can provide a 0 or a 1 during operation. However, a logic gate "configured to" provide an enable signal to a clock does not include every potential logic gate that can provide a 1 or a 0. Rather, the logic gate is a logic gate that is coupled in such a way that it outputs a 1 or a 0 during operation to enable the clock. Note again that the use of the term "configured to" does not require operation, but rather focuses on the potential state of the device, hardware, and / or element, where in the potential state the device, hardware, and / or element is designed to perform a particular task when the device, hardware, and / or element is operating.
[0121] Further, in one embodiment, the use of the phrases "for," "to," and / or "operable to" refers to a device, logic, hardware, and / or element that is designed in such a manner that the device, logic, hardware, and / or element can be used in a specified manner. As noted above, in one embodiment, the use of for, to, or operable to refers to a potential state of the device, logic, hardware, and / or element, where the device, logic, hardware, and / or element is not operating, but is designed in such a manner that the device can be used in a specified manner.
[0122] Values, as used herein, include any known representation of a number, state, logic state, or binary logic state. Typically, use of a logic level, logic value, or logical value also refers to use of 1 and 0, which simply represent binary logic states. For example, 1 refers to a high logic level, and 0 refers to a low logic level. In one embodiment, a storage unit, such as a transistor or a flash memory cell, is capable of holding a single logic value or multiple logic values. However, other representations of values in computer systems have been used. For example, the decimal number ten can also be represented as the binary value 1010 and the hexadecimal letter A. Thus, values include any representation of information capable of being maintained in a computer system.
[0123] Furthermore, a state can be represented by a value or portion of a value. As an example, a first value, such as a logic 1, can represent a default or initial state, while a second value, such as a logic 0, can represent a non-default state. Additionally, in one embodiment, the terms reset and set refer to a default value or state and an updated value or state, respectively. For example, a default value potentially includes a high logic value (i.e., reset), while an updated value potentially includes a low logic value (i.e., set). Note that any combination of values can be used to represent any number of states.
[0124] Embodiments of the methods, hardware, software, firmware or code set forth above can be implemented via instructions or code stored on a machine-accessible, machine- readable, computer-accessible, or computer-readable medium. Non-transitory machine- accessible / readable media include any mechanism that provides (i.e., stores and / or transmits) information in a form accessible by a machine (e.g., a computing device, electronic system, etc.). For example, the non-transitory, machine- accessible medium includes random-access memory (RAM) (e.g., static RAM (SRAM) or dynamic RAM (DRAM)), ROM, magnetic or optical storage medium, flash memory devices, electrical storage devices, optical storage devices, acoustical storage devices, other form of storage devices for holding information
[0125] Instructions used to program logic to perform embodiments of the present application can be stored within a memory in the system, such as DRAM, cache, flash memory, or other storage. Furthermore, instructions can be distributed via a network or by way of other computer readable media. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), but is not limited to, soft disks, optical disks, magnetic
[0126] The following examples relate to embodiments in accordance with the present specification. Example 1 is an apparatus comprising a retimer including an elastic buffer, buffer logic, a receiver, and a controller. The buffer logic can add or subtract data in the elastic buffer to compensate for different bit rates of two devices to be connected by a link, where the retimer is located between the two devices on the link. The receiver can receive a data stream to be transmitted between the two devices on the link. The controller can determine a modification to one or more characteristics of the link from the data stream, and cause a size of the elastic buffer to change from a first size to a second size based on the modification.
[0127] Example 2 can include the subject matter of Example 1, wherein the one or more characteristics include a link width of the link.
[0128] Example 3 can include the subject matter of Example 2, wherein the size of the elastic buffer decreases in response to an increase in the link width, and decreases in response to a decrease in the link width.
[0129] Example 4 can include the subject matter of any of Examples 1-3, wherein the modification includes a change to a clocking pattern applied at a clock of the two devices.
[0130] Example 5 can include the subject matter of Example 4, wherein the change to the clocking pattern includes enabling or disabling spread spectrum modulation at the clock.
[0131] Example 6 can include the subject matter of any of Examples 1-5, wherein the controller is further to determine a second modification to another characteristic of the link based on a sideband message sent from one of the two devices to the retimer; and cause another change to the size of the elastic buffer based on the second modification.
[0132] Example 7 can include the subject matter of Example 6, wherein the second modification comprises a modification to a maximum packet size supported on the link.
[0133] Example 8 can include the subject matter of any one of Examples 1-7, wherein the size of the elastic buffer is to be changed in a link training phase of the link.
[0134] Example 9 can include the subject matter of Example 8, wherein the data stream is received during the link training phase, the data stream comprises one or more ordered sets, and the modification is determined from the one or more ordered sets.
[0135] Example 10 can include the subject matter of Example 8, wherein the data stream is received during an active link state of the link, and causing the size of the elastic buffer to change from the first size to the second size comprises forcing a reset of the link to cause the link training phase, and the change to the size of the elastic buffer is completed during the link training phase.
[0136] Example 11 can include the subject matter of Example 10, wherein forcing the reset of the link comprises changing a particular data in the data stream received from a first device of the two devices to a particular value; and sending the changed particular data to a second device of the two devices, wherein the particular value artificially causes an error condition on the link.
[0137] Example 12 can include the subject matter of any one of Examples 1-11, wherein the buffer logic is to add or subtract data from a particular ordered set that is sent in a loop in the data stream and buffered at the elastic buffer.
[0138] Example 13 can include the subject matter of Example 12, wherein the particular ordered set is defined according to a particular protocol specification for compensating for different bit rates of devices on the link.
[0139] Example 14 can include the subject matter of Example 13, wherein the particular ordered set comprises a SKP ordered set.
[0140] Example 15 can include the subject matter of any one of Examples 13-14, wherein the particular protocol specification comprises a specification of one of a Universal Serial Bus (USB) based protocol or a Peripheral Component Interconnect Express (PCIe) based protocol.
[0141] Example 16 can include the subject matter of any one of Examples 1-15, wherein the first size comprises a size corresponding to a worst case elastic buffer size of the link.
[0142] Example 17 is a method comprising: receiving a data stream at a re-timer, where the data stream is to be transmitted between two devices connected over a link, the re-timer is located between the two devices over the link, and the re-timer includes an elastic buffer to add or subtract data in the elastic buffer to compensate for a clock frequency variation between clocks of the two devices; detecting an indication of a modification to one or more characteristics of the link in the data stream; determining a particular amount of change to be made to a size of the elastic buffer based on the modification; and changing the size of the elastic buffer by the particular amount.
[0143] Example 18 can include the subject matter of Example 17, wherein the one or more characteristics include a link width of the link.
[0144] Example 19 can include the subject matter of Example 18, wherein the size of the elastic buffer is decreased in response to an increase in the link width, and is decreased in response to a decrease in the link width.
[0145] Example 20 can include the subject matter of any of Examples 17-19, wherein the modification includes a change to a timing mode applied at the clocks of the two devices.
[0146] Example 21 can include the subject matter of Example 20, wherein the change to the timing mode includes enabling or disabling spread spectrum modulation at the clocks.
[0147] Example 22 can include the subject matter of any of Examples 17-21, further comprising: determining a second modification to another characteristic of the link based on a sideband message transmitted from one of the two devices to the re-timer; and causing another change to the size of the elastic buffer based on the second modification.
[0148] Example 23 can include the subject matter of Example 22, wherein the second modification includes a modification to a maximum packet size supported over the link.
[0149] Example 24 can include the subject matter of any of Examples 17-23, wherein the size of the elastic buffer is to be changed in a link training phase of the link.
[0150] Example 25 can include the subject matter of Example 24, wherein the data stream is received during the link training phase, the data stream includes one or more ordered sets, and the modification is determined from the one or more ordered sets.
[0151] Example 26 can include the subject matter of Example 24, wherein the data stream is received during an active link state of the link, and causing the size of the elastic buffer to change from a first size to a second size includes: forcing a reset of the link to cause the link training phase; and completing the change to the size of the elastic buffer during the link training phase.
[0152] Example 27 can include the subject matter of Example 26, wherein forcing a reset of the link comprises changing a particular data in a data stream received from a first device of the two devices to a particular value that artificially causes an error condition on the link; and sending the changed particular data to a second device of the two devices.
[0153] Example 28 can include the subject matter of any of Examples 17-27, wherein the elastic buffer is to add or subtract data from a particular ordered set that is sent in a loop in the data stream and buffered at the elastic buffer.
[0154] Example 29 can include the subject matter of Example 28, wherein the particular ordered set is defined according to a particular protocol specification for compensating for different bit rates of devices on the link.
[0155] Example 30 can include the subject matter of Example 29, wherein the particular ordered set comprises a SKP ordered set.
[0156] Example 31 can include the subject matter of any of Examples 29-30, wherein the particular protocol specification comprises a specification of one of a Universal Serial Bus (USB) based protocol or a Peripheral Component Interconnect Express (PCIe) based protocol.
[0157] Example 32 can include the subject matter of any of Examples 17-31, wherein the first size comprises a size corresponding to a worst case elastic buffer size of the link.
[0158] Example 33 is a system comprising means for performing the method of any of Examples 17-32.
[0159] Example 34 is a machine-accessible storage medium having stored thereon instructions, which, when executed on a machine, cause the machine to: detect, in a data stream received by a re-timer, an indication of a modification to one or more characteristics of a link, wherein the data stream is to be transmitted between two devices connected on the link, the re-timer is located between the two devices on the link, and the re-timer comprises an elastic buffer to add or subtract data in the elastic buffer to compensate for a clock frequency variation between clocks of the two devices; determine, based on the modification, a particular amount of change to be made to a size of the elastic buffer; and cause the size of the elastic buffer to change by the particular amount.
[0160] Example 35 can include the subject matter of Example 34, wherein the one or more characteristics comprise a link width of the link.
[0161] Example 36 can include the subject matter of Example 35, wherein the size of the elastic buffer is decreased in response to an increase in the link width and is decreased in response to a decrease in the link width.
[0162] Example 37 can include the subject matter of any one of Examples 34-36, wherein the modification comprises a change to a timing mode applied at a clock of the two devices.
[0163] Example 38 can include the subject matter of Example 37, wherein the change to the timing mode comprises enabling or disabling spread spectrum modulation at the clock.
[0164] Example 39 can include the subject matter of any one of Examples 34-38, wherein the instructions, when executed, further cause the machine to determine a second modification to another characteristic of the link based on a sideband message sent from one of the two devices to the re-timer, and cause another change to the size of the elastic buffer based on the second modification.
[0165] Example 40 can include the subject matter of Example 39, wherein the second modification comprises a modification to a maximum packet size supported on the link.
[0166] Example 41 can include the subject matter of any one of Examples 34-40, wherein the size of the elastic buffer is to be changed in a link training phase of the link.
[0167] Example 42 can include the subject matter of Example 41, wherein the data stream is received during the link training phase, the data stream comprises one or more ordered sets, and the modification is determined from the one or more ordered sets.
[0168] Example 43 can include the subject matter of Example 41, wherein the data stream is received during an active link state of the link, and causing the size of the elastic buffer to change from the first size to the second size comprises forcing a reset of the link to cause the link training phase, and the change to the size of the elastic buffer is completed during the link training phase.
[0169] Example 44 can include the subject matter of Example 43, wherein forcing the reset of the link comprises changing a particular data in the data stream received from a first device of the two devices to a particular value, and sending the changed particular data to a second device of the two devices, wherein the particular value artificially causes an error condition on the link.
[0170] Example 45 can include the subject matter of any one of Examples 34-44, wherein the elastic buffer is to add or subtract data from a particular ordered set that is sent in a loop in the data stream and buffered at the elastic buffer.
[0171] Example 46 can include the subject matter of Example 45, wherein the particular ordered set is defined according to a particular protocol specification for compensating for different bit rates of devices on the link.
[0172] Example 47 can include the subject matter of Example 46, wherein the particular ordered set comprises a SKP ordered set.
[0173] Example 48 can include the subject matter of any one of Examples 46-47, wherein the specific protocol specification comprises a specification of one of a Universal Serial Bus (USB) based protocol or a Peripheral Component Interconnect Express (PCIe) protocol.
[0174] Example 49 can include the subject matter of any one of Examples 34-48, wherein the first size comprises a size corresponding to a worst-case elastic buffer size of the link.
[0175] Example 50 is a system comprising: a first device comprising a first clock; a second device comprising a second clock, wherein the first device is connected to the second device through a link; and a re-timer device located between the first device and the second device on the link. The re-timer comprises: an elastic buffer to add or subtract a specific data buffered in the elastic buffer to compensate for different frequencies of the first clock and the second clock; a receiver to receive a data stream to be transmitted on the link between the first device to the second device, and a controller. The controller can determine a modification to a characteristic of the link according to the data stream, and cause a size of the elastic buffer to change based on the modification.
[0176] Example 51 can include the subject matter of Example 50, wherein the re-timer device further comprises re-timer logic to regenerate the data stream received from the first device, and transmit the regenerated data stream to the second device.
[0177] Example 52 can include the subject matter of any one of Examples 50-51, wherein the controller comprises a controller according to a Physical Interface for PCI Express (PIPE) based interface.
[0178] Example 53 can include the subject matter of any one of Examples 50-52, wherein the characteristic comprises a link width of the link.
[0179] Example 54 can include the subject matter of Example 53, wherein the size of the elastic buffer is decreased in response to an increase in the link width, and decreased in response to a decrease in the link width.
[0180] Example 55 can include the subject matter of any one of Examples 50-52, wherein the characteristic comprises a change to a clocking mode applied at the clocks of the two devices.
[0181] Example 56 can include the subject matter of Example 55, wherein the change to the clocking mode comprises enabling or disabling spread spectrum modulation at the clocks.
[0182] Example 57 can include the subject matter of any one of Examples 50-56, wherein the controller is further to determine a second modification to another characteristic of the link based on a sideband message transmitted from one of the two devices to the re-timer, and cause another change to the size of the elastic buffer based on the second modification.
[0183] Example 58 can include the subject matter of Example 57, wherein the second modification comprises a modification to a maximum packet size supported on the link.
[0184] Example 59 can include the subject matter of any one of Examples 50-58, wherein the size of the elastic buffer is to be changed in a link training phase of the link.
[0185] Example 60 can include the subject matter of Example 59, wherein the data stream is received during the link training phase, the data stream comprises one or more ordered sets, and the modification is determined from the one or more ordered sets.
[0186] Example 61 can include the subject matter of Example 59, wherein the data stream is received during an active link state of the link, and causing the size of the elastic buffer to change from the first size to the second size comprises forcing a reset of the link to cause the link training phase, and the change to the size of the elastic buffer is completed during the link training phase.
[0187] Example 62 can include the subject matter of Example 61, wherein forcing the reset of the link comprises changing a particular data in the data stream received from a first device of the two devices to a particular value, and sending the changed particular data to a second device of the two devices, wherein the particular value artificially causes an error condition on the link.
[0188] Example 63 can include the subject matter of any one of Examples 50-62, wherein the elastic buffer is to add or subtract data from a particular ordered set that is sent in a loop in the data stream and buffered at the elastic buffer.
[0189] Example 64 can include the subject matter of Example 63, wherein the particular ordered set is defined according to a particular protocol specification for compensating for different bit rates of devices on the link.
[0190] Example 65 can include the subject matter of Example 64, wherein the particular ordered set comprises a SKP ordered set.
[0191] Example 66 can include the subject matter of any one of Examples 64-65, wherein the particular protocol specification comprises a specification of one of a Universal Serial Bus (USB) based protocol or a Peripheral Component Interconnect Express (PCIe) based protocol.
[0192] Example 67 can include the subject matter of any one of Examples 50-66, wherein the first size comprises a size corresponding to a worst case elastic buffer size of the link.
[0193] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily referring to the same embodiment. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0194] In the foregoing specification, specific embodiments have been described in some detail. It will be apparent, however, that various modifications and changes can be made to these details without departing from the broader spirit and scope of the application as set forth in the appended claims. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. Furthermore, the foregoing use of terminology in the detailed description of embodiments should not be taken to limit the scope of the present application as claimed.
Claims
1. A retimer device comprising: a retimer logic to: receive a first signal from a first device and regenerate the first signal for transmission to a second device; and receive a second signal from the second device and regenerate the second signal for transmission to the first device, wherein the first device comprises a processor device; a resilient buffer to compensate for different bit rates of the first device and the second device, wherein the retimer device is located between the first device and the second device on a link; and a controller to: monitor the first signal; determine a modification to one or more characteristics of the link from the first signal; and based on the modification, cause a size of the resilient buffer to change from a first size to a second size.
2. The retimer device of claim 1, wherein protocol activity comprises equalization of the link comprising the first device, the second device, and the retimer device, and the retimer device is located between the first device and the second device in the link.
3. The retimer device of claim 2, wherein the first signal identifies that equalization of the link is to be performed, and the controller is further to: receive an equalization signal from the second device; determine whether an error is present in the equalization signal; and generate a metric based on whether an error is present in the equalization signal, wherein a sideband interface is to provide the metric to the first device.
4. The retimer device of claim 3, wherein the equalization signal comprises a first instance of the equalization signal, and the controller is further to: generate a second instance of the equalization signal, wherein the first instance of the equalization signal is compared to the second instance of the equalization signal to determine whether an error is present in the first instance of the equalization signal; transmit the second instance of the equalization signal to the first device over the link; obtain feedback data from the first device; and adjust transmitter parameters of the retimer device based on the feedback data.
5. The retimer device of claim 4, wherein the feedback data is received over the sideband interface.
6. The retimer device of claim 4, wherein the feedback data is received during a slow mode.
7. The retimer device of claim 4, wherein the controller comprises a signal generator to generate instances of the equalization signal.
8. The retimer device of claim 7, wherein the equalization signal comprises a particular pseudo-random bit sequence (PRBS), and the signal generator comprises a linear feedback shift register (LFSR).
9. The retimer device of claim 4, wherein the controller is further to: receive a third instance of the equalization signal from the first device; determine whether an error is present in the third instance of the equalization signal; and generate a second metric based on whether an error is present in the third instance of the equalization signal. 10. The retimer device of claim 9, wherein the transmitter parameter corresponds to a first transmitter of the retimer device used to transmit data to the first device, a second transmitter of the retimer device is used to transmit data to the second device, and the controller is further to: transmit a fourth instance of the equalized signal to the second device; obtain a second metric generated by the second device based on the fourth instance of the equalized signal; and adjust the transmitter parameter corresponding to the second transmitter based on the second metric.
11. The re-timer device of claim 1, further comprising: receiver detection logic to: detect that another device is connected to a transmitter of the retimer device; and selectively enable termination of a receiver of the retimer device, wherein termination of the receiver is used to indicate to the first device that the other device is connected to the transmitter of the retimer device, and the other device comprises the second device.
12. The retimer device of claim 11, wherein receiver detection logic is further to enter a receiver detection state defined in a state machine of a protocol, and detect that the other device is connected to the transmitter during the receiver detection state. protocol logic to support a plurality of protocol activities for a plurality of different protocols.
13. The re-timer device of claim 1, further comprising:
14. The retimer device of claim 13, wherein the protocol logic comprises less than an entire protocol stack of one protocol of the plurality of different protocols.
15. The retimer device of claim 13, wherein the plurality of protocol activities comprises a speed change, and the retimer device is to operate at a different speed as a result of the speed change.
16. A method for link extension, comprising: receiving a signal from a processor at a retimer on a link, wherein the retimer comprises a resilient buffer used to compensate for different bit rates of devices on the link; regenerating the signal; sending the regenerated signal to another device on the link; identifying a modification to one or more characteristics of the link from the signal; and changing the resilient buffer from a first size to a second size based on the modification.
17. The method of claim 16, wherein a protocol activity comprises equalization of the link comprising a first device, a second device, and the retimer, and the retimer is located between the first device and the second device in the link.
18. The method of claim 17, wherein a first signal identifies that equalization of the link is to be performed, and the method further comprises: receiving an equalization signal from the second device; determining whether an error is present in the equalization signal; and generating a metric based on whether an error is present in the equalization signal, wherein a sideband interface is used to provide the metric to the first device.
19. The method of claim 18, wherein the equalization signal comprises a first instance of the equalization signal, and the method further comprises: generating a second instance of the equalized signal, wherein the first instance of the equalized signal is compared to the second instance of the equalized signal to determine whether an error exists in the first instance of the equalized signal; sending the second instance of the equalized signal to the first device over the link; obtaining feedback data from the first device; and adjusting a transmitter parameter of the re-timer based on the feedback data.
20. A system for link extension comprising means for performing the method of any of claims 16-19.
21. The system of claim 20, wherein the means comprises a machine accessible storage medium having instructions stored thereon, wherein the instructions, when executed on a machine, cause the machine to perform the method of any of claims 16-19.
22. An apparatus for link extension comprising: a processor device; a second device coupled to the processor device over a link; and a re-timer located between the processor device and the second device over the link, the re-timer comprising: a protocol stack comprising physical layer logic for performing one or more protocol specific events; a resilience buffer for compensating for different bit rates of the processor device and the second device; and a controller for: determining a modification to one or more characteristics of the link; and changing the resilience buffer from a first size to a second size based on the modification.
23. The apparatus of claim 22, wherein the second device comprises a second processor device.
24. The apparatus of claim 22, wherein the re-timer comprises a subset of the physical layer logic of the protocol stack, and auxiliary logic for supplementing the subset of the physical layer logic of the re-timer.
Citation Information
Patent Citations
Reducing Latency OF Unified Memory Transactions
US20150067433A1
Phase detector and retimer for clock and data recovery circuits
US20160028537A1
Dynamic data-link selection over common physical interface
US20170039162A1
Multi-function bypass port and port bypass circuit
US7474612B1