High-performance physical coupling structure layer

The HPI architecture addresses the challenges of high-performance computing by implementing a layered protocol stack and point-to-point links to enhance communication efficiency and power management in computer systems, achieving scalable and efficient performance across diverse computing platforms.

DE112013002880B4Active Publication Date: 2026-02-26INTEL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE112013002880
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-10-22
Filing Date
2013-03-27
Publication Date
2026-02-26
Estimated Expiration
2033-03-27

AI Technical Summary

Technical Problem

Existing coupling structure architectures in computer systems face challenges in meeting the increasing demands for high-performance computing and energy efficiency, particularly in multi-socket configurations, as they struggle to handle the communication and power requirements of advanced processors.

Method used

A high-performance coupling structure (HPI) architecture is introduced, featuring a layered protocol stack with a transaction layer, data link layer, and physical layer, utilizing point-to-point links and credit-based flow control to enhance communication efficiency and manage power consumption.

Benefits of technology

The HPI architecture improves communication bandwidth and reduces latency, enabling efficient power management and scalability across various computing platforms, including servers and mobile devices, while supporting diverse market segments with tailored performance and energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Device comprising: an interface logic for differential signaling on a plurality of paths, where the interface logic is configured to transmit a plurality of flits, wherein the plurality of flits has a plurality of nibbles, and wherein a first of the plurality of flits is transmitted with a clean flit boundary such that each of the plurality of paths is used to transmit initial nibbles of the first flit in a first unit interval (UI), and wherein initial nibbles of a second flit of the plurality of flits, together with final nibbles of the first flit, are transmitted on the plurality of paths during a subsequent UI.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The present disclosure relates generally to the field of computer development and in particular to software development, which involves the coordination of mutually dependent constrained systems. BACKGROUND

[0002] Advances in semiconductor processing and logic design have allowed for an increase in the amount of logic that can be present in integrated circuit devices. Consequently, computer system configurations have evolved from a single or multiple integrated circuits within a system to multiple cores, multiple hardware threads, and multiple logical processors present in individual integrated circuits, as well as other interfaces integrated within such processors. A processor or integrated circuit typically comprises a single physical processor chip layer, which can include any number of cores, hardware threads or logical processors, interfaces, memory, controller hubs, and so on.

[0003] As a result of the increased ability to pack more computing power into smaller packages, smaller computing devices have gained popularity. Smartphones, tablets, ultra-thin notebooks, and other subscriber terminals have proliferated exponentially. However, these smaller devices rely on servers for both data storage and complex processing that exceeds their form factor. Consequently, demand in the high-performance computing market (i.e., server storage space) has also increased. For example, modern servers typically contain not just a single multi-core processor, but multiple physical processors (also known as multi-sockets) to boost computing power. But as computing power increases along with the number of devices in a computer system, communication between sockets and other components becomes more critical.

[0004] In fact, coupling structures have evolved from more traditional multipoint interconnect buses, which primarily handle electrical communications, into fully mature coupling structure architectures that facilitate high-speed communication. Unfortunately, the corresponding demands are being placed on the capabilities of existing coupling structure architectures, while the demand for future processors with even higher power consumption rates is increasing.

[0005] US 2012 / 0079156A1 describes a method and apparatus for implementing the Intel QuickPath Interconnect® (QPL) protocol over a PCIe interface. The upper layers of the QPL protocol are implemented across a PCIe physical layer by applying QPL data bit mappings to corresponding PCIe x16, x8, and x4 lane configurations. A QPL interconnect layer to the PCIe physical layer interface is used to abstract the QPL interconnect, routing, and protocol layers from the underlying PCIe physical layer (and corresponding PCIe interface circuitry), thereby enabling QPL protocol messages to be used across PCIe hardware.

[0006] US 2009 / 0265472A1 describes initialization techniques for systems and components in a point-to-point architecture. These techniques enable flexible system / socket layer parameters that can be tailored to the requirements of the platform, such as desktop, mobile devices, small servers, large servers, etc., as well as to the component types, such as IA32 / IPF processors, memory controllers, I / O hubs, etc. Furthermore, the techniques facilitate booting with the correct point-of-care (POC) values, thereby avoiding multiple warm resets and improving boot time. In one embodiment, registers for storing new values, such as reset-controlled configuration values ​​(CVDR) and reset-captured configuration values ​​(CVCR), can be eliminated.

[0007] US 2011 / 0307577A1 describes a method for processing network data that involves a network interface controller (NIC) collecting a multitude of transmit buffer indicators (TX buffer indicators) in a multitude of priority lists of connections. Each of the multitude of TX buffer indicators identifies data ready to transmit that is outside the NIC and has not previously been received by the NIC. One or more of the multitude of TX buffer indicators can be selected. The identified data ready to transmit can be retrieved into the NIC based on the selected one or more TX buffer indicators. At least a portion of the identified data ready to transmit can be transmitted.

[0008] The invention is defined in the main claim and in dependent claims 12 and 17. BRIEF DESCRIPTION OF THE DRAWINGS Fig.Figure 1 illustrates a simplified block diagram of a system that includes a serial point-to-point coupling structure to connect I / O devices in a computer system according to one embodiment; Fig. Figure 2 illustrates a simplified block diagram of a shift log stack according to one embodiment; Fig. Figure 3 illustrates one embodiment of a transaction describer. Fig. Figure 4 illustrates an embodiment of a serial point-to-point link. Fig. Figure 5 illustrates embodiments of potential high-performance coupling structure (HPI) system configurations. Fig. Figure 6 illustrates an embodiment of a layer protocol stack connected to an HPI. Fig. Figure 7 illustrates a representation of an exemplary state machine. Fig. Figure 8 illustrates exemplary control supersequences. Fig. Figure 9 illustrates a flowchart that represents an exemplary entry into a partial width transmission state. Fig. Figure 10 illustrates a representation of an exemplary flit sent over an exemplary twenty-lane data link. Fig. Figure 11 illustrates a representation of an example flit sent over an example eight-lane data link. Fig. Figure 12 illustrates an embodiment of a block diagram for a computer system that includes a multi-core processor. Fig. Figure 13 illustrates another embodiment of a block diagram for a computer system that includes a multi-core processor. Fig. Figure 14 illustrates an embodiment of a block diagram for a processor. Fig.Figure 15 illustrates another embodiment of a block diagram for a computer system that includes a processor. Fig. Figure 16 illustrates an embodiment of a block for a computer system that includes a multiprocessor socket. Fig. Figure 17 illustrates another embodiment of a block diagram for a computer system.

[0009] Identical reference numbers and designations in the different drawings refer to similar elements. DETAILED DESCRIPTION

[0010] The following description presents numerous specific details, such as examples of certain types of processors and system configurations, specific hardware arrangements, specific details about architecture and microarchitecture, special register configurations, special instruction types, special system components, special processor pipeline stages, special coupling structure layers, special packet / transaction configurations, special transaction names, special protocol exchange operations, special link widths, special implementations and operations, etc., to ensure a thorough understanding of the present invention. However, it is obvious to a person skilled in the art that these specific details are not necessarily required to implement the subject matter of the present disclosure. In other cases, the well-detailed description of known components or methods, such as...Special and alternative processor architectures, special logic circuits / special code for described algorithms, special firmware code, special low-level interconnection operations, special logic configurations, special manufacturing techniques and materials, special compiler implementations, special implementation of algorithms in code, special shutdown and gating techniques / logic, and other special operating details of computer systems are not described in detail to avoid unnecessary obfuscation of the present invention.

[0011] Although the following embodiments may be described with reference to energy saving, energy efficiency, processing efficiency, and so forth, in relation to specific integrated circuits such as computer platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic assemblies. Similar techniques and teachings of embodiments described herein may be applied to other types of circuits or semiconductor devices that may also benefit from these features. For example, the disclosed embodiments are not limited to server computer systems, desktop computer systems, laptops, and Ultrabooks™, but may also be used in other devices such as handheld devices, smartphones, tablets, other thin notebooks, systems-on-a-chip (SoC) devices, and embedded applications. Some examples of handheld devices include, but are not limited to, the following:Mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Similar techniques for a high-performance coupling structure can be applied here to increase performance (or even save energy) in a low-energy coupling structure. Embedded applications typically include, among other things, a microcontroller, digital signal processor (DSP), system-on-a-chip, network computer (NetPC), set-top boxes, network hubs, wide area network (WAN) switches, or other systems capable of performing the functions and operations described below. Furthermore, the devices, methods, and systems described here are not limited to physical computing equipment but can also involve software optimizations for energy saving and efficiency.As is readily apparent from the following description, the embodiments of the methods, devices and systems described herein (whether referring to hardware, firmware, software or a combination thereof) can be considered crucial for a “green technology” future, balanced with performance considerations.

[0012] As computer systems evolve, their components become more complex. The coupling architecture used to connect and communicate between these components has also increased in complexity to ensure that bandwidth requirements are met for optimal component operation. Furthermore, different market segments require different aspects of coupling architectures to suit their respective markets. For example, servers require higher performance, while the mobile ecosystem is sometimes able to sacrifice overall performance for energy savings. Yet, the singular purpose of most coupling structures is to deliver the highest possible performance with maximum energy efficiency. Therefore, a wide variety of different coupling structures can potentially benefit from the subject matter described here.

[0013] The Peripheral Component Interconnect (PCI) Express (PCIe) interconnect architecture and the QuickPath Interconnect (QPI) interconnect architecture, among others, can potentially be improved according to one or more of the principles described here. For example, a primary goal of PCIe is to enable components and devices from different vendors to interoperate in an open architecture spanning multiple market segments: clients (desktop and mobile), servers (standard and enterprise), and embedded and communications devices. PCI Express is a universal, high-performance I / O interconnect architecture for a wide variety of computing and communications platforms.Some PCI attributes, such as its usage model, load-store architecture, and software interfaces, have been retained across revisions, while previous parallel bus implementations have been replaced by a highly scalable, fully serial interface. Newer versions of PCI Express leverage advancements in point-to-point coupling structures, switch-based technology, and packaged protocol to deliver new levels of performance and features. Power management, Quality of Service (QoS), hot-plug / hot-swap support, data integrity, and error handling are some of the advanced features supported by PCI Express.Although the primary discussion herein refers to a new HPI architecture, aspects of the invention described herein may be applied to other coupling structure architectures, such as a PCIe-compliant architecture, a QPI-compliant architecture, a MIPI-compliant architecture, a high-performance architecture, or any other known coupling structure architecture.

[0014] With reference to Fig.Figure 1 illustrates an embodiment of a structure consisting of point-to-point links connecting a set of components. The system 100 includes a processor 105 and system memory 110 coupled to the controller hub 115. The processor 105 can include any processing elements, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or another processor. The processor 105 is coupled to the controller hub 115 via the front-side bus (FSB) 106. In one embodiment, the FSB 106 is a serial point-to-point coupling structure, as described below. In another embodiment, the link 106 includes a serial differential coupling structure architecture conforming to a different coupling structure standard.

[0015] System memory 110 comprises any memory unit, such as random access memory (RAM), non-volatile (NV) memory, or other memory accessible to the components of system 100. System memory 110 is coupled to the controller hub 115 via memory interface 116. Examples of memory interfaces include dual-rate DDR memory interface, dual-channel DDR memory interface, and dynamic RAM (DRAM) memory interface.

[0016] In one embodiment, the controller hub 115 can include a root hub, root complex, or root controller, as in a PCIe connection hierarchy. Examples of a controller hub 115 include a chipset, memory controller hub (MCH), northbridge, coupling structure controller hub (ICH), southbridge, and a root controller / hub. Often, the term chipset refers to two physically separate controller hubs, such as a memory controller hub (MCH) coupled to a coupling structure controller hub (ICH). It should be noted that current systems often include the MCH integrated within the processor 105, while the controller 115 communicates with I / O devices in a manner similar to that described below. In some embodiments, peer-to-peer routing is optionally supported by the root complex 115.

[0017] Here, the controller hub 115 is coupled to the switch / bridge 120 via the serial link 119. The I / O modules 117 and 121, which can also be referred to as interfaces / ports 117 and 121, can include / implement a layered protocol stack to provide communication between the controller hub 115 and the switch 120. In one embodiment, multiple devices are capable of being coupled to the switch 120.

[0018] Switch / bridge 120 routes packets / messages from device 125 upstream, i.e., a hierarchy upwards towards a root complex to the controller hub 115, and downstream, i.e., a hierarchy downwards away from a root controller of processor 105 or system memory 110 to device 125. In one embodiment, switch 120 is referred to as a logical assembly of multiple virtual PCI-to-PCI bridge devices. Device 125 includes any internal or external device or component that is coupled to an electronic system, such as an I / O device, a network interface controller (NIC), an add-in card, an audio processor, a network processor, a hard disk drive, a storage device, a CD / DVD-ROM drive, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a FireWire device, a Universal Serial Bus (USB) device, a scanner, and other input / output devices.In PCIe terminology, such a device is often referred to as an endpoint. Although not specifically shown, Device 125 may include a bridge (e.g., a PCIe-to-PCI / PCI-X bridge) to support legacy or other versions of devices, or coupling structure assemblies supported by such devices.

[0019] A graphics accelerator 130 can also be coupled to the controller hub 115 via a serial link 132. In one embodiment, the graphics accelerator 130 is coupled to an MCH, which is coupled to an ICH. The switch 120, and consequently the I / O device 125, is then coupled to the ICH. The I / O modules 131 and 118 also implement a layered protocol stack for communication between the graphics accelerator 130 and the controller hub 115. Similar to the MCH discussion above, a graphics controller or the graphics accelerator 130 itself can be integrated into the processor 105.

[0020] With current reference to Fig. Figure 2 illustrates an embodiment of a layered protocol stack. The layered protocol stack 200 can include any form of layered communication stack, such as a QPI stack, a PCle stack, a next-generation HPI stack, or any other layered stack. In one embodiment, the protocol stack 200 can include the transaction layer 205, the link layer 210, and the physical layer 220. An interface such as interfaces 117, 118, 121, 122, 126, and 131 in Fig. 1 can be represented as a communication protocol stack 200. The representation as a communication protocol stack can also be described as a module or interface that implements / includes a protocol stack.

[0021] Packets can be used to communicate information between components. Packets can be formed in the transaction layer (205) and the data link layer (210) to transport information from the sending component to the receiving component. As the transmitted packets flow through the other layers, they are augmented with additional information necessary for processing packets at those layers. On the receiving side, the reverse process occurs, and the packets are transformed from their physical layer (220) representation to the data link layer (210) representation and finally (for transaction layer packets) into a form that can be processed by the receiving device's transaction layer (205).

[0022] In one embodiment, the transaction layer 205 can provide an interface between a device's processor core and the coupling structure architecture, such as the data link layer 210 and the physical layer 220. In this respect, a primary responsibility of the transaction layer 205 can include merging and splitting packets (i.e., transaction layer packets, or TLPs). The translation layer 205 can also manage credit-based flow control for TLPs. Some implementations can use split transactions, i.e., transactions where the request and response are separated by time, allowing one link to carry other traffic while the destination device, among other things, gathers data for the response.

[0023] Credit-based flow control can be used to implement virtual channels and networks that utilize the coupling structure. For example, a device might offer an initial amount of credit for each of the receive buffers in transaction layer 205. An external device at the opposite end of the link, such as controller hub 115, would then receive credit. Fig. 1. The system can count the number of credits consumed by each TLP. A transaction can be sent as long as it does not exceed the credit limit. Upon receiving a response, a credit amount is restored. One example of an advantage, among other potential benefits of such a credit scheme, is that the latency of credit return does not impact performance, provided the credit limit is not reached.

[0024] In one embodiment, four transaction address spaces can include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more read and write requests to transfer data to / from a memory-allocated location. In one embodiment, memory space transactions are capable of using two different instruction types, such as a short address format like a 32-bit address or a long address format like a 64-bit address. Configuration space transactions can be used to access the configuration space of various devices connected to the coupling structure. Configuration space transactions can include read and write requests.Message space transactions (or simply messages) can also be defined to support in-band communication between coupling structure agents. Therefore, in an exemplary embodiment, the transaction layer 205 can concatenate packet headers / payloads 206.

[0025] With brief reference to Fig.Figure 3 illustrates an exemplary embodiment of a transaction layer packet describer. In one embodiment, the transaction describer 300 can be a mechanism for transporting transaction information. In this respect, the transaction describer 300 supports the identification of transactions in a system. Other possible uses include tracking modifications of standard transaction order and associating transactions with channels. For example, the transaction describer 300 can include the global identifier field 302, attribute field 304, and channel identifier field 306. In the illustrated example, the global identifier field 302 is comprehensively represented as the local transaction identifier field 308 and the source identifier field 310. In one embodiment, the global transaction identifier 302 is unique for all pending requests.

[0026] According to one implementation, the local transaction identifier field 308 is generated by a requesting agent and can be unique for all pending requests requiring completion for that requesting agent. Furthermore, in this example, the source identifier 310 uniquely identifies the requesting agent within a coupling structure hierarchy. Therefore, the local transaction identifier field 308, together with the source ID 310, provides the global identification of a transaction within a hierarchy domain.

[0027] Attribute field 304 specifies properties and relationships of the transaction. In this respect, attribute field 304 is potentially used to provide additional information that allows modification of the default transaction handling. In one implementation, attribute field 304 includes the priority field 312, reserved field 314, order field 316, and the no-snoop field 318. Here, the priority subfield 312 can be modified by an initiator to assign a priority to the transaction. The reserved attribute field 314 remains reserved for future or vendor-defined use. Possible usage models that utilize priority or security attributes can be implemented using the reserved attribute field.

[0028] In this example, the order attribute field 316 is used to provide optional information that conveys the type of order that can modify the default ordering rule. According to one example implementation, an order attribute of "0" indicates that default ordering rules should be applied, while an order attribute of "1" indicates relaxed ordering, where writes can pass through other writes in the same direction and read operations can pass through other writes in the same direction. The snoop attribute field 318 is used to determine whether transactions are queried via snooping. As shown, the channel identifier field 306 determines a channel to which a transaction is associated.

[0029] To discuss Fig.Returning to section 2, a link layer 210, also referred to as data link layer 210, can act as an intermediate stage between the transaction layer 205 and the physical layer 220. In one embodiment, the responsibility of the data link layer 210 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components at a link. One side of the data link layer 210 accepts TLPs concatenated by the transaction layer 205, applies the packet sequence identifier 211 (i.e., an identification number or packet number), calculates and applies an error detection code (CRC 212), and submits the modified TLPs to the physical layer 220 for transmission over a physical layer to an external device.

[0030] In one example, the physical layer 220 includes the logical sub-block 221 and the electrical sub-block 222 to physically send a packet to an external device. Here, the logical sub-block 221 is responsible for the "digital" functions of the physical layer 221. In this respect, the logical sub-block can include a transmit section to prepare outgoing information for transmission by the physical sub-block 222, and a receive section to determine and prepare received information before passing it on to the link layer 210.

[0031] The physical block 222 includes a transmitter and a receiver. The transmitter is supplied with symbols by the logical subblock 221, which the transmitter serializes and sends to a peripheral device. The receiver is supplied with serialized symbols from a peripheral device and transforms the received signals into a bitstream. The bitstream is deserialized and provided to the logical subblock 221. In one exemplary embodiment, an 8b / 10b transmission code is used, with ten-bit symbols being sent / received. Here, special symbols are used to form a packet containing the frames 223. In another example, the receiver also provides a symbol clock recovered from the incoming serial stream.

[0032] Although the transaction layer 205, link layer 210, and physical layer 220 are described above with reference to a specific embodiment of a protocol stack (such as a PCle protocol stack), a layered protocol stack is not restricted in this respect. In fact, any layered protocol may be included / implemented and adopt features described herein. As an example, a port / interface represented as a layered protocol may include: (1) a first layer to concatenate packets, i.e., a transaction layer; (2) a second layer to sequentialize packets, i.e., a link layer; and (3) a third layer to send the packets, i.e., a physical layer. As a specific example, an HPI layered protocol as described below is used.

[0033] With current reference to Fig.Figure 4 illustrates an exemplary embodiment of a serial point-to-point structure. A serial point-to-point link can include any transmission path for sending serial data. In the embodiment shown, a link can include two differentially driven low-voltage signal pairs: a transmit pair 406 / 411 and a receive pair 412 / 407. Accordingly, device 405 includes the transmit logic 406 to send data to device 410 and the receive logic 407 to receive data from device 410. In other words, some implementations of a link include two transmit paths, i.e., paths 416 and 417, and two receive paths, i.e., paths 418 and 419.

[0034] A transmission path refers to any route for sending data, such as a transmit line, a copper wire, an optical line, a wireless communication channel, an infrared communication link, or any other communication path. A connection between two devices, such as Device 405 and Device 410, is called a link, for example, Link 415. A link can support one lane—each lane represents a set of differential signal pairs (one pair for transmitting, one pair for receiving). To scale bandwidth, a link can aggregate multiple lanes, denoted by xN, where N is each supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.

[0035] A difference pair can refer to two transmission paths, such as lines 416 and 417, to transmit differential signals. For example, when line 416 switches from a low voltage level to a high voltage level (i.e., a rising edge), line 417 transitions from a high logic level to a low logic level (i.e., a falling edge). Differential signals, among other advantages, potentially exhibit better electrical characteristics, such as improved signal integrity (i.e., reduced cross-coupling, over- / under-voltage, and ringing). This allows for a narrower timing window, enabling faster transmission frequencies.

[0036] In one embodiment, a new HPI is provided. The HPI can include a next-generation, cache-coherent, link-based coupling structure. As an example, the HPI can be used in high-performance computing platforms such as workstations or servers, including systems where PCIe or another coupling structure protocol is typically used to connect processors, accelerators, I / O devices, and the like. However, the HPI is not limited to this. Instead, the HPI can be used in any of the systems or platforms described here. Furthermore, the individually developed ideas can be applied to other coupling structures and platforms such as PCIe, MIPI, QPI, etc.

[0037] To support multiple devices in an exemplary implementation, the HPI can incorporate instruction set architecture (ISA) agnosticism (i.e., the HPI can be implemented on several different devices). In another scenario, the HPI can also be used to connect high-performance I / O devices, not just processors or accelerators. For example, a high-performance PCle device can be coupled to the HPI via a suitable translation bridge (i.e., HPI to PCle). Furthermore, the HPI links of many HPI-based devices, such as processors, can be used in various configurations (e.g., star topologies, rings, meshes, etc.). Fig.Figure 5 illustrates exemplary implementations of several potential multi-socket configurations. A two-socket configuration, 505, can include two HPI links as shown; however, other implementations may use only one HPI link. For larger topologies, any configuration can be used as long as an identifier (ID) can be assigned and there is some form of virtual path, along with other additional or substitute features. As shown in one example, a four-socket configuration, 510, has an HPI link from each processor to every other processor. But in the eight-socket implementation shown in configuration 515, not every socket is directly connected to each other by an HPI link. However, if a virtual path or channel exists between the processors, the configuration is supported. A range of supported processors includes 2-32 in a native domain.Higher numbers of processors can be achieved, among other examples, by using multiple domains or other coupling structures between node controllers.

[0038] The HPI architecture includes a definition of a layered protocol architecture, which in some examples includes protocol layers (coherent, incoherent, and optionally other memory-based protocols), a routing layer, a link layer, and a physical layer. Furthermore, the HPI can include additional extensions related to power managers (such as power control units (PCUs)), design for testing and debugging (DFT), error handling, registers, and security, among other examples. Fig. Figure 5 illustrates an embodiment of an exemplary HPI layer protocol stack. In some implementations, at least some of the features shown in Figure 5 can be used. Fig.The five illustrated layers are optional. Each layer deals with its own level of granularity or amount of information (the protocol layer 605a,b with the packets 630, the link layer 610a,b with the flits 635, and the physical layer 605a,b with the phits 640). Note that in some embodiments, a packet may include partial flits, a single flit, or multiple flits, depending on the implementation.

[0039] As a first example, the width of a Phit 640 implies a one-to-one mapping of the link width to bits (e.g., a 20-bit link width implies a Phit of 20 bits, and so on). Flits can have larger sizes, such as 184, 192, or 200 bits. It's important to note that if the Phit 640 is 20 bits wide and the size of the Flit 635 is 184 bits, then a fraction of the Phits 640 is required to send a Flit 635 (e.g., 9.2 Phits at 20 bits to send a 184-bit Flit 635, or 9.6 at 20 bits to send a 192-bit Flit, among other examples). It should be noted that the widths of the elementary link at the physical layer can vary. For example, the number of lanes per instruction can include 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, etc. In one embodiment, the link layer 610a,b is capable of embedding multiple parts of different transactions in a single flit and one or more headers (e.g.,1, 2, 3, 4) can be embedded within the flit. In one example, HPI divides the headers into corresponding slots to allow multiple messages in the flit intended for different nodes.

[0040] In one embodiment, physical layer 605a,b can be responsible for the rapid transmission of information over the physical medium (electrical or optical, etc.). The physical link can be point-to-point between two link layer entities, such as layers 605a and 605b. Link layer 610a,b can abstract physical layer 605a,b from the upper layers and provides the capability to reliably transmit data (as well as requests) and manage flow control between two directly connected entities. The link layer can also be responsible for virtualizing the physical channel into multiple virtual channels and message classes. Protocol layer 620a,b relies on link layer 610a,b to assign protocol messages to the appropriate message classes and virtual channels before passing them to physical layer 605a,b for transmission over the physical links.The link layer 610a,b can support multiple messages such as a request, snoop response, write-back, and incoherent data, among other examples.

[0041] The physical layer 605a,b (or PHY) of the HPI can be implemented above the electrical layer (i.e., electrical conductors connecting two components) and below the link layer 610a,b, as illustrated in Fig.6. The physical layer and corresponding logic can reside at each agent and connect the link layers separately at two agents (A and B) (e.g., devices on either side of a link). The local and remote electrical layers are connected by physical media (e.g., wires, conductors, optical, etc.). In one embodiment, the physical layer 605a,b has two essential phases: initialization and operation. During initialization, the connection to the link layer is opaque, and signaling can involve a combination of timed states and handshake events. During operation, the connection to the link layer is transparent, and signaling occurs at a rate with all pathways operating together as a single link. During the operation phase, the physical layer transports flits from agent A to agent B and from agent B to agent A.The connection, also referred to as a link, abstracts some physical aspects, including media, width, and speed, from the link layers, while flits and control / status of the current configuration (e.g., width) are exchanged with the link layer. The initialization phase includes subordinate phases, such as querying and configuration. The operational phase includes subordinate phases (e.g., link power management states).

[0042] In one embodiment, the link layer 610a,b can be implemented to provide reliable data transmission between two protocol or routing entities. The link layer can abstract the physical layer 605a,b from the protocol layer 620a,b and can be responsible for flow control between two protocol agents (A, B) and provide virtual channel services to the protocol layer (message classes) and routing layer (virtual networks). The interface between the protocol layer 620a,b and the link layer 610a,b can typically be at the packet layer. In one embodiment, the smallest transfer unit at the link layer is called a flit with a specific number of bits, such as 192 bits, or by another name.Link layer 610a,b relies on physical layer 605a,b to transform the physical layer 605a,b transmission unit (phit) into the link layer 610a,b transmission unit (flit). Furthermore, link layer 610a,b can be logically divided into two parts: a sender and a receiver. A sender / receiver pair at one entity can be connected to a receiver / sender pair at another entity. Flow control is often performed on both a flit and a packet basis. Error detection and correction are also potentially performed on a flit-layer basis.

[0043] In one embodiment, routing layer 615a,b can provide a flexible and distributed method for routing HPI transactions from a source to a destination. The scheme is flexible because routing algorithms for multiple topologies can be specified by programmable routing tables at each router (the programming is performed by firmware, software, or a combination thereof in one embodiment). The routing functionality can be distributed; routing can be performed through a series of routing steps, each defined by a lookup of a table at either the source, intermediate, or destination routers. The lookup at a source can be used to introduce an HPI packet into the HPI fabric. The lookup at an intermediate router can be used to route an HPI packet from an input port to an output port.Looking up a destination port can be used to address the destination HPI protocol agent. It's important to note that the routing layer can be thin in some implementations because the routing tables, and therefore the routing algorithms, are not specifically defined by specification. This allows for flexibility and a variety of usage models, including flexible architectural platform topologies, which are defined by the system implementation. Routing layer 615a,b relies on link layer 610a,b to provide the use of up to three (or more) virtual networks (VNs)—in one example, two VNs without deadlocks, VN0 and VN1, with multiple message classes defined in each virtual network.A shared adaptive virtual network (VNA) may be defined in the link layer, but this adaptive network may not be directly exposed in routing concepts, as each message class and virtual network may have fixed resources allocated and guaranteed progress, among other features and examples.

[0044] In some implementations, HPI can use an embedded clock. A clock signal can be embedded in data transmitted using the coupling structure. With the clock signal embedded in the data, distinct and associated clock paths can be omitted. This can be useful, for example, because it allows more pins of a device to be allocated for data transmission, especially in systems where pin space is at a premium.

[0045] A link can be established between two agents on either side of a coupling structure. An agent sending data can be a local agent, and the agent receiving the data can be a remote agent. State machines can be used by both agents to manage various aspects of the link. In one embodiment, the physical layer data path can send FLITS from the link layer to the electrical front end. The control path in one implementation includes a state machine (also called a link-training state machine or similar). The actions and state exits of the state machine can depend on internal signals, timers, external signals, or other information. In fact, some states, such as certain initialization states, can have timers to provide a timeout value for exiting a state.It should be noted that in some embodiments, "detect" refers to the detection of an event on both segments of a path, but not necessarily simultaneously. In other embodiments, "detect" refers to the detection of an event by a reference agent. "Debouncing," for example, refers to a sustained assertion of a signal. In one embodiment, the HPI supports operation in the case of malfunctioning paths. Here, paths can be dropped under specific conditions.

[0046] States defined in the state machine can include reset states, initialization states, and operating states, among other categories and subcategories. For example, some initialization states may have a secondary timer used to exit the state on a timeout (essentially an abort due to failure to progress within the state). An abort may involve updating registers such as status registers. Some states may also have primary timers used to time the primary functions within the state. Other states may be defined such that internal or external signals (such as handshake protocols), among other examples, cause the transition from one state to another.

[0047] A state machine can also support debugging through single-step debugging, freezing on initialization failure, and the use of checkers. State exits can be deferred / held until the debugging software is ready. In one instance, the exit can be deferred / held until the secondary timeout. Actions and exits can, in one embodiment, be based on the exchange of training sequences. In another embodiment, the link state machine runs in the clock domain of the local agent, and the transition from one state to the next coincides with a training sequence limit of the sender. State registers can be used to reflect the current state.

[0048] Fig.Figure 7 illustrates a representation of at least part of a state machine used by agents in an exemplary HPI implementation. It is important to understand that the states represented in the state table of Fig. The seven included states represent a non-exhaustive list of possible states. For example, some transitions are omitted to simplify the diagram. Furthermore, some states may be combined, split, or omitted, while others may be added. Such states might include: Event reset state: entered during a warm or cold start event. Restores default values. Initializes counters (e.g., synchronization counters). Can exit to a different state, such as another reset state. Timed Reset State: Timed state for in-band reset. Can trigger a predefined electrical ordered set (EOS) so that remote receivers can detect the EOS and also enter the timed reset. The receiver has traces that hold electrical settings. Can exit to an agent to calibrate the reset state. Calibration Reset State: Calibrate without signaling on the track (e.g., receiver calibration state) or turning off drivers. May remain in the state for a predetermined duration based on a timer. May set an operating speed. May act as a wait state when a port is not enabled. May include a minimum residence time. Receiver conditioning or staggered off may occur based on the design. May exit to a receiver detection state after a timeout and / or completion of calibration. Receiver detection state: detects the presence of a receiver on a track or tracks. Can see after a receiver completion (e.g., receiver pulldown insertion). Can exit to the calibration reset state after a specified value is set, or if another specified value is not set. Can exit to the transmitter calibration state when a receiver is detected or a timeout is reached. Transmitter calibration state: for transmitter calibrations. Can be a timed state assigned for transmitter calibrations. Can include signaling on a path. Can continuously drive an EOS such as an electrically inactive Exit-Ordered-Set (EIEOS). Can exit to the compliance state upon completion of calibration or after a timer expires. Can exit to the transmitter detection state when a counter expires or a secondary timeout occurs. Sender detection state: qualifies valid signaling. Can be a handshake state where an agent completes actions and exits to the next state based on remote agent signaling. The receiver can qualify valid signaling from the sender. In one embodiment, the receiver looks for a wake-up detection, and if debouncing occurs on one or more lanes, it looks for it on the other lanes. The sender drives a detection signal. Can exit to a query state in response to debouncing that is complete for all lanes and / or a timeout, or if debouncing is not complete on all lanes and there is a timeout. Here, one or more monitoring lanes can be kept awake to debounce a wake-up signal. And if debouncing occurs, then the other lanes are potentially debounced. This can enable power savings in low-power states. Query state: The receiver adapts, initializes the drift buffer, and locks at bits / bytes (e.g., determining symbol boundaries). Tranches can be deskewed. A remote agent can initiate an exit to the next state (e.g., a link width state) in response to an acknowledgment message. Queries can additionally include a training sequence lock by locking to an EOS and a training sequence header. Tranche-to-tranche bit offset at the remote sender can be capped at a first length for high speed and a second length for low speed. Deskew can be performed in a slow mode as well as an operating mode. The receiver can have a special maximum for deskewing tranche-to-tranche bit offset, such as 8, 16, or 32 intervals of bit offset. Receiver actions can include latency fixing.In one embodiment, receiver actions can be completed upon successful deskew of a valid path assignment. A successful handshake can be achieved, for example, by receiving a number of consecutive training sequence headers containing acknowledgments and sending a number of training sequences with an acknowledgment after the receiver has completed its actions. Link Width State: The agent communicates with the final path mapping to the remote sender. The receiver receives and decodes the information. The receiver can record a configured path mapping in one structure after the checkpoint of a previous path mapping value in a second structure. The receiver can also respond with an acknowledgment ("ACK"). It can initiate an in-band reset. As an example, the first state is used to initiate an in-band reset. In one embodiment, exiting to a subsequent state, such as a flit configuration state, is performed in response to the ACK. A reset signal can also be generated before entering the low-energy state if the frequency of an alarm detection signal falls below a setpoint (e.g., 1 every number of unit intervals (Uls), such as 4K Ul). The receiver can retain current and previous path mappings.The transmitter can use different groups of paths based on training sequences that have different values. Path assignment may not modify some status registers in some implementations. Flitlock Configuration State: Entry by a sender, but the state is considered exited (i.e., secondary timeout disputed) when both sender and receiver have exited to a Blocking Link state or another link state. Sender exit to a link state, in one embodiment, involves starting from a Data Sequence (SDS) and Training Sequence (TS) boundary after receiving a planetary synchronization signal. Here, receiver exit can be based on receiving an SDS from a remote sender. This state can be a bridge from the agent to the link state. Receiver determines the SDS. Receiver can exit to the Blocking Link (BLS) state (or a control window) when SDS received after a descrambler is initialized. If a timeout occurs, exit can be to the reset state. Sender controls orbits with a configuration signal.Transmitter exit can occur due to reset, BLS, or other states based on conditions or timeouts. Send Link State: A link state. Flits are sent to a remote agent. Entry can occur from a Blocking Link State, and return to a Blocking Link State can occur upon an event, such as a timeout. Sender sends Flits. Receiver receives Flits. Can also exit to a Low Energy Link State. In some implementations, the Send Link State (TLS) may be referred to as the L0 state. Block Link State: A link state in which the sender and receiver operate in a unified manner. It can be a timed state during which link layer flits are paused while physical layer information is communicated to the remote agent. It can exit to a low-energy link state (or another link state based on the design). In one embodiment, a Block Link State (BLS) occurs periodically. The period is called a BLS interval and can be timed and may vary between slow speed and operating speed. Note that the link layer can be periodically blocked with respect to sending flits, allowing a physical layer control sequence of a length similar to that transmitted during a Send Link State or a Part-Width Send Link State.In some implementations, the Block Link State (BLS) may be referred to as an L0 control or an L0c state. Partial-width transmit link state: Link state. Can save energy by entering a partial-width state. In one embodiment, asymmetric partial-width refers to each direction of a two-way link, which has different widths that can be supported in some designs. An example of an initiator, such as a transmitter, that sends a partial-width cue to enter the partial-width transmit link state is shown in the example of Fig.Figure 14 shows that a partial-width cue is sent while transmitting on a link with a first width to cause the link to transmit at a second, new width. A mismatch can result in a reset. Note that speeds cannot be changed, but widths can. Therefore, flits are potentially transmitted at different widths. This can be logically similar to a transmit link state; however, because a smaller width is present, it can take longer to transmit flits. It can exit to other link states, such as a low-energy link state based on certain received and transmitted messages, or exit the partial-width transmit link state, or a link-blocking state based on other events. In one embodiment, a transmit port can shut down idle pathways in a staggered manner to improve signal integrity (i.e.,(noise reduction). Here, non-retrying flits, such as null flits, can be used during periods when the link width is changing. A suitable receiver can drop these null flits and shut down inactive lanes in a staggered manner, as well as record the current and previous lane assignments in one or more structures. Note that status and associated status registers can remain unchanged. In some implementations, the partial-width transmit link state may be referred to as a partial L0 or LOp state. Partial-width transmit link state exit: Exit the partial-width state. May or may not use a blocking link state in some implementations. In one embodiment, the sender initiates the exit by sending partial-width exit patterns on inactive lanes to train and deskew them. As an example, an exit pattern starts with an EIEOS that is detected and debounced to signal that the lane is ready to enter a full transmit link state, and it may end with an SDS or a fast training sequence (FTS) on inactive lanes. Any failure during the exit sequence (receiver actions such as not completing a deskew before the timeout) stops flit transmissions to the link layer and asserts a reset, which is handled by resetting the link the next time a blocking link state occurs.The SDS can also initialize the scrambler / descrambler on the lanes to suitable values. Low-energy link state: This is a low-energy state. In one embodiment, it is lower energy than the partial-width link state because signaling is stopped on all paths and in both directions. Transmitters can use a block link state to request a low-energy link state. Here, the receiver can decode the request and respond with an ACK or NAK; otherwise, a reset can be triggered. In some implementations, the low-energy link state may be referred to as an L1 state.

[0049] In some implementations, state transitions can be simplified to allow states to be bypassed, for example, when state actions, such as certain calibrations and configurations, have already been completed. Previous state results and configurations of a link can be stored and reused in subsequent initializations and configurations of that link. Instead of repeating such configurations and state actions, the corresponding states can be bypassed. However, traditional systems that implement state bypasses often employ complex designs and costly validation escapes. Instead of using a traditional bypass, one example, HPI, can use short timers in certain states where, for instance, the state actions do not need to be repeated.This could, among other potential advantages, enable potentially more consistent and synchronized state machine transitions.

[0050] In one example, a software-based controller (e.g., via an external control point for the physical layer) can activate a short-time timer for one or more specific states. For a state where actions have already been executed and stored, the state can be timed short, for instance, to facilitate a quick exit to the next state. However, if the preceding state action fails or cannot be applied within the short-time timer's duration, a state exit can occur. Furthermore, the controller can deactivate the short-time timer, for example, if the state actions are due to be executed again. A long-time or default timer can be set for each corresponding state. If configuration actions for the state cannot be completed within the long-time timer's duration, a state exit can occur.The long-term timer can be set to an appropriate duration to allow the completion of state actions. In contrast, the short-term timer can be considerably shorter, which in some cases makes it impossible to execute state actions without referencing previously executed state actions, among other examples.

[0051] In some HPI implementations, supersequences can be defined, with each supersequence corresponding to a specific state or entry / exit into / out of that state. A supersequence can include a repetition sequence of data records and symbols. In some cases, the sequences can repeat until the completion of a state or state transition, or the communication of a corresponding event, among other examples. In some cases, the repetition sequence of a supersequence can repeat at a defined frequency, such as a defined number of unit intervals (Uls). A unit interval (UI) can correspond to the time interval for transmitting a single bit on a lane of a link or system. In some implementations, the repetition sequence can begin with an EOS (End of System).Accordingly, an instance of the EOS can be expected to repeat itself according to a predefined frequency. Such ordered sets can be implemented as defined 16-byte codes, which can be represented in hexadecimal format, among other examples. In one example, the EOS of a supersequence can be an electrically inactive ordered set (or EIEIOS). In another example, an EIEOS can resemble a low-frequency clock signal (e.g., a predefined number of repeating hexadecimal symbols FF00 or FFF000, etc.). A predefined data set can follow the EOS, such as a predefined number of training sequences or other data. Such supersequences can be used in state transitions, which include, among other examples, link state transitions and initialization.

[0052] As introduced above, initialization in one embodiment can initially occur at a slow speed, followed by a high-speed initialization. The slow-speed initialization uses the default values ​​for the registers and timers. Software then uses the slow-speed link to set up the registers, timers, and electrical parameters, and clears the calibration semaphores to prepare for a high-speed initialization. As an example, the initialization could consist of states or tasks such as reset, detect, query, and configure, among other potentially different states.

[0053] In one example, a link-layer blocking control sequence (i.e., a blocking link state (BLS) or L0c state) can include a timed state during which link-layer flits are paused while PHY information is communicated to the remote agent. Here, the sender and receiver can initiate a blocking control sequence timer. Upon expiration of the timer, the sender and receiver can exit the blocking state and perform other actions, such as exiting to reset, exiting to a different link state (or any other state), including states that allow sending flits over the link.

[0054] In one embodiment, link training can be provided and include the transmission of one or more encrypted training sequences, ordered sets, and control sequences, as well as those associated with a defined supersequence. A training sequence symbol can include one or more headers, reserved portions, a target latency, a pair number, physical path mapping code reference paths, or a group of paths and an initialization state. In one embodiment, the header can be sent with an ACK or NAK, among other examples. As an example, training sequences can be sent as part of supersequences and can be encrypted.

[0055] In one embodiment, ordered sets and control sequences are neither encrypted nor staggered and are transmitted simultaneously and completely on all lanes in an identical manner. Valid reception of an ordered set may include verification of at least a portion of the ordered set (or the entire ordered set for partial ordered sets). Ordered sets may include an EOS such as an electrically inactive ordered set (EIOS) or an EIEOS. A supersequence may include a data sequence start (SDS) or a fast training sequence (FTS). These sets and control supersequences may be predefined and may have any pattern or hexadecimal representation and any length. For example, ordered sets and supersequences may be 8 bytes, 16 bytes, or 32 bytes long, etc.For example, FTS can also be used for fast bit locking during exit from a partial-width transmit link state. It should be noted that the FTS definition can be per lane and can use a rotated version of the FTS.

[0056] In one implementation, supersequences can include the introduction of an EOS, such as an EIEOS, into a training sequence stream. When signaling starts, pathways can be switched on in a staggered manner. However, this can result in initial supersequences being truncated on some pathways at the receiver. Supersequences can, however, be repeated over short intervals (e.g., approximately one thousand unit intervals (or ~1 KUI)). The training supersequences can additionally be used for one or more functions, including deskew, configuration, and communicating an initialization target, pathway assignment, etc. The EIEOS can be used for one or more functions, including switching a pathway from the inactive to the active state, screening for good pathways, and identifying symbol and TS boundaries.

[0057] With current reference to Fig.Figure 8 shows illustrations of exemplary supersequences. For example, an exemplary recognition supersequence 805 may be defined. The recognition supersequence 805 may include a repetition sequence of a single EIEOS (or another EOS) followed by a predefined number of instances of a special training sequence (TS). In one example, the EIEOS may be sent, immediately followed by seven repeated instances of TS. When the last of the seven TS is sent, the EIEOS may be sent again, followed by seven additional instances of TS, and so on. This sequence may be repeated according to a special predefined frequency. In the example of Fig.8. The EIEOS can reappear on the lanes approximately once every thousand ULs (~1KUI), followed by the rest of the detection supersequence 805. A receiver can monitor lanes for the presence of a repeating detection supersequence 805 and, after validating supersequence 705, conclude that a remote agent is present, has been added to the lanes (e.g., connected while in operation), has woken up, or has been reinitialized, etc.

[0058] In another example, a different supersequence 810 can be defined to indicate a query, configuration, or loopback condition or state. As with the exemplary detection supersequence 805, pathways of a link through a receiver can be monitored for such a query / config / loop supersequence 810 to determine a query state, configuration state, or loopback state or condition. In one example, a query / config / loop supersequence 810 can begin with an EIEOS followed by a predefined number of repeated instances of a TS. For example, in one example, the EIEOS can be followed by thirty-one (31) instances of TS, with the EIEOS being repeated approximately every four thousand UI (e.g., ~4 KUI).

[0059] Furthermore, in another example, a Partial Width Transmit State (PWTS) exit supersequence 815 can be defined. In one example, a PWTS exit supersequence can include an initial EIEOS to repeat the lane pre-preparation before transmitting the first full sequence in the supersequence. For example, the sequence to be repeated in supersequence 815 can begin with an EIEOS (to repeat approximately once every 1 KUI). Furthermore, Fast Training Sequences (FTS) can be used instead of other Training Sequences (TS), with the FTS configured to assist with faster bit locking, byte locking, and deskew. In some implementations, an FTS can be decrypted to further assist in reactivating inactive lanes as quickly and unobtrusively as possible.As with other supersequences that precede entry into a link transmit state, supersequence 815 can be interrupted and terminated by sending an SDS. Furthermore, a partial FTS (FTSp) can be sent to assist in synchronizing the new lanes with the active lanes, for example by allowing bits to be subtracted from (or added to) the FTSp, among other examples.

[0060] Supersequences such as the detection supersequence 705 and the query / config / loop supersequence 710, etc., can potentially be sent essentially during the initialization or re-initialization of a link. In some cases, after receiving and detecting a specific supersequence, a receiver can respond by echoing the same supersequence back to the sender via the pathways. The receiving and validation of a specific supersequence by the sender and receiver can serve as a handshake to confirm a state or condition communicated by the supersequence. For example, such a handshake (e.g., using a detection supersequence 705) can be used to determine the re-initialization of a link. In another example, such a handshake can be used to indicate the end of an electrical reset or low-energy state, which is in...This results in corresponding pathways that are reactivated, among other examples. The end of the electrical reset can be determined, for example, by a handshake between the sender and receiver, in which each sends a recognition supersequence 705.

[0061] In another example, pathways can be monitored for supersequences, and these supersequences can be used, among other things, in conjunction with pathway screening for detecting, waking, state exits, and entries. The predefined and predictable nature and form of supersequences can be further used to perform initialization tasks such as bit locking, byte locking, debouncing, descrambling, deskew, matching, latency fixing, negotiated delays, and other possible uses. In essence, pathways can be continuously monitored for such events to enhance the system's ability to respond to and process these conditions.

[0062] In the case of debouncing, transients can be introduced onto tracks due to a variety of conditions. For example, adding or turning on a device can introduce transients onto the track. Additionally, voltage irregularities on a track can occur due to poor track quality or an electrical fault. In some cases, track bouncing can produce false positive results, such as a false EIEOS. However, in some implementations, defined supersequences can include further additional sequences of data as well as a defined frequency with which the EIEOS is repeated, while supersequences can begin with an EIEOS. Even where a false EIEOS appears on a track, a logic analyzer at the receiver can determine that the EIEOS is a false positive by validating data that follows the false EIEOS.For example, if an expected Transient Signal (TS) or other data does not follow the EIEOS, or if the EIEOS does not repeat within a specific predefined frequency of one of the predefined supersequences, the receiver logic analyzer may fail to validate the received EIEOS. Since bounce can occur during initiation while a device is being added to a link, false negatives can also result. For example, after being added to a set of lanes, a device may begin transmitting a detection supersequence 705 to alert the other side of the link to its presence and begin initializing the link. However, transients introduced on the lanes can corrupt the initial EIEOS, the TS instances, and the other data in the supersequence.However, a logic analyzer at the receiving device can continue to monitor the paths and determine the next EIEOS sent by the new device in the repeating recognition supersequence 705, among other examples.

[0063] In some implementations, an HPI link is capable of operating at multiple speeds, facilitated by the embedded clock. For example, a slow mode may be defined. In some cases, the slow mode can be used to assist in initializing the link. Link calibration can involve software-based controllers that provide logic to set various calibrated properties of the link. These properties include, among others, which pathways the link is to be used for, the pathway configuration, the link's operating speed, pathway and agent synchronization, deskew, and target latency. These software-based tools can utilize external control points to add data to physical layer registers to control various aspects of the physical layer facilities and logic.

[0064] The operating speed of a link can be considerably higher than the effective operating speed of software-based controllers used during link initialization. A slow mode can be used to enable the use of such software-based controllers in other situations, such as during link initialization or reinitialization. Slow mode can also be applied to pathways connecting a receiver and transmitter, for example, when a link is powered on, initialized, reset, etc., to facilitate link calibration.

[0065] In one embodiment, the clock can be embedded in the data, eliminating the need for separate clock lanes. Flits can be sent according to the embedded clock. Furthermore, the flits sent over the lanes can be encrypted to facilitate clock recovery. For example, the receiver clock recovery unit can supply sampling clocks to a receiver (i.e., the receiver extracts the clock from the data and uses it to sample the incoming data). In some implementations, receivers continuously adapt to an incoming bitstream. Embedding the clock can potentially reduce pin usage. Embedding the clock in the in-band data can also change how an in-band reset is handled. In one embodiment, a block-link state (BLS) can be used after initialization.Furthermore, electrical ordered-set supersequences can be used during initialization to facilitate reset, among other considerations. The embedded clock can be shared between devices on a link, and the common operating clock can be set during link calibration and configuration. For example, HPI links can reference a common clock using drift buffers. Such an implementation can achieve lower latency than elastic buffers used in non-shared reference clocks, among other potential advantages. Furthermore, the reference clock distribution segments can be adjusted within specified limits.

[0066] As previously mentioned, an HPI link can operate at multiple speeds, including a "slow mode" for standard power-on, initialization, and other processes. The operating (or "fast") speed or mode of each device can be statically set by the BIOS. The common clock on the link can be configured based on the respective operating speeds of each device on either side of the link. For example, the link speed can be based on the slower of the two device operating speeds, among other possibilities. Each change in operating speed can be accompanied by a warm or cold boot.

[0067] In some examples, the link initializes into slow mode at a throughput rate of, for example, 100 MT / s upon power-up. Software then configures the two sides for the link's operating speed and begins initialization. In other cases, a sideband mechanism can be used to establish a link that incorporates the common clock on the link, for example, in the absence or unavailability of a slow mode.

[0068] In one implementation, a slow-mode initialization phase may use the same encoding, encryption, training sequences (TS), states, etc., as the operating speed, but with potentially fewer features (e.g., no electrical parameter setup, no adjustments, etc.). Similarly, the slow-mode operating phase may potentially use the same encoding, encryption, etc. (although other implementations may not), but may have fewer states and features compared to the operating speed (e.g., no low-energy states).

[0069] Furthermore, the slow mode can be implemented using the device's native phase-locked loop (PLL) clock frequency. For example, the HPI can support an emulated slow mode without changing the PLL clock frequency. While some designs may use separate PLLs for slow and fast speeds, in some HPI implementations, the emulated slow mode can be achieved by allowing the PLL clock to run at the same fast operating speed during slow mode. For example, a transmitter can emulate a slower clock signal by repeating bits multiple times to simulate a slow high clock signal and then a slow low clock signal. The receiver can then oversample the received signal to locate edges emulated by the repeating bits and determine the corresponding bit.In such implementations, ports that share a PLL can coexist at slow and fast speeds.

[0070] Some HPI implementations can support path matching at a link. The physical layer can support both receiver matching and transmitter matching. With receiver matching, the transmitter on a path sends sample data to the receiver, which the receiver logic can process to determine deficiencies in the path's electrical properties and signal quality. The receiver can then adjust path calibration settings to optimize the path based on the analysis of the received sample data. In the case of transmitter matching, the receiver can again receive sample data and develop a metric that describes the path quality, but in this case, the metric is linked to the transmitter (e.g.,using a return channel such as a software, hardware, embedded, sideband or other channel) to enable the transmitter to make adjustments to the track based on the feedback.

[0071] Since both devices on a link can operate at the same reference clock (e.g., ref clk), elasticity buffers can be omitted (any elastic buffers can be bypassed or used as drift buffers with the lowest possible latency). However, phase-matching or drift buffers can be used on each lane to transfer the appropriate receiver bitstream from the remote clock domain to the local clock domain. The latency of the drift buffers can be sufficient to handle the sum of drift from all sources in the electrical specification (e.g., voltage, temperature, residual SSC introduced by reference clock routing mismatches, and so on), but it can be as small as possible to reduce transport delay. If the drift buffer is too shallow, drift errors can result and manifest as a series of CRC errors.Therefore, some implementations can provide a drift alarm that can initiate a reset of the physical layer before an actual drift error occurs, among other examples.

[0072] Some HPI implementations can support two sides running at the same nominal reference clock frequency but with a ppm difference. In this case, frequency matching (or elasticity) buffers may be required and may need to be re-adjusted during an extended BSL window or during specific sequences that would occur periodically, among other examples.

[0073] Some systems and devices that use HPI can be deterministic, so their transactions and interactions with other systems, including communications over an HPI link, are synchronized with specific events at the system or device. Such synchronization can occur according to a planetary synchronization point or signal that corresponds to the deterministic events. For example, a planetary synchronization signal can be used to synchronize state transitions, including entering a link transmit state, with other events at the device. In some cases, synchronization counters can be used to maintain alignment with a device's planetary alignment. For example, each agent can include a local synchronization counter that is initialized by a planetarily aligned signal (i.e.,, jointly and simultaneously (apart from fixed bit offset) with all agents / layers that are synchronized). This synchronization counter can correctly count alignment points even in power-down or low-power states (e.g., L1 state) and can be used to time the initialization process (after reset or L1 exit), including the boundaries (i.e., starting or ending) of an EIEOS (or any other EOS) enclosed in a supersequence used during initialization. These supersequences can be of fixed size and larger than the maximum possible latency on a link. EIEOS TS boundaries in a supersequence can therefore be used as a proxy for a remote synchronization counter value.

[0074] Furthermore, HPI can support master-slave models, where a deterministic master device or system can time the interaction with another device according to its own planetary alignment moments. Additionally, master-master determinism can be supported in some examples. Master-master or master-slave determinism can ensure that two or more link pairs can be in lock-step at the link layer and above. In master-master determinism, exit in either direction can be controlled by the initialization of the corresponding sender. In the case of master-slave determinism, a master agent can control the determinism of the link pair (i.e., in both directions) by having a slave sender initialization exit wait until, for example, its receiver exits initialization, among other potential examples and implementations.

[0075] In some implementations, a synchronization (or "sync") counter can be used in conjunction with maintaining determinism within an HPI environment. For example, a synchronization counter can be implemented to count a defined amount, such as 256 or 512 UI. This synchronization counter can be reset by an asynchronous event and can then count continuously (with rollover), potentially even during a low-energy link state. Pin-based resets (e.g., power-on reset, warm start), among other examples, can synchronize events that reset a synchronization counter. In one embodiment, these events can occur on two sides, with the bit offset being smaller (and in many cases, much smaller) than the synchronization counter value. During initialization, the start of the sent exit-ordered set (e.g.,EIEOS), which precedes a training sequence of a training supersequence, can be aligned with the reset value of the synchronization counter (e.g., synchronization counter rollover). Such synchronization counters can be maintained for each agent on a link to preserve determinism by maintaining constant latency of Flit transfers over a dedicated link.

[0076] Control sequences and codes, among other signals, can be synchronized with a planetary synchronization signal. For example, EIEOS sequences, BLS or L0c windows (and enclosed codes), SDSes, etc., can be configured to synchronize with a planetary alignment. Furthermore, synchronization counters can be reset by an external signal, such as a planetary synchronization signal from a device, so that it is itself synchronized with the planetary alignment, among other examples.

[0077] Synchronization counters of both agents on a link can be synchronized. Resetting, initializing, or reinitializing a link may include resetting the synchronization counters to realign them with each other and / or an external control signal (such as a planetary synchronization signal). In some implementations, synchronization counters can only be reset by entering a reset state. In some cases, determinism can be maintained, such as by returning to an L0 state without resetting the synchronization counter. Instead, other signals already aligned to a planetary alignment, or another deterministic event, can be used as a proxy for a reset. In some implementations, an EIEOS can be used upon entering a deterministic state.In some cases, the EIEOS boundary and an initial TS of a supersequence can be used to determine a synchronization moment and synchronize synchronization counters from one of the agents on a link. The end of an EIEOS can be used, for example, to prevent transients from corrupting the start boundary of the EIEOS, among other uses.

[0078] Latency fixing can also be provided in some HPI implementations. Latency can include not only the latency introduced by the transmit line used for communication between flits, but also the latency resulting from processing by the agent on the other side of the link. The latency of a lane can be determined during link initialization. Furthermore, changes in latency can also be determined. From the determined latency, latency fixing can be initiated to compensate for such changes and restore the expected latency for the lane to a constant, expected value. Maintaining consistent latency on a lane can be critical for maintaining determinism in some systems.

[0079] In some implementations that use a latency buffer in conjunction with determinism, the latency at a receiver link layer can be fixed to a programmed value and activated by initiating a detection (e.g., by sending a detection supersequence) on a synchronization counter rollover. Accordingly, in one example, a sent EIEOS (or another EOS) can be polled and configured on a synchronization counter rollover. In other words, the EIEOS can be precisely aligned with the synchronization counter, so that a synchronized EIEOS (or another EOS) can, in some cases, act as a proxy for the synchronization counter value itself, at least in conjunction with certain latency-fixing activities.For example, a receiver can add enough latency to a received EIEOS to meet the dictated target latency at the physical layer-link layer interface. As an example, if the target latency is 96 UI and the receiver's EIEOS is deskew at a synchronization count of 80 UI, 16 UI of latency can be added. Essentially, given the synchronization of an EIEOS, the latency of a lane can be determined based on the delay between when it is known that the EIEOS was sent (e.g., at a specific synchronization count) and when the EIEOS was received. Furthermore, the latency can be fixed using the EIEOS (e.g., by adding latency to the transmission of an EIEOS to maintain a target latency, etc.).

[0080] Latency fixing can be used within the context of determinism to allow an external entity (such as an entity providing a planetary synchronization signal) to synchronize the physical state of two agents bidirectionally across the link. Such a feature can be used, for example, for debugging problems in the field and for supporting lock-step behavior. Accordingly, such implementations can include external control of one or more signals that can cause the physical layer of two agents to transition to a send-link (TLS) state. Agents possessing deterministic capabilities can exit initialization at a TS bound, which is potentially also the clean flit bound, if, or after, the signal is asserted.Master-slave determinism allows a master to synchronize the physical layer state of master and slave agents bidirectionally across the link. When enabled, the slave sender's exit from initialization can depend on its receiver's exit from initialization (e.g., follow or be coordinated with it) (in addition to other considerations based on determinism). Agents exhibiting deterministic capability can possess additional functionality to enter a BLS or L0c window on a clean flit, among other examples.

[0081] Determinism can also be referred to as automatic test setup (ATE) when it is used to synchronize test patterns in ATE with a device under test (DUT), which controls the physical and link layer state by fixing the latency at the receiver link layer to a programmed value using a latency buffer.

[0082] In some implementations, determinism in HPI can include facilitating an agent's ability to determine and apply a delay based on a deterministic signal. A master can send a hint of a target latency to a remote agent. The remote agent can determine the actual latency along an orbit and apply a delay to adjust the latency to achieve the target latency (e.g., determined in a TS). Adjusting the delay or latency can aid in facilitating the eventual synchronized entry into a link transmit state at a planetary alignment point. A delay value can be communicated from a master to a slave, for example, in TS payload data of a supersequence. The delay can specify a particular number of ULs (ultra-increments) dedicated to the delay.The slave can delay entry into a state based on a specified delay. Such delays can be used, for example, to facilitate checking, to stagger L0c intervals on the paths of a link, among other examples.

[0083] As previously mentioned, a state exit can occur according to a planetary alignment point. For example, a state exit signal (SDS) can be sent to interrupt a state supersequence to effect the transition from one state to another. Sending the SDS can be timed to coincide with a planetary alignment point, and in some cases, in response to a planetary synchronization signal. In other cases, sending an SDS can be synchronized with a planetary alignment point based on a synchronization counter value or some other signal synchronized to planetary alignment. An SDS can be sent at any point in a supersequence, in some cases interrupting a specific state transition (TS), end-of-state exit (EIEOS), etc., within the supersequence.This can ensure that the state changes with little delay while maintaining alignment with a planetary alignment point, among other examples.

[0084] In some implementations, the HPI can support flits with a width that, in some cases, is not a multiple of the nominal lane width (e.g., using a flit width of 192 bits and 20 lanes as a purely illustrative example). Indeed, in implementations that allow partial-width transmit states, the number of lanes over which flits are transmitted can fluctuate, even during the link's lifetime. In some cases, for example, the flit width might be a multiple of the number of active lanes at one instant but not a multiple of the number of active lanes at another instant (e.g., while the link is changing state and lane width). In cases where the number of lanes is not a multiple of a current lane width (e.g.,For example, with a flit width of 192 bits across 20 lanes, in some embodiments successive flits can be configured to be sent overlapping across lanes in order to conserve bandwidth (e.g., sending five consecutive 192-bit flits overlapping across the 20 lanes).

[0085] Fig. Figure 10 illustrates a representation of the transmission of successive flits that overlap on a number of paths. For example, it shows Fig. 10 is a representation of five overlapping 192-bit flits sent over a 20-lane link (the lanes represented by columns 0-19). Each cell of Fig.10 represents a corresponding "4-bit unit" or grouping of four bits (e.g., the bits 4n+3:4n) enclosed in a flit that is sent over a 4-UI span. For example, a 192-bit flit can be divided into 48 four-bit units. In one example, 4-bit unit 0 includes bits 0-3, 4-bit unit 1 includes bits 4-7, and so on. The bits in the 4-bit units can be sent so that they overlap or are interleaved (e.g., "swizzled"), so that higher-priority fields of the flit are presented earlier and error detection properties (e.g., CRC) are preserved, among other considerations. Indeed, a swizzling scheme can also provide that some 4-bit units (and their corresponding bits) are sent out of order (as in the examples of the Fig. 10 and Fig.11) In some implementations, a swizzling scheme may depend on the architecture of the link layer and the format of the flits used in the link layer.

[0086] The bits (or 4-bit units) of a flit with a length that is not a multiple of the active paths can be swizzled, as in the example of Fig.10. For example, during the first 4 UI, the 4-bit units 1, 3, 5, 7, 9, 12, 14, 17, 19, 22, 24, 27, 29, 32, 34, 37, 39, 42, 44, and 47 can be sent. The 4-bit units 0, 2, 4, 6, 8, 11, 13, 16, 18, 21, 23, 26, 28, 31, 33, 36, 38, 41, 43, and 46 can be sent during the next 4 UI. In UIs 8-11, only eight 4-bit units from the first flit remain. These final 4-bit units (i.e., 10, 15, 20, 25, 30, 40, 45) of the first flit can be sent simultaneously with the first 4-bit units (i.e., the 4-bit units 2, 4, 7, 9, 12, 16, 20, 25, 30, 35, 40, 45) of the second flit, so that the first and second flits overlap or are swizzled. Using such a technique, in the present example, five complete flits can be sent in 48 UI, with each flit being sent over a fractional 9.6 UI period.

[0087] In some cases, swizzling can result in periodic "clean" flit boundaries. For example, in the example of Fig.10. The initiating 5-flit boundary (the header of the first flit) can also be referred to as a clean flit boundary, since all lanes starting with the 4-bit unit transmit from the same flit. Agent link-layer logic can be configured to determine lane swizzling and can reconstruct the flit from the swizzled bits. Additionally, physical layer logic can include functionality to determine when and how to swizzle a stream of flit data based on the number of lanes currently in use. In fact, upon transitioning from one link-width state to another, agents can configure themselves to determine how to apply swizzling to the data stream. Indeed, both sides of the link can determine the scheme to be used for swizzling a data stream, thus determining how a link-width state transition affects the stream.In some implementations, among other examples, the length of a partial FTS (FTSp) can be tailored so that the signal exit is synchronized to facilitate a link-width state transition on a jagged edge of a flit. Furthermore, the physical layer logic can be configured to maintain determinism despite jagged flit boundaries resulting from swizzling, among other features.

[0088] As mentioned previously, links can switch between web widths, in some cases operating at an initial or full width and later switching to (and from) a partial width that uses smaller webs. In some cases, the defined width of a flit can be divisible by the number of webs. For example, the example of Fig. 11. Such an example, where the 192-bit flit from the previous examples is sent over an 8-lane link. As shown in Fig.11. Four-bit units of a 192-bit flit can be evenly distributed and transmitted over 8 lanes (i.e., because 192 is a multiple of 8). In fact, a single flit can be transmitted over 24 UI when working with an 8-lane partial width. Furthermore, any flit boundary in the example of Fig. 11. Be clean. While clean flit boundaries can simplify state transitions, determinism, and other features, allowing swizzling and occasional jagged flit boundaries can minimize wasted bandwidth on a link.

[0089] While the example of Fig.Figure 11 shows that lanes 0-7 are the lanes that have remained actively in a partial-width state. Furthermore, any set of 8 lanes can potentially be used. It should be noted that the examples above are for illustrative purposes only. Flits can potentially be defined to have any width. Links can also potentially have any link width. Additionally, the swizzling scheme of a system can be flexibly designed according to the formats and fields of the flits and the preferred lane widths in a system, among other considerations and examples.

[0090] The operation of the logical HPI-PHY layer can be independent of the underlying transmission media, provided that the latency does not result in latency fixing errors or timeouts at the link layer, among other considerations.

[0091] External interfaces can be provided in the HPI to assist in managing the physical layer. For example, external signals (from pins, fuses, other layers), timers, control registers, and status registers can be provided. The input signals can change relative to the PHY state at any given time, but the physical layer must consider them at specific points in time when in a corresponding state. For example, a changing synchronization signal (as introduced below) can be received but have no effect after the link enters a transmit link state, among other examples. Similarly, instruction register values ​​can be considered by physical layer entities only at specific times. For example, the physical layer logic can take a snapshot of the value and use it in subsequent operations.Therefore, in some implementations, updates to instruction registers may be associated with a limited subset of special time periods (e.g., in a send-link state or when held in reset calibration, in send-link slow mode) to avoid abnormal behavior.

[0092] Because status values ​​track hardware changes, the values ​​read can depend on when they are read. However, some status values, such as link allocation, latency, speed, etc., cannot change after initialization. For example, a re-initialization (or a low-energy link state (LPLS) or L1 state exit) is the only thing that can cause these to change (e.g., a critical path fault in a TLS cannot result in link reconfiguration until re-initialization is triggered, among other examples).

[0093] Interface signals can include signals that are external to the physical layer but affect its behavior. Examples of such interface signals include encoding and clock signals. Interface signals can be design-specific. These signals can represent an input or an output. Some interface signals, such as semaphores and those prefixed with EO, can be active once per assertion edge, meaning they can be deasserted and then reasserted to be effective again. For example, Table 1 includes an exemplary list of such functions. TABLE 1 function Input PIN reset (also known as warm start) Input pin reset (also known as cold start) Input in-band reset pulse; causes semaphore to be set; semaphore is cleared when in-band reset occurs. Input enables low-energy states Input loopback parameter; applied to loopback pattern Entrance to enter PWLTS Entrance to exit PWLTS Entrance to enter LPLS Entrance to exit LPLS Input from inactive exit detection (also known as squelch break) Input enables the use of CPylnitBegin Input of local or planetary alignment for the transmitter to exit initialization. Output when remote agent NAKs LPLS request Exit when agent enters LPLS Exit to the link layer to avoid forcing further flits attempts. Output to the link layer to force zero flits Output when transmitter is in Partial Width Link Transmitting State (PWLTS) Output when receiver is in PWLTS

[0094] CSR timer default values ​​can be provided in pairs—one for slow mode and one for operating speed. In some cases, a value of 0 disables the timer (i.e., a timeout never occurs). Timers can include those shown in Table 2 below. Primary timers can be used to time expected actions in a state. Secondary timers are used to abort initializations that are not progressing or to perform progress state transitions in an ATE mode at precise times. In some cases, secondary timers can be much larger than the primary timers in a state. Exponential timer sets can be designated with the suffix exp, and the timer value is 2 raised to the power of the field value. For linear timers, the timer value is the field value. Each timer could use different levels of precision.Additionally, some timers in the power management section may be found in a set called a timing profile. These may be linked to a timing diagram of the same name. TABLE 2 Timer Table Tpriexp Set Reset residence to navigate to EIEOS Minimum receiver calibration time; for staggered transmitter from Minimum transmitter calibration time; for staggered one Tsecexp Set Timed receiver calibration Timed transmitter calibration Detect / debounce squelch exit DetectAtRx overhang for handshake Adapt+bit lock / byte lock / deskew Configuring link widths Waiting for planetarily aligned clean Flit boundary Re-byte lock / Deskew Tdebugexp Set For hot-plugging; non-zero value regarding debug hangs. TBL Sentry set BLS boarding delay - fine BLS boarding delay - rough TBLS set BLS duration for the transmitter BLS duration for the recipient BLS clean Flit interval for the transmitter TBLS clean flit interval for the receiver

[0095] Command and control registers can be provided. Control registers can be late-action and, in some cases, can be read or written by software. Values ​​of a late action can take effect continuously during reset (e.g., passing through from the software-facing to the hardware-facing section). Control semaphores (prefixed CP) are RW1S and can be cleared by hardware. Control registers can be used to perform any of the elements described here. They can be modifiable and accessible by hardware, software, firmware, or a combination thereof.

[0096] Status registers can be provided to track hardware changes (written and used by hardware) and can be read-only (but debugging software may also be able to write to them). Such registers do not impair interoperability and can usually be augmented with many private status registers. Status semaphores (prefixed with SP) can be mandated, as they can be cleared by software to repeat the actions that set the status. "Standard" means that initial (on reset) values ​​can be provided as a subset of these initialization-related status bits. On an aborted initialization, this register can be copied to a memory structure.

[0097] Toolbox registers can be provided. For example, testability toolbox registers in the physical layer can provide pattern generation, pattern verification, and loopback control mechanisms. Higher-level applications can use these registers, along with electrical parameters, to determine clearances. For instance, a test integrated into the coupling structure can use this toolbox to determine clearances. For transmitter tuning, these registers can be used in conjunction with the special registers described in previous sections, among other examples.

[0098] In some implementations, the HPI supports Reliability, Availability, and Maintainability (RAS) capabilities using the physical layer. In one embodiment, the HPI supports hot-plugging and hot-removal with one or more layers, which may include software. Hot-removal can include shutting down the link, and an initialization start state / signal can be cleared for the agent being removed. A remote agent (i.e., the one not being removed, such as the host agent) can be set to a slow speed, and its initialization signal can also be cleared. An in-band reset (e.g., by BLS) can cause both agents to wait in a reset state, such as a Calibration Reset State (CRS); and the agent to be removed can be removed (or kept in the addressed pin reset, shut down), among other examples and features.In fact, some of the aforementioned events can be omitted and additional events can be added.

[0099] Hot-add can include a default initialization speed of slow, and an initialization signal can be set on the agent being added. Software can set the speed to slow and clear the initialization signal on the remote agent. The link can emerge in slow mode, and software can determine an operating speed. In some cases, no PLL relock of a remote agent is performed at this point. Operating speed can be set on both agents, and an activation can be set for customization (if not previously done). The initialization start indicator can be cleared on both agents, and an in-band BSL reset can cause both agents to wait in CRS. Software can assert a warm start (e.g., an addressed or self-reset) from an agent (to be added), which can cause a PLL to relock.Software can also set the initialization start signal using any known logic and further set it at Remote (thus progressing it to the Receiver Detect State (RDS)). Software can disable a warm start of the agent being added (thus progressing it to RDS). The link can then initialize to a Send Link State (TLS) at operating speed (or to Loopback if the matching signal is set), among other examples. Indeed, some of the events mentioned above can be omitted, and additional events can be added.

[0100] Data lane recovery after a failure can be supported. In one embodiment, a link in the HPI can be resilient against a critical fault on a single lane by self-configuring to less than the full width (e.g., less than half the full width), thus eliminating the faulty lane. For example, the configuration can be performed by the link state engine, and unused lanes can be deactivated in the configuration state. As a result, the Flit can be transmitted at a narrower width, among other possibilities.

[0101] In some HPI implementations, trace reversal can be supported on certain links. For example, trace reversal can refer to traces 0 / 1 / 2... of a transmitter that are connected to traces n / n-1 / n-2... of a receiver (e.g., n can be 19 or 7, etc.). Trace reversal can be detected at the receiver as defined in a field of a TS header. The receiver can handle the trace reversal by starting in a query state using the physical trace n...0 for the logical trace 0..n. Therefore, references to a trace can point to a logical trace number. This allows board designers to more effectively design the physical or electrical layout, and the HPI can work with virtual trace assignments as described here. Furthermore, in one embodiment, the polarity can be inverted (i.e., when a differential transmitter + / - is connected to a receiver - / +).Polarity can also be detected by a receiver of one or more TS header fields and, in one embodiment, handled in the query state.

[0102] With reference to Fig.Figure 12 shows an embodiment of a block diagram for a computer system that includes a multi-core processor. The processor 1200 includes any processor or processing device such as a microprocessor, an integrated processor, a digital signal processor (DSP), a network processor, a handheld processor, an application processor, a coprocessor, a system-on-a-chip (SoC), or any other device for executing code. In one embodiment, the processor 1200 includes at least two cores—cores 1201 and 1202—which may be asymmetric cores or symmetric cores (the illustrated embodiment). However, the processor 1200 can include any number of processing elements, which may be symmetric or asymmetric.

[0103] In one embodiment, a processing element refers to hardware or logic to support a software thread. Examples of hardware processing elements include: a thread unit, a thread slot, a thread window, a process unit, a context, a context unit, a logical processor, a hardware thread, a core, and / or any other element that can contain a state for a processor, such as an execution state or architectural state. In other words, in one embodiment, a processing element refers to any hardware that can be independently associated with code, such as a software thread, operating system, application, or other code. A physical processor (or processor socket) typically refers to an integrated circuit that potentially includes any number of other processing elements, such as cores or hardware threads.

[0104] A core often refers to logic within an integrated circuit capable of maintaining independent architectural states, each with its own dedicated execution resources. In contrast to cores, a hardware thread typically refers to any logic on an integrated circuit capable of maintaining independent architectural states, where these independently maintained states share access to execution resources. It is evident that the line between the nomenclature of a hardware thread and a core overlaps when certain resources are shared and others are tightly allocated to a specific architectural state.Nevertheless, a core and a hardware thread are often regarded by an operating system as individual logical processors, whereby the operating system can schedule operations on each logical processor individually.

[0105] The physical processor 1200, as in Fig.Figure 12 illustrates this, including two cores, core 1201 and 1202. Here, cores 1201 and 1202 are considered symmetric cores, that is, cores with the same configurations, functional units, and / or logic. In another embodiment, core 1201 includes an out-of-order processor core, while core 1202 includes an in-order processor core. Cores 1201 and 1202 can be individually selected by any type of core, such as a native core, a software-managed core, a core adapted to run a native instruction set architecture (ISA), a core adapted to run a translated instruction set architecture (ISA), a co-designed core, or any other known core. In a heterogeneous core environment (i.e., asymmetric cores), a form of translation, such as binary translation, can be used to schedule or execute code on one or both cores.To advance the discussion, the functional units illustrated in core 1201 are described in further detail below, since the units in core 1202 operate in a similar manner in the embodiment shown.

[0106] As shown, core 1201 includes two hardware threads, 1201a and 1201b, which can also be referred to as hardware thread slots 1201a and 1201b. Therefore, software entities, such as an operating system, potentially view the 1200 processor in one embodiment as four separate processors—that is, four logical processors or processing elements capable of executing four software threads concurrently. As noted above, a first thread is associated with architecture state registers 1201a, a second thread is associated with architecture state registers 1201b, a third thread is associated with architecture state registers 1202a, and a fourth thread is associated with architecture state registers 1202b. Here, each of the architecture state registers (1201a, 1201b, 1202a and 1202b) can be referred to as processing elements, thread slots or thread units as described above.As illustrated, the architecture state registers 1201a are repeated in the architecture state registers 1201b, allowing individual architecture states / contexts to be stored for logical processor 1201a and logical processor 1201b. In core 1201, other smaller resources, such as instruction pointers and rename logic in the assign and rename block 1230, can also be repeated for threads 1201a and 1201b. Some resources, such as reorder buffers in reorder / reorder unit 1235, ILTB 1220, load / store buffers, and queues, can be shared through partitioning. Other resources, such as... B. Internal universal registers, page table base registers, subordinate data cache and data TLB 1215, execution unit(s) 1240 and parts of out-of-order unit 1235 are potentially fully shared.

[0107] The 1200 processor often includes additional resources that can be fully shared, shared through partitioning, or permanently allocated to / by processing elements. Fig.Figure 12 illustrates an embodiment of a purely exemplary processor with illustrative logical units / resources of a processor. It should be noted that a processor may include or omit any of these functional units, as well as any other known functional units, logic, or firmware not shown. As illustrated, core 1201 includes a simplified, representative out-of-order (OOO) processor core. However, different embodiments may use an in-order processor. The OOO core includes a branch target buffer 1220 to predict branches to be executed / taken and an instruction translation buffer (I-TLB) 1220 to store address translation entries for instructions.

[0108] The 1201 core further includes the 1225 decoding module, which is coupled to the 1220 retrieval unit to decode retrieved elements. In one embodiment, the retrieval logic includes individual sequencers connected to the thread slots 1201a and 1201b, respectively. Typically, the 1201 core is connected to a first ISA that specifies / defines instructions executable by the 1200 processor. Often, machine instructions that are part of the first ISA include a portion of the instruction (referred to as an instruction code) that references / specifies an instruction or operation to be performed. The 1225 decoding logic includes circuitry that recognizes these instructions from their instruction codes and passes the decoded instructions into the pipeline for processing, as defined by the first ISA.For example, as described in more detail below, in one embodiment, the Decoder 1225 can include logic designed or adapted to recognize special instructions, such as a transaction instruction. As a result of recognition by Decoder 1225, the Architecture or Core 1201 performs special, predefined actions to execute tasks associated with the corresponding instruction. Crucially, some of the tasks, blocks, operations, and procedures described here can be executed in response to one or more instructions, some of which may be new or existing instructions. It should be noted that in one embodiment, the Decoder 1226 recognizes the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, the Decoder 1226 recognizes a second ISA (either a subset of the first ISA or a different ISA).

[0109] In one example, the assigner and renamer block 1230 includes an assigner to reserve resources such as register files to store instruction processing results. However, threads 1201a and 1201b are potentially capable of out-of-order execution, with the assigner and renamer block 1230 also reserving additional resources, such as reorder buffers, to track instruction results. Unit 1230 may also include a register renamer to rename program / instruction reference registers to other registers within processor 1200. The reorder / reorder unit 1235 includes components such as the aforementioned reorder buffers, load buffers, and memory buffers to support out-of-order execution and subsequent in-order reordering of out-of-order instructions.

[0110] In one embodiment, the scheduler and execution unit block 1240 includes a scheduler unit for scheduling instructions / operations at execution units. For example, a floating-point instruction is scheduled on a port of an execution unit that has an available floating-point execution unit. Register files associated with the execution units are also included for storing processing results of information instructions. Example execution units include a floating-point execution unit, an integer execution unit, a jump execution unit, a load execution unit, a memory execution unit, and other known execution units.

[0111] The subordinate data cache and data translation buffer (D-TLB) 1250 are coupled to the execution unit(s) 1240. The data cache is intended to store recently used / operated elements, such as data operands, which are potentially maintained in memory coherence states. The D-TLB stores recent virtual / linear to physical address translations. As a specific example, a processor might include a page table structure to divide physical memory into a multitude of virtual pages.

[0112] Here, cores 1201 and 1202 share access to higher-level or more distant caches, such as a second-level cache, connected to the on-chip interface 1210. It's important to note that "higher-level" or "more distant" refers to cache levels that increase in or are located further away from the execution unit(s). In one embodiment, the higher-level cache 1200 is a last-level data cache—the last cache in the memory hierarchy on processor 1200—such as a second-level or third-level data cache. However, "higher-level cache" is not limited in this respect, as it can be connected to or include an instruction cache. A trace cache—a type of instruction cache—can instead be coupled behind decoder 1225 to store recently decoded traces. Here, an instruction potentially refers to a macro instruction (i.e., a macro instruction)., a non-privileged instruction that is recognized by the decoders), which can decode into a number of micro-instructions (micro-operations).

[0113] In the configuration shown, the processor 1200 also includes the interface module on the chip 1210. Historically, a memory controller, described in more detail below, was included in a computer system located external to the processor 1200. In this scenario, the chip's internal interface 121 communicates with devices located outside of the processor 1200, such as the system memory 1275, a chipset (which often includes a memory controller hub to connect to memory 1275 and an I / O controller hub to connect peripherals), a memory controller hub, a northbridge, or another integrated circuit. And in this scenario, the bus 1205 can include any known coupling structure, such as a multipoint link bus, a point-to-point coupling structure, a serial coupling structure, a parallel bus, a coherent (e.g.,cache-coherent) bus, a layered protocol architecture, a differential bus and a GTL bus.

[0114] Memory 1275 can be dedicated to processor 1200 or shared with other devices in a system. Conventional examples of memory types 1275 include DRAM, permanent storage (NV memory), and other familiar storage devices. It should be noted that device 1280 can include a graphics accelerator, processor or card coupled with a memory controller hub, data storage coupled with an I / O controller hub, a wireless transceiver, a flash memory device, an audio controller, a network controller, or any other familiar device.

[0115] While more logic and components can be integrated onto a single chip layer such as a SoC, each of these components can be included in the 1200 processor. For example, in one embodiment, a memory controller hub is located on the same package and / or chip layer as the 1200 processor. Here, part of the core (a part on the core) 1210 includes one or more controllers to connect to other devices such as memory 1275 or a graphics assembly 1280. The configuration that includes a coupling structure and controllers to connect to such devices is often referred to as an on-core (or non-core) configuration. As an example, the on-chip interface 1210 includes a ring coupling structure for on-chip communication and a high-speed serial point-to-point link 1205 for off-chip communication.Nevertheless, in the SOC environment, even more devices such as the network interface, coprocessors, memory 1275, graphics processor 1280 and any other known computer device / interface can be integrated on a single chip layer or integrated circuit to provide a small form factor with high functionality and low power consumption.

[0116] In one embodiment, the processor 1200 is capable of executing compiler, optimizer, and / or translator code 1277 to compile, translate, and / or optimize application code 1276 to support or connect with the devices and procedures described herein. A compiler often includes a program or set of programs to translate source text / code into target text / code. Typically, the compilation of program / application code with a compiler is performed in multiple stages and passes to transform high-level programming language code into low-level machine or assembly code. However, single-pass compilers can still be used for simple compilation.A compiler can use any known compilation techniques and perform any known compiler operations, such as lexical analysis, preprocessing, parsing, semantic analysis, code generation, code translation, and code optimization.

[0117] Larger compilers often include multiple phases, but most commonly these phases are enclosed within two general categories: (1) a front end, which is generally where syntactic processing, semantic processing, and some transformation / optimization can occur, and (2) a back end, which is generally where analysis, transformations, optimizations, and code generation can occur. Some compilers refer to a middle ground, illustrating the blurring of the boundaries between a compiler's front end and back end. As a result, reference to introduction, connection, generation, or any other compiler operation can occur in any of the phases or passes mentioned above, as well as any other known compiler phases or passes. As an illustrative example, a compiler potentially adds operations, calls, functions, and so on.Dynamic compilation involves one or more phases of the compilation process, such as inserting calls / operations in a front-end phase and then transforming those calls / operations into lower-level code during a transformation phase. It's important to note that during dynamic compilation, compiler code or dynamic optimization code can insert such operations / calls and optimize the code for runtime execution. As a specific illustrative example, binary code (already compiled code) can be dynamically optimized at runtime. Here, the program code can include the dynamic optimization code, the binary code, or a combination thereof.

[0118] Similar to a compiler, a translator, such as a binary translator, translates code either statically or dynamically to optimize and / or translate code. Therefore, references to code execution, application code, program code, or another software environment can refer to: (1) the execution of a compiler program or...(1) Compiler programs, optimization code optimizers, or translators, either dynamically or statically, to compile program code, to maintain software structures, to perform other operations, to optimize code, or to translate code; (2) the execution of the main program code, including operations / calls such as application code that has been optimized / compiled; (3) the execution of other program code, such as libraries associated with the main program code, to maintain software structures, to perform other software-related operations, or to optimize code; or (4) a combination thereof.

[0119] Referring to Fig. Figure 13 shows a block diagram of an embodiment of a multi-core processor. As shown in the embodiment of Fig.The processor 1300 includes several domains. Specifically, a core domain 1330 includes a multitude of cores 1330A-1330N, a graphics domain 1360 includes one or more graphics engines, which include a media engine 1365 and a system agent domain 1310.

[0120] In various embodiments, the System Agent domain 1310 handles power control events and power management, allowing individual units in domains 1330 and 1360 (e.g., cores and / or graphics engines) to be controlled independently. This enables them to dynamically operate at an appropriate power mode / level (e.g., active, turbo, sleep, hibernation, deep sleep, or another state similar to those found in extended configuration and power management interfaces) based on the activity (or inactivity) occurring in the given unit. Each of domains 1330 and 1360 can operate at different voltages and / or power levels, and furthermore, individual units within each domain can potentially operate at independent frequencies and voltages.It should be noted that, although it is shown with only three domains, the scope of the present invention is not limited in this respect and additional domains may be present in other embodiments.

[0121] As shown, each 1330 core further includes low-level caches in addition to various execution units and additional processing elements. Here, the different cores are coupled with each other and with a shared cache memory formed from a multitude of units or segments of a 1340A-1340N Last Level Cache (LLC); these LLCs often include memory and cache controller functionality and are shared among the cores and potentially also among the graphics engine.

[0122] As seen, a ring coupling structure 1350 couples the cores together and provides the connection between the core domain 1330, the graphics domain 1360, and the system agent circuits 1310 via a multitude of ring stops 1352A-1352N, each at a coupling between a core and an LLC segment. As can be seen in Fig.In section 13, the coupling structure 1350 is used to transport various pieces of information, including address information, data information, acknowledgment information, and snoop / invalid information. Although a ring coupling structure is illustrated, any known coupling structure or on-chip lattice can be used. As an illustrative example, some of the lattices discussed above (e.g., another on-chip coupling structure, the OSF (On-Chip System Fabric), an Advanced Microcontroller Bus Architecture (AMBA) coupling structure, a multidimensional mesh fabric, or another known coupling structure architecture) can be used in a similar manner.

[0123] As further described, the system agent domain 1310 includes the display engine 1312, which provides control of and an interface to a connected display. The system agent domain 1310 may include other units such as: an integrated memory controller 1320, which provides an interface to system memory (e.g., a DRAM implemented with multiple DIMMs); ​​and coherence logic 1322 to perform memory coherence operations. Multiple interfaces may be present to enable communication between the processor and the other circuitry. In one embodiment, at least one Direct Media Interface (DMI) 1316 interface and one or more PCIe™ interfaces 1314 are provided. The display engine and these interfaces typically couple to memory via a PCIe™ bridge 1318.To provide communication between other agents such as additional processors or other circuits, one or more other interfaces can be provided.

[0124] Referring to Fig. Figure 14 shows a block diagram of a representative kernel; specifically, logical building blocks of a back-end of a kernel such as kernel 1330. Fig. 13. In general, the in Fig. The structure shown in Figure 14 incorporates an out-of-order processor that includes a front-end unit 1470, which is used to retrieve incoming instructions, perform various processing operations (e.g., caching, decoding, branch prediction, etc.), and pass instructions / operations to an out-of-order (OOO) engine 1480. The OOO engine 1480 performs further processing on the decoded instructions.

[0125] Especially in the embodiment of Fig.The out-of-order engine 1480 includes an allocation unit 1482 to receive decoded instructions, which may be in the form of one or more micro-instructions or µOps from the front-end unit 1470, and to allocate them to the appropriate resources, such as registers. The instructions are then provided to a reservation station 1484, which reserves resources and schedules them for execution by a variety of execution units 1486A-1486N. Various types of execution units may be present, including, for example, arithmetic logic units (ALUs), load and store units, vector processing units (VPUs), and floating-point execution units. Results from these different execution units are provided to a reorder buffer (ROB) 1488, which takes any out-of-order results and reassembles them into the correct program order.

[0126] With further reference to Fig.It should be noted that both the front-end unit 1470 and the out-of-order engine 1480 are coupled to different levels of a memory hierarchy. Specifically, an instruction-level cache 1472 is shown, which in turn is coupled to a middle cache 1476, which is in turn coupled to a last-level cache 1495. In one embodiment, the last-level cache 1495 is implemented in an on-chip (sometimes referred to as non-core) unit 1490. As an example, unit 1490 is assigned to the system agent 1310 of Fig.13 similarly. As described above, non-core 1490 communicates with system memory 1499, which in the illustrated embodiment is implemented via ED RAM. It should also be noted that the various execution units 1486 within the out-of-order engine 1480 are in communication with a level 1 cache 1474, which is also in communication with the mid-level cache 1476. It should also be noted that the additional cores 1430N-2 - 1430N can couple with LLC 1495. Although at this high level in the embodiment of Fig. As shown in Figure 14, it is obvious that various modifications and additional components may be present.

[0127] With current reference to Fig.Figure 15 illustrates a block diagram of an exemplary computer system formed with a processor that includes execution units for executing an instruction, wherein one or more of the coupling structures implement one or more features according to an embodiment of the present invention. System 1500 includes a component, such as a processor 1502, to employ execution units that include logic for executing algorithms on process data according to the present invention as described herein. System 1500 is representative of processing systems based on the PENTIUM III™, PENTIUM 4™, Xeon™, Itanium, XScale™, and / or StrongARM™ microprocessors, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, and the like) can also be used.In one embodiment, the example system 1500 runs a version of the WINDOWS™ operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (for example, UNIX and Linux), embedded software, and / or graphical user interfaces can also be used. Thus, the embodiments of the present invention are not limited to a specific combination of hardware circuits and software.

[0128] Embodiments are not limited to computer systems. Alternative embodiments of the present invention can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications can include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of executing one or more instructions according to at least one embodiment.

[0129] In this illustrated embodiment, the processor 1502 comprises one or more execution units 1508 for implementing an algorithm that executes at least one instruction. One embodiment can be described in the context of a desktop or server system with a single processor, but alternative embodiments can be included in a multiprocessor system. The System 1500 is an example of a "hub" system architecture. The Computer System 1500 includes a processor 1502 for processing data signals.The Processor 1502 includes, as an illustrative example, a Complex Instruction Set Computer (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, a processor with a combination of instruction sets, or another processing unit, such as a digital signal processor. The Processor 1502 is coupled to a Processor Bus 1510, which transmits data signals between the Processor 1502 and other components in the System 1500. The System 1500 elements include, for example, the Graphics Accelerator 1512, Memory Controller Hub (MCH) 1516, Memory 1520, I / O Controller Hub (ICH) 1524, Wireless Transceiver 1526, Flash BIOS 1528, Network Controller 1534, and Audio Controller. 1536, serial expansion port 1538, I / O controller 1540, etc.) fulfill their conventional functions, which are well known to a professional.

[0130] In one embodiment, the processor 1502 includes an internal Level 1 (L1) cache memory 1504. Depending on the architecture, the processor 1502 may have a single internal cache or multiple levels of internal cache. Other embodiments include a combination of both internal and external caches, depending on the specific implementation and requirements. The register file 1506 stores various data types in different registers, including integer registers, floating-point registers, vector registers, banked registers, shadow registers, checkpoint registers, status registers, and instruction pointer registers.

[0131] The execution unit 1508, containing the logic for executing integer and floating-point operations, is also resident in the processor 1502. In one embodiment, the processor 1502 includes a microcode read-only memory (ucode) for storing microcode that, when executed, performs algorithms for specific macro instructions or handles complex scenarios. Here, the microcode may be updatable to address logical errors / corrections for the processor 1502. In another embodiment, the execution unit 1508 includes logic for processing a packed instruction set 1509. By incorporating the packed instruction set 1509 into the instruction set of a general-purpose processor 1502, along with the associated circuitry for executing the instructions, operations used by many multimedia applications can be performed using the packed data in a general-purpose processor 1502.This allows many multimedia applications to run faster and more efficiently by utilizing the full bus width of a processor to perform operations on packed data. This potentially eliminates the need to transfer smaller data units across the processor data bus to perform one or more operations, processing one data element at a time.

[0132] Other embodiments of an execution unit 1508 can also be used in microcontrollers, embedded processors, graphics processing units, DSPs, and other types of logic circuits. The system 1500 includes a memory 1520. The memory 1520 includes dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, or another memory element. The memory 1520 stores instructions and / or data represented by data signals to be executed by processor 1502.

[0133] It should be noted that any of the above-mentioned features or aspects of the invention may be present in one or more of the following: Fig.The 15 illustrated coupling structures can be used. For example, an on-chip coupling structure (QDI) not shown, used to couple internal units of the processor 1502, implements one or more aspects of the invention described above. Or the invention is connected to a processor bus 1510 (e.g., another known high-performance computing coupling structure), a high-bandwidth memory path 1518 to memory 1520, a point-to-point link to graphics accelerator 1512 (e.g., a Peripheral Component Interconnect Express (PCIe) compliant structure), a controller hub coupling structure 1522, an I / O or other coupling structure (e.g., USB, PCI, PCIe) to couple the other illustrated components.Some examples of such components include the audio controller 1536, the firmware hub (flash BIOS) 1528, the wireless transceiver 1526, the data storage device 1524, the legacy I / O controller 1510 with interfaces for user input and keyboard 1542, the serial expansion port 1538 such as a Universal Serial Bus (USB), and the network controller 1534. The data storage device 1524 can include a hard disk drive, a floppy disk drive, a CD-ROM drive, a flash memory unit, or a mass storage device.

[0134] Referring to Fig. Figure 16 shows a block diagram of a second system 1600 according to an embodiment of the present invention. As in Fig.As shown in Figure 16, the multiprocessor system 1600 is a system with a point-to-point coupling structure and comprises a first processor 1670 and a second processor 1680, which are coupled via a point-to-point coupling structure 1650. Each of the processors 1670 and 1680 can be a variant of a processor. In one embodiment, 1652 and 1654 are part of a serial, coherent point-to-point coupling structure, such as a high-performance architecture. As a result, the invention can be implemented within the QPL architecture.

[0135] Although only two processors, 1670 and 1680, are shown, it is understood that the scope of the present invention is not limited in this way. In other embodiments, one or more additional processors may be present in a given processor.

[0136] Processors 1670 and 1680 are each shown with integrated memory controller units 1672 and 1682, respectively. Processor 1670 also includes point-to-point (PP) interfaces 1676 and 1678 as part of its bus controller units; similarly, processor 1680 includes PP interfaces 1686 and 1688. Processors 1670 and 1680 can exchange information via a PP interface 1650 using PP interface circuits 1678 and 1688. As shown in Fig. As shown in Figure 16, the IMCs 1672 and 1682 couple the processors to the respective memories, namely to a memory 1632 and a memory 1634, which can be parts of a main memory that is locally connected to the respective processors.

[0137] The 1670 and 1680 processors each exchange data with a 1690 chipset via individual PP interfaces 1652 and 1654, using PP interface circuits 1676, 1694, 1686, and 1698. The 1690 chipset also exchanges information with a high-performance graphics circuit 1638 via an interface circuit 1692 along a high-performance graphics coupling structure 1639.

[0138] Each processor may contain a shared cache (not shown) either within each processor or outside of the two processors, but which is connected to the processors via a PP coupling structure such that one (or both) of the local cache information of the processors can be stored in the shared cache when a processor is put into a power-saving mode.

[0139] The chipset 1690 can be coupled to a first bus 1616 via the interface 1696. In one embodiment, the first bus 1616 can be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third-generation I / O interconnect bus, although the scope of the present invention is not limited in this way.

[0140] As in Fig.As shown in Figure 16, various I / O devices 1614 are coupled to the first bus 1616, together with a bus bridge 1618 that couples the first bus 1616 to a second bus 1620. In one embodiment, the second bus 1620 includes a Low Pin Count (LPC) bus. Various devices are coupled to the second bus 1620, including, for example, a keyboard and / or mouse 1622, communication devices 1627, and a storage unit 1628 such as a disk drive or other mass storage device, which in one embodiment often includes instructions / code and data 1630. Furthermore, an audio I / O 1624 is shown coupled to the second bus 1620. It should be noted that other architectures are possible in which the included components and coupling structure architectures vary. For example, a system may use a different architecture than the point-to-point architecture of Fig. 16. Implement a multidrop bus or other such architecture.

[0141] With current reference to Fig. Figure 17 shows an embodiment of a system-on-a-chip (SOC) design according to the inventions. As a specific illustrative example, SOC 1700 is enclosed in the subscriber terminal equipment (STE). In one embodiment, STE refers to any device used by an end-user for communication, for example, a handheld telephone, smartphone, tablet, ultra-thin notebook, notebook with broadband adapter, or any other similar communication device. Frequently, a STE connects to a base station or node, which by its nature may correspond to a mobile station (MS) in a GSM network.

[0142] Here, the SOC 1700 includes two cores – 1706 and 1707. Similar to the description above, cores 1706 and 1707 can correspond to an instruction set architecture, such as an Intel® Architecture Core™-based processor, an Advanced Micro Devices, Inc. (AMD) processor, a MIPS-based processor, or an ARM-based processor design, or a customer thereof, as well as their licensees or users. Cores 1706 and 1707 are coupled to the cache controller 1708, which is connected to the bus interface unit 1709 and the L2 cache 1711 to communicate with other parts of the system 1700. The coupling structure 1710 includes an on-chip coupling structure such as an IOSF, AMBA, or another coupling structure discussed above, which may implement one or more of the aspects described here.

[0143] The interface 1710 provides communication channels to other components such as a Subscriber Identity Module (SIM) 1730 as an interface to a SIM card, a boot ROM 1735 for storing boot code for execution by cores 1706 and 1707 for initializing and booting the SOC 1700, an SDRAM controller 1740 as an interface to external memory (e.g., DRAM 1760), a flash controller 1745 as an interface to non-volatile memory (e.g., Flash 1765), a peripheral controller 1750 (e.g., serial peripheral interface) as an interface to peripheral devices, video codecs 1720 and video interface 1725 for displaying and receiving input (e.g., touch-enabled input), GPU 1715 for performing graphics-related calculations, etc. Each of these interfaces can incorporate aspects of the invention described herein.

[0144] Furthermore, the system illustrates communication peripherals such as a Bluetooth module (1770), 3G modem (1775), GPS (1785), and Wi-Fi (1785). As mentioned above, UE includes radio communication capabilities. Consequently, not all of these peripheral communication modules are required. However, some form of radio communication for external communication must be present in a UE.

[0145] Although the present invention has been described with regard to a limited number of embodiments, those skilled in the art are aware that many further modifications and variants are possible. The appended claims are intended to cover all such modifications and variants that correspond to the purpose and scope of protection of the present invention.

[0146] A design can go through various stages, from creation to simulation to manufacturing. Data representing a design can represent it in several ways. First, as is useful in simulations, the hardware can be represented using a hardware description language or another functional description language. Additionally, a circuit-level model with logic and / or transistor gates can be created at some stages of the design process. Furthermore, most designs eventually reach a data layer that represents the physical arrangement of various devices within the hardware model.When conventional semiconductor fabrication techniques are used, the data representing the hardware model can be data specifying the presence or absence of various features on different mask layers for masks used to fabricate the integrated circuit. In a design representation, the data can be stored in the form of a machine-readable medium. A memory or magnetic or optical storage medium, such as a disk, can be the machine-readable medium for storing information transmitted by means of an optical or electrical wave that is modulated or otherwise generated to transmit such information.When an electrical carrier wave, which displays or carries the code or design, is transmitted to such an extent that copying, buffering, or retransmission of the electrical signal occurs, a new copy is created. Thus, a communications service provider or network service provider can at least temporarily store an item, such as information encoded in a carrier wave, on a specific machine-readable medium, embodying the techniques of embodiments of the present invention.

[0147] A module as used herein refers to any combination of hardware, software, and / or firmware. For example, a module includes hardware, such as a microcontroller, connected to a non-volatile medium to store code adapted for execution by the microcontroller. Therefore, in one embodiment, the use of a module refers to the hardware specifically configured to recognize and / or execute the code stored on a non-volatile medium. In another embodiment, the use of a module further refers to the non-volatile medium containing the code specifically adapted for execution by the microcontroller to perform predetermined operations.And as can be inferred from yet another embodiment, the term module (in this example) can refer to the combination of microcontroller and non-volatile medium. Module boundaries, illustrated as separate, conventionally often vary and can potentially overlap. For example, a first and a second module may share hardware, software, firmware, or a combination thereof, while potentially keeping some independent hardware, software, or firmware separate. In one embodiment, the use of the term logic includes hardware such as transistors, registers, or other hardware such as programmable logic assemblies.

[0148] The use of the phrase "is configured" in an embodiment refers to the arranging, assembling, manufacturing, offering for sale, importing, and / or constructing of a device, hardware, logic, or element to perform an intended or specific task. In this example, a device or element thereof that is not operating is still "configured" to perform an intended task if it is designed, coupled, and / or connected to perform that intended task. As a purely illustrative example, a logic gate can provide a 0 or 1 during operation. But a logic gate that is "configured" to provide a enable signal to a clock does not include every potential logic gate that can provide a 1 or 0. Instead, the logic gate is one that is coupled in such a way that, during operation, the 1 or 0 output enables the clock.Once again, it should be noted that the use of the term "is configured" does not require operation, but instead focuses on the latent state of a device, hardware and / or element, where in the latent state the device, hardware and / or element is designed to perform a specific task when the device, hardware and / or element is operating.

[0149] Furthermore, the use of the terms "to," "capable of," and / or "operational to" in an embodiment refers to a device, logic, hardware, and / or element that is designed in such a way as to enable the use of the device, logic, hardware, and / or element in a specified manner. As above, it should be noted that the use of "to," "capable of," and / or "operational to" in an embodiment refers to the latent state of a device, logic, hardware, and / or element where the device, logic, hardware, and / or element is not operational but is designed in such a way as to enable the use of a device in a specified manner.

[0150] A value, as used herein, includes any known representation of a number, a state, a logical state, or a binary logical state. The use of logic levels, logic values, or logical values ​​is also referred to as 1s and 0s, which simply represent binary logical states. For example, a 1 refers to a high logic level, and 0 refers to a low logic level. In one embodiment, a memory cell, such as a transistor or a flash cell, may contain a single logical value or multiple logical values. However, other representations of values ​​have been used in computer systems. The decimal number ten, for example, can also be represented as a binary value, 1010, and a hexadecimal letter, A. Therefore, a value includes any representation of information that can be contained in a computer system.

[0151] Furthermore, states can be represented by values ​​or parts of values. For example, a first value, such as a logical one, can represent a default or initial state, while a second value, such as a logical zero, can represent a non-default state. Additionally, in one embodiment, the terms reset and set refer to a default and an updated value or state, respectively. A default value, for instance, potentially includes a high logical value, i.e., reset, while an updated value potentially includes a low logical value, i.e., set. It should be noted that any combination of values ​​can be used to represent any number of states.

[0152] The embodiments of the aforementioned methods, hardware, software, firmware, or code may be implemented by instructions or code stored on a machine-accessible, machine-readable, or computer-readable medium and capable of being executed by a processing element. A non-volatile machine-accessible / machine-readable medium includes any mechanism that provides (i.e., stores and / or transmits) information in a form readable by a machine such as a computer or electronic device. For example, a non-volatile machine-accessible medium includes random-access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; a magnetic or optical storage medium; flash memory devices; an electrical storage device, optical storage devices, acoustic storage devices; or any other form of storage device to store information transmitted by volatile (propagated) signals (e.g.,carrier waves, infrared signals, digital signals) etc. were received, which are to be distinguished from the non-volatile media that can receive information from them.

[0153] Commands used to program logic to execute embodiments of the invention can be stored in a memory within the system, such as DRAM, cache, flash memory, or other storage devices. Furthermore, the commands can be disseminated over a network or using other machine-readable media. Thus, a machine-readable medium can represent any mechanism for storing or transmitting information in a (e.g.,Computer-readable medium includes, but is not limited to, floppy disks, optical drives, CDs, read-only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or any non-volatile, machine-readable memory used in the transmission of information over the internet using electrical, optical, acoustic, or other forms of propagating signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, computer-readable medium includes any type of non-volatile, machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0154] The following examples relate to embodiments according to this specification. One or more embodiments can provide a device, a system, a machine-readable memory, a machine-readable medium, and a method for providing a synchronization counter and a layer stack that includes physical layer logic, link layer logic, and protocol layer logic, wherein the physical layer logic synchronizes a reset of the synchronization counter to an external deterministic signal and synchronizes entry into a link transmit state with the deterministic signal.

[0155] In at least one example, the physical layer logic further initializes a data link using one or more supersequences.

[0156] In at least one example, the entry into the link sending state coincides with the start of the data sequence (SDS) that was sent to complete the initialization of the data link.

[0157] In at least one example, the SDS is sent according to the deterministic signal.

[0158] In at least one example, each supersequence includes a corresponding repetition sequence that includes an electrically inactive exit-ordered set and a corresponding number of training sequences.

[0159] In at least one example, the SDS interrupts the super sequences.

[0160] In at least one example, the supersequences each include a corresponding repetition sequence that includes at least one electrically inactive exit-ordered set (EIEOS) and a corresponding number of training sequences.

[0161] In at least one example, the EIEOS of a supersequence is sent so that it coincides with the synchronization counter.

[0162] In at least one example, the physical layer logic continues to synchronize with a deterministic interval based on a received EIEOS.

[0163] In at least one example, synchronizing to a deterministic interval based on a received EIEOS involves determining an end boundary of the received EIEOS.

[0164] In at least one example, the end boundary is used to synchronize entry into the link sending state.

[0165] In at least one example, the end limit is used to synchronize the exit from a partial width link transmit state.

[0166] In at least one example, the physical layer logic further generates a special supersequence and sends the special supersequence, which is to be synchronized with the deterministic signal.

[0167] In at least one example, the physical layer logic specifies a target latency to a remote agent, where the remote agent uses the target latency to apply a delay to match the actual latency to the target latency.

[0168] In at least one example, the target latency is communicated in the user data of a training sequence.

[0169] In at least one example, the deterministic signal includes a planetary synchronization signal for a device.

[0170] In at least one example, the physical layer logic further synchronizes a periodic control window embedded in a link layer data stream with the deterministic signal via a serial data link, with the control window being configured for the exchange of physical layer information during a link send state.

[0171] In at least one example, the physical layer information includes information for use in initiating state transitions at the data link.

[0172] In at least one example, the control windows are embedded according to a defined control interval, and the control interval is based at least partially on the deterministic signal.

[0173] One or more examples can further provide the sending of the supersequences to a remote agent connected to the data link during the initialization of the data link, and at least one element of the supersequence is synchronized with the deterministic signal.

[0174] In at least one example, the element includes an EIEOS.

[0175] In at least one example, each supersequence includes a corresponding repetition sequence that includes at least EIEOS and a corresponding number of training sequences.

[0176] One or more examples can further provide the sending of a stream of link layer flits in the link send state.

[0177] One or more examples can further provide the synchronization of a regular control window to be embedded in the stream with the deterministic signal, with the control window configured for the exchange of physical layer information during the link-send state.

[0178] One or more examples can further provide the sending of delay information to a remote agent connected to the data link, where the delay corresponds to the deterministic signal.

[0179] One or more embodiments can provide a device, a system, a machine-readable memory, a machine-readable medium and a method for determining a target latency for a serial data link, receiving a data sequence over the data link that is synchronized with a synchronization counter connected to the data link, and maintaining the target latency using the data sequence.

[0180] In at least one example, the data sequence includes a supersequence to enclose a repetition sequence, where the sequence is repeated with a defined frequency.

[0181] In at least one example, the sequence includes an electrically inactive exit-ordered set (EIEOS).

[0182] In at least one example, the sequence begins with the EIEOS followed by a predefined number of training sequences.

[0183] In at least one example, at least one of the training sequences includes data that determines the target latency.

[0184] In at least one example, at least part of the sequence is encrypted using a pseudorandom binary sequence (PRBS).

[0185] One or more examples can further provide the determination of an actual latency of the data link based on the reception of the data sequence.

[0186] One or more examples can further provide information on determining a deviation of the actual latency from the target latency.

[0187] One or more examples can further provide the impetus for the deviation to be corrected.

[0188] One or more embodiments can provide a device, a system, a machine-readable memory, a machine-readable medium and a method for determining whether the width of flits to be sent over a serial data link enclosing a number of lanes is a multiple of the number of lanes, and sending the flits over the serial data link, sending two flits so that they overlap on the lanes, if the width of the flits is not a multiple of the number of lanes.

[0189] In at least one example, the overlap involves sending one or more bits of a first of the two flits over a first part of the number of tracks simultaneously with sending one or more bits of a second of the two flits over a second part of the number of tracks.

[0190] In at least one example, at least some bits of the flits are not sent in sequence.

[0191] In at least one example, flits do not overlap when the width of the flits is a multiple of the number of lanes.

[0192] In at least one example, the width of the flits includes 192 bits.

[0193] In at least one example, the number of lanes includes 20 lanes in at least one link transmission state.

[0194] One or more examples can further provide the transition to a different new link width, which includes a second number of lanes.

[0195] One or more examples can further provide the means to determine whether the width of the flits is a multiple of the second number of lanes.

[0196] In at least one example, the transition is aligned with a non-overlapping flit boundary.

[0197] One or more embodiments can provide a device, a system, a machine-readable memory, a machine-readable medium and a method for providing physical layer logic to receive a bitstream including a set of flits over a serial data link, wherein corresponding parts of at least two of the set of flits are simultaneously sent on lanes of the data link, and link layer logic to reconstruct the set of flits from the received bitstream.

[0198] In at least one example, part of Flits' theorem has overlapping boundaries.

[0199] In at least one example, overlapping boundaries include sending one or more final bits of a first of the two flits over a first part of the number of tracks simultaneously with sending one or more starting bits of a second of the two flits over a second part of the number of tracks.

[0200] In at least one example, the width of the flits is not a multiple of the number of lanes of the data link.

[0201] In at least one example, the width of the flits includes 192 bits and the number of lanes includes 20 lanes.

[0202] In at least one example, at least some of the bits of the flits are not sent in sequence.

[0203] One or more examples can further provide a physical layer (PHY) configured to be coupled to a link, wherein the link includes an initial number of traces and the PHY enters a loopback state, and wherein, when the PHY is memory-resident in the loopback state, it introduces a specialized pattern on the link.

[0204] One or more examples may further include a physical layer (PHY) configured to be coupled to a link, wherein the link includes an initial number of lanes, and wherein the PHY includes a synchronization (sync) counter, and wherein the PHY sends an electrically inactive exit order set (EIEOS) aligned with the synchronization counter associated with a training sequence.

[0205] In at least one example, a synchronization counter value is not exchanged by the synchronization counter during each training sequence.

[0206] In at least one example, the EIEOS alignment with the synchronization counter acts as a proxy for exchanging the synchronization counter value from the synchronization counter during each training sequence.

[0207] One or more examples can further provide a physical layer (PHY) configured to be coupled to a link, wherein the PHY includes a PHY state machine for transitioning between a plurality of states, and wherein the PHY state machine is capable of transitioning from a first state to a second state based on a handshake event and of transitioning the PHY from a third state to a fourth state based on a primary timer event.

[0208] In at least one example, the PHY state machine is capable of transitioning the PHY from a fifth state to a sixth state based on a primary time event in combination with a secondary timer event.

[0209] References in this description to “an embodiment” mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, uses of the phrase “in an embodiment” at various points in this entire description do not necessarily all refer to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0210] The foregoing description provides a detailed account with reference to specific exemplary embodiments. However, it is obvious that various modifications and changes can be made to it without deviating from the broader meaning and scope of the invention as set forth in the appended claims. The description and drawings should therefore be viewed in an illustrative rather than a limiting sense. Furthermore, the foregoing use of "embodiment" and other exemplary language does not necessarily refer to the same embodiment or example, but may refer to different and distinct embodiments, as well as potentially to the same embodiment.

Claims

[1] Device comprising: an interface logic for differential signaling on a plurality of paths, where the interface logic is configured to transmit a plurality of flits, wherein the plurality of flits has a plurality of nibbles, and wherein a first of the plurality of flits is transmitted with a clean flit boundary such that each of the plurality of paths is used to transmit initial nibbles of the first flit in a first unit interval (UI), and wherein initial nibbles of a second flit of the plurality of flits, together with final nibbles of the first flit, are transmitted on the plurality of paths during a subsequent UI. [2] Device according to claim 1, wherein the plurality of flits comprises at least five flits which are to be transmitted over 48 UIs before a subsequent clean flit boundary is to be transmitted. [3] Device according to claim 1, wherein a flit has 192 bits and a nibble has 4 bits. [4] Device according to claim 3, wherein the bits of each flit are to be transmitted in a sequence starting with the bit 4n+3. [5] Device according to claim 1, wherein the plurality of paths comprises eight paths or twenty paths. [6] Device according to claim 1, wherein the interface logic comprises a physical level logic, a link level logic and a protocol level logic. [7] Device according to claim 6, wherein the protocol-level logic is said to support cache coherence transactions. [8] Device according to claim 1, wherein the interface logic is contained in a processor which is connected in a socket of a server having at least two sockets. [9] Device according to claim 1, wherein the interface logic is contained in a system on a chip (SoC). [10] Device according to claim 1, wherein the SoC is coupled with a plurality of other SoCs in a microserver. [11] Device according to claim 9, further comprising a radio. [12] Device comprising: A controller for connecting via an interface between at least a first processor for detecting a first instruction set and a second processor for detecting a second instruction set that is different from the first instruction set, wherein the controller has interface logic for being connected to a serial differential interconnect having a plurality of paths, wherein the interface logic is configured to transmit a plurality of flits, each of the plurality of flits having a corresponding plurality of nibbles, wherein a first of the plurality of flits is transmitted with a clean flit boundary, with initial nibbles of the first flit being sent on each path of the plurality of paths in a first unit interval (UI), and initial nibbles of a second of the plurality of flits being sent with end nibbles of the first flit on the plurality of paths during a subsequent UI. [13] Device according to claim 12, wherein each of the plurality of Flit comprises 192 bits and 48 nibbles. [14] Device according to claim 12, wherein the plurality of paths comprises eight paths or twenty paths. [15] Device according to claim 12, wherein the first and the second processor are coupled to the controller. [16] Device according to claim 15, wherein the first instruction set comprises an Intel-based instruction set. [17] Device comprising: an interface logic for differential signaling on a plurality of paths, where the interface logic is intended to logically group data into a plurality of information quanta, where the interface logic at least a first part of a first information quantum of the plurality of information quanta should be transferred to all of the plurality of paths during a first unit interval (UI); and at least a second part of the first information quantum, together with a part of a second information quantum of the plurality of information quanta, should be transferred along the plurality of paths during a subsequent UI. [18] Device according to claim 17, wherein the first part of the first information quantum is to be transmitted during at least the first four unit intervals and wherein the second part of the first information quantum and the part of the second information quantum together are to be transmitted on the plurality of paths during the following four successive UIs. [19] Device according to claim 17, wherein the first information quantum comprises 192 bits. [20] Device according to claim 17, wherein the interface logic includes protocol logic to support cache-coherent transactions.

Citation Information

Patent Citations

  • Method, System, and Apparatus for System Level Initialization

    US20090265472A1

  • Method and system for transmit scheduling for multi-layer network interface controller (NIC) operation

    US20110307577A1

  • IMPLEMENTING QUICKPATH INTERCONNECT PROTOCOL OVER A PCIe INTERFACE

    US20120079156A1