High-performance wiring bit transmission layer

The HPI architecture addresses the limitations of existing wiring systems by implementing a layered protocol stack for efficient communication and power management, enhancing performance and power efficiency in high-performance computing environments.

DE112013005104B4Active Publication Date: 2026-02-26INTEL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE112013005104
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2012-10-22
Filing Date
2013-03-27
Publication Date
2026-02-26
Estimated Expiration
2033-03-27

AI Technical Summary

Technical Problem

As computing systems evolve with increased complexity and demand for higher performance and power efficiency, existing wiring architectures are pushed to their limits, particularly in high-performance computing environments, necessitating improved communication between components to meet bandwidth requirements and optimize power consumption.

Method used

A high-performance wiring (HPI) architecture is introduced, featuring a layered protocol stack with a transaction layer, data link layer, and physical layer, utilizing point-to-point connections and credit-based flow control to enhance communication efficiency and power management, applicable to various computing platforms including servers, mobile devices, and embedded systems.

Benefits of technology

The HPI architecture enhances communication performance and power efficiency by optimizing bandwidth and reducing power consumption, supporting flexible topologies and power management, while being adaptable to different market segments' specific needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A periodic control window is embedded in a data link layer data stream transmitted over a serial data link. This control window is configured to provide data link layer information, including details for initiating state transitions on the data link. The data link layer data can be sent during a link transmission state, and the control window can interrupt the sending of flits. In one aspect, the information includes link width transition data, indicating an attempt to change the number of active traces on the link.
Need to check novelty before this filing date? Find Prior Art

Description

AREA

[0001] The present disclosure relates generally to the field of computer development and in particular to software development, which includes the coordination of interdependent constrained systems. BACKGROUND

[0002] Advances in semiconductor processing and logic design allow for an increase in the amount of logic that can be contained on integrated circuit devices. As a result, computer system configurations have evolved from a single integrated circuit or multiple integrated circuits in a system to multiple cores, multiple hardware threads, and multiple logical processors contained within a single integrated circuit, as well as other interfaces integrated within such processors. Typically, a processor or integrated circuit comprises a single physical processor chip, which may contain any number of cores, hardware threads, logical processors, interfaces, memory, controller hubs, and so on.

[0003] As a result of the increased ability to pack more processing power into smaller components, the popularity of smaller computing devices has grown. Smartphones, tablets, ultrathin notebooks, and other consumer devices have proliferated exponentially. However, these smaller devices rely on servers for both data storage and complex processing that exceeds their form factor. Consequently, the demand in the high-performance computing market (i.e., server room) has also increased. For example, modern servers typically contain not just a single multi-core processor, but also multiple physical processors (also known as multi-sockets) to increase computing power. But as computing power grows along with the number of devices in a computer system, communication between sockets and other devices becomes more critical.

[0004] In fact, wiring systems have evolved from rather conventional multipoint buses, which primarily handled electrical communications, into fully featured wiring architectures that enable high-speed communication. As the demand for power consumption increases with future processors operating at even higher rates, the capabilities of existing wiring architectures are unfortunately being pushed to their limits.

[0005] US 7089485 B2 discloses a device comprising a layer stack that includes a data link layer logic and a physical layer logic.

[0006] US patent 2006 / 0034295A1 discloses that the number of tracks used by a physical layer can be increased or decreased. It also discloses that physical layer logic can be brought into a low-power state.

[0007] From US 2006 / 004 1696 A1 it is known that a bit transmission layer is brought into a reset state. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 illustrates a simplified block diagram of a system that includes serial point-to-point wiring for connecting I / O devices in a computer system in accordance with one embodiment; Fig. Figure 2 illustrates a simplified block diagram of a shift protocol stack in accordance with one embodiment; Fig. Figure 3 illustrates an embodiment of a transaction descriptor; Fig. Figure 4 illustrates an embodiment of a serial point-to-point connection. Fig. Figure 5 illustrates embodiments of potential high-performance wiring system configurations (HPI system configurations). Fig.Figure 6 illustrates an embodiment of a shift log stack associated with an HPI. Fig. Figure 7 illustrates a representation of an exemplary state machine. Fig. Figure 8 illustrates exemplary control supersequences. Fig. Figure 9 illustrates a representation of an example control window embedded in a data stream. Fig. Figure 10 illustrates a flowchart of an example receipt exchange. Fig. Figure 11 illustrates a flowchart of an exemplary transition to a partial width state. Fig. Figure 12 illustrates an exemplary transition from a partial width state. Fig. Figure 13 illustrates an embodiment of a block diagram for a computer system that includes a multi-core processor. Fig.Figure 14 illustrates another embodiment of a block diagram for a computer system that includes a multi-core processor. Fig. Figure 15 illustrates an embodiment of a block diagram for a processor. Fig. Figure 16 illustrates another embodiment of a block diagram for a computer system that includes a processor. Fig. Figure 17 illustrates an embodiment of a block for a computer system that contains multiple processor sockets. Fig. Figure 18 illustrates another embodiment of a block diagram for a computer system.

[0008] The same reference symbols and designations indicate the same elements in the different drawings. DETAILED DESCRIPTION

[0009] To provide a thorough understanding of the present invention, numerous specific details are set forth in the following description, such as examples of specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitecture details, specific register configurations, specific instruction types, specific system components, specific processor pipeline stages, specific wiring layers, specific packet / transaction configurations, specific transaction names, specific protocol exchanges, specific link widths, specific implementations, and operation, etc. However, it may be obvious to the person skilled in the art that these specific details need not necessarily be used to realize the subject matter of the present disclosure.In other cases, a very detailed description of known components or procedures, such as specific or alternative processor architectures, specific logic circuits / specific code for described algorithms, specific firmware code, a low-level wiring operation, specific logic configurations, specific manufacturing techniques and materials, specific compiler implementations, a specific expression of the algorithm in code, specific shutdown and gate control techniques / shutdown and gate control logic, and other specific operational details of a computer system has been avoided in order to prevent unnecessary obfuscation of the present disclosure.

[0010] Although the following embodiments, with respect to energy saving, energy efficiency, processing efficiency, etc., can be described in specific integrated circuits such as computer platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of the embodiments described herein can be applied to other types of circuits or semiconductor devices that can also utilize these features. For example, the disclosed embodiments are not limited to server computer systems, desktop computer systems, laptops, or Ultrabooks™, but can also be used in other devices such as handheld devices, smartphones, tablets, other thin notebooks, single-chip devices (SoC devices), and embedded applications.Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Similar techniques for high-performance wiring can be applied here to increase performance (or even save power) in low-power wiring. Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of performing the functions and operations taught herein. Furthermore, the devices, methods, and systems described here are not limited to physical computing devices but can also relate to software optimizations for power saving and energy efficiency.As will become clearer from the following description, the embodiments of the methods, devices and systems described herein (whether they relate to hardware, firmware, software or a combination thereof) can be considered indispensable for a future with a “green technology” that is aligned with performance considerations.

[0011] As computing systems evolve, their components become more complex. To ensure that bandwidth requirements for optimal component operation are met, the complexity of the wiring architecture for coupling and communication between components has also increased. Furthermore, different market segments demand different aspects of wiring architectures to suit their specific needs. For example, servers require higher performance, while the mobile ecosystem may occasionally sacrifice overall performance for power savings. Nevertheless, a common goal of most fabrics is to deliver the highest possible performance with maximum power savings. Moreover, the subject matter described here can potentially utilize a variety of different wiring configurations.

[0012] Among other examples, the Peripheral Component Interconnect (PCI) Express Wiring Fabric Architecture (PCIe Wiring Fabric Architecture) and the QuickPath Interconnect Fabric Architecture (QPI Fabric Architecture) can potentially be improved in accordance with one or more of the principles described here. For example, a primary goal of PCIe is to enable components and devices from different vendors to work together in an open architecture spanning multiple market segments: clients (desktop and mobile), servers (standard and enterprise), and embedded and communication devices. PCI Express is a high-performance, general-purpose I / O wiring standard defined for a wide variety of future computing and communication platforms.Some PCI attributes, such as its usage model, load-store architecture, and software interfaces, have been maintained across revisions, while earlier implementations using a parallel bus have been replaced by a highly scalable, fully serial interface. More recent versions of PCI Express leverage point-to-point wiring, switch-based technology, and a packet protocol to deliver new levels of performance and features. Power management, quality of service (QoS), hot-plug support, data integrity, and error handling are just some of the advanced features supported by PCI Express.Although the main discussion here relates to a new high-performance wiring architecture (HPI architecture), aspects of the invention described herein may be applied to other wiring architectures such as a PCIe-compatible architecture, a QPI-compatible architecture, a MIPI-compatible architecture, a high-performance architecture, or any other known wiring architecture.

[0013] In Fig.Figure 1 shows an embodiment of a fabric consisting of point-to-point connections that wire a set of components together. The system 100 includes a processor 105 and a system memory 110 coupled to a controller hub 115. The processor 105 can include any processing element, such as a microprocessor, a host processor, an embedded processor, a coprocessor, or another processor. The processor 105 is coupled to the controller hub 115 via a front-side bus (FSB) 106. In one embodiment, the FSB 106 is a serial point-to-point connection as described below. In another embodiment, the connection 106 includes a serial differential wiring architecture compatible with a different wiring standard.

[0014] System memory 110 contains some storage device such as read / write memory (RAM), non-volatile memory (NV memory), or other memory that devices in system 100 can access. System memory 110 is coupled to the controller hub 115 via a memory interface 116. Examples of memory interfaces include a double data rate memory interface (DDR memory interface), a dual-channel DDR memory interface, and a dynamic RAM (DRAM) memory interface.

[0015] In one embodiment, the controller hub 115 can contain a root hub, root complex, or root controller, similar to a PCIe interconnect hierarchy. Examples of the controller hub 115 include a chipset, a memory controller hub (MCH), a northbridge, a wiring controller hub (ICH), a southbridge, and a root controller / hub. Often, the term chipset refers to two physically separate controller hubs, such as a memory controller hub (MCH) coupled to a wiring controller hub (ICH). It is noted that current systems often include the MCH integrated with the processor 105, while the controller 115 is intended to communicate with I / O devices in a manner similar to that described below. In some embodiments, the root complex 115 optionally supports peer-to-peer routing.

[0016] The controller hub 115 is connected to the switch / bridge 120 via a serial connection 119. The input / output modules 117 and 121, which can also be referred to as interfaces / ports 117 and 121, can contain / implement a layered protocol stack to provide communication between the controller hub 115 and the switch 120. In one embodiment, multiple devices can be connected to the switch 120.

[0017] The switch / bridge 120 routes packets / messages upwards from the device 125, i.e., in a hierarchy towards a root complex to the controller hub 115, and downwards, i.e., in a hierarchy from a root controller, from the processor 105, or from the system memory 110 to the device 125. In one embodiment, the switch 120 is referred to as a logical arrangement of several virtual PCI-to-PCI bridge devices.Device 125 includes any internal or external device or component intended to be coupled to an electronic system, such as an I / O device, a network interface controller (NIC), an add-in card, an audio processor, a network processor, a hard disk drive, a storage device, a CD / DVD-ROM drive, a monitor, a printer, a mouse, a keyboard, a router, a portable storage device, a FireWire device, a universal serial bus (USB) device, a scanner, and other input / output devices. In PCIe terminology, such a device is often referred to as an endpoint. Although not specifically shown, Device 125 may include a bridge (e.g., a PCIe-to-PCI / PCI-X bridge) to support legacy devices or other versions of devices, or wiring fabrics supported by such devices.

[0018] The graphics accelerator 130 can also be coupled to the controller hub 115 via a serial connection 132. In one embodiment, the graphics accelerator 130 is coupled to an MCH, which is coupled to an ICH. The switch 120, and consequently the I / O device 125, is then coupled to the ICH. The I / O modules 131 and 118 are also intended to implement a layer protocol stack for communication between the graphics accelerator 130 and the controller hub 115. Similar to the MCH discussion above, a graphics controller or the graphics accelerator 130 itself can be integrated into the processor 105.

[0019] Transition to Fig.Figure 2 shows an embodiment of a layer protocol stack. The layer protocol stack 200 can include any form of layer communication stack, such as a QPI stack, a PCIe stack, a next-generation high-performance computer wiring (HPI) stack, or another layer stack. In one embodiment, the protocol stack 200 can include a transaction layer 205, a data link layer 210, and a physical layer 220. The communication protocol stack 200 can be an interface such as interfaces 117, 118, 121, 122, 126, and 131 in Fig. 1. The representation as a communication protocol stack can also be described as a module or as an interface that implements / contains a protocol stack.

[0020] Packets can be used to transmit information between components. Packets can be formed in the transaction layer (205) and the data link layer (210) to transmit information from the sending component to the receiving component. As the sent packets flow through the other layers, they are augmented with additional information used to process packets in those layers. On the receiving side, the reverse process takes place, and the packets are transformed from their physical layer (220) representation to the data link layer (210) representation and finally (for transaction layer packets) to a form that can be processed by the transaction layer (205) of the receiving device.

[0021] In one embodiment, the transaction layer 205 can provide an interface between a processing core of the device and the wiring structure, such as the data link layer 210 and the physical layer 220. In this regard, a primary responsibility of the transaction layer 205 can include merging and splitting packets (i.e., transaction layer packets or TLPs). The transaction layer 205 can also manage credit-based flow control for TLPs. In some implementations, among other examples, split transactions—that is, transactions with a request and response separated by a time interval—can be used, allowing one connection to transmit other traffic while the destination device gathers data for the response.

[0022] To implement virtual channels and networks using the wiring fabric, credit-based flow control can be employed. For example, a device for each of the receive buffers in transaction layer 205 can announce an initial amount of credit. An external device at the opposite end of the connection, such as the controller hub 115 in Fig. One can count the number of credits consumed by each TLP. A transaction can be sent if it does not exceed a credit limit. Upon receiving a response, a credit amount is restored. One example of an advantage of such a credit scheme, among other potential benefits, is that the latency of credit return does not affect performance as long as the credit limit is not reached.

[0023] In one embodiment, four transaction address spaces can contain a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions contain one or more read and write requests to transfer data to / from a memory-image-based location. In one embodiment, memory space transactions can use two different address formats, such as a short address format like a 32-bit address or a long address format like a 64-bit address. Configuration space transactions can be used to access the configuration space of various devices connected to the wiring. Transactions into the configuration space can contain read and write requests.To support in-band communication between wiring agents, message space transactions (or simply messages) can also be defined. Thus, in an exemplary embodiment, transaction 205 can concatenate the packet start block / payload information 206.

[0024] Briefly based on Fig.Figure 3 shows an exemplary embodiment of a transaction layer packet descriptor. In one embodiment, the transaction descriptor 300 can be a mechanism for transmitting transaction information. In this respect, the transaction descriptor 300 supports the identification of transactions in a system. Other potential uses include tracking changes to the default transaction order and mapping a transaction to channels. The transaction descriptor 300 can, for example, contain a global identifier field 302, an attribute field 304, and a channel identifier field 306. In the example shown, the global identifier field 302 is shown to include a local transaction identifier field 308 and a source identifier field 310. In one embodiment, the global transaction identifier 302 is unique for all pending requests.

[0025] In accordance with one implementation, the local transaction identifier field 308 is a field generated by a requesting agent and can be unique for all pending requests requiring completion by that requesting agent. Furthermore, in this example, the source identifier 310 uniquely identifies the requesting agent within a wiring hierarchy. Accordingly, the local transaction identifier field 308, together with the source ID 310, provides a global identification of a transaction within a hierarchy domain.

[0026] Attribute field 304 specifies properties and relationships of the transaction. In this regard, attribute field 304 is potentially used to provide additional information that allows for modification of the default transaction handling. In one embodiment, attribute field 304 contains a priority field 312, a reserved field 314, an ordering field 316, and a do-not-listen field 318. The priority subfield 312 can be modified by an initiator to assign a priority to the transaction. The reserved attribute field 314 is left reserved for future use or for a provider-defined usage. Using the reserved attribute field, possible usage models can be implemented using priority or security attributes.

[0027] The order attribute field 316 is used in this example to provide optional information that communicates the type of order, which can modify the default ordering rules. In accordance with an example implementation, an order attribute of "0" indicates that default ordering rules should be applied, while an order attribute of "1" indicates a more relaxed order in which writes can skip writes in the same direction and read completions can skip writes in the same direction. The listen-in attribute field 318 is used to determine whether transactions are listened in. As shown, the channel ID field 306 identifies a channel to which a transaction is associated.

[0028] Returning to the discussion from Fig.2. A data link layer 210, also referred to as data link layer 210, can act as an intermediate layer between the transaction layer 205 and the physical layer 220. In one embodiment, the data link layer 210 is responsible for providing a reliable mechanism for exchanging transaction layer packets (TLPs) between two components on a connection. One side of the data link layer 210 accepts TLPs concatenated by the transaction layer 205, applies a packet sequence identifier 211, i.e., an identification number or packet number, calculates and applies an error detection code, i.e., CRC 212, and passes the modified TLPs to the physical layer 220 for transmission via a physical device to an external device.

[0029] In one example, the physical layer 220 contains a logical subblock 221 and an electrical subblock 222 to physically send a packet to an external device. The logical subblock 221 is responsible for the "digital" functions of the physical layer 221. In this respect, the logical subblock can contain a send section for preparing outgoing information for transmission by the physical subblock 222 and a receive section for identifying and preparing received information before it is passed to the data link layer 210.

[0030] The physical block 222 contains a transmitter and a receiver. The logical subblock 221 supplies symbols to the transmitter, which it serializes and sends to an external device. The receiver receives serialized symbols from an external device and transforms the received signals into a bitstream. The bitstream is deserialized and supplied to the logical subblock 221. In one exemplary embodiment, an 8b / 10b transmit code is used, with ten-bit symbols being sent / received. Here, special symbols are used to frame a packet with frame 223. In one example, the receiver also provides a symbol clock, which is recovered from the incoming serial stream.

[0031] Although the transaction layer (205), data link layer (210), and physical layer (220) are discussed in relation to a specific implementation of a protocol stack (such as a PCIe protocol stack), a layered protocol stack as mentioned above is not limited to this. In fact, any layered protocol stack can be included / implemented and adopt the features discussed here. As an example, a port / interface represented as a layered protocol may include: (1) a first layer for assembling packets, i.e., a transaction layer; (2) a second layer for sequencing packets, i.e., a data link layer; and (3) a third layer for transmitting the packets, i.e., a physical layer. A high-performance wiring layered protocol, as described here, is used as a specific example.

[0032] The following is based on Fig.Figure 4 shows an exemplary embodiment of a serial point-to-point fabric. A serial point-to-point link can include any transmission path for transmitting serial data. In the embodiment shown, a link can include two differentially driven low-voltage signal pairs: a transmit pair 406 / 411 and a receive pair 412 / 407. Accordingly, the device 405 includes transmit logic 406 for sending data to the device 410 and receive logic 407 for receiving data from the device 410. In other words, one implementation of a link includes two transmit paths, i.e., paths 416 and 417, and two receive paths, i.e., paths 418 and 419.

[0033] A transmission path refers to any route used to send data, such as a transmission line, copper wire, optical line, wireless communication channel, infrared communication link, or other communication path. A connection between two devices, such as a Device 405 and a Device 410, such as Link 415, is called a link. A link can support one track—each track representing a set of different signal pairs (one pair for transmitting, one pair for receiving). To scale bandwidth, a link can group multiple tracks, denoted by xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.

[0034] A differential pair can refer to two transmission paths, such as lines 416 and 417, for transmitting differential signals. For example, line 417 switches from a high logic level to a low logic level (i.e., a falling edge), while line 416 switches from a low voltage level to a high voltage level (i.e., a rising edge). Differential signals exhibit, among other advantages, potentially better electrical characteristics such as improved signal integrity (i.e., reduced cross-coupling, voltage overshoot / undershoot, and damped oscillation). This allows for a better timing window, enabling higher transmission frequencies.

[0035] In one embodiment, a novel high-performance wiring (HPI) is provided. The HPI can incorporate next-generation cache-coherent, connection-based wiring. As an example, the HPI can be used in high-performance computing platforms such as workstations or servers, including systems where PCIe or another wiring protocol is commonly used to connect processors, accelerators, I / O devices, and the like. However, the HPI is not limited to these applications. Instead, the HPI can be used in any of the systems or platforms described herein. Furthermore, the individual concepts developed can be applied to other wiring protocols and platforms such as PCIe, MI-PI, QPI, and so on.

[0036] To support multiple devices, the HPI in an exemplary implementation can be instruction set architecture agnostic (ISA agnostic) (i.e., the HPI can be implemented in several different devices). In another scenario, the HPI can also be used to connect high-performance I / O devices, not just processors or accelerators. For example, a high-performance PCIe device can be coupled to an HPI via a suitable translation bridge (i.e., HPI to PCIe). Furthermore, the HPI connections can be used by many HPI-based devices, such as processors, in various configurations (e.g., star, ring, mesh, etc.). Fig.Figure 5 illustrates exemplary implementations of several potential multi-socket configurations. A two-socket configuration 505, as shown, can contain two HPI connections; however, other implementations may use a single HPI connection. For larger topologies, any configuration can be used as long as an identifier (ID) can be assigned and as long as some form of virtual path exists among other additional or substitute features. As shown, in one example, a four-socket configuration 510 has an HPI connection from one processor to another. In contrast, in the eight-socket implementation shown in configuration 515, not every socket is directly connected to every other via an HPI connection. However, the configuration is supported if a virtual path or channel exists between the processors. A range of supported processors in a native domain includes 2-32.Higher numbers of processors can be achieved, among other examples, by using multiple domains or other wiring between node controllers.

[0037] The HPI architecture includes a definition of a layered protocol architecture, which in some examples contains protocol layers (coherent, incoherent, and optionally other memory-based protocols), a routing layer, a data link layer, and a physical layer. Furthermore, the HPI may include extensions relating to power managers (such as power control units (PCUs)), design for testing and verification (DFT), fault handling, registers, and security, among other examples. Fig. Figure 5 illustrates an embodiment of an exemplary HPI layer protocol stack. In some implementations, at least some of the features shown in Fig.The five layers shown are optional. Each layer handles its own granularity or information quantum level (the protocol layer 605a,b with packets 630, the data link layer 610a,b with flits 635, and the physical layer 605a,b with phits 640). It is noted that in some embodiments, based on the implementation, a packet may contain partial flits, a single flit, or multiple flits.

[0038] As a first example, a Phit 640 contains a 1:1 mapping of a link width to bits (e.g., a link width of 20 bits contains a Phit of 20 bits, etc.). Flits can be of a larger size, such as 184, 192, or 200 bits. It is noted that a fractional number of Phits 640 are required to transmit a Flit 635 if the Phit 640 is 20 bits wide and the size of the Flit 635 is 184 bits (e.g., among other examples, 9.2 Phits at 20 bits to transmit a 184-bit Flit 635, or 9.6 at 20 bits to transmit a 192-bit Flit). It is noted that the widths of the underlying link in the physical layer can vary. For example, the number of tracks per direction can be 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, etc. In one embodiment, the data link layer 610a,b can embed multiple pieces of different transactions into a single flit and can contain one or more initial blocks (e.g., 1, 2, 3, 4, 5, 6, 6, 7, 8, 9, 10 ...B. 1, 2, 3, 4) are embedded in the flit. In one example, the HPI separates the initial blocks into corresponding slots to allow multiple messages in the flit to be destined for different nodes.

[0039] In one embodiment, the physical layer 605a,b can be responsible for the rapid transmission of information in the physical medium (electrical or optical, etc.). The physical connection between two data link layer entities, such as between layer 605a and layer 605b, can be point-to-point. The data link layer 610a,b can abstract the physical layer 605a,b from the upper layers and provides the capability to reliably transmit data (and requests) and manage flow control between two directly connected entities. Furthermore, the data link layer can be responsible for virtualizing the physical channel into multiple virtual channels and message classes.The protocol layer 620a,b relies on the data link layer 610a,b to map protocol messages to appropriate message classes and virtual channels before passing them to the physical layer 605a,b for transmission over the physical links. The data link layer 610a,b can support multiple messages, such as request, eavesdrop, response, write-back, and incoherent data, among other examples.

[0040] As in Fig.As shown in Figure 6, the physical layer 605a,b (or PHY) of the HPI can be implemented above the electrical layer (i.e., the electrical conductors connecting two components) and below the data link layer 610a,b. The physical layer and its associated logic can reside in each agent and connect the data link layers in two separate agents (A and B) (e.g., in devices on either side of a connection). The local and remote electrical layers are connected by physical media (e.g., wires, conductors, optical, etc.). In one embodiment, the physical layer 605a,b has two main phases: initialization and operation. During initialization, the link is opaque to the data link layer, and signaling can include a combination of timed states and acknowledgment exchange events.During operation, the link is permeable to the data link layer, and signaling occurs at a single rate, with all traces working together as a single connection. During the operational phase, the physical layer transports flits from agent A to agent B and from agent B to agent A. The link is also referred to as a connection and abstracts some physical aspects, including media, width, and speed, from the data link layers, while exchanging flits and control / status information (e.g., the width) with the data link layer. The initialization phase includes minor phases, such as queries and configuration. The operational phase also includes minor phases (e.g., link power management states).

[0041] In one embodiment, the data link layer 610a,b can be implemented to provide reliable data transmission between two protocols or routing entities. The data link layer can abstract the physical layer 605a,b from the protocol layer 620a,b and can be responsible for flow control between two protocol agents (A, B), providing virtual channel services for the protocol layer (message classes) and the routing layer (virtual networks). The interface between the protocol layer 620a,b and the data link layer 610a,b can typically be at the packet level. In one embodiment, the smallest transmission unit of the data link layer is called a flit, which is a specified number of bits, such as 192 bits or some other nominal value.To frame the transmission unit (phit) of physical layer 605a,b within the transmission unit (flit) of data link layer 610a,b, data link layer 610a,b relies on physical layer 605a,b. Furthermore, data link layer 610a,b can be logically decomposed into two parts: a sender and a receiver. A sender / receiver pair in one entity can be connected to a receiver / sender pair in another entity. Flow control is often performed on both a flit and a packet basis. Potentially, error detection and correction are also performed on a flit-level basis.

[0042] In one embodiment, the routing layer 615a,b can provide a flexible and distributed method for routing HPI transactions from a source to a destination. Since routing algorithms for multiple topologies can be specified by programmable routing tables at each router (where, in one embodiment, the programming is performed by firmware, software, or a combination thereof), the scheme is flexible. The routing functionality can be distributed; routing can be performed through a series of routing steps, each routing step being defined by looking up a table at either the source router, an intermediate router, or the destination router. Looking up a table at a source can be used to inject an HPI packet into the HPI fabric. Looking up a table at an intermediate router can be used to route an HPI packet from an input port to an output port.Looking up a destination port can be used for targeting the HPI protocol agent. It is noted that the routing layer may be thin in some implementations because the routing tables, and therefore the routing algorithms, are not specifically defined by the specification. This allows for flexibility and enables a variety of usage models, including flexible platform architecture topologies, to be defined through system implementation. Routing layer 615a,b relies on data link layer 610a,b to provide the use of up to three (or more) virtual networks (VNs)—in one example, two deadlock-free VNs, VN0 and VN1, with multiple message classes defined in each virtual network.In the data link layer, a shared adaptive virtual network (VNA) can be defined, but this adaptive network is not directly exposed in routing concepts, since each message class and each virtual network can have dedicated resources and guaranteed forward progress under other characteristics and examples.

[0043] In some implementations, the HPI can utilize an embedded clock. A clock signal can be embedded within data transmitted via wiring. When the clock signal is embedded in the data, several dedicated clock tracks can be omitted. This can be useful, for example, in systems where pin space is limited, as it allows more pins to be dedicated to data transmission.

[0044] A connection can be established between two agents on either side of a wiring connection. An agent that sends data can be a local agent, and the agent that receives the data can be a remote agent. Both agents can use state machines to manage various aspects of the connection. In one embodiment, the physical layer data path can send FLITS from the data link layer to the electrical front end. In another implementation, the control path includes a state machine (also called a link-training state machine or similar). The actions and exits of the state machine's states can depend on internal signals, timers, external signals, or other information. In fact, some of the states, such as a few initialization states, can have timers to provide a timeout value for exiting a state.It is noted that in some embodiments, detection refers to the detection of an event on both branches of a track, though not necessarily simultaneously. However, in other embodiments, detection refers to the detection of an event by a reference agent. As an example, debouncing refers to the sustained declaration of a signal. In one embodiment, HPI supports operation in the case of non-functional tracks. In specific states, tracks can be dropped.

[0045] States defined in the state machine can include, among other categories and subcategories, reset states, initialization states, and operating states. For example, some initialization states may have a secondary timer used to exit the state when a certain time expires (essentially an abort, as no progress occurs within the state). An abort may involve updating registers such as the status register. Some states may also have one or more primary timers used to time the primary functions within the state. Other states may be defined such that, among other examples, internal or external signals (such as acknowledgment exchange protocols) trigger the transition from one state to another.

[0046] A state machine can also support single-step testing, freezing upon aborted initialization, and the use of testers. State exits can be deferred / held until the testing software is ready. In one instance, the exit can be deferred / held until the secondary timing expires. In one embodiment, actions and exits can be based on the exchange of training sequences. In another embodiment, the link state machine is executed in the clock domain of the local agent, and the transition from one state to the next occurs such that it coincides with a sender training sequence boundary. Status registers can be used to reflect the current state.

[0047] Fig.Figure 7 illustrates a representation of at least one section of a state machine used by agents in an exemplary implementation of the HPI. It should be noted that the state table consists of Fig. The seven states included in the diagram represent a non-exhaustive list of possible states. For example, some transitions have been omitted to simplify the diagram. Furthermore, some states can be combined, split, or omitted, while others could be added. These states can include the following: Event reset state: This state is entered during a hot or cold reset event. It restores default values. It initializes counters (e.g., synchronous counters). It can exit to a different state, such as another reset state. Timed Reset State: A timed state for in-band reset. It can target a predefined electrically ordered set (EOS) so that remote receivers can detect the EOS and also enter the timed reset. The receiver has traces that hold electrical settings. It can exit to an agent to calibrate the reset state. Calibration reset state: Calibration without signaling on the track (e.g., receiver calibration state) or driver shutdown. It can be a predetermined duration in the state based on a timer. It can set an operating speed. If a port is not enabled, it can act as a wait state. It can include a minimum dwell time. Based on the design, receiver conditioning or staggering can occur. After a time elapsed and / or after completion of a calibration, it can exit a receiver detection state. Receiver Detection State: Detects the presence of a receiver on one or more tracks. It can search for a receiver termination (e.g., after a receiver abort insertion). If a specified value is set, or if another specified value is not set, it can exit to a calibration reset state. If a receiver is detected, or if a timeout is reached, it can exit to the transmitter calibration state. Transmitter calibration state: for transmitter calibrations. It can be a time-based state assigned for transmitter calibrations. It can contain a signal on a track. It can continuously drive an EOS such as an EIEOS. If a calibration is performed or a timer expires, it can exit to a compatibility state. If a counter has expired or a secondary timeout has occurred, it can exit to a transmitter detection state. Sender detection state: Denotes a valid signal. It can be an acknowledgment exchange state, where an agent, based on a signal from a remote agent, completes actions and exits to the next state. The receiver can denote a valid signal from a sender. In one embodiment, the receiver looks for a wake-up detection, searching for it on other tracks if it has been debounced on one or more tracks. The sender triggers a detection signal. In response to the completion of debouncing for all tracks and / or a timeout, or if debouncing is not complete on all tracks and there is a timeout, it can exit to a query state. One or more monitoring tracks can be kept awake. To debounce a wake-up signal, the other tracks can potentially be debounced as well.This can enable energy savings in low-power states. Query state: The receiver adapts, initializes the drift buffer, and locks onto bits / bytes (identifying, for example, symbol boundaries). The tracks can be skew-compensated. A remote agent can, in response to an acknowledgment message, initiate an exit to the next state (e.g., a link width state). The query can additionally include a training sequence lock by locking onto an EOS and a training sequence start block. Track-to-track skew at the remote sender can be limited to a first length for high speed and a second length for low speed. Skew compensation can be performed in both slow and operating modes. The receiver can have a specific maximum skew-to-track compensation limit, such as 8, 16, or 32 skew intervals.The receiver actions can include latency correction. In one embodiment, the receiver actions can be completed upon successful skew compensation of a valid track card. For example, a successful acknowledgment exchange can be achieved by receiving a number of consecutive training sequence start blocks with acknowledgments and by sending a number of training sequences with an acknowledgment after the receiver has completed its actions. Link Width State: The agent communicates with the last track card to the remote sender. The receiver receives and decodes the information. The receiver can record a configured track card in one structure after the checkpoint of a previous track card value in a second structure. The receiver can also respond with an acknowledgment ("ACK"). It can initiate an in-band reset. As an example, it is a first state for initiating the in-band reset. In one embodiment, exit to the next state, such as a flit configuration state, is performed in response to the ACK. Furthermore, a reset signal can also be generated before entering a low-power state if the frequency with which a wake-up detection signal occurs falls below a specified value (e.g., 1 for every number of unit intervals (UIs), such as 4 kUI).The receiver can hold current and previous track maps. The transmitter can use different groups of tracks based on training sequences with varying values. In some embodiments, the track map cannot modify certain status registers. Flit-lock configuration state: This state is entered by a sender, but is considered exited (i.e., secondary timing is questionable) when both the sender and receiver have exited to a link lock state or another link state. In one embodiment, the sender exit to a link state includes the beginning of a data sequence boundary (SDS boundary) and a training sequence boundary (TS boundary) after receiving a planetary alignment signal. The receiver exit can be based on receiving an SDS from a remote sender. This state can be a bridge from an agent to a link state. The receiver identifies the SDS. If an SDS is received after a dewritter has been initialized, the receiver can exit to a link lock state (BLS) (or into a control window).If a timeout occurs, the output can be a reset state. The transmitter controls the tracks with a configuration signal. Based on conditions or timeouts, the transmitter output can be a reset, a BLS (Battery Life System), or other states. Link Transfer State: A link state. Flits are sent to a remote agent. It can be entered from a link lock state and can revert to a link lock state upon an event such as a timeout. The sender transmits flits. The receiver receives flits. It can also exit a low-power link state. In some implementations, the Link Transfer State (TLS) may be referred to as the L0 state. Link lock state: a link state. The sender and receiver operate in a unified manner. It can be a time-limited state during which data link layer flits are withheld while physical layer information is transmitted to the remote agent. It can exit to a low-power link state (or another link state based on the design). In one embodiment, a link lock state (BLS) occurs periodically. The period is called a BLS interval and can be time-limited and may differ between low speed and operating speed.It is noted that the data link layer can be periodically blocked from sending flits, such as during a link-transmit state or a partial-width link-transmit state, so that a physical layer control sequence of length can be sent. In some implementations, the link-lock state (LLS) may be referred to as an L0 control state or an L0c state. Partial-width link transmission state: Link state. It can save power by entering a partial-width state. In one embodiment, the asymmetric partial width refers to each direction of a bidirectional link with different widths, which can be supported in some designs. An example of an initiator, such as a transmitter, that sends a partial-width indication to enter a partial-width link transmission state is shown in the example from Fig.Figure 9 shows that a partial width is sent while the connection is being transmitted on a link with a first width, so that the link switches to transmitting at a second, new width. A mismatch can lead to a reset. It is noted that speeds cannot be changed, but the width can be changed. Thus, potentially, flits are sent at different widths. This can be logically similar to a link transmission state; however, it may take longer to transmit flits because of the smaller width. It can exit to other link states, such as a low-power link state, based on certain received and sent messages, or from a partial width link transmission state, or from a link lock state based on other events.In one embodiment, a transmitter port, as shown in the timing diagram, can stagger the phasing of unused tracks to provide improved signal integrity (e.g., noise reduction). During periods when the link width is changing, non-repeatable flits, such as null flits, can be used. In one or more structures, a corresponding receiver can drop these null flits and stagger the phasing of unused tracks, recording the current and previous track map. It is noted that the status and associated status register can remain unchanged. In some implementations, the partial-width link transmission state may be referred to as the partial L0 state or L0p state. Exit from partial-width link transmission state: Exiting the partial-width state. In some implementations, this may or may not be a link lock state. In one embodiment, the transmitter initiates an exit by sending partial-width exit patterns on the unused lanes to train them and compensate for their skew. As an example, an exit pattern begins with an EIEOS, which is detected and debounced to signal that the lane is ready to begin entering a full-width link transmission state, and may end with an SDS or a fast-training sequence (FTS) on unused lanes.Any disturbance during the exit sequence (receiver actions such as misalignment compensation not completed before the timer expires) halts Flit transmissions to the data link layer and triggers a reset, which is handled by resetting the link the next time the link lock state occurs. Additionally, the SDS can initialize the shuffler / unshuffler on the tracks to appropriate values. Low-power link state: This is a lower-power state. In one embodiment, it is lower power than the partial-width link state because signaling is stopped on all lanes and in both directions. Transmitters can use a link-lock state to request a low-power link state. The receiver can decode the request and respond with an ACK or a NAK; otherwise, a reset can be triggered. In some implementations, the low-power link state may be referred to as an L1 state.

[0048] In one embodiment, two types of pin resets can be supported: power-on reset (or "cold" reset) and warm reset. A reset initiated by software or originating from one agent (in the physical layer or another layer) can be transmitted in-band to the other agent. However, due to the use of an embedded clock, an in-band reset can be handled by communicating with another agent using an ordered set, such as a specific electrical ordered set or EIEOS, as introduced above. Such ordered sets can be implemented, among other examples, as defined 16-byte codes that can be represented in hexadecimal format. The ordered set can be sent during initialization, and a PHY control sequence (or a "link lock state") can be sent after initialization.The link lock state can prevent the data link layer from sending flits. As another example, data link layer traffic can be blocked from sending new NULL flits, which can be discarded at the receiver.

[0049] In some HPI implementations, supersequences can be defined, where each supersequence corresponds to a specific state or the entry into / exit from a specific state. A supersequence can contain a repeating sequence of data sets and symbols. Among other examples, the sequences can repeat until the completion of a state or state transition, or until the transmission of a corresponding event. In some cases, the repeating sequence of a supersequence can repeat at a defined frequency, such as a defined number of unit intervals (UIs). A unit interval (UI) can correspond to the time interval for transmitting a single bit on a track of a connection or system. In some implementations, the repeating sequence can begin with an Electrically Ordered Set (EOS).Accordingly, an instance of the EOS can be expected to repeat itself according to a predefined frequency. Such ordered sets can be implemented as defined 16-byte codes, which, among other examples, can be represented in hexadecimal format. In one example, the EOS of a supersequence can be an Electrically Ordered Electrical Idle Ordered Set (EIEIOS). In another example, an EIEOS can resemble a low-frequency clock signal (e.g., a predefined number of repeating FF00 or FFF000 hexadecimal symbols, etc.). The EOS can be followed by a predefined set of data, such as a predefined number of training sequences or other data. Among other examples, such supersequences can be used in state transitions, including link state transitions, as well as in initialization.

[0050] As introduced above, in one embodiment, initialization can initially occur at low speed, followed by high-speed initialization. The low-speed initialization uses default values ​​for the registers and timings. Software then uses the low-speed connection to set up the registers, timings, and electrical parameters, and clears the calibration semaphore to prepare for the high-speed initialization. As an example, the initialization could consist of, among potentially other, such states or tasks as resetting, detection, querying, and configuration.

[0051] In one example, a link-layer lock control sequence (i.e., a link lock state (BLS) or L0c state) can contain a timed state during which link-layer flits are withheld while PHY information is transmitted to the remote agent. The sender and receiver can initiate a control sequence lock timer here. Additionally, when the timer expires, the sender and receiver can exit the lock state and take other actions, such as exiting to reset, exiting to another link state (or state at all), including states that allow flits to be sent over the link.

[0052] In one embodiment, link training may be provided and may involve sending one or more jumbled training sequences, ordered sets, and control sequences, such as in conjunction with a defined supersequence. A training sequence symbol may contain an initial block and / or reserved sections and / or a target latency and / or a pair number and / or code reference tracks or a group of code reference tracks from a map of physical tracks and / or an initialization state. In one embodiment, the initial block may be sent with an ACK or a NAK, among other examples. As an example, training sequences may be sent as part of supersequences and may be jumbled.

[0053] In one embodiment, ordered sets and control sequences are not scrambled or staggered and are sent equally, simultaneously, and completely on all tracks. A valid reception of an ordered set may include checking at least one section of the ordered set (or the entire ordered set for ordered subsets). Ordered sets may contain an Electrically Ordered Set (EOS), such as an Electrical Idle Ordered Set (EIOS) or an EIEOS. A supersequence may contain the beginning of a Data Sequence (SDS) or a Fast Training Sequence (FTS). Such sets and control supersequences may be predefined and may have any pattern or hexadecimal representation and any length. For example, ordered sets and supersequences may be 8 bytes, 16 bytes, or 32 bytes long, etc.As an example, the AGV can also be used for fast bit locking during exit from a partial-width link transmission state. It is noted that the AGV definition can be per lane and can utilize a rotated version of the AGV.

[0054] In one embodiment, supersequences can include the insertion of an EOS, such as an EIEOS, into the training sequence stream. When signaling begins, tracks in one implementation are switched on in a staggered manner. However, this can result in initial supersequences being seen truncated on some tracks at the receiver. Supersequences can be repeated over short intervals (e.g., approximately 1,000-unit intervals (or 1 kUI)). Furthermore, the training supersequences can be used for skew compensation, configuration, and / or the transmission of an initialization target, track map, etc. The EIEOS can be used, among other examples, for transitioning a track from the inactive to the active state, for selecting good tracks, and / or for identifying symbol and TS boundaries.

[0055] Transition to Fig.Figure 8 shows illustrations of exemplary supersequences. For example, an exemplary detection supersequence 805 can be defined. The detection supersequence 805 can contain a repeating sequence of a single EIEOS (or another EOS), followed by a predefined number of instances of a specific training sequence (TS). In one example, the EIEOS can be sent, immediately followed by seven repeated instances of the TS. When the last of the seven TS has been sent, the EIEOS can be sent again, followed by seven additional instances of the TS, and so on. The sequence can be repeated according to a specific predefined frequency. In the example from Fig.8. The EIEOS can reappear on the tracks approximately once every thousand UI (□1 kUI), followed by the remainder of the detection supersequence 805. A receiver can monitor tracks for the presence of a repeating detection supersequence 805 and, upon validating the supersequence 705, infer that a distant agent is present, has been added to the tracks (e.g., by hot-plugging), has woken up, or has been reinitialized, etc.

[0056] In another example, a different supersequence 810 can be defined to specify a query, configuration, or loopback condition or state. As with the exemplary detection supersequence 805, traces of a connection can be monitored by a receiver for such a query / configuration / loop supersequence 810 to identify a query state, configuration state, or loopback condition. In one example, a query / configuration / loopback supersequence 810 can begin with an EIEOS followed by a predetermined number of repeated instances of a TS. For example, the EIEOS in one example can be followed by thirty-one (31) instances of a TS, with the EIEOS repeating approximately every four thousand UI (e.g., 4kUI).

[0057] Furthermore, in another example, a partial-width transmit-state exit supersequence (PWTS exit supersequence) 815 can be defined. In one example, a PWTS exit supersequence can contain an initial EIEOS to pre-prepare tracks before transmitting the first full sequence in the supersequence. For example, the sequence to be repeated in supersequence 815 can begin with an EIEOS (to be repeated approximately once per 1 kUI). Additionally, instead of other training sequences (TS), fast training sequences (FTS) can be used, with the FTS configured to assist with faster bit locking, byte locking, and skew compensation. In some implementations, an FTS can be de-dicated to further assist in reactivating unused tracks as quickly and seamlessly as possible.As with other supersequences that precede entry into a link transfer state, supersequence 815 can be interrupted and terminated by sending a data sequence start (SDS). Furthermore, a partial FTS (FTSp) can be sent to assist in synchronizing the new tracks with the active tracks, for example, by allowing bits to be removed from (or added to) the FTSp.

[0058] Potentially, supersequences such as a detection supersequence 705 and a query / configuration / loop supersequence 710, etc., can be sent essentially throughout the entire initialization or reinitialization of a connection. In some cases, upon receiving and detecting a particular supersequence, a receiver can respond by identically relaying the same supersequence back to the sender via the traces. The reception and validation of a particular supersequence by a sender and receiver can serve as an acknowledgment exchange to confirm a state or condition transmitted via the supersequence. For example, such an acknowledgment exchange (e.g., using a detection supersequence 705) can be used to identify a connection reinitialization.In another example, such an acknowledgment exchange can be used, among other examples, to indicate the end of an electrical reset state or low-power state, which causes the corresponding tracks to be restarted. The end of the electrical reset can be identified, for example, by an acknowledgment exchange between a sender and a receiver, each transmitting a detection supersequence 705.

[0059] In another example, tracks can be monitored for supersequences and can utilize these supersequences in conjunction with track screening for detection and wake-up state exits and entries, among other events. Furthermore, the predefined and predictable nature and form of supersequences can be used to perform initialization tasks such as bit locking, byte locking, debouncing, drafting, skew compensation, adaptation, latency correction, tuned delays, and other potential uses. In fact, tracks can be monitored essentially continuously for such events to accelerate the system's ability to respond to and process these conditions. In the case of debouncing, transients can be introduced onto tracks as a result of a variety of conditions.For example, adding or turning on a device can introduce transients onto the track. Additionally, voltage irregularities can be represented on a track due to poor track quality or electrical interference. Such irregularities can be easily detected on supersequences with predictable values, such as when EIEOS values ​​unexpectedly deviate in conjunction with transients or other bit errors.

[0060] In one example, a transmitting device might attempt to enter a specific state. For instance, the transmitting device might attempt to activate the link and enter an initialization state. In another example, among others, the transmitting device might attempt to exit a low-power state, such as an L1 state. In some cases, an L1 state might serve as a power-saving, idle, or standby state. In fact, in some examples, main power supplies might remain active in the L1 state. Upon exiting an L1 state, a first device might send a supersequence associated with the transition from the L1 state to a specific other state, such as an L0 link transfer state (L0-TLS).As in other examples, the supersequence can be a repeating sequence of an EOS followed by a predetermined number of TS, such that the EOS is repeated at a specific, predefined frequency. A receiving device can receive and validate the data, identify the supersequence, and complete the acknowledgment exchange with the transmitting device by sending the supersequence back to the transmitting device.

[0061] If both the transmitting and receiving devices receive the same supersequence, each device can perform additional initialization tasks that utilize the supersequences. For example, each device can perform debouncing, bit locking, byte locking, drafting, and skew compensation using the supersequences. Additional initialization information can be transmitted via the start blocks and payload of the Transaction Signatures (TS) contained in the supersequences. When the link is initialized, in some cases a Data Send Start Sequence (SDS sequence) can be sent that interrupts the supersequence (for example, by sending it in the middle of a TS or EIEOS), allowing the respective devices on both sides of the link to prepare for synchronized entry into the TLS.In TLS or in an "S0" state, supersequences can be terminated and flits can be transmitted using the data link layer of the protocol stack.

[0062] Within the TLS, further limited opportunities can be provided for the physical layer to perform control tasks. For example, during an L0 state, bit errors and other errors on one or more tracks can be identified. In one implementation, a control state L0c can be provided. The L0c state can be provided as a periodic window within the TLS to allow physical layer control messages to be sent between streams of FLITS transmitted over the data link layer. As in the Fig.As illustrated in the 9th example, an L0 state can be divided into L0c intervals. Each L0c interval can begin with an L0c state or L0c window (e.g., 905) in which physical layer control codes and other data can be sent. The remainder (e.g., 910) of the L0c interval can be dedicated to sending flits. The length of the L0c interval and the L0c state within each interval can be defined by programming, for example, by the BIOS of one or more devices or by another software-based controller. The L0c state can be exponentially shorter than the remainder of an L0c interval. For example, in one example, the L0c might be 8 UI, while the remainder of the L0c interval is on the order of 4 kUI.This can enable windows in which relatively short, predefined messages can be sent without significantly interrupting or wasting connection data bandwidth.

[0063] An L0c state message can convey a variety of conditions at the physical layer. For example, a device might initiate a link or track reset based on bit errors or other faults exceeding a certain threshold. Such faults can also be conveyed in L0c windows (such as previous L0c windows). Furthermore, the L0c state can be effectively used to implement other in-band signaling, such as signaling to support or trigger transitions between other link states. For example, L0c messages can be used to transition a link from an active L0 state to a standby or low-power state, such as an L1 state. As shown in the simplified flowchart from Fig.As shown in Figure 10, a specific L0c state can be used to transmit an L1 entry request (e.g., 1010). While the device (or an agent within the device) waits for an acknowledgment of request 1010, further flits (e.g., 1020, 1030) can be sent. The other device on the link can then send the acknowledgment (e.g., 1040). In some examples, the acknowledgment can also be sent within an L0c window. In some cases, the acknowledgment can be sent in the next L0c window, followed by the receipt / transmission of the L1 request 1010. Timers can be used to synchronize the L0c intervals in each device, and the requesting device can recognize acknowledgment 1040, among other examples, based on an identification that acknowledgment 1040 was sent in the next L0c window, as an acknowledgment of request 1010 (e.g.,instead of an independent L1 entry request). In some cases, an acknowledgment may be transmitted via an L0c code different from the one used in L1 entry request 1010. In other cases, acknowledgment 1040, among other examples, may contain the identical reproduction of the L1 entry request code used in request 1010. Furthermore, in alternative examples, a non-acknowledgment signal or NAK may be transmitted in the L0c window.

[0064] In addition to (or as an alternative to) acknowledgment exchange using L0c codes, supersequences, such as a detection supersequence, can be sent in conjunction with resetting and reinitializing the link. While the supersequences are sent by a first device and identically retransmitted by the second, receiving device, further acknowledgment exchange can occur between the devices. As described above, supersequences can be used to assist in reinitializing the link, including debouncing, bit locking, byte locking, drafting, and skew compensation of the link traces. Furthermore, the devices can use the timer (which, for example, embodies the L0c interval) to synchronize the devices and the link entering the requested L1 state.For example, the receipt of acknowledgment 1040 for the devices can, among other examples, indicate that they should both enter (or begin to enter) the L1 state at the end of the L0c interval corresponding to the L0c window in which the acknowledgment was sent. For example, the data sent in an L0c window included in or otherwise associated with acknowledgment 1040 can, among other potential examples, specify the time at which the devices should enter the L1 state. In some cases, additional flits (e.g., 1050) can be sent while the devices wait for the time elapsed corresponding to the transition to the L1 state.

[0065] In some HPI implementations, connections can be established on any number of two or more tracks. Furthermore, a connection can be initialized with an initial number of tracks and later transition to a partial-width state, so that only a portion of the tracks are used. The partial-width state can be defined as a lower-performance state, such as an L0p state. For example, an L0c state can be used to transition from an L0 state, where the initial number of tracks are active, to an L0p state, where a smaller number of tracks are active. As in the example from Fig.As shown in Figure 11, a connection with an initial width of 1110 can be active. In some cases, the initial width can be the full width (e.g., at L0). In other cases, the connection can transition from an initial L0p state using a first number of tracks to another L0p state using a different number (or quantity) of tracks. During an L0c window of the tracks with the initial width, an L0p entry code 1120 can be sent. The L0p entry request 1120 can identify which new width should be applied. In some cases, the new connection width can be predetermined and easily identified from receiving the L0p request 1120. Furthermore, in conjunction with the L0p request 1120, the specific tracks to be dropped in the partial width state can be specified, identified, or preconfigured in some other way.

[0066] Continuing with the example from Fig. 11. Flits or other data can continue to be transmitted across the full width of the tracks while the link waits to transition to the L0p state. For example, synchronized timers on the devices connected via the link can specify a duration t to synchronize entry into the L0p state. In one example, the duration t might correspond to a remainder of an L0c interval corresponding to request 1120. At the end of the interval, some tracks remain active, while others enter the inactive or idle state. The link then operates at the new width (e.g., 1140) at least until an L0p exit request or another link width transition request is received.

[0067] The HPI can utilize one or more power control units (PCUs) to assist in timing transitions between an L0 state and lower-power states such as L0p and L1. Furthermore, the HPI can support master-slave, master-master, and other architectures. For example, a PCU may be present on only one of the devices connected via a link or otherwise associated with it, with the device containing the PCU being considered the master. Master-master configurations can be implemented, for instance, when both devices are associated with a PCU that can initiate a link state transition. Some implementations may be configured for a specific low-power state, such as...Among other examples, L0p or L1 specify a minimum duration to try to minimize transitions between states and to try to maximize power savings within a low-power state that has been entered.

[0068] Exiting from a low-power section width state can be adjusted to be efficient and fast, minimizing the impact and interruption of active lanes. In some implementations, L0c windows and L0c codes can also be used to trigger an exit from an L0p state or another state to reactivate unused lanes. For example, moving on to the examples from Fig. Figure 12 shows a simplified flowchart illustrating an exemplary exit from an L0p state. In the specific example from Fig.Flits (e.g., 1205) can be sent when an L0c window 1210 is detected containing an L0 entry request (or an L0p exit request). Additional flits 1215 can be sent before the point at which the L0p exit is to occur. As in other examples, an L0c code 1210 can identify a time at which a state transition is to begin / end, as well as specific events of the state transition, or implicitly identify them. To maximize data transmission while the devices wait to enter the state transition, further flits (e.g., 1215) can be sent.

[0069] In one example, an EIEOS 1220 (or other data such as another EOS) can be sent on the inactive tracks to begin track reconditioning. In some cases, such inactive tracks (e.g., tracks "n+1" through "z") may have been inactive for a period of time, and waking the tracks can introduce electrical transients and other instability. Accordingly, the EIEOS 1220, as well as partial-width supersequences sent in conjunction with exiting the L0p state, can be used to debounce the tracks as they wake up. Furthermore, in some cases, transients on the waking tracks (e.g., on tracks "n+1" through "z") can potentially affect the active tracks (e.g., tracks "0" through "n").To prevent irregularities originating from the reawakening of unused tracks from negatively affecting the active tracks, the active tracks can be synchronized to send null flits at or immediately before the initial signals (e.g., 1220) sent via the awakening tracks (e.g., at 1225).

[0070] In some implementations, the reinitialization of unused tracks can be scheduled to begin at a time similar to the completion of a corresponding L0c interval. In other cases, an alternative time can be used to initiate the reinitialization early. In such cases, a sender of the L0p exit request can cause the used tracks to be pre-processed, for example, by sending one or more individual EIEOS signals. Among other examples, the sending of such pre-processing signals can be coordinated with the active tracks such that, coinciding with the initial transmission of the EIEOS and the protection of the active tracks from interfering transients during the start of the unused tracks, temporary null flits are sent on the active tracks. Other techniques include, for example,Alternatively or additionally, data link layer buffers can be used to protect against bit loss resulting from such transients in reawakening unused tracks.

[0071] Furthermore, in some implementations, after sending an initial EIEOS (or initial supersequence), a partial-width state exit supersequence (e.g., 1230) can be sent. At least part of the supersequence can be repeated on the active tracks (e.g., at 1225). Furthermore, the device receiving the supersequence 1225 can, among other examples, identically replay the supersequence to perform an acknowledgment exchange of the state transition. Additionally, sending the supersequence (e.g., 1230) can be used to perform bit locking, byte locking, debouncing, drafting, and skew compensation. For example, the reactivated tracks can be skew compensated relative to the active tracks.In some cases, the initial configurations designated for the unused tracks in the original initialization of the connection can be accessed and applied, although in other cases the unused nature of the tracks may lead to changes in skew and other track properties, resulting in the effective reinitialization of the unused tracks.

[0072] Briefly returning to Fig. Figure 8 shows an example of sequences that can be sent in conjunction with a partial-width transmit state exit (e.g., a transition from an L0p state to an L0 state). Since tracks should remain active before and after such a transition, the focus can be placed on accelerating the state transition to provide minimal interruption for the active tracks. In one example (e.g., as in Figure 1220), Fig.12) A partial supersequence can be sent without the subsequent training sequences to speed up debouncing. For example, one might attempt to resolve transients within the first EIEOS without waiting another 1 kUI for a second full EIEOS to be sent to begin bit locking, byte locking, skew compensation, and other tasks. Furthermore, the full partial-width transmit-state exit supersequence can contain a repeating sequence of an EOS (e.g., EIEOS) followed by a predefined number of training sequences. In the example from Fig.8. An EIEOS followed by a series of training sequences (e.g., seven consecutive training sequences) can be sent. In one implementation, an abbreviated "rapid training sequence" (FTS) can be sent instead of a full training sequence (such as the "TS" used in supersequences 805 and 810). Among other features, the symbols of the FTS can be optimized to aid in fast bit and byte locking and skew compensation of reactivated tracks. In one example, the FTS can have a length of less than 150 UI (e.g., 128 UI). Furthermore, FTS can be left unscrambled to further aid in the fast recovery of unused tracks.

[0073] As shown in the third row of Item 815, a partial-width transmit state exit supersequence can also be interrupted by an SDS when a controller has determined that the reactivated tracks have been effectively initialized. In one example, a partial FTS (or FTSp) can follow the SDS to help synchronize the reactivated track (e.g., when bit locking, byte locking, and skew compensation have completed) with the active tracks. For example, the bit length of the FTSp can be set according to a clean flit boundary for the final width between the receiving tracks and the active tracks. To facilitate fast track synchronization, bits can be added to or removed from a track at the receiver before or during the FTSp to account for skew.Alternatively or additionally, bits can also be added to or removed from the track at the receiver before or during the SDS, among other examples to facilitate the skew compensation of a newly activated track.

[0074] Returning to the discussion from Fig.In some examples, data flit transmission on the active tracks (e.g., tracks 0 to n) can be resumed (e.g., at 1225) while the initialization of the waking-up tracks is being completed. For example, data link layer transmissions can be resumed once debouncing has been resolved. In some cases, flit transmission can be temporarily interrupted in connection with the final reactivation and synchronization of the previously unused tracks (e.g., tracks n + 1 to z) (e.g., in conjunction with transmitting an FTSp 1235) (e.g., at 1240). Once the tracks have been restored, flit data transmission can then be resumed on all tracks (1245).

[0075] In one embodiment, the clock can be embedded in the data, eliminating the need for separate clock tracks. To facilitate clock recovery, flits sent over the tracks can be scrambled. For example, the receiver clock recovery unit can supply sampling clocks to a receiver (i.e., the receiver recovers the clock from the data and uses it to sample the incoming data). In some implementations, the receivers continuously adapt to an incoming bitstream. Embedding the clock can potentially reduce the connector pin layout. However, embedding the clock in the in-band data can change how in-band reset is handled. In one embodiment, a link lock state (BLS) can be used after initialization.Furthermore, among other considerations, supersequences of an electrically ordered set can be used during initialization to enable resetting (e.g., as described above). The embedded clock can be shared between devices on a link, and the common operating clock can be set during link calibration and configuration. For example, HPI links can reference a common clock using drift buffers. Among other potential advantages, such an implementation can achieve lower latency than elastic buffers used in non-shared reference clocks. Furthermore, the reference clock distribution segments can be adjusted within specified limits.

[0076] As noted above, an HPI connection can operate at multiple speeds, including a "slow mode" for standard power-on, initialization, and so on. The operating speed or mode (or "fast" speed or "quick" mode) of each device can be statically set by the BIOS. The shared clock on the connection can be configured based on the respective operating speeds of each device on either end of the connection. For example, the connection speed can be based on the slower of the two devices' operating speeds. A change in operating speed may be accompanied by a warm or cold reset.

[0077] In some examples, the link initializes to the slow operating mode with a transmission rate of, for example, 100 MT / s upon power-up. Software then configures the two sides for the link's operating speed and begins the initialization process. In other cases, a sideband mechanism can be used to configure a link, including the shared clock signal, for example, in the absence or unavailability of a slow operating mode.

[0078] In one embodiment, an initialization phase of the slow operating mode can use the same encoding, scrambling, training sequences (TS), states, etc., as the operating speed mode, but with potentially fewer features (e.g., no electrical parameter adjustment, no adaptation, etc.). The operating phase of the slower mode can also potentially use the same encoding, scrambling, etc. (although this may not be the case for other implementations), but compared to the operating speed mode, it can have fewer states and features (e.g., no low-power states).

[0079] Furthermore, a slow operating mode can be implemented using the device's native phase-locked loop (PLL) clock frequency. For example, the HPI can support an emulated slow operating mode without changing the PLL clock frequency. Although some designs may use separate PLLs for low and high speed, in some HPI implementations, an emulated slow operating mode can be achieved by maintaining the same high operating speed for the PLL clock during the slow mode. For example, a transmitter can emulate a slower clock signal by repeating bits multiple times to simulate a slow high clock signal followed by a slow low clock signal. The receiver can then oversample the received signal to detect edges emulated by the repeating bits and identify the corresponding bit.In such implementations, ports that share a PLL can coexist at low and high speeds.

[0080] A common speed of a slow operating mode can be initialized between two devices. For example, the two devices on a link can have different operating speeds. For instance, a common slow operating speed can be configured on the link during a discovery phase or while in a discovery state. In one example, an emulation multiple can be set as an integer (or non-integer) ratio of the high speed to the low speed, and the different high speeds can be down-converted to operate at the same low speed. For example, two device agents that support at least one common frequency can be hot-connected, regardless of the speed at which the host port is running.The software discovery can then use the connection in slow mode to identify and set the optimal connection operating speeds. Where the multiple is an integer ratio of a high speed to a low speed, different high speeds can operate at the same slow speed, which can be used during the discovery phase (e.g., when connecting while the system is running).

[0081] In some HPI implementations, track adaptation on a connection can be supported. The physical layer can support both receiver and transmitter / sender adaptation. In receiver adaptation, the transmitter on a track can send sample data to the receiver, which the receiver logic can process to identify deficiencies in the electrical characteristics of the track and the signal quality. The receiver can then make adjustments to the track calibration based on the analysis of the received sample data to optimize the track. In the case of sender adaptation, the receiver can again receive sample data and develop metrics that describe the quality of the track, but in this case, the metrics are not directly related to the track's electrical characteristics (e.g., the signal strength, frequency, and overall signal quality).using a return channel (such as a software channel, hardware channel, embedded channel, sideband channel, or other channel) to the sender, enabling the sender to make track adjustments based on the feedback. Receiver adaptation can be initialized at the start of the query state using the query supersequence sent by the remote sender. Similarly, sender adaptation can be performed by repeating the following for each sender parameter. Both agents can enter a loopback pattern state as the master and transmit a specified pattern. Both receivers can measure the metric (e.g., BER) for this particular sender setting at a remote agent. Both agents can enter and then reset the loopback mark state and use return channels (e.g., slow-mode TLS or sideband) to exchange metrics.Based on these metrics, the next transmitter setting can be identified. Finally, the optimal transmitter setting can be identified and saved for subsequent use.

[0082] Since both devices can operate on a single link with the same reference clock (e.g., ref clk), elastic buffers can be omitted (any elastic buffers can be bypassed or used as drift buffers with the lowest possible latency). However, phase adjustment or drift buffers can be used on each lane to transfer the respective receiver bitstream from the remote clock domain to the local clock domain. The latency of the drift buffers can be sufficient to handle the sum of drift from all sources in an electrical specification (e.g., voltage, temperature, residual SSC introduced by reference clock routing mismatches, etc.), but should be as small as possible to minimize transmission delay. If the drift buffer is too shallow, drift errors can occur and manifest as a series of CRC errors.Consequently, some implementations, among other examples, may include a drift alarm that can initiate a reset of the physical layer before an actual drift error occurs.

[0083] Some HPI implementations can support the two sides running at the same nominal reference clock frequency but with a ppm difference. In this case, frequency adjustment buffers (or elasticity buffers) may be necessary, which, among other examples, can be recalibrated during an extended BLS window or during specific sequences that would occur periodically.

[0084] Among other considerations, the operation of the logical layer of the HPI PHY can be independent of the underlying transmission media, provided that the latency does not lead to latency correction errors or timeouts at the data link layer.

[0085] In HPI, external interfaces can be provided to assist in managing the physical layer. For example, external signals (from connector pins, fuses, other layers), timers, control registers, and status registers can be provided. The input signals can change relative to the PHY state at any time, but the physical layer must consider them at specific points in a given state. For example, among other examples, a changing alignment signal (as introduced later) can be received but have no effect after the link has entered a link transmission state. Similarly, instruction register values ​​of physical layer entities can be considered only at specific times. For example, the physical layer logic can take a snapshot of the value and use it in subsequent operations.Consequently, in some implementations, updates to instruction registers can be allocated to a limited subset of specific periods (e.g., in a link transfer state or when held in a reset calibration, in a link transfer state of slow operation) to avoid anomalous behavior.

[0086] Because status values ​​track hardware changes, the values ​​read can depend on when they are read. However, some status values, such as the link card, latency, speed, etc., cannot change after initialization. For example, a reinitialization (or exiting a low-power link state (LPLS) or L1 state) is the only thing that can cause these to change (e.g., a hard track disruption in a TLS, among other examples, might only lead to a link reconfiguration after a reinitialization has been triggered).

[0087] Interface signals can contain signals that are external to the physical layer behavior but affect it. Examples of such interface signals include encoding and timing signals. Interface signals can be design-specific. These signals can be inputs or outputs. Some interface signals, such as semaphores and prefix EO, can be active once per declaration edge; that is, their declaration can be canceled and then re-declared to become effective again. For example, Table 1 provides an exemplary list of such functions: TABLE 1 function Input connector reset (aka warm reset) Input connector reset (aka cold reset) Input of an in-band reset pulse; causes a semaphore to be set; if an in-band reset occurs, the semaphore is cleared. Input releases low-power states Input of loopback parameters; applied to loopback patterns Entry to a PWLTS Input for exiting a PWLTS Entry for admission to an LPLS Input for exiting an LPLS Input from the detection of an unused output (aka squelch interruption) Input that enables the use of CPhyInitBegin Input of local or planetary orientation for transmitter to exit initialization Output if, furthermore, agent LPLS request is answered with NAKs Output when agent enters an LPLS Output to the data link layer to force non-repeatable flits Output to the data link layer to enforce null flits Output when the transmitter is in a partial width link transmission (PWLTS) state Output when receiver is in a PWLTS

[0088] The CSR timer default values ​​can be provided in pairs—one for the slow operating mode and one for the operating speed. In some cases, a value of 0 disables the timer (i.e., the timer never occurs). The timers can include those shown in Table 2 below. Primary timers can be used to time expected actions in a state. Secondary timers are used to abort initializations that are not progressing or to perform forward state transitions at precise times in an operating mode of an automated test equipment (ATE). In some cases, secondary timers in a state can be much larger than primary timers. Exponential timers may have the suffix exp, and the timer value is 2 raised to the power of the field value. For linear timers, the timer value is the field value. Each timer could use different granularities.Furthermore, some timekeepers in the performance management section may be grouped into a set called a timing profile. These may be assigned to a timing diagram with the same name. TABLE 2 Timer Set table Tpriexp Resetting the location to navigate to EIEOS Receiver calibration minimum time; for transmitter staggering Transmitter calibration minimum time; for staggering Set Tsecexp time-definite receiver calibration time-based transmitter calibration Squelch exit detection / squelch exit debouncing DetectAtRx overhang for receipt exchange Adaptation and bit locking / byte locking / skew compensation Configuring connection widths Waiting for a planetarily aligned, clean Flit boundary Re-locking of bytes / skew compensation set Tdebugexp For plugging in during operation; non-zero value for testing hangs Setting up TBL Sentry BLS entry delay - fine BLS entry delay - coarse Set up TBLS BLS duration for the transmitter BLS duration for the recipient BLS clean-flit interval for the transmitter TBLS clean-flit interval for the receiver

[0089] Command and control registers may be provided. Control registers may be late-action registers and, in some cases, may be read or written by software. Late-action values ​​may take effect continuously upon reset (e.g., passing from a software-based stage to a hardware-based stage). Control semaphores (prefix CP) are RW1S and may be cleared by hardware. Control registers may be used to execute any of the positions described here. They may be modifiable by hardware, software, firmware, or a combination thereof, and may be accessed by these means.

[0090] Status registers can be used to track hardware changes (written and used by hardware) and can be read-only (although test software may also be able to write to them). Such registers do not affect interoperability and can typically be augmented by many private status registers. They may be prefixed with status semaphores (SP), as these can be cleared by software to undo the actions that set the status. As a subset of these status bits related to initialization, default mean initial values ​​(upon reset) may be provided. If initialization is aborted, these registers can be copied to a storage structure.

[0091] Toolbox registers may be provided. For example, the test capability toolbox registers in the physical layer may provide pattern generation, pattern verification, and loopback control mechanisms. Higher-level applications can use these registers, along with electrical parameters, to determine limits. For example, a tester embedded in the wiring may use this toolbox to determine limits. Among other examples, these registers may be used in conjunction with the specific transmitter adaptation registers described in previous sections.

[0092] In some implementations, HPI supports reliability, availability, and usability (RAS) capabilities using the physical layer. In one embodiment, HPI, with one or more layers that may contain software, supports hot-plugging and hot-plugging. Hot-plugging may involve putting the connection into a sleep state and clearing an initialization start state / signal for the agent being removed. A remote agent (i.e., the one not to be removed, such as the host agent) may be set to low speed, and its initialization signal may also be cleared. Among other examples and features, in-band reset (e.g.,(via BLS) cause both agents to wait in a reset state, such as a calibration reset state (CRS); and that the agent to be removed can be removed (or kept powered down in a targeted pin reset). In fact, some of the above events can be omitted, and additional events can be added.

[0093] Hot-adding agents can include setting the initialization speed to slow by default and setting an initialization signal on the agent being added. Software can set the speed to low and clear the initialization signal on the remote agent. The link can come up in slow mode, and software can determine an operating speed. In some cases, no PLL relocking of a remote facility is performed at this point. The operating speed can be set on both agents, and an adaptation enable can be set (if not already done). The initialization start indicator can be cleared on both agents, and an in-band BLS reset can cause both agents to wait in the CRS. Software can perform a warm reset (e.g.,Software can declare a (to be added) targeted reset or self-reset of an agent, which can cause a PLL to relock. Additionally, software can set the initialization start signal using some known logic and further set it to remote (thus switching it to the receiver detection state (RDS)). Software can also override the warm reset declaration of the added agent (thus switching it to RDS). Among other examples, the link can then initialize at operating speed to a link transfer state (TLS) (or to loopback if the adaptation signal is set). In fact, some of the above events can be omitted, and additional events can be added.

[0094] Data track corruption recovery can be supported. In one embodiment, a link in the HPI can be elastic against a hard fault in a single track by configuring itself to less than the full width (e.g., less than half the full width), thereby excluding the corrupted track. As an example, the configuration can be performed by a link state machine, and unused tracks can be disabled in the configuration state. As a result, the data can be transmitted over a narrower width, among other examples.

[0095] In some HPI implementations, track reversal can be supported on certain connections. The track reversal can refer, for example, to tracks 0 / 1 / 2... of a transmitter connected to tracks n / n-1 / n-2... of a receiver (where n can be, for example, 19 or 7, etc.). The track reversal can be detected at the receiver because it is identified in a field of a TS start block. The receiver can handle the track reversal by starting in a query state, using physical track n...0 for logical track 0...n. Thus, references to a track can refer to the number of a logical track. This allows board designers to define the physical or electrical layout more efficiently, and enables the HPI to work with virtual track assignments as described here. Furthermore, in one embodiment, the polarity can be reversed (i.e.,(when a differential transmitter (+ / -) is connected to a receiver (+ / -). The polarity can also be detected at a receiver from one or more TS start block fields and, in one embodiment, handled in the query state.

[0096] In Fig.Figure 13 shows an embodiment of a block diagram for a computer system containing a multi-core processor. The processor 1300 contains any processor or processing device, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, a handheld device processor, an application processor, a coprocessor, a system-on-a-chip (SoC), or any other code-executing device. In one embodiment, the processor 1300 contains at least two cores—core 1301 and core 1302—which may be asymmetric cores or symmetric cores (the embodiment shown). However, the processor 1300 may contain any number of processing elements, which may be symmetric or asymmetric.

[0097] In one embodiment, a processing element refers to hardware or logic for supporting a software thread. Examples of hardware processing elements include: a thread unit, a thread slot, a thread, a process unit, a context, a context unit, a logical processor, a hardware thread, a core, and / or any other element that can maintain state for a processor, such as an execution state or an architectural state. In other words, in one embodiment, a processing element refers to any hardware that can be independently associated with code, such as a software thread, an operating system, an application, or other code. Typically, a physical processor (or processor socket) refers to an integrated circuit that potentially contains any number of processing elements, such as cores or hardware threads.

[0098] A kernel often refers to logic within an integrated circuit that can maintain independent architectural states, with each independently maintained architectural state having at least some dedicated execution resources. In contrast to kernels, a hardware thread usually refers to any logic within an integrated circuit that can maintain independent architectural states, with the independently maintained architectural states sharing access to execution resources. As can be seen, the line between the nomenclature of a hardware thread and a kernel overlaps when certain resources are shared and others are dedicated to a single architectural state.Nevertheless, a core and a hardware thread are often viewed by an operating system as individual logical processors, with the operating system being able to schedule operations on each logical processor individually.

[0099] The physical processor 1300, as it is in Fig.Figure 13 shows two cores—core 1301 and core 1302. Cores 1301 and 1302 are considered here to be symmetric cores, i.e., cores with the same configurations, functional units, and / or logic. In another embodiment, core 1301 contains an out-of-order processor core, while core 1302 contains an in-order processor core. However, cores 1301 and 1302 can be individually selected from any core type, such as a native core, a software-managed core, a core adapted to execute a native instruction set architecture (I-SA), a core adapted to execute a translated instruction set architecture (ISA), a co-designed core, or any other known core. In a heterogeneous core environment (i.e.,(In asymmetric cores) a form of translation, such as binary translation, can be used to schedule or execute code in one or both cores. However, to further the discussion, the functional units shown in core 1301 are described in more detail below, since the units in core 1302 operate similarly in the embodiment shown.

[0100] As shown, the 1301 core contains two hardware threads, 1301a and 1301b, which can also be referred to as hardware thread slots 1301a and 1301b. Thus, software entities such as an operating system can potentially view the 1300 processor in one embodiment as four separate processors, i.e., four logical processors or processor elements capable of executing four software threads simultaneously. As mentioned above, a first thread is assigned to architecture state registers 1301a, a second thread is assigned to architecture state registers 1301b, a third thread can be assigned to architecture state registers 1302a, and a fourth thread can be assigned to architecture state registers 1302b. Each of the architecture state registers (1301a, 1301b, 1302a and 1302b) can be referred to here as processing elements, thread slots or thread units, as described above.As shown, the architecture state registers 1301a are repeated in the architecture state registers 1301b, so that separate architecture states / architecture contexts can be stored for logical processor 1301a and logical processor 1301b. In core 1301, other small resources such as instruction pointers and rename logic can also be repeated in an assigner / renamer block 1330 for threads 1301a and 1301b. Some resources, such as reorder buffers in the reorder / removal unit 1335, the I-TLB 1320, load / store buffers, and queues, can be shared through partitioning. Other resources such as internal general-purpose registers, one or more page table base registers, a lower-level data cache and a data TLB 1315, one or more execution units 1340 and sections of the out-of-order unit 1335 are potentially fully shared.

[0101] The processor often contains 1300 other resources that may be fully shared, shared through partitioning, or dedicated to / for processing elements. Fig.Figure 13 is an embodiment of a purely exemplary processor with illustrative logic units / logic resources of a processor shown. It is noted that a processor may include any of these functional units or omit them, and may include any other known functional units, logic, or firmware not shown. As shown, core 1301 contains a simplified representative out-of-order processor core (OOO processor core). However, an in-order processor may be used in other embodiments. The OOO core contains a branch target buffer 1320 to predict branches to be executed / taken, and an instruction translation buffer (I-TLB) 1320 to store address translation entries for instructions.

[0102] Furthermore, the core 1301 contains a decoding module 1325, which is coupled to a retrieval unit 1320 to decode retrieved elements. In one embodiment, the retrieval logic includes individual sequence controls assigned to three slots 1301a and 1301b, respectively. Typically, the core 1301 is associated with a first ISA that defines / specifies instructions executable in the processor 1300. Often, machine code instructions that are part of the first ISA contain a section of the instruction (referred to as an opcode) that refers to / specifies an instruction or operation to be executed. The decoding logic 1325 contains a circuit arrangement that recognizes these instructions from their opcodes and passes the decoded instructions into the pipeline for processing as defined by the first ISA. As discussed in more detail below, the decoders 1325 in one embodiment contain, for example,Logic designed or intended to recognize specific instructions, such as a transaction instruction. As a result of recognition by the 1325 decoders, the 1301 architecture or core takes specific, predefined actions to execute tasks associated with the corresponding instruction. It is important to note that any of the tasks, blocks, operations, and procedures described here can be executed in response to one or more instructions, some of which may be new or old instructions. It is noted that in one embodiment, the 1326 decoders recognize the same ISA (or a subset thereof). Alternatively, in a heterogeneous core environment, the 1326 decoders recognize a second ISA (either a subset of the first ISA or a different ISA).

[0103] In one example, the assigner and renamer block 1330 contains an assigner for reserving resources such as register files to store instruction processing results. However, threads 1301a and 1301b are potentially capable of out-of-order execution, and the assigner and renamer block 1330 also reserves other resources, such as reassigner buffers, to track instruction results. Additionally, unit 1330 may contain a register renamer for renaming program / instruction reference registers to other registers internal to processor 1300. The reassigner / elimination unit 1335 contains components such as the aforementioned reassigner buffers, load buffers, and memory buffers to support out-of-order execution and subsequent in-order elimination of instructions that are executed out of order.

[0104] In one embodiment, the scheduler and execution unit(s) block 1340 includes a scheduler unit for scheduling instructions / operation in execution units. For example, a floating-point instruction is scheduled on a port of an execution unit that has an available floating-point execution unit. Also included are register files associated with the execution units for storing information instruction processing results. Exemplary execution units include a floating-point execution unit, an integer execution unit, a jump execution unit, a load execution unit, a memory execution unit, and other known execution units.

[0105] The one or more execution units 1340 are coupled with a lower-level data cache and a data translation buffer (D-TLB) 1350. The data cache is intended to store recently used / processed elements, such as data operands, which are potentially held in memory coherence states. The D-TLB is intended to store recent translations of virtual / linear addresses to physical addresses. As a specific example, a processor may contain a page table structure to partition physical memory into multiple virtual pages.

[0106] Cores 1301 and 1302 share access to a higher-level cache or a more distant cache, such as a second-level cache associated with the chip-integrated interface 1310. It is noted that higher-level or more distant cache refers to cache levels that increase in or are located further away from the execution unit(s). In one embodiment, the higher-level cache is a last-level data cache—the last cache in the memory hierarchy in the 1300 processor—such as a second-level or third-level data cache. However, the higher-level cache is not limited to this, as it can be associated with or contain an instruction cache. According to the 1325 decoder, a trace cache—a type of instruction cache—can instead be coupled to store recently decoded traces.Here, an instruction potentially refers to a macro instruction (i.e., a general instruction recognized by the decoders) that can be decoded into a number of micro instructions (micro-operations).

[0107] Furthermore, the 1300 processor in the configuration shown includes a chip-integrated interface module 1310. Historically, a memory controller, as described in more detail below, is contained in a computer system external to the 1300 processor. In this scenario, the chip-integrated interface 1310 is intended to communicate with devices external to the 1300 processor, such as the system memory 1375, a chipset (which often includes a memory controller hub for connecting to the 1375 memory and an I / O controller hub for connecting to peripheral devices), a memory controller hub, a northbridge, or another integrated circuit. Additionally, in this scenario, the 1305 bus can use any known wiring, such as a multipoint bus, point-to-point wiring, serial wiring, a parallel bus, a coherent (e.g.,a cache coherent bus, a layer protocol architecture, a differential bus and a GTL bus.

[0108] The memory 1375 can be assigned to the processor 1300 or shared with other devices in a system. Common examples of memory 1375 types include DRAM, SRAM, non-volatile memory (NV memory), and other known storage devices. It is noted that the device 1380 may include a graphics accelerator, a processor or card coupled to a memory controller hub, a data storage device coupled to an I / O controller hub, a wireless transceiver, a flash device, an audio controller, a network controller, or any other known device.

[0109] However, recently, as more logic and devices are integrated onto a single chip such as a SoC, each of these devices can be integrated into the 1300 processor. For example, in one embodiment, a memory controller hub is in the same assembly and / or on the same single chip as the 1300 processor. A section of the core (a core-integrated section) 1310 here contains one or more controllers as an interface to other devices such as the 1375 memory or a 1380 graphics device. The configuration that includes wiring and controllers as an interface to such devices is often referred to as a core-integrated (or un-core) configuration. As an example, the 1310 chip-integrated interface includes ring wiring for chip-integrated communication and a high-speed serial point-to-point link 1305 for chip-off communication.Nevertheless, in the SOC environment, even more devices such as the network interface, coprocessors, memory 1375, a graphics processor 1380 and any other known computer devices / any other known computer interface can be integrated on a single chip or integrated circuit to provide a small form factor with high functionality and low power consumption.

[0110] In one embodiment, the processor 1300 can execute compiler, optimizer, and / or translator code 1377 to compile, translate, and / or optimize application code 1376 to support or interface with the devices and procedures described herein. A compiler often contains a program or set of programs for translating source text / code into target text / code. Typically, the compilation of program code / application code with a compiler is performed in multiple stages and passes to convert code from a higher-level programming language into low-level machine code or assembly language code. However, single-pass compilers can still be used for simple compilation.A compiler can use any known compilation techniques and perform any known compiler operations such as lexical analysis, preprocessing, parsing, semantic analysis, code generation, code transformation, and code optimization.

[0111] Larger compilers often contain multiple phases, though most of these phases fall within two general categories: (1) a frontend, which is generally where syntax processing, semantic processing, and some transformation / optimization take place, and (2) a backend, which is generally where parsing, transformations, optimizations, and code generation occur. Some compilers employ a middle ground, blurring the lines between a compiler's frontend and backend. As a result, references to insertion, mapping, generation, or any other compiler operation can occur in any of the phases or passes mentioned above, as well as in any other known stages or passes of a compiler. For example, a compiler might potentially insert operations, calls, functions, and so on.Dynamic compilation involves one or more phases of the compilation process, such as the insertion of calls / operations in a frontend phase and the subsequent transformation of these calls / operations into lower-level code during a transformation phase. It is noted that during dynamic compilation, compiler code or dynamic optimization code can insert such operations / calls and optimize the code for execution at runtime. As a specific illustrative example, binary code (already compiled code) can be dynamically optimized at runtime. The program code can contain the dynamic optimization code, the binary code, or a combination thereof.

[0112] A translator, such as a binary translator, translates code similarly to a compiler, either statically or dynamically, in order to optimize and / or translate code.Thus, reference to the execution of code, application code, program code, or any other software environment may refer to: (1) the execution of one or more compiler programs, optimization code optimizers, or translators, either dynamically or statically, to compile program code, to maintain software structures, to perform other operations, to optimize code, or to translate code; (2) the execution of main program code, including operations / calls such as application code that has been optimized / compiled; (3) the execution of other program code, such as libraries associated with the main program code, to maintain software structures, to perform other software-related operations, or to optimize code; or (4) a combination thereof.

[0113] Now in Fig. Figure 14 shows a block diagram of an embodiment of a multi-core processor. As in the embodiment of Fig. As shown in Figure 14, the 1400 processor contains multiple domains. More precisely, a core domain 1430 contains multiple cores 1430A-1430N, a graphics domain 1460 contains one or more graphics machines with a media machine 1465, and a system agent domain 1410.

[0114] In various embodiments, the system agent domain 1410 handles power control events and power management, allowing individual units of domains 1430 and 1460 (e.g., cores and / or graphics engines) to be independently controllable to operate dynamically in an appropriate power mode / at an appropriate power level (e.g., active, turbo, sleep, hibernate, deep sleep, or another state similar to an advanced configuration power interface) based on the activity (or inactivity) occurring in the given unit. Each of domains 1430 and 1460 can operate at a different voltage and / or power, and furthermore, the individual units within the domains can each potentially operate at an independent frequency and voltage.Although only three domains are shown, it is noted that the scope of protection of the present invention is of course not limited in this respect and that additional domains may be present in other embodiments.

[0115] As shown, each 1430 core contains, in addition to various execution units and additional processing elements, lower-level caches. The different cores are coupled to each other and to a shared cache memory formed from multiple units or slices of a last-level cache (LLC) 1440A-1440N; these LLCs often include storage and cache controller functionality and are shared among the cores and potentially also with the graphics engine.

[0116] As can be seen, a ring wiring 1450 couples the cores to each other via several ring holders 1452A-1452N, each at a coupling between a core and an LLC slice, and provides wiring between the core domain 1430, the graphics domain 1460, and the system agent circuit arrangement 1410. As shown in Fig.As shown in Figure 14, wiring 1450 is used to transmit various pieces of information, including address information, data information, acknowledgment information, and eavesdropping / invalid information. Although a ring wiring configuration is shown, any known chip-integrated wiring or chip-integrated fabric can be used. As an illustrative example, some of the fabrics discussed above (e.g., another chip-integrated wiring configuration, a chip-integrated system fabric (OSF), an advanced microcontroller bus architecture (AMBA) wiring configuration, a multidimensional mesh fabric, or another known wiring architecture) can be used in a similar manner.

[0117] As further shown, the system agent domain 1410 includes a display machine 1412, which provides control of an associated display and an interface with it. The system agent domain 1410 may include other units such as: an integrated memory controller 1420, which provides an interface to system memory (e.g., to DRAM implemented with multiple DIMMs); ​​and coherence logic 1422 for performing memory coherence operations. Multiple interfaces may be provided to enable communication between the processor and another circuit arrangement. For example, in one embodiment, at least one Direct Media Interface (DMI) interface 1416 and one or more PCIe™ interfaces 1414 are provided. Typically, this display machine and these interfaces couple to the memory via a PCIe™ bridge 1418.Furthermore, one or more additional interfaces may be provided to enable communication between other agents, such as additional processors or a different circuit arrangement.

[0118] Now in Fig. Figure 15 is a block diagram of a representative kernel, specifically of logic blocks of a backend of a kernel such as kernel 1430. Fig. 14, shown. In general, the one in Fig. The structure shown in Figure 15 depicts an out-of-order processor comprising a front-end unit 1570, which is used to retrieve incoming instructions, perform various processing operations (e.g., cache processing, decoding, branch prediction, etc.), and pass instructions / operations along to an out-of-order machine (OOO machine) 1580. The OOO machine 1580 performs further processing on the decoded instructions.

[0119] More precisely, the out-of-order machine 1580 in the embodiment shown contains Fig. 15. An allocation unit 1582 receives decoded instructions, which may be in the form of one or more micro-instructions or µops, from the front-end unit 1570 and allocates them to appropriate resources such as registers. The instructions are then provided to a reservation station 1584, which reserves resources and schedules them for execution in one of several execution units 1586A-1586N. Various types of execution units may be present, including, for example, arithmetic logic units (ALUs), load and store units, vector processing units (VPUs), floating-point execution units, and others. The results from these various execution units are provided to a reorder buffer (ROB) 1588, which takes any unordered results and reassembles them into the correct program order.

[0120] Further based on Fig. It is noted in Figure 15 that both the front-end unit 1570 and the out-of-order machine 1580 are coupled to different levels of a memory hierarchy. More precisely, an instruction-level cache 1572 is shown, which in turn is coupled to a middle-level cache 1576, which in turn is coupled to a top-level cache 1595. In one embodiment, the top-level cache 1595 is implemented in a chip-integrated unit (occasionally referred to as an un-core unit) 1590. As an example, the unit 1590 is similar to the system agent 1410 from [reference missing]. Fig.14. As discussed above, the Un-Core 1590 communicates with system memory 1599, which in the illustrated embodiment is implemented via ED-RAM. It is also noted that the various execution units 1586 within the out-of-order machine 1580 are connected to a first-level cache 1574, which is also connected to a middle-level cache 1576. It is further noted that the LLC 1595 can couple additional cores 1530N-2 - 1530N. Although this is not the case in the embodiment shown. Fig. As shown at this high level, 15 indicates that various changes and additional components may be present.

[0121] Transition to Fig.Figure 16 shows a block diagram of an exemplary computer system formed with a processor containing execution units for executing an instruction, wherein one or more of the wiring implements one or more features in accordance with an embodiment of the present invention. System 1600 includes a component, such as a processor 1602, for utilizing execution units containing logic for executing algorithms for processing data in accordance with the present invention, such as in the embodiment described herein. System 1600 represents processing systems based on PENTIUM III™, PENTIUM 4™, Xeon™, Itanium, XScale™, and / or StrongARM™ microprocessors, although other systems (including PCs with other microprocessors, engineering workstations, set-top boxes, and the like) can also be used.In one embodiment, the prototype system 1600 runs a version of the WINDOWS™ operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical user interfaces can also be used. Thus, embodiments of the present invention are not limited to any specific combination of hardware circuitry and software.

[0122] Embodiments are not limited to computer systems. Alternative embodiments of the present invention can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include mobile phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications can include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of executing one or more instructions in accordance with at least one embodiment.

[0123] In this illustrated embodiment, the processor 1602 includes one or more execution units 1608 for implementing an algorithm that is to execute at least one instruction. One embodiment can be described in the context of a single-processor desktop system or a single-processor server system, although alternative embodiments may be included in a multi-processor system. The system 1600 is an example of a 'hub' system architecture. The computer system 1600 includes a processor 1602 for processing data signals.As an illustrative example, the Processor 1602 contains a computer's complex instruction set (CISC) microprocessor, a computer's reduced instruction set (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other processing device such as a digital signal processor. The Processor 1602 is coupled to a Processor Bus 1610, which transmits data signals between the Processor 1602 and other components in the System 1600. The elements of the System 1600 (e.g., the graphics accelerator 1612, the memory controller hub 1616, the memory 1620, the I / O controller hub 1624, the wireless transceiver 1626, the flash BIOS 1628, the network controller 1634, the audio controller 1636, the serial expansion port 1638, the I / O controller 1640, etc.)They perform their conventional functions, which are well known to experts in the field.

[0124] In one embodiment, the processor 1602 includes an internal Level 1 (L1) cache memory 1604. Depending on the architecture, the processor 1602 may have a single internal cache or multiple levels of internal caches. Other embodiments, depending on the specific implementation and requirements, include a combination of both internal and external caches. The register file 1606 is intended to store various data types in different registers, including integer registers, floating-point registers, vector registers, bank registers, shadow registers, checkpoint registers, status registers, and instruction pointer registers.

[0125] The execution unit 1608, which contains logic for performing integer and floating-point operations, is also located in the processor 1602. In one embodiment, the processor 1602 includes a microcode ROM (µcode ROM) for storing microcode which, when executed, is intended to run algorithms for specific macro instructions or to handle complex scenarios. The microcode is potentially updatable to handle logic errors / corrections for the processor 1602. In another embodiment, the execution unit 1608 includes logic for handling a packed instruction set 1609. By incorporating the packed instruction set 1609 into the instruction set of a general-purpose processor 1602, and with the associated circuitry for executing the instructions, the operations used by many multimedia applications can be performed using packed data in a general-purpose processor 1602.Thus, many multimedia applications are accelerated and run more efficiently by utilizing the full width of a processor's data bus to perform operations on packed data. This potentially eliminates the need to transfer smaller data units, only one data element at a time, across the processor's data bus to perform one or more operations.

[0126] Alternative embodiments of an execution unit 1608 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. The system 1600 includes a memory 1620. The memory 1620 includes a dynamic read / write memory device (DRAM device), a static read / write memory device (SRAM device), a flash memory device, or another storage device. The memory 1620 stores instructions and / or data represented by data signals to be executed by the processor 1602.

[0127] It is noted that any of the above-mentioned features or any of the above-mentioned aspects of the invention may be incorporated into one or more of the inventions in Fig.The invention can be used in the 16 wiring configurations shown. For example, an on-chip wiring (ODI) configuration, not shown, implements one or more aspects of the invention described above for coupling internal units of the processor 1602. Alternatively, the invention is associated with a processor bus 1610 (e.g., another known high-performance computer wiring configuration), a high-bandwidth memory path 1618 to the memory 1620, a point-to-point connection to the graphics accelerator 1612 (e.g., a Peripheral Component Interconnect Express compatible fabric (PCIe compatible fabric)), a controller hub wiring configuration 1622, an I / O interface, or other wiring (e.g., USB, PCI, PCIe) for coupling the other components shown.Some aspects of such components include the audio controller 1636, the firmware hub (flash BIOS) 1628, the wireless transceiver 1626, the data storage device 1624, the alt I / O controller 1610, which includes user input and keyboard interfaces 1642, a serial expansion port 1638 such as Universal Serial Bus (USB), and a network controller 1634. The data storage device 1624 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or another mass storage device.

[0128] Now in Fig. Figure 17 shows a block diagram of a second system 1700 in accordance with an embodiment of the present invention. As in Fig.As shown in Figure 17, the multiprocessor system 1700 is a point-to-point wiring system and comprises a first processor 1770 and a second processor 1780, which are coupled via a point-to-point wiring connection 1750. Each of the processors 1770 and 1780 can be a specific version of a processor. In one embodiment, 1752 and 1754 are parts of a serial coherent point-to-point wiring fabric, such as a high-performance architecture. As a result, the invention can be implemented within the QPI architecture.

[0129] Although only two processors, 1770 and 1780, are shown, it should be noted that the scope of protection of the present invention is not limited thereto. In other embodiments, one or more additional processors may be present in a given processor.

[0130] The 1770 and 1780 processors are shown to contain integrated memory controller units 1772 and 1782, respectively. The 1770 processor also includes point-to-point (PP) interfaces 1776 and 1778 as part of its bus controller units; similarly, the 1780 processor includes PP interfaces 1786 and 1788. The 1770 and 1780 processors can exchange information via a point-to-point (PP) interface 1750 using PP interface circuits 1778 and 1788. As shown in Fig. As shown in Figure 17, IMCs 1772 and 1782 can couple the processors with their respective memories, i.e., with a memory 1732 and a memory 1734, which can be sections of main memory that are locally connected to the respective processors.

[0131] The 1770 and 1780 processors each exchange information with a 1790 chipset via individual PP interfaces 1752 and 1754, using point-to-point interface circuits 1776, 1794, 1786, and 1798. Additionally, the 1790 chipset exchanges information with a 1738 high-performance graphics circuit via an interface circuit 1792 together with a high-performance graphics wiring 1739.

[0132] Each processor, or separate from both processors but connected to the processors via PP wiring, may contain a shared cache (not shown) so that the local cache information of one or both processors can be stored in the shared cache if one processor is placed in a low-power operating mode.

[0133] The chipset 1790 can be coupled to a first bus 1716 via an interface 1796. In one embodiment, the first bus 1716 can be a peripheral component interconnect bus (PCI bus) or a bus such as a PCI Express bus or another third-generation I / O wiring bus, although the scope of protection of the present invention is not limited thereto.

[0134] As in Fig.As shown in Figure 17, various I / O devices 1716 are coupled to the first bus 1716 together with a bus bridge 1718, which couples the first bus 1716 to a second bus 1720. In one embodiment, the second bus 1720 includes a low-pin-count bus (LPC bus). In another embodiment, various devices are coupled to the second bus 1720, including, for example, a keyboard and / or mouse 1722, communication devices 1727, and a storage unit 1728 such as a disk drive or other mass storage device, which often contains commands / code and data 1730. Furthermore, an audio I / O 1724 is shown coupled to the second bus 1720. It is noted that other architectures are possible, with varying components and wiring architectures. For example, a system may use a point-to-point architecture instead of a... Fig. 17. Implement a multipoint bus or other such architecture.

[0135] Moving on to Fig. Figure 18 shows an embodiment of a single-chip design (SOC design) in accordance with the inventions. As a specific illustrative example, the SOC 1800 is contained in a user device (UE). In one embodiment, UE refers to any device that an end user uses to communicate, such as a mobile phone, a smartphone, a tablet, an ultrathin notebook, a notebook with a broadband adapter, or any other similar communication device. Frequently, a UE connects to a base station or node, which is potentially equivalent to the nature of a mobile station (MS) in a GSM network.

[0136] The SOC 1800 contains two cores – 1806 and 1807. Similar to the discussion above, cores 1806 and 1807 can be in accordance with an instruction set architecture such as that of an Intel® Architecture Core™-based processor, an Advanced Microdevices, Inc. (AMD) processor, a MIPS-based processor, an ARM-based processor design, or that of a customer, licensee, or user thereof. Cores 1806 and 1807 are coupled to a cache controller 1808, to which the bus interface unit 1809 and the L2 cache 1811 are associated for communication with other parts of the System 1800. The wiring 1810 contains on-chip wiring such as IOSF, AMBA, or other wiring discussed above, which potentially implements one or more of the aspects described here.

[0137] The 1810 wiring provides communication channels to other components such as a Subscriber Identity Module (SIM) 1830 as an interface with a SIM card, a Boot ROM 1835 for holding boot code for execution by the cores 1806 and 1807 to initialize and boot the SOC 1800, an SDRAM controller 1840 as an interface with external memory (e.g., with the DRAM 1860), a Flash controller 1845 as an interface with non-volatile memory (e.g., with the Flash 1865), a peripheral controller 1850 (e.g., a Serial Peripheral Interface) as an interface with peripheral devices, video codecs 1820 and a video interface 1825 for displaying and receiving input (e.g., touch input), a GPU 1815 for performing graphics-related calculations, etc. Any of these interfaces may contain aspects of the invention described herein.

[0138] Furthermore, the system illustrates communication peripherals such as a Bluetooth module 1870, a 3G modem 1875, GPS 1885, and WLAN 1885. It is noted that, as mentioned above, a UE contains a radio communication device. Consequently, not all of these peripheral communication modules are required. However, a UE must contain some form of radio for external communication.

[0139] Although the present invention has been described with respect to a limited number of embodiments, those skilled in the art will recognize numerous modifications and deviations thereof. The appended claims are intended to include all such modifications and deviations as they are inherent in the true concept and scope of protection of this present invention.

[0140] A design can go through various phases, from creation to simulation to manufacturing. Data representing a design can represent it in a number of ways. First, the hardware, as is useful in simulations, can be represented using a hardware description language or another functional description language. Additionally, at some stages of the design process, a circuit-level model with logic and / or transistor gates can be created. Furthermore, most designs, at some stage, reach a level of data that represents the physical arrangement of various devices within the hardware model.If conventional semiconductor fabrication techniques are used, the data representing the hardware model can be the data specifying the presence or absence of various features in different mask layers for masks used to fabricate the integrated circuit. In any representation of the design, the data can be stored on some form of machine-readable medium. A storage medium, such as a magnetic or optical disk, can be a machine-readable medium for storing information transmitted by an optical or electrical wave that is modulated or otherwise generated to transmit such information. When an electrical carrier wave is transmitted that specifies or conveys the code or design, a new copy is created to the extent that copying, buffering, or retransmission of the electrical signal is performed.Thus, a communications provider or network provider can at least temporarily store an item, such as information encoded in a carrier wave, which embodies techniques of embodiments of the present invention, in a specific machine-readable medium.

[0141] As used here, a module refers to any combination of hardware, software, and / or firmware. For example, a module contains hardware such as a microcontroller associated with a non-temporary medium for storing code designed to be executed by the microcontroller. Thus, in one embodiment, the reference to a module refers to the hardware specifically configured to recognize and / or execute the code to be held in a non-temporary medium. Furthermore, in another embodiment, the use of a module refers to a non-temporary medium containing the code specifically designed to be executed by the microprocessor to perform predefined operations.As can be deduced, the term module (in this example) can, in yet another embodiment, refer to the combination of the microcontroller and the non-temporary medium. Module boundaries, depicted as separate, typically vary frequently and potentially overlap. For example, a first and a second module may share hardware, software, firmware, or a combination thereof, while potentially retaining some independent hardware, software, or firmware. In one embodiment, the use of the term logic includes hardware such as transistors, registers, or other hardware such as programmable logic devices.

[0142] In one embodiment, the use of the term 'configured to' refers to arranging, assembling, manufacturing, offering for sale, importing, and / or designing a device, piece of hardware, logic, or element to perform a desired or specific task. In this example, a device or element thereof that is not operating is still 'configured' to perform a specific task if it is designed, coupled, and / or connected to perform that specific task. As a purely illustrative example, a logic gate can provide a 0 or a 1 during operation. However, a logic gate that is 'configured' to provide an enable signal for a clock signal does not include every potential logic gate that can provide a 1 or a 0.Instead, the logic gate is one that is coupled in such a way that during operation, a 1 or 0 is output to release the clock signal. It is noted again that the use of the term 'configured to' does not require operation, but instead focuses on the latent state of a device, piece of hardware, and / or element, wherein the device, hardware, and / or element in the latent state is designed to perform a specific task when the device, hardware, and / or element is operating.

[0143] Furthermore, in an embodiment, the use of the terms 'to', 'capable of', and / or 'operable of' refers to a device, logic, hardware, and / or element that is designed to enable the use of the device, logic, hardware, and / or element in a specified manner. It is noted that in an embodiment as above, the use of 'to', 'capable of', and 'operable of' refers to the latent state of a device, logic, hardware, and / or element, wherein the device, logic, hardware, and / or element is not operating but is designed to enable the use of a device in a specified manner.

[0144] A value, as used here, contains any known representation of a number, a state, a logic state, or a binary logic state. Logic levels, logic values, or logical values ​​are often referred to as 1s and 0s, thus simply representing binary logic states. For example, a 1 refers to a high logic level, and a 0 refers to a low logic level. In one embodiment, a storage cell, such as a transistor or flash cell, may be able to hold a single logic value or multiple logic values. However, other representations of values ​​are used in computer systems. For example, the decimal number ten can also be represented as the binary value 1010 and as the hexadecimal letter A. Thus, a value contains any representation of information that can be held in a computer system.

[0145] Furthermore, states can be represented as values ​​or segments of values. For example, a first value, such as a logical one, can represent a default or initial state, while a second value, such as a logical zero, can represent a non-default state. Additionally, in one embodiment, the terms reset and set refer to a default value and an updated value, respectively. For instance, a default value potentially contains a high logic value, i.e., reset, while an updated value potentially contains a low logic value, i.e., set. It should be noted that any combination of values ​​can be used to represent any number of states.

[0146] The embodiments of methods, hardware, software, firmware, or code described above can be implemented by means of instructions or code that can be implemented in a machine-accessible, machine-readable, computer-accessible, or computer-readable medium executable by a processing element. A non-temporary machine-accessible / machine-readable medium contains any mechanism that provides (i.e., stores and / or transmits) information in a form readable by a machine, such as a computer or electronic system.For example, a non-temporary, machine-accessible medium contains a read / write memory (RAM) such as static RAM (SRAM) or dynamic RAM (DRAM); a ROM; a magnetic or optical storage medium; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; another form of storage device for holding information received from transient (propagating) signals (e.g., carrier waves, infrared signals, digital signals); etc., which are to be distinguished from the non-temporary media that can receive information from them.

[0147] Instructions that can be used to program logic for executing embodiments of the invention can be stored within a memory in the system, such as in a DRAM, a cache, flash memory, or other storage medium. Furthermore, the instructions can be distributed over a network or through other computer-readable media. Thus, a machine-readable medium can be any mechanism for storing or transmitting information in a way that is readable by a machine (e.g., a computer).Computer-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer). This includes floppy disks, optical discs, compact disc read-only memory (CD-ROMs) and magneto-optical disks, read-only memory (ROMs), read / write memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical disks, flash memory, or any specific machine-readable storage medium used in, but not limited to, the transmission of information over the Internet via electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, computer-readable medium includes any type of specific machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0148] The following examples relate to embodiments in accordance with this description. One or more embodiments can provide a device, a system, a machine-readable repository, and a method for embedding a periodic control window in a data link layer data stream to be transmitted over a serial data link, wherein the control window is configured to provide physical layer information, including information for use in initiating state transitions on the data link.

[0149] In at least one example, the data stream includes a series of flits.

[0150] In at least one example, the data link layer data stream is sent during a link transfer state of the data link.

[0151] One or more examples may further provide the identification of a particular control window in the data stream and the sending of reset data to a device connected to the data link during the particular control window, wherein the reset data is intended to transmit an attempt to enter a reset state from the link transmission state.

[0152] One or more examples may further provide the generation of a supersequence associated with the reset state and the sending of the supersequence to the device.

[0153] One or more examples may further provide the identification of a particular control window in the data stream and the sending of link width transition data to a device connected to the data link during the particular control window, wherein the link width transition data is intended to convey an attempt to change the number of active tracks on the link.

[0154] In at least one example, the number of tracks is to be reduced from an original number to a new number, whereby the reduction in the number of active tracks is associated with entering a partial width link transmission state.

[0155] One or more examples may further provide the identification of a subsequent control window in the data stream and the sending of section-width exit data to the device during the subsequent control window, wherein the section-width exit data is intended to convey an attempt to reset the number of active lanes to the original number.

[0156] One or more examples may further include identifying a particular control window in the data stream and sending low-power data to a device connected to the data link during the particular control window, the low-power data being intended to convey an attempt to transition from the link transmission state to a low-power state.

[0157] In at least one example, control windows are embedded in accordance with a defined control interval, and devices connected via the data link are intended to synchronize the state transition with the end of a corresponding control interval.

[0158] One or more embodiments can provide a device, a system, a machine-readable storage, and a method for receiving a data stream, wherein the data stream is to contain alternating transmission intervals and control intervals, wherein data link layer flits are to be sent during the transmission intervals, and wherein the control intervals are to provide opportunities for sending physical layer control information, wherein control data to be included in a particular control interval is to be identified, wherein the control data is to indicate an attempted entry from a first state into a particular state, wherein the data stream is to be received in the first state, and wherein the transition to the particular state is to be enabled.

[0159] In at least one example, the definite state includes a reset state.

[0160] In at least one example, enabling the transition to the specified state involves sending an acknowledgment of the attempted entry into the specified state.

[0161] In at least one example, the acknowledgment is sent within the control interval.

[0162] In at least one example, the data stream is sent over a serial data link containing multiple active tracks, and the specified state includes a partial width state, where at least a subset of the tracks contained in the multiple active tracks are to remain unused in the partial width state.

[0163] One or more examples may further provide the identification of subsequent data contained in a subsequent control interval, wherein the subsequent data indicate an attempt to exit the partial width state and reactivate the unused tracks.

[0164] In at least one example, the specific state includes a low-power transfer state.

[0165] In at least one example, the data stream is received via a serial data link containing multiple active tracks, and the determined state includes a partial width state, where at least a subset of the tracks contained in the multiple active tracks are to remain unused in the partial width state.

[0166] In at least one example, the definite state includes a reset state.

[0167] In at least one example, the physical layer control information describes a data link error.

[0168] One or more embodiments can provide a device, a system, a machine-readable storage device, and a method for embedding a clock signal in data to be transmitted from a first device via a serial data link containing multiple tracks, and for transitioning from a first link transmission state using a first number of the multiple tracks to a second link transmission state using a second number of the multiple tracks.

[0169] In at least one example, the second number of tracks is greater than the first number of tracks.

[0170] In at least one example, the transition from a first link transfer state to the second link transfer state involves sending a partial width state exit supersequence that includes one or more instances of a sequence containing an Electrical Ordered Set (EOS) and multiple instances of a training sequence.

[0171] In at least one example, the transition from the first link transmission state to the second link transmission state further includes sending an initial EOS that precedes the partial width state exit supersequence.

[0172] In at least one example, null flits should be sent on the active tracks during the transmission of the initial EOS.

[0173] In at least one example, the training sequence includes an unjumbled rapid training sequence (FTS).

[0174] In at least one example, the transition from the first link transfer state to the second link transfer state further involves using the partial width state exit supersequence to initialize at least some of the unused tracks contained in the multiple tracks.

[0175] In at least one example, the transition from the first link transfer state to the second link transfer state also includes sending a data sequence start (SDS) after initializing the portion of the unused tracks.

[0176] In at least one example, the transition from the first link transfer state to the second link transfer state also includes sending a partial FTS (FTSp) after sending the SDS.

[0177] In at least one example, the transition from the first link transfer state to the second link transfer state further includes receiving an acknowledgment of the transition, wherein the acknowledgment contains the partial width state exit supersequence.

[0178] In at least one example, the transition from the first link transmission state to the second link transmission state involves sending an in-band signal over the data link to the second device.

[0179] In at least one example, the first number of tracks is greater than the second number of tracks.

[0180] In at least one example, the data comprises a data stream containing alternating transmission intervals and control intervals, and the signal is sent within a specific control interval, indicating the transition from the first link transmission state to the second link transmission state.

[0181] In at least one example, the transition from the first connection transmission state to the second connection transmission state should be synchronized with the end of a specific transmission interval immediately after the specific control interval.

[0182] In at least one example, the transition is based on a requirement from a power control unit.

[0183] One or more embodiments may include a device, a system, a machine-readable storage device, and a method for receiving a data stream, wherein the data stream is to contain alternating transmission intervals and control intervals, wherein the control intervals are to provide opportunities for sending physical layer control information, and wherein the data stream is to be transmitted over a serial data link, which is to contain active tracks and inactive tracks, for identifying control data contained in a particular control interval, wherein the data is to indicate an attempt to activate at least a portion of the inactive tracks of the link, and for enabling the activation of the portion of the inactive tracks.

[0184] In at least one example, the data stream is received while the data link is in a partial width state, and the control data is intended to indicate an attempt to exit the partial width state.

[0185] In at least one example, enabling the activation of the portion of inactive tracks should include receiving a supersequence that specifies an attempt to activate the portion of inactive tracks.

[0186] In at least one example, the supersequence should include one or more instances of a sequence containing an Electric Idle Exit Ordered Set (EIEOS) and multiple instances of a training sequence.

[0187] In at least one example, enabling the activation of the portion of inactive tracks involves sending at least one initial EIEOS that immediately precedes the supersequence.

[0188] In at least one example, null flits should be sent on the active tracks during the transmission of the initial EIEOS.

[0189] In at least one example, the training sequence includes an unjumbled rapid training sequence (FTS).

[0190] In at least one example, enabling the activation of the portion of inactive tracks also involves using the supersequence to initialize that portion of inactive tracks.

[0191] In at least one example, enabling the activation of the portion of inactive tracks also includes receiving a data sequence start (SDS) after initializing the portion of inactive tracks.

[0192] In at least one example, enabling the activation of the portion of the inactive tracks also includes receiving a partial FTS (FTSp) after the SDS.

[0193] In at least one example, enabling the activation of the portion of the inactive tracks also includes acknowledging the attempt by identically reproducing the supersequence.

[0194] One or more embodiments may provide a device, a system, a machine-readable storage device, and a method for receiving a data stream, wherein the data stream is to contain alternating transmission intervals and control intervals, wherein data link layer flits are to be sent during the transmission intervals, and wherein the control intervals are to provide opportunities for sending physical layer control information, for identifying control data indicating an attempted entry from a link transmission state to a low-power state, wherein the data stream is to be received in the link transmission state, and for transitioning to the low-power state.

[0195] In at least one example, the tax data includes a predefined code.

[0196] In at least one example, the transition to the low-power state involves the identical reproduction of the predefined code in a subsequent control interval.

[0197] In at least one example, the transition to the low-power state involves receiving a supersequence that indicates the transition to the low-power state.

[0198] In at least one example, the transition to the low-power state also involves the identical replay of the supersequence.

[0199] In at least one example, the supersequence comprises one or more instances of a sequence containing an Electrical Ordered Set (EOS) followed by a predetermined number of instances of a training sequence.

[0200] In at least one example, the EOS includes an Electrical Idle Electrical Ordered Set (EIEOS).

[0201] One or more embodiments may include a device, a system, a machine-readable repository, and a method for identifying a specific instance of a periodic control interval to be embedded in a data stream on a serial data link during a link transmission state, for sending state transition data during the specific instance of the control interval to a device, wherein the state transition data is to indicate an attempt to enter a low-power state and provide for transitioning to the low-power state.

[0202] One or more examples can further provide the receipt of an acknowledgment from the device, wherein the acknowledgment includes the state transition data.

[0203] In at least one example, the acknowledgment should coincide with the next periodic control interval.

[0204] In at least one example, the transition to the low-power state involves sending a supersequence, indicating the transition to the low-power state, to the device.

[0205] In at least one example, the transition to the low-power state also involves receiving a repeated instance of the supersequence from the device.

[0206] In at least one example, the supersequence comprises one or more instances of a sequence containing an Electrical Ordered Set (EOS) followed by a predetermined number of instances of a training sequence.

[0207] In at least one example, the EOS includes an Electric Idle Exit Ordered Set (EIEOS).

[0208] In at least one example, the transition to the low-power state is based on a request from a power control unit.

[0209] One or more examples can also provide the initiation of a transition from the low-power state to the link transfer state.

[0210] One or more examples may further provide a physical layer (PHY) configured to be coupled to a serial differential link, wherein the PHY is to periodically issue a link lock state (BLS), wherein the BLS request is to cause an agent to enter a BLS to hold the link layer flit transmission for a duration, wherein the PHY is to use the serial differential link during the duration for tasks assigned to the PHY.

[0211] In at least one example, the PHY is to utilize the serial differential connection during the duration for tasks assigned to the PHY, including sending one or more messages from a priority message list containing NOP, reset, in-band reset, entering a low-power state, entering a partial-width state, entering another PHY state, etc.

[0212] One or more examples may further provide a physical transmission layer (PHY) configured to be coupled to a link, wherein the link contains a first number of tracks, wherein the PHY is to send flits over the first number of tracks in a full-width link transmission state, and wherein the PHY is to send flits over a second number of tracks, which is smaller than the first number of tracks, in a partial-width link transmission state.

[0213] In at least one example, the PHY should use a link lock state to enter the partial width link transfer state from the link lock state.

[0214] In at least one example, the flits have the same size when transmitted over the first number of tracks and over the second number of tracks.

[0215] In at least one example, the PHY uses an embedded clock for transmission over the first number of tracks and over the second number of tracks.

[0216] In at least one example, the PHY uses an embedded clock for transmission over the first number of tracks and a forwarded clock for transmission over the second number of tracks.

[0217] Throughout this description, references to "one embodiment" or "an embodiment" mean that a particular feature, structure, or property described in connection with the embodiment is included in at least one embodiment of the present invention. Therefore, occurrences of the expressions "in one embodiment" or "in an embodiment" at various points throughout this description do not necessarily all refer to the same embodiment. Furthermore, the particular features, structures, or properties may be combined in one or more embodiments in any way.

[0218] The foregoing description provides a detailed account with reference to specific exemplary embodiments. However, various modifications and alterations can obviously be made to these without deviating from the broader inventive concept and scope of protection as set forth in the appended claims. Accordingly, the description and the drawings should be viewed in an illustrative rather than a limiting sense. Furthermore, the foregoing use of an embodiment and other exemplary language does not necessarily refer to the same embodiment or the same example, but may refer to different and distinct embodiments, as well as potentially to the same embodiment.

Claims

[1] Device comprising the following: a layer stack comprising a data link layer logic and a physical layer logic, wherein the physical layer logic is at least partially implemented in a hardware circuit and is intended to embed a periodic control window in a data link layer data stream to be sent over a serial data link, wherein the control window is configured to provide physical layer information including information for use in initiating state transitions on the data link, wherein the control window is to be embedded periodically according to a defined interval, a series of successive flits in the data link layer data stream are to be sent between control windows embedded in the data link layer data stream, and the length of the control window is to be shorter than the interval between control windows. [2] Device according to claim 1, wherein the data link layer data stream is sent during a link transmission state of the data link. [3] Device according to claim 1, wherein the bit transmission layer logic is further said to identify a specific control window in the data stream and, during the specific control window, to send reset data to a device connected to the data link, wherein the reset data is said to transmit an attempt to enter a reset state from the link transmission state. [4] Device according to claim 3, wherein the bit transmission layer logic is further said to generate a supersequence associated with the reset state and to send the supersequence to the device. [5] Device according to claim 2, wherein the physical layer logic is further said to identify a specific control window in the data stream and, during the specific control window, to send link width transition data to a device connected to the data link, wherein the link width transition data is said to transmit an attempt to change the number of active tracks on the link. [6] Device according to claim 5, wherein the number of tracks is to be reduced from an original number to a new number, wherein the reduction of the number of active tracks is associated with entering a partial width link transmission state. [7] Device according to claim 6, wherein the bit layer logic further identifies a subsequent control window in the data stream and sends subwidth state exit data to the device during the subsequent control window, wherein the subwidth state exit data transmits an attempt to reset the number of active tracks to the original number. [8] Device according to claim 2, wherein the bit transmission layer logic is further said to identify a specific control window in the data stream and during the specific control window is said to send low-power data to a device connected to the data link, wherein the low-power data is said to transmit an attempt to enter a low-power state from the link transmission state. [9] Device according to claim 1, wherein the control windows are embedded in accordance with a defined control interval and devices connected to the data link are intended to synchronize the state transition with an end of a corresponding control interval. [10] Device comprising the following: a layer stack comprising physical layer logic, data link layer logic, and protocol layer logic, wherein the physical layer logic is at least partially implemented in a hardware circuit and is intended to perform the following: Receiving a data stream, wherein the data stream is to contain alternating transmission intervals and control intervals, wherein during the transmission intervals a series of successive data link layer flits are to be sent, the control intervals are to provide opportunities to send physical layer control information, and the transmission intervals are to be longer than the control intervals; Identifying control data to be included in a specific control interval, wherein the control data is to indicate an attempted entry from a first state into a specific state, and wherein the data stream is to be received in the first state; and Enabling the transition to the specified state. [11] Device according to claim 10, wherein the determined state comprises a reset state. [12] Device according to claim 11, wherein enabling the transition to the specified state includes sending an acknowledgment of the attempted entry into the specified state. [13] Device according to claim 12, wherein the acknowledgment is sent within the control interval. [14] Device according to claim 10, wherein the data stream is sent via a serial data link containing multiple active tracks and wherein the determined state comprises a partial width state, wherein at least a subset of the tracks contained in the multiple active tracks are to be left unused in the partial width state. [15] Device according to claim 14, wherein the bit layer logic is further said to identify subsequent data contained in a subsequent of the control intervals, wherein the subsequent data indicate an attempt to exit the partial width state and reactivate the unused tracks. [16] Device according to claim 10, wherein the specified state comprises a low-power transmission state. [17] Method comprising the following: Receiving a data stream at a device, wherein the data stream contains alternating transmission intervals and control intervals, wherein during the transmission intervals a series of successive data link layer flits are sent, the control intervals provide opportunities to send physical layer control information, and the transmission intervals are intended to be longer than the control intervals; Identifying, using hardware of the device, control data contained in a specific control interval, wherein the control data indicates an attempted entry from a first state into a specific state, the data stream being received in the first state; and Enabling the transition to the specified state. [18] Method according to claim 17, wherein the data stream is received via a serial data link containing multiple active tracks, and wherein the determined state comprises a partial width state, wherein at least a subset of the tracks contained in the multiple active tracks is to be left unused in the partial width state. [19] Method according to claim 17, wherein the determined state comprises a reset state. [20] Method according to claim 17, wherein the bit layer control information describes a fault in the data link. [21] System comprising the following: a first device; and a second device which is communicatively coupled to the first device using a serial data link, wherein the second device contains a bit driver transmission layer module which is executed by at least one processor which is to perform the following: Receiving a data stream, wherein the data stream contains alternating transmission intervals and control intervals, wherein during the transmission intervals a series of successive data link layer flits are sent, the control intervals provide opportunities to send physical layer control information, and the transmission intervals are intended to be longer than the control intervals; Identifying control data contained in a particular control interval, wherein the control data indicates an attempted entry from a first state into a particular state, the data stream being received in the first state; and Enabling the transition to the specified state. [22] System according to claim 21, wherein the first device comprises a microprocessor. [23] System according to claim 21, wherein the second device comprises a second microprocessor. [24] System according to claim 22, wherein the second device comprises a graphics accelerator. [25] System according to claim 21, wherein the first device includes a bit transmission layer logic designed to perform the following: Sending the tax data within the specified tax interval; and Transition to the specified state. [26] System according to claim 21, wherein the determined state is contained in several states of a link training state machine.

Citation Information

Patent Citations

  • Dynamically modulating link width

    US20060034295A1

  • Methods and apparatuses for the physical layer initialization of a link-based system interconnect

    US20060041696A1

  • Simple link protocol providing low overhead coding for LAN serial and WDM solutions

    US7089485B2