TRAINING A VALID LANE
The MCPL addresses the challenge of high-speed communication in complex computing systems by implementing a protocol-independent interconnect solution with efficient data transfer and reduced power consumption, enhancing data rates and reliability across multiple protocols.
Patent Information
- Application Number
- DE112015006953
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2015-09-26
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2035-09-26
AI Technical Summary
As computing systems evolve with increased processing power and complexity, the demand for high-speed communication between components in interconnect architectures has outpaced the capacity of existing interconnects, particularly in high-performance computing environments like servers, where multiple processors and devices require efficient data transfer without significant power consumption.
The implementation of a multi-chip module link (MCPL) with a physical layer (PHY) and logical layer (LPHY) that supports multiple protocols, enabling high-bandwidth, low-latency communication through regulated mid-rail termination, active crosstalk suppression, and clock straightening, along with a modular common physical layer that is protocol-independent, facilitating efficient data transfer across multiple protocols.
The MCPL provides high data rates exceeding 8-10 Gb/s with energy efficiency, supporting multiple protocols and reducing power consumption while maintaining reliable communication between integrated circuits, even over longer channel lengths.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
REGIONThis disclosure relates to computing system and, more particularly (but not exclusively), point-to-point connections.BACKGROUNDAdvances in semiconductor processing and logic design have enabled an increase in the amount of logic that may be present on integrated circuit devices. As a concomitant feature, computing system configurations have evolved from systems having one or more integrated circuits to those having multiple cores, multiple hardware threads, and multiple logic processors that may be present on individual integrated circuits, as well as other interfaces integrated with such processors. A processor or integrated circuit typically comprises a single physical processor die, where the processor die may have any number of cores, hardware threads, logic processors, interfaces, memories, controller hubs, etc.As a result of the increased ability to accommodate more processing power in smaller packages, the popularity of smaller computing devices has increased. Smartphones, tablets, ultra-flat notebooks, and other user devices have grown exponentially. However, these smaller devices rely on servers for both data storage and complex processing operations that exceed the form factor. Consequently, demand on the market for high performance data processing (i.e., server site) has also been increasing. In modern servers, for example, there is generally not only a single processor with multiple cores, but multiple physical processors (also referred to as multiple sockets) in order to increase the computing power. However, as computing power increases along with the number of devices in a computing system, communication between sockets and other devices becomes more critical.Indeed, connections from more traditional multi-drop buses (multi-drop buses) that mainly handle electrical communications have evolved to fully mature connection architectures that facilitate fast communication. Unfortunately, as the demand for future processors that consume at even higher speeds increases, the demand for capacities of existing interconnect architectures correspondingly increases.WO 2014 / 164004 A1 relates to techniques for de-correlating training pattern sequences for high speed links and high speed interconnects that use multiple lanes in each direction to send and receive data and are physically implemented via signal paths in an inter-plane board via a cable. During link training, a training pattern with a pseudo-random bit sequence (PBRS) is sent over each track. The PBRS for each track is generated by a PBRS generator based on a PRB S polynomial unique to that track. Since a different PRBS polynomial is used for each track, the training patterns between the tracks are substantially uncorrelated. A link negotiation may be performed between the endpoints of the link to ensure that the PBRS polynomials used for all lanes of the high speed link are unique.US 2009 / 0019326 A1 relates to a self-synchronizing bit error analyzer and a circuit therefor capable of determining the bit error rate of single data rate (SDR) and double data rate (DDR) data. The self-synchronizing data bus analyzer includes a generator linear feedback shift register (LFSR) for generating a first data set and a receiver LFSR for generating a second data set. The data bus analyzer also includes a bit sampler to sample the first data set received via a data bus coupled to the generator LFSR and output a sampled first data set. A comparator is included to compare the sampled first data set to the second data set generated by the receiver LFSR and provide a signal to the receiver LFSR to adjust a phase of the receiver LFSR until the second data set substantially matches the first data set.SUMMARY OF THE INVENTIONThe solution according to the invention described herein relates to apparatuses according to main claim 1 and dependent claim 15, methods according to dependent claims 25 and 27 and a system according to dependent claim 29.BRIEF DESCRIPTION OF THE DRAWINGSFIG. 1 illustrates an embodiment of a computing system including an interconnect architecture. FIG. 2 illustrates an embodiment of an interconnect architecture including a layer stack. FIG. 3 illustrates an embodiment of a request or packet to be generated or received within a connection architecture. FIG. 4 illustrates an embodiment of a transmitter and receiver pair for a link architecture. FIG. 5 illustrates an embodiment of a multi-chip module. FIG. 6 is a simplified block diagram of a multichip module link (MCPL). FIG. 7 is an illustration of an example signal transmission on an example MCPL. Figure 8 is a simplified block diagram of an MCPL. FIG. 9 is an illustration of signal transmission to enter a power saving state. FIG. 10 is a diagram of a portion of an example link state machine. FIG. 11 is an illustration of an example link state machine. FIG. 12 is an illustration of example signal transmission on an example MCPL. FIG. 13 is a block diagram of an example offset balanced clock tree. FIG. 14 is an illustration of an example signaling on an example MCPL during link training, according to one embodiment. FIG. 15 is a block diagram of an example linear feedback shift register (LFSR). FIG. 16 is an illustration of an example signaling on an example MCPL during link training, according to one embodiment. FIG. 17 is a block diagram of another example linear feedback shift register (LFSR). FIG. 18 is a table illustrating an example implementation of a non-correlated pseudo-random binary sequence generated from an example LFSR. FIG. 19 is a block diagram illustrating a bumpout layout of two devices connected by an example MCPL. FIG. 20 is a table illustrating an example implementation of crosstalk stress tests. FIG. 21 is a table illustrating an example implementation of crosstalk stress tests to include testing a valid lane of an example MCPL. FIGS. 22A-22B are simplified flow diagrams illustrating techniques for training a valid lane of an MCPL. FIG. 23 illustrates an embodiment of a block for a computing system that includes multiple processors.Like reference numerals and labels in the various drawings indicate like elements.DETAILED DESCRIPTIONIn the following description, numerous specific details are set forth, such as examples of specific types of processors and system configurations, specific hardware structures, specific architectural and microarchitectural details, specific register configurations, specific instruction types, specific system components, specific dimensions / heights, specific processor pipeline stages and operation, etc., to provide a thorough understanding of the present invention. It will be apparent, however, to one skilled in the art that these specific details need not necessarily be employed to practice the present invention. In other instances, well-known components or methods, such as specific and alternative processor architectures, specific logic circuitry / code for described algorithms, specific firmware code, specific interconnect operation, specific logic configurations, specific fabrication techniques and materials, specific compiler implementations, specific conversion of algorithms into code, specific shutdown and gating techniques / logic, and other specific operational details of computing system(s), have not been described in detail to avoid unnecessarily obscuring the present invention.Although the following embodiments may be described with respect to power saving and power efficiency in specific integrated circuits such as computer platforms or microprocessors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of embodiments described herein may be applied to other types of circuits or semiconductor devices that may also benefit from higher energy efficiency and energy saving. For example, the disclosed embodiments are not limited to desktop computing systems or Ultrabooks™ but may also be applied in other devices such as portable devices, tablets, other flat notebooks, systems on a chip (SOC) devices, and embedded applications. Some examples of wearable devices include cellular phones, Internet protocol devices, digital cameras, personal digital assistants (PDAs), and wearable PCs. Embedded applications typically include a microcontroller, a digital signal processor (DSP), a system on a chip, network computers (NetPC), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of performing the functions and operations taught below. Moreover, the apparatuses, methods, and systems described herein are not limited to physical computing devices, but may also relate to software optimizations for energy conservation and efficiency. As will be readily apparent from the description below, the embodiments of the methods, devices, and systems described herein (whether in terms of hardware, firmware, software, or a combination thereof) are critical to a future "green technology" and are balanced with performance considerations.As computer systems continue to develop, the components included therein become more and more complex. As a result, the complexity of the interconnect architecture for coupling the components and communicating therebetween also increases to ensure that the demand for bandwidth is satisfied for optimum operation of the components. Further, different market segments require different aspects of interconnect architectures to meet the needs of the market. For example, servers require higher power while the mobile ecosystem is sometimes able to sacrifice overall power for power savings. However, a singular purpose of most fabrics is to provide the highest possible performance with maximum power saving. Discussed below is a number of connections that could potentially benefit from aspects of the invention described herein.One interconnect fabric architecture includes the Peripheral Component Interconnect (PCI) Express (PCIe) architecture. A primary goal of PCIe is to enable components and devices from different suppliers to operate in an open architecture spanning multiple market segments; clients (desktop and mobile), servers (standard and enterprise), and embedded and communication devices. PCI Express is a high performance universal I / O link defined for a wide range of future computing and communication platforms. Some PCI attributes, such as their usage model, load store architecture, and software interfaces, have been maintained at their revisions, while previous parallel bus implementations have been replaced with a highly scalable, full serial interface. The more recent versions of PCI Express make use of advances in point-to-point connections, switch-based technology, and packetized protocol to provide new levels of performance and functions. More advanced functions supported by PCI Express include power management, quality of service (QoS), hot plug / hot swap support, data integrity, and error handling.Referring to FIG. 1, an embodiment of a fabric composed of point-to-point links that connect a set of components is illustrated. The system 100 includes a processor 105 and a system memory 110 coupled to a controller hub 115. The processor 105 includes any processing element such as a microprocessor, a host processor, an embedded processor, a co-processor, or another processor. The processor 105 is coupled to the controller hub 115 through a front side bus (FSB) 106. In one embodiment, the FSB 106 is a serial point-to-point connection, as described below. In another embodiment, link 106 includes a serial differential interconnect architecture that conforms to different interconnect standards. One or more components of system 100 may be provided with logic to implement the features described herein.System memory 110 includes any storage device such as random access memory (RAM), nonvolatile (NV) memory, or other memory accessible from devices in system 100. System memory 110 is coupled to controller hub 115 via memory interface 116. Examples of a memory interface include a dual data rate (DDR) memory interface, a dual channel DDR memory interface, and a dynamic RAM (DRAM) memory interface.In one embodiment, controller hub 115 is a root hub, root complex, or root controller in a peripheral component interconnect express (PCIe or PCIE) connection hierarchy. Examples of a controller hub 115 include a chipset, a memory controller hub (MCH), a north bridge, an interconnect controller hub (ICH), a south bridge, and a root controller / hub. The term chipset often refers to two physically separate controller hubs, i.e. a memory controller hub (MCH) coupled to an interconnect controller hub (ICH). It should be appreciated that current systems often include the MCH integrated with the processor 105, while the controller 115 is to communicate with I / O devices in a similar manner as described below. In some embodiments, peer-to-peer (peer-to-peer) routing is optionally supported by a root complex 115.Here, the controller hub 115 is coupled to a / a switch / bridge 120 by a serial link 119. The input / output modules 117 and 121, which may also be referred to as interfaces / ports 117 and 121, include / implement a layered protocol stack to provide communication between the controller hub 115 and the switch 120. In one embodiment, multiple devices are capable of being coupled to the switch 120.The switch / bridge 120 directs packets / messages from the device 125 upstream, i.e., an up hierarchy to a root complex, to the controller hub 115, and downstream, i.e., a down hierarchy away from a root controller, from the processor 105, or from the system memory 110, to the device 125. The switch 120 is referred to as a logical assembly of multiple PCI-to-PCI virtual bridge devices, in one embodiment. The device 125 includes any internal or external device or component to be coupled to an electronic system, such as an I / O device, a network interface controller (NIC), an add-in card (add-in card), an audio processor, a network processor, a hard disk drive, a storage device, a CD / DVD ROM, a screen, a printer, a mouse, a keyboard, a router, a portable storage device, a firewire device, a universal serial bus (USB) device, a scanner, and other input / output devices. In the PCIe art jargon, such a device is often referred to as an endpoint. Although not specifically shown, device 125 may include a PCIe-to-PCI / PCI-X bridge to support legacy PCI devices or other versions of PCI devices. Endpoint devices in PCIe are often referred to as legacy, PCIe, or root complex integrated endpoints.A graphics accelerator 130 is also coupled to the controller hub 115 through the serial link 132. In one embodiment, graphics accelerator 130 is coupled to an MCH, which is coupled to an ICH. The switch 120, and accordingly the I / O device 125, is then coupled to the ICH. I / O modules 131 and 118 also implement a layered protocol stack to communicate between graphics accelerator 130 and controller hub 115. Similar to the discussion of the MCH above, a graphics controller or graphics accelerator 130 itself may be integrated with processor 105.Referring to FIG. 2, an embodiment of a layered protocol stack is illustrated. The layer protocol stack 200 includes any form of layer communication stack, such as a quick path interconnect (QPI) stack, a PCIe stack, a next generation high performance data processing interconnect stack, or another layer stack. Although the discussions immediately below with reference to FIGS. 1-4 relate to a PCIe stack, the same concepts may be applied to other interconnect stacks. In one embodiment, protocol stack 200 is a PCIe protocol stack including a transaction layer 205, a link layer 210, and a physical layer 220. An interface, such as interfaces 117, 118, 121, 122, 126, and 131 in FIG. 1, may be depicted as communication protocol stack 200. A representation as a communication protocol stack may also be referred to as a module or interface that implements / includes a protocol stack.PCI Express uses packets to communicate information between components. The packets are formed in the transaction layer 205 and the data link layer 210 to transport the information from the sending component to the receiving component. As the transmitted packets flow through the other layers, they are augmented by additional information necessary to handle packets at these layers. At the receiving end, the opposite process occurs and packets are converted from the representation of their physical layer 220 to the representation of the data link layer 210 and finally (for transaction layer packets) to the form that can be processed by the transaction layer 205 of the receiving device.Transaction LayerIn one embodiment, transaction layer 205 is to provide an interface between a device's processor core and the interconnect architecture such as data link layer 210 and physical layer 220. In this regard, the major responsibility of the transaction layer 205 is assembly and disassembly of packets (i.e., transaction layer packets or TLPs). The translation layer 205 typically manages credit-based flow control for TLPs. PCIe implements split transactions, i.e., transactions where request and response are separated by time, which allows a link to carry other traffic while the target device collects data for the response.In addition, PCIe uses credit-based flow control (PCIe). In this scheme, a device offers an initial set of credit for each of the receive buffers in the transaction layer 205. An external device at the opposite end of the link, such as controller hub 115 in FIG. 1, counts the number of credits consumed by each TLP. A transaction may be sent when the transaction does not exceed a credit limit. Upon receiving a response, a credit amount is recovered. An advantage of such a credit scheme is that the latency of credit return does not affect performance provided that the credit limit is not reached.In one embodiment, four transaction address spaces include a configuration address space, a memory address space, an input / output address space, and a message address space. Memory space transactions include one or more of read requests and write requests to transfer data to or from a memory-allocated location. In one embodiment, memory space transactions are capable of using two different address formats, e.g., a short address format such as a 32-bit address or a long address format such as a 64-bit address. Configuration space transactions are used to access the configuration space of the PCIe devices. Transactions to configuration space include read requests and write requests. Message space transactions (or simply messages) are defined to support in-band communication between PCIe agents.Thus, in one embodiment, transaction layer 205 assembles packet header / payload 206. Formats for current packet headers / payloads can be found in the PCIe specification on the PCIe specification web page.Referring now briefly to FIG. 3, an embodiment of a PCIe transaction descriptor is illustrated. In one embodiment, transaction descriptor 300 is a mechanism for carrying transaction information. In this regard, transaction descriptor 300 supports the identification of transactions in a system. Other possible uses include tracking modifications of standard transaction ordering and assigning the transaction to channels.Transaction descriptor 300 includes global identifier field 302, attribute field 304, and channel identifier field 306. In the illustrated example, the global identifier field 302 is depicted as including the local transaction identifier field 308 and the source identifier field 310. In one embodiment, the global transaction identifier 302 is unique to all outstanding requests.According to one implementation, the local transaction identifier field 308 is a field generated by a requesting agent and is unique to all outstanding requests that require completion for that requesting agent. Further, in this example, the source identifier 310 uniquely identifies the requesting agent within a PCIe hierarchy. Accordingly, the local transaction identifier field 308 along with the source ID 310 provides a global identification of a transaction within a hierarchy domain.The attribute field 304 specifies properties and relationships of the transaction. In this regard, the attribute field 304 is potentially used to provide additional information that allows modification of the default handling of transactions. In one embodiment, attribute field 304 includes priority field 312, reserved field 314, ordering field 316, and non-listening field 318. Here, priority sub-field 312 may be modified by an initiator to assign a priority to the transaction. Reserved attribute field 314 remains reserved for future or supplier-defined use. Possible usage models using priority or security attributes may be implemented using the reserved attribute field.In this example, the ordering attribute field 316 is used to provide optional information conveying the type of ordering that can modify the default ordering rules. According to an example implementation, an order attribute "0" denotes that default order rules are to be applied, whereas an order attribute "1" denotes a relaxed order, wherein writes may pass writes in the same direction and read access executions may pass writes in the same direction. The intercept free attribute field 318 is used to determine whether transactions are intercepted. As shown, the channel ID field 306 identifies a channel to which a transaction is associated.Link LayerLink layer 210, also referred to as data link layer 210, acts as an intermediate between transaction layer 205 and physical layer 220. In one embodiment, a responsibility of the data link layer 210 is to provide a reliable mechanism for exchanging transaction layer packets (TLPs) between two components of a link. A data link layer 210 side accepts TLPs that have been merged by the transaction layer 205, applies the packet sequence identifier 211, i.e., an identification number or a packet number, calculates and applies an error detection code, i.e., CRC 212, and submits the modified TLPs to the physical layer 220 for transmission over a physical to an external device.Physical layerIn one embodiment, physical layer 220 includes a logical sub-block 221 and an electrical sub-block 222 to physically send a packet to an external device. Here, the logical sub-block 221 is responsible for the "digital" functions of the physical layer 221. In this regard, the logical sub-block includes a transmit portion to prepare outgoing information for transmission by the physical sub-block 222, and a receiver portion to identify and prepare received information before forwarding it to the link layer 210.The physical block 222 includes a transmitter and a receiver. Symbols are supplied to the transmitter from logical sub-block 221 which the transmitter serializes and transmits to an external device. The receiver is supplied with serialized symbols from an external device and converts the received signals into a bit stream. The bitstream is deserialized and provided to logical sub-block 221. In one embodiment, an 8b / 10b transmission code is employed, wherein ten-bit symbols are transmitted / received. Here, specific symbols are used to form a packet with the frames 223. Additionally, in one example, the receiver also provides a symbol clock recovered from the incoming serial stream.Although the transaction layer 205, link layer 210, and physical layer 220 are discussed with respect to a specific embodiment of a PCIe protocol stack as noted above, a layer protocol stack is not so limited. Indeed, any layer protocol may be included / implemented. As an example, a port / interface, represented as a layer protocol, includes: (1) a first layer to assemble packets, i.e., a transaction layer; a second layer to order packets, i.e., a link layer; and a third layer to send the packets, i.e., a physical layer. As a specific example, a common system interface (CSI) layer protocol is used.Referring next to FIG. 4, an embodiment of a PCIe serial point-to-point fabric is illustrated. Although an embodiment of a PCIe point-to-point serial link is illustrated, a point-to-point serial link is not so limited as it includes any transmission path for transmitting serial data. In the embodiment shown, a single PCIe link includes two differentially driven low voltage signal pairs: a transmitter pair 406 / 411 and a receiver pair 412 / 407. Accordingly, device 405 includes transmit logic 406 to transmit data to device 410 and receive logic 407 to receive data from device 410. In other words, a PCIe link includes two transmit paths, i.e., paths 416 and 417, and two receive paths, i.e., paths 418 and 419.A transmission path refers to each path for transmitting data, such as a transmission line, a copper line, an optical line, a wireless communication channel, an infrared communication link, or another communication path. A connection between two devices, such as device 405 and device 410, is referred to as a link, such as link 415. A link may support a lane - each lane representing a set of differential signal pairs (a pair for transmission, a pair for reception). To scale bandwidth, a link may accumulate multiple lanes denoted xN, where N is any supported link width, such as 1, 2, 4, 8, 12, 16, 32, 64, or wider.A differential pair refers to two transmission paths, such as lines 416 and 417, to transmit differential signals. As an example, when line 416 switches from a low voltage level to a high voltage level, i.e., a rising edge, line 417 is controlled from a high logic level to a low logic level, i.e., a falling edge. Differential signals potentially exhibit better electrical properties, such as better signal integrity, i.e., cross coupling, voltage overshoot / undershoot, ringing, etc. This allows for a better time window that allows for faster transmission frequencies.FIG. 5 is a simplified block diagram 500 illustrating an example multi-chip module 505 including one or more chips or dies (e.g., 510, 515) communicatively interconnected using an example multi-chip module link (MCPL) 520. While FIG. 5 illustrates an example of two (or more) dies interconnected using an example MCPL 520, it should be appreciated that the principles and features described herein regarding implementations of an MCPL may be applied to any connection or link connecting a die (e.g., 510) and other components, including connecting two or more dies (e.g., 510, 515), connecting a die (or chip) to another component outside the die, connecting a die to another device or die outside the package (e.g., 505), connecting a die to a BGA package, implementation of a patch-on-interposer (POINT), among potentially other examples.Generally, a multi-chip module (e.g., 505) may be an electronic module in which multiple integrated circuits (ICs), semiconductor dies, or other discrete components (e.g., 510, 515) are packaged onto an interconnecting substrate (e.g., silicone or other semiconductor substrate), which allows the combined components to be used as a single component (e.g., as by a larger IC). In some cases, the larger components (e.g., dies 510, 515) may themselves be IC systems, such as systems on chip (SoC), multiprocessor chips, or other components that include multiple components (e.g., 525- 530 and 540- 545) on the device, for example, on a single die (e.g., 510, 515). Multi-chip modules 505 may provide flexibility to build complex and varied systems from potentially multiple discrete components and systems. For example, each of the dies 510, 515 may be manufactured or otherwise provided by two different devices, with the silicone substrate of the package 505 provided by yet a third device, among many other examples. Further, dies and other components within a multi-chip module 505 may themselves include a connection or other communication fabric (e.g., 535, 550), thereby providing the infrastructure for inter-component communication (e.g., 525-530 and 540-545) within the device (e.g., 510, 515, respectively). The various components and connections (e.g., 535, 550) may potentially support or use multiple different protocols. Further, communication between dies (e.g., 510, 515) may potentially include transactions between the various components on the dies via multiple different protocols. Designing mechanisms for providing inter-chip (or die) communication on a multi-chip module may present a challenge, where conventional solutions that employ highly specialized, expensive, and package specific solutions rely on the specific combinations of components (and desired transactions) to be connected.The examples, systems, algorithms, devices, logic, and functions described within this specification may address at least some of the problems identified above, including potentially many others not expressly mentioned herein. For example, in some implementations, a high bandwidth, low power, and low latency interface may be provided to connect a host device (e.g., a CPU) or other device to a companion chip residing in the same module as the host. Such a multichip module link (MCPL) may support multiple module options, multiple I / O protocols, and reliability, availability, and serviceability (RAS) functions. Further, the physical layer (PHY) may include an electrical layer and a logical layer and may support longer channel lengths, including channel lengths up to approximately 45 mm and in some cases beyond. In some implementations, an example MCPL may operate at high data rates, including data rates above 8-10Gb / s.In an example implementation of an MCPL, a PHY electrical layer may improve traditional multi-channel connection solutions (e.g., multi-channel DRAM I / O), thereby expanding data rate and channel configuration, for example, by a number of functions, for example, regulated mid-rail termination, active low-energy crosstalk suppression, circuit redundancy, per-bit duty cycle correction and clock straightening, row coding, and transmitter compensation, among potentially other examples.In an example implementation of an MCPL, a PHY logical layer may be implemented that may further contribute to data rate expansion and channel configuration expansion (e.g., electrical layer functions) and also activates the connection to route multiple protocols across the electrical layer. Such implementations may provide and define a modular common physical layer that is protocol independent and configured to operate with potentially any existing or future connection protocol.Referring to FIG. 6, a simplified block diagram 600 is shown illustrating at least a portion of a system illustrating an example implementation of a multi-chip module link (MCPL). An MCPL may be implemented using physical electrical connections (e.g., wires implemented as lanes) that connect a first device 605 (e.g., a first die including one or more subcomponents) to a second device 610 (e.g., a second die including one or more subcomponents). In the particular example shown in the overview of diagram 600, all signals (in channels 615, 620) may be unidirectional and lanes may be provided for the data signals to include both upstream and downstream data transmission. While the block diagram 600 of FIG. 6 refers to the first component 605 as the upstream component and the second component 610 as the downstream component, and to the physical lanes of the MCPL used to send data as a downstream channel 615 and lanes used to receive data (from component 610) as an upstream channel 620, it should be understood that the MCPL between the devices 605, 610 may be used by each device to both send and receive data between the devices.In an example implementation, an MCPL may provide a physical layer (PHY) including MCPL PHY 625a,b (or shared 625) and executable logic implementing MCPL PHY 630a,b (or shared 630). The electrical or physical PHY 625 may provide the physical connection over which data is communicated between the devices 605, 610. Signal conditioning components and logic may be implemented in conjunction with physical PHY 625 to establish a high data rate and channel configuration capability of the link, which in some applications may include closely grouped physical connections of lengths of approximately 45 mm or more. The logical PHY 630 may include logic to enable timing, link state management (e.g., for link layers 635 a, 635 b), and protocol multiplexing between potentially multiple, different protocols used for communication over the MCPL.In an example implementation, the physical PHY 625 may include, for each channel (e.g., 615, 620), a set of data lanes over which in-band data may be transmitted. In this particular example, 50 data lanes are provided in each of the upstream and downstream channels 615, 620, although any other number of lanes may be used, such as permitted by layout and performance constraints, desired applications, device constraints, etc. Each channel may further include one or more dedicated lanes for a strobe or clock signal for the channel, one or more dedicated lanes for a valid signal for the channel, one or more dedicated lanes for a stream signal, and one or more dedicated lanes for link state machine management or a sideband signal. The physical PHY may further include a sideband link 640, which in some examples may be a lower frequency bidirectional control signal link used to coordinate state transitions and other attributes of the MCPL connecting the devices 605, 610, among other examples.In some implementations, in-band data (and other data) sent over the MCPL may be encrypted. In one example, the data on each lane may be encrypted using a pseudo random binary sequence (PRBS). In some implementations, the PRBS may be generated to be scrambled with outgoing data using a linear feedback shift register (LFSR). A receiving device may decrypt the data to clearly see the data, among other examples.As noted above, multiple protocols may be supported using an implementation of an MCPL. Indeed, multiple independent transaction layers 650a, 650b may be provided on each device 605, 610. For example, each device 605, 610 may support and use two or more protocols such as, but not limited to, PCI, PCIe, QPI, Intel In-Die Interconnect (IDI). IDI is a coherent protocol used on the die for communicating between cores, last level caches (LLCs), memory, graphics, and IO controllers. Other protocols, including Ethernet protocol, infiniband protocols, and other PCIe fabric-based protocols may also be supported. The combination of the logical PHY and the physical PHY may also be used as a die-to-die link to connect a SerDes PHY (PCIe, Ethernet, Infiniband, or other high speed SerDes) on one die to their upper layers implemented on the other die, among other examples.The logical PHY 630 may support multiplexing between these multiple protocols on an MCPL. For example, its own stream lane may be used to activate an encoded stream signal that identifies which protocol to apply to data transmitted substantially in parallel on the channel's data lanes. Further, the logical PHY 630 may be used to negotiate the various types of link state transitions that the different protocols may support or request. In some cases, LSM_SB signals transmitted over the channel's own LSM_SB lane may be used in conjunction with sideband link 640 to communicate and negotiate link state transitions between devices 605, 610. Further, link training, fault detection, skew detection, clock straightening, and other functionalities of conventional connections may be replaced or regulated, in part using logical PHY 630. For example, valid signals transmitted over one or more dedicated valid signal lanes in each channel may be used to signal link activity, detect clock offsets and link errors, and implement other functions, among other examples. In the specific example of FIG. 6, multiple valid lanes are provided per channel. For example, data lanes may be bundled or grouped (physical and / or logical) within a channel, and a valid lane may be provided for each group. Further, multiple strobe lanes may be provided, in some cases also to provide a separate strobe signal for each group in multiple data lane groups in a channel, among other examples.As mentioned above, the logical PHY 630 may be used to negotiate and manage link control signals sent between devices connected by the MCPL. In some implementations, the logical PHY 630 may include link layer packet (LLP) generation logic 660, which may be used to send link layer control messages over the MCPL (i.e., in-band). Such messages may be sent over data lanes of the channel, where the data stream lane identifies that the data is a link layer-to-link layer notification, such as link layer control data, among other examples. Link layer messages enabled using the LLP module 660 may aid in negotiation and execution of link layer transitions, power management, loops, disabling, re-centering, encrypting, among other link layer functions between the link layers 635 a, 635 bof the respective devices 605, 610.Referring to FIG. 7, a diagram 700 is shown illustrating an example signaling using a set of lanes (e.g., 615, 620) in a particular channel of an example MCPL. In the example of FIG. 7, two groups of twenty five (25) data lanes are provided for a total of fifty (50) data lanes in the channel. A portion of the lanes is shown, while others (e.g., DATA[4-46] and a second strobe signal lane (STRB)) have been omitted (e.g., as redundant signals) for ease of illustrating the specific example. When the physical layer is in an active state (e.g., not powered down or in a power saving mode (e.g., an L1 state)), strobe lanes (STRB) may be provided with a synchronous clock signal. In some implementations, data may be sent on both the rising and falling edges of the strobe. Each edge (or half clock cycle) may delimit a unit interval (UI). Accordingly, in this example, one bit (e.g., 705) may be sent on each lane, allowing one byte to be sent every 8UI. A byte time period 710 may be defined as 8UI, or the time to send a byte on a single one of the data lanes (e.g., DATA[0-49]).In some implementations, a valid signal transmitted on one or more dedicated valid signal channels (e.g., VALID0, VALID1) may serve as a leading indicator for the receiving device to identify, when enabled (high), for the receiving device or sink, that data is transmitted from the transmitting device or source on data lanes (e.g., DATA[0-49]) during the following time period, such as byte time period 710. Alternatively, if the valid signal is low, the source indicates to the sink that the sink will not send data on the data lanes in the following time period. Accordingly, if the logical PHY of the sink detects that the valid signal is not asserted (e.g., on lanes VALID0 and VALID1), the sink may ignore any data detected on the data lanes (e.g., DATA[0-49]) during the following time period. For example, cross-talk sounds or other bits may occur on one or more of the data lanes when the source is not actually sending data. Due to a low or non-asserted valid signal during the previous time period (e.g., the previous byte time period), the sink may determine that the data lanes are to be ignored during the following time period.Data transmitted on one of the lanes of the MCPL may be aligned exactly with the strobe signal. A time period may be defined based on the strobe, such as a byte time period, and each of these periods may correspond to a defined window in which signals are to be transmitted on the data lanes (e.g., DATA[0-49]), the valid lanes (e.g., VALID1, VALID2), and the stream lane (e.g., STREAM). Accordingly, alignment of these signals may allow identification that a valid signal in a previous time period window applies to data in the following time period window and that a stream signal applies to data in the same time period window. The stream signal may be an encoded signal (e.g., 1 byte of data for a byte time period window) encoded to identify the protocol applicable to data transmitted during the same time period window.For illustrative purposes, in the specific example of Figure 7, a byte time period window is defined. A valid is activated at time period window n (715) before any data is applied to data lanes DATA[0-49]. In the following time period window n+1(720), data is sent on at least some of the data lanes. In this case, data is transmitted on all fifty data lanes during n+1(720). Because a valid has been activated for the duration of the previous time period window n (715), the sink device may validate the data received on the data lanes DATA[0-49] during time period window n+1 (720). In addition, the leading nature of the valid signal during the time period window n (715) allows the receiving device to prepare for the incoming data. Continuing with the example of FIG. 7, the valid signal (on VALID 1 and VALID 2) remains asserted during the duration of time period window n+1(720), causing the sink device to expect the data sent over data lanes DATA[0-49] during time period window n+2(725). If the valid signal remained activated during time period window n+2(725), the sink device could still expect to receive (and process) additional data that would be sent during an immediately subsequent time period window n+3(730). However, in the example of FIG. 7, the valid signal is deasserted during the duration of time period window n+2(725), indicating to the sink device that no data is being sent during time period window n+3(730), and that any bits detected on data lanes DATA[0-49] are to be ignored during time period window n+3(730).As noted above, multiple valid lanes and strobe lanes may be maintained per channel. This may assist in maintaining circuit simplicity and synchronization among the groups of relatively long physical lanes connecting the two devices, among other advantages. In some implementations, a set of data lanes may be divided into groups of data lanes. For example, in the example of FIG. 7, the data lanes DATA[0-49] may be divided into two groups of twenty five lanes each, and each group may have its own valid and strobe lanes. For example, valid lane VALID1 may be assigned to data lanes DATA[0-24], and valid lane VALID2 may be assigned to data lanes DATA[25-49]. The signals on each "copy" of the valid and strobe lanes may be identical for each group.As introduced above, data on the stream lane STREAM may be used to indicate to the receiving logical PHY which protocol is to apply to corresponding data transmitted on the data lanes DATA[0-49]. In the example of FIG. 7, a stream signal is sent on STREAM during the same time period window as data on the data lanes DATA[0-49] to indicate the protocol of the data on the data lanes. In alternative implementations, the stream signal may be transmitted during a previous time period window, such as corresponding valid signals, among other possible modifications. However, continuing with the example of FIG. 7, a stream signal 735 is transmitted during time period window n+1 (720) encoded to indicate the protocol (e.g., PCIe, PCI, IDI, QPI, etc.) that applies to the bits transmitted over data lanes DATA[0-49] during time period window n+1 (720). Similarly, another stream signal 740 may be sent during the subsequent time period window n+2 (725) to indicate the protocol applicable to the bits sent over the data lanes DATA[0-49] during time period window n+2 (725), and so on. In some cases, such as the example of FIG. 7 (where both stream signals 735, 740 have the same binary FF coding), data in successive time period windows (e.g., n+1(720) and n+2(725)) may belong to the same protocol. However, in other cases, data in successive time period windows (e.g., n+1(720) and n+2(725)) may originate from different transactions for which different protocols are to be applicable, and stream signals (e.g., 735, 740) may be appropriately encoded to identify the different protocols that are applicable to the successive bytes of data on the data lanes (e.g., DATA[0-49]), among other examples.In some implementations, a power save or sleep state may be defined for the MCPL. For example, if no device on the MCPL is transmitting data, the physical layer (electrical and logic) of the MCPL may transition to a sleep or power saving state. For example, in the example of FIG. 7, the MCPL in time period window n-2 (745) is in a sleep or sleep state and the strobe is disabled to save power. The MCPL may exit from the power save or idle state and reactivate the strobe at time period windows n-1 (e.g., 705). The strobe may complete a transmit preamble (e.g., to assist in re-activating and synchronizing each lane of the channel as well as the sink device) by starting the strobe signal before every other signal transmission on the other non-strobe lanes. After this time period window n-1 (705), as discussed above, the valid signal on time period window n (715) may be activated to notify the sink that data is pending in the following time period window n+1 (720).The MCPL may reenter a low power or idle state (e.g., an L1 state) after detecting idle conditions on the valid lanes, the data lanes, and / or other lanes on the MCPL channel (or, simply speaking, "MCPL"). For example, no signal transmission can be detected starting at time period window n+3 (730) and then proceeding forward. Among other examples and principles (including those discussed later herein), logic on either the source or target sink devices may initiate a transition back to a low power state, which in turn (e.g., in time period window n+5 (755)) results in the strobe transitioning to a low power state.The electrical features of the physical PHY may include one or more of asymmetric signal transmission, forward half-frequency clocking, matching delays on the interconnect channel, and on-chip transport of transmitter (source) and receiver (sink), optimized electrostatic discharge (ESD) protection, pad capacitance, among other functions. Further, an MCPL may be implemented to achieve a higher data rate (e.g., approximately 16 Gb / s) and energy efficiency features than conventional module I / O solutions.Referring to FIG. 8, a simplified block diagram 800 is shown illustrating an example logical PHY of an example MCPL. A physical PHY 805 may be connected to a die that includes a logical PHY 810 and additional logic that supports a link layer of the MCPL. The die may further include logic to support multiple different protocols on the MCPL in this example. For example, in the example of FIG. 8, PCIe logic 815 and IDI logic 820 may be provided so that the dies may communicate using either PCIe or IDI over the same MCPL connecting the two dies, among potentially many other examples, including examples in which more than two protocols or protocols different from PCIe and IDI are supported over the MCPL. Different protocols supported between the dies may provide different levels of service and functionality.The logical PHY 810 may include link state machine management logic 825 to negotiate link state transitions associated with requests from the die's higher logic (e.g., received via PCIe or IDI). The logical PHY 810 may further include link test and debug logic (e.g., 830) in some implementations. As noted above, an example MCPL may support control signals sent between dies over the MCPL to enable protocol independence, high performance, and energy efficiency functions (among other example functions) of the MCPL. For example, logical PHY 810 may support the generation and transmission, as well as the reception and processing of valid signals, stream signals, and LSM sideband signals in conjunction with the transmission and reception of data over its own data lanes, as described in the examples above.In some implementations, multiplexer (e.g., 835) and demultiplexer (e.g., 840) logic may be included in or otherwise accessible to the logical PHY 810. For example, multiplexer logic (e.g., 835) may be used to identify data (e.g., embodied as packets, messages, etc.) to be sent on the MCPL. Multiplexer logic 835 may identify the protocol that regulates the data and generate a stream signal encoded to identify the protocol. For example, in an example implementation, the stream signal may be encoded as a byte of two hexadecimal digits (e.g., IDI: FFh; PCIe: F0h; LLP: AAh; sideband: 55h; etc.), and may be transmitted during the same window (e.g., a byte time period window) of the data regulated by the identified protocol. Similarly, demultiplexer logic 840 may be employed to interpret incoming data stream signals to decode the data stream signal and identify the protocol to apply to data received in parallel with the data stream signal on the data lanes. Demultiplexer logic 840 may then apply (or ensure) protocol specific processing by the link layer and cause the data to be processed by the corresponding protocol logic (e.g., PCIe logic 815 or IDI logic 820).The logical PHY 810 may further include link layer packet logic 850 that may be used to handle various link control functions, including power management tasks, loop switching (loopback), disabling, re-centering, encrypting, etc. LLP logic 850 may enable link layer-to-link layer messages via the MCLP, among other functions. Data corresponding to the LLP signaling may also be identified from a stream signal transmitted on a separate stream signal lane encoded to identify that the data lanes are LLP data. Multiplexer and demultiplexer logic (e.g., 835, 840) may also be used to generate and interpret the data stream signals corresponding to LLP traffic as well as to cause such traffic to be processed by the appropriate die logic (e.g., LLP logic 850). Similarly, in further implementations, an MCLP may include its own sideband (e.g., sideband 855 and supporting logic), such as a lower frequency asynchronous and / or sideband channel, among other examples.The logical PHY logic 810 may further include link state machine management logic that may generate and receive (and use) link state management messaging over its own LSM sideband lane. For example, an LSM sideband lane may be used to perform handshake to advance the link training state, leave power management states (e.g., an L1 state), among other possible examples. The LSM sideband signal may be an asynchronous signal in that it is not aligned with the link's data, valid, and stream signals, but instead corresponds to signal transfer state transitions, and among other examples, the link state machine aligns between the two dies or chips connected by the link. Providing a dedicated LSM sideband lane may allow, among other example advantages, conventional noise blocking and received detection circuits of an analog front end (AFE) to be eliminated, among other example advantages.FIG. 9 is a simplified block diagram 900 illustrating an example flow in a transition between an active state (e.g., L0) and a low power or quiescent state (e.g., L1). In this specific example, a first device 905 and a second device 910 are communicatively coupled using an MCPL. In the active state, data is transmitted over the MCPL lanes (e.g., DATA, VALID, STREAM, etc.). Link layer packets (LLPs) may be communicated over the lanes (e.g., data lanes, where the stream signal indicates that the data is LLP data) to assist in facilitating link state transitions. For example, LLPs may be sent between the first and second devices 905, 910 to negotiate entry of L0 into L1. For example, higher layer protocols supported by the MCPL may communicate that entry into L1 (or other state) is desired, and the higher layer protocols may cause LLPs to be sent over the MCPL to enable a link layer handshake that causes the physical layer to enter L1. For example, FIG. 9 shows at least a portion of transmitted LLPs including an "ingress-to-L1" request LLP transmitted from the second (upstream) device 910 to the first (downstream) device 905. In some higher layer implementations and protocols, the downstream port does not initiate entry into L1. The receiving first device 905 may transmit a change-to-L1" request LLP in response, which the second device 910 may acknowledge, by a change-to-L1" acknowledgement (ACK) LLP, among other examples. Upon detecting the completion of the handshake, the logical PHY may cause a sideband signal on its own sideband link to be activated to confirm that the ACK has been received and that the device (e.g., 905) is ready to enter and expect L1. For example, the first device 905 may activate a sideband signal 915 sent to the second device 910 to acknowledge receipt of the last ACK in the link layer handshake. Additionally, the second device 910 may also activate a sideband signal in response to the sideband signal 915 to notify the first device 905 of the sideband ACK 905 of the first device. When link layer control and sideband handshakes are complete, the MCPL PHY may be forced to the L1 state, causing all lanes of the MCPL to be set to the idle / low power state mode, including respective MCPL strobes of 920, 925 of devices 905, 910. The L1 may be exited when the higher layer logic of one of the first and second devices 905, 910 requests re-entry into L0, for example, in response to detecting data to be sent to the other device via the MCPL.As noted above, in some implementations, an MCPL may enable communication between two devices that potentially supports multiple different protocols, and the MCPL may enable communications according to potentially any of the multiple protocols over the MCPL lanes. However, facilitating multiple protocols may complicate entry and re-entry into at least some link states. For example, while some conventional connections have a single higher layer protocol that takes the role of the master during state transitions, in an implementation of MCPL with multiple different protocols, multiple masters are effectively present. As an example, as shown in FIG. 9, each of PCIe and IDI may be supported between two devices 905, 910 via an implementation of an MCPL. For example, setting the physical layer to a sleep or low power state may be conditioned on the first-obtained permission from each of the supported protocols (e.g., both PCIe and IDI).In some examples, entry into L1 (or another state) may be requested by only one of the many supported protocols supported for an implementation of an MCPL. While there may be a probability that the other protocols may also request entry into the same state (e.g., due to identifying similar conditions (e.g., little or no traffic) on the MCPL), the logical PHY may wait until permission or commands are received from each higher layer protocol before actually allowing state transition. The logical PHY may track which higher layer protocols requested the state change (e.g., performed a corresponding handshake) and trigger the state transition by identifying that each of the protocols requested that particular state change, such as a transition from L0 to L1 or another transition that would interfere with or interfere with communication with other protocols. In some implementations, protocols may be blinded for their at least partial dependence on other protocols in the system. Further, in some cases, a protocol may expect a response (e.g., from the PHY) to a request to enter a particular state, such as an acknowledgement or rejection of the requested state transition. Accordingly, in such cases, while waiting for permission from other supported protocols to enter a idle link state, the logical PHY may generate synthetic responses to a request to enter the idle state to "excrete" the requesting higher protocol so that it assumes that a particular state has been entered (if, in fact, the lanes are still active, at least until the other protocols also request entry into the idle state). Among other possible advantages, this may facilitate coordinating entry into the low power state between multiple protocols, among other examples.Referring to FIG. 10, a simplified link state machine transition diagram 1000 is shown, along with sideband handshaking used between the state transitions. For example, a reset.idle state (e.g., where phase lock loop (PLL) lock calibration is performed) may transition to a reset.cal state (e.g., where the link is further calibrated) by a sideband handshake. Reset.Cal may transition to a reset.ClockDCC state by a sideband handshake (e.g., where duty cycle correction (DCC) and delay-locked loop (DLL) locking may be performed). An additional handshake may be performed to transition from the reset.clockDC to a reset.quiet state (e.g., to disable the valid signal). To assist in directing signaling on the MCPL lanes, the lanes may be transmitted through a center. Pattern state can be centered.In some implementations, during the Center.Pattern state, the transmitter may generate training patterns or other data. The training patterns may be used, for example, by data lanes of the MCPL to assist in aligning and optimizing scanning on the MCPL lanes, such as the data lanes. For example, the lanes may be centered by a Center.Pattern state defined by the link state machine (LSM). In one example, during centering, the receiver device may condition its receiver circuit to receive such training patterns, for example, by adjusting the phase interpolator position and the vref position and adjusting the comparator. The receiver may continually compare the received patterns to expected patterns and store the result in a register. After a set of patterns is completed, the receiver may increment the phase interpolator setting and thus maintain the reference voltage (vref) equal. The test pattern generation and comparison process may proceed step by step and new comparison results may be stored in the register, the process repeatedly passing through all phase interpolator values and all vref values. The results may be compared and analyzed, for example, by the logical PHY and / or system management software, to determine the optimal conditions for operation of the multiple lanes of the link. Such conditions may include optimal values for the phase interpolator (e.g., defining a delay between the strobe and the data sampling clock) and vref. Following this analysis or comparison process, in one example, a Center.Quiet state may be entered when the pattern generation and comparison process is complete. After centering the lanes by the Center.Pattern and Center.Quiet link states (and any other example link training states), sideband handshake (e.g., using an LSM sideband signal over the own LSM sideband lane of the link) may be enabled to transition to a link.init state to initialize the MCPL according to the results of the link training and cause the MCPL to enter an active, L0, or transmit state. In the active link state, the transmission of (basic) data on the MCPL is enabled.As mentioned above, sideband handshakes may be used to enable the link state machine to transition between dies or chips in a multi-chip module. For example, signals on the LSM sideband lanes of an MCPL may be used to synchronize the state machine transitions across the die. For example, if the conditions to exit a state (e.g., reset.idle) are met, the page that met these conditions may activate an LSM sideband signal on its outgoing LSM_SB lane and wait for the other external die to meet the same condition and activate an LSM sideband signal on its LSM_SB lane. When both LSM_SB signals are asserted, the link state machine of the respective die may transition to the next state (e.g., a reset.cal state). A minimum overlap time may be defined during which both LSM_SB signals should remain activated before transitioning to a state. Further, a minimum dwell time may be defined after LSM_SB is deactivated to allow accurate detection of the swing. In some implementations, each link state machine transition may be conditioned on and enabled by such LSM_SB handshakes.FIG. 11 is a more detailed diagram of a link state machine 1100 illustrating at least some of the additional link states and link state transitions that may be included in an example MCPL. In some implementations, an example link state machine may include, in addition to the other states and state transitions illustrated in FIG. 11, a directed loop operation ("directed loop back" transition may be provided to put the lanes of an MCPL into a digital loop. For example, the receiver lanes of an MCPL may be looped back to the transmitter lanes after the clock recovery circuits. An "LB_Center" state may also be provided in some cases, which may be used to align the data characters. Additionally, as shown in FIG. 11, an MCPL may support multiple link states including an active L0 state and low power states such as an L1 sleep state and an L2 sleep state, among potentially other examples. As another example, configuration or centering states (e.g., CENTER) may be extended or system to assist in reconfiguring a link while being powered to allow the links lanes to be remapped to route data around one or more links of the link that, among other examples, are determined to be faulty or malicious.While the above examples generally illustrate the techniques and example benefits for training at least some of the lanes of an example link, such as an MCPL, training some lanes may be problematic. For example, in the case of an MCPL (or similar connections), it may be difficult not only to train a valid lane, but to synchronize with the data lanes that it effectively controls. For example, as described above, a valid lane may be used to signal a receiver that data is expected on the data lanes of an MCPL. This may cause the receiver to provide power, go out of sleep, or otherwise prepare the receiver circuitry for receiving the expected data and starting its processing.In an MCPL, a dedicated valid lane is used to allow for very fast power state transition. For example (as shown and discussed in connection with the example of FIG. 7 ), the transmitting device (or transmitter) may activate the valid lane (e.g., by generating the valid signal) to indicate the beginning of data traffic on the data lanes in a next signal transmission window (e.g., 8UI later). The receiver can continually scan the valid lane and detect valid activations. From the first UI (or signal transmission window boundary) in which valid was enabled, the receiver may then count a UI duration (e.g., 8 UI) before scanning the corresponding multiple data lanes of the MCPL. As a result, it is critical that the rising and falling edges of valid be accurately sampled. For example, among other example problems, failure to correctly sample the valid lane (and / or data lanes) may result in errors in the event of an offset between sampling or clocking the valid and data lanes, such as arrival of data before it is expected, clipping of data before it is completed, incorrectly interpreting transients on data lanes as data, etc.Training and centering of the data lanes, although desirable, may complicate synchronized scanning of the valid and data lanes. For illustrative purposes, with reference to FIG. 12, a diagram 1200 is shown illustrating an example signal transmission on an MCPL. In this example, training patterns 1205, 1210, 1215 are sent to the MCPL lanes to be trained. The lanes may include the MCPL's data lanes (e.g., DATA[0], etc.), and potentially also one or more control lanes including the stream lane. In this example, as in the example of FIG. 7, the valid lane is used to indicate to a receiver that data (e.g., 1205, 1210, 1215) is on the data lanes. The valid signal (e.g., 1220) is to be transmitted in the signal transmission window (beginning at n) that precedes the signal transmission window (beginning at n+1) in which the training data (e.g., 1205) is to be transmitted. The valid signal is activated and held for a duration (e.g. a number of UIs or signal transmission windows) corresponding to the direction of the data to be transmitted, in this case two signal transmission windows or 16UI. In this example, the clock is implemented as a 2UI clock, which means that a UI is clocked on both the rising edge and the falling edge of the clock (or strobe) cycle. Further, training patterns or data in this example may be stored at a center, among other conditions. Pattern state can be transmitted.While the valid lane is useful in the example of FIG. 12 by supporting fast signal transmission transitions, using valid (and providing its own valid Tx and Rx logic) to support training the data lanes by indicating incoming training patterns of the valid lane makes it difficult to train itself. This can also complicate the synchronization of the scanning of the valid and data lanes. Indeed, in cases where the valid lane is an active party in training the data lanes (as in the example of FIG. 12 ), it is not possible to train itself. However, not training and not centering the link may result in an offset between scanning the valid and data lanes, as the data lanes may include centering (i.e., interpolating) the scan clock in a suboptimal setting for the valid lane by centering, for example, as mentioned above.To synchronize or unify the sampling of the data lanes with the valid lane (on which the timing of the data lanes may depend), some implementations may provide solutions such as using a complementary clock tree of the DLL clock (corresponding to the strobe) to sample control lanes as valid. However, using such a complementary clock tree may be expensive from the perspective of system power and not feasible for all systems. In other implementations, a multi-phase link training protocol may be defined, where the control lanes (e.g., valid) and data lanes are trained in two separate processes (i.e., where the control lanes synchronize the data lanes for training, and vice versa). However, among other example disadvantages, two-phase training may increase training time and increase complexity of system and circuit design to support the multiple phases of link training.As noted above, synchronizing the scanning of the valid lane and the data lanes may be critical to successful operation of the MCPL. In one example, training the valid lane may be kept simple (while supporting more complex training of the data lanes (as shown in the example of FIG. 12 )) by providing an offset balanced clock tree to provide the sampling clock for sampling both the data lanes and the valid lanes. FIG. 13 is a simplified block diagram 1300 illustrating an example implementation using an offset balanced clock tree design. Such a clocking scheme may be used, for example, where it remains desirable to use a valid lane in its typical context to assist in training the data lanes by controlling (or displaying) pending transmission of training patterns on the data lanes. However, in such cases, the valid lane is not included in the centering process that defines the optimal sampling point (e.g., at the receiver) for sampling the data lanes. This also leads to a circular problem, as the sampling of the valid data lanes is not necessarily synchronized with the sampling of the data lanes during the training of the data lanes (e.g., shown in FIG. 12 ) to allow accurate bitwise comparison of received training patterns at the receiver during the training, among other example problems.To address this, an offset balanced clock tree may be provided as shown in FIG. 13. Because the data must be sampled at an accurate UI or signal transmission window boundary corresponding to the beginning of a valid signal (e.g., 8 UI) to ensure that the sampling clock offsets for valid and data lanes match, the offset-compensated clock tree may provide a delay circuit 1305 that provides a sampling clock offset that ensures that the sampling of the valid lane is within the range of sampling the data lanes. In the example of FIG. 13, a phase interpolator shifts the DLL clock (1310) of the strobe 1315, thereby providing an interpolated version of the strobe signal to the data lane clock domain (e.g., at 1320). Indeed, the interpolated sampling clock of data lanes 1320 is optimized to center the sampling of the data lane at the center of the "eye" of the signal, such that the sampling occurs near the center of the UI (or the x,5 UI interval (where x is an integer from 0 to ∞)). The same strobe signal may be provided to the clock domain 1325 of the valid lane via a clock tree.The delay circuit 1305 may cause the clock tree to be offset balanced by implementing the delay m that remains greater than zero for each expected process, voltage, or temperature, but is equal to or less than 0.5UI during operation (i.e., m<=0.5UI). This condition may ensure that there is sufficient sample margin relative to the offset on the data lanes resulting from training the data lanes (e.g., 1320). In one example, the delay circuit 1305 may be implemented as a chain of inverter elements, where the number of inverter elements corresponds to the desired delay value m. In another implementation, a network of inverters (or other elements) may be provided that is configurable to selectively enable / disable this combination of elements in delay circuit 1305 to provide the desired delay value m. This may allow the delay value to be dynamically determined or even modified. Regardless of implementation, delay circuit 1305 may provide a phase shift of the sampling clock for the valid lane, resulting in an offset x,5UI+m. This delay may ensure that the sampling of the data lanes and the valid lanes remains substantially synchronized to guarantee that valid signals accurately indicate the arrival of data on the data lanes.In another example implementation, the link training methods applied to the data lanes (such as in the example of FIG. 12 ) may also be applied to the valid lane in a single phase or stage such that the valid lane is centered (and otherwise trained) along with the data lanes. This can ensure that scanning of the valid and data lanes is consistent, as well as allow somewhat higher ranked link training techniques (and corresponding advantages) from which the valid lane may benefit. In such implementations, the valid lane (and supporting Rx and Tx circuitry on the two devices to be connected by the corresponding MCPL) may operate in two modes. In a first mode or training mode, the valid lane may operate like any other lane and does not send valid signals to queue the arrival of data on corresponding data lanes (e.g., as in the example of FIG. 12 ). Likewise, data lanes (and supporting Rx and Tx logic) may be configured in this implementation to support training and transmission / reception of corresponding training patterns on the data lanes without the request and help of a valid lane. In a second or active mode following successful training of the link, the valid lane may operate as a control lane as described above, activating valid signals to indicate further arrival of data on corresponding data lanes. Accordingly, devices may send / receive data on both valid lanes and data lanes during link training other than in active states. Indeed, in one implementation, there may effectively be no difference between the role of the valid lane and the data lane during link training.FIG. 14 shows a diagram 1400 illustrating example signaling on an MCPL during enhanced link training mode. The example of FIG. 14 may represent single phase centering of both the valid lane and the data lanes. In one example, as discussed in more detail below, single phase centering may include training patterns generated using self-initialized Fibonacci linear feedback shift registers (LFSRs) that generate a pseudo-random binary sequence (PRBS) to synchronize a transmitting device with a remote receiving device via the MCPL. Such a solution may be achieved without a complementary clock tree for control signals and may allow centering of data and control lanes in a single phase such that centering of the data lanes and valid lanes (and potentially other control lanes such as the stream lane) are completed at the same time and together such that environmental factors (e.g., cross-talk) and other factors are better accounted for.For example, as shown in Figure 14, an alternative implementation of the link training shown in the example of Figure 12 is illustrated. Rather than the valid lane providing valid signals to indicate the arrival of training patterns (e.g., 1205, 1210, 1215) on the data lanes (e.g., DATA[0]), the valid lane is included in the training and treated as one of the data lanes, with corresponding training patterns (e.g., 1405, 1410, 1415) being sent on the valid lane. Indeed, the valid lane and set of data lanes (and potentially other lanes, such as a stream lane) are trained in parallel together in a single training procedure. In this example, the training patterns (e.g., 1205, 1210, 1215, 1405, 1410, 1415) may be sent to center and synchronize the set of lanes. For example, the training patterns may be sent in a Center.Pattern (or other) link training state. Centering can be used to determine an optimal interpolation phase and comparison voltage to apply to sampling the collection of lanes (valid and data lanes). Arriving at the optimal setting may include the transmitter transmitting training patterns (e.g., 1205, 1210, 1215, 1405, 1410, 1415) on the lanes that are checked bit by bit by the remote receiver, respectively. In one example, the centering training patterns may be versions of a PRBS. Training patterns are sent and evaluated at each of a (large) set of potential receiver phase interpolation and voltage settings (PI, vref) to determine which combination yields the best results. These settings are then adopted for use in scanning the collection of lanes when the MCPL transitions to an active state (and the valid lane transitions to its operational mode (e.g., as shown in the example of FIG. 7 )).As noted above, training patterns used in training lanes of an MCPL (including the valid lane) may include or be based on instances of a PRBS. In some implementations, an LFSR, such as a 23-bit Fibonacci LFSR (as shown in FIG. 15 ), can be used to generate PRBSs for training a collection of MCPL lanes. In the example of FIG. 15, the LFSR may be implemented as a chain of flip-flops (e.g., x1-x23) that are connected as a linear shift register to an XOR gate 1505, which provides feedback to the shift register through a multiplexer (mux) 1510.In some implementations, prior to centering (or other link training activities that include PRBS-based training patterns), all of the LFSRs used to generate the PRBSs for training may be initially initialized with all 0s at the transmitter, such that valid and data are kept low on the way high to the training patterns (as shown in the example signal transmission diagram 1600 of FIG. 16 ). At the receiver, the LFSR mux controller 1510 clocks incoming transmitter values into bit 1 of the LFSR. When the transmitter LFSR is initially initialized with all 1's (e.g., 23 consecutive bits for the 23 serial flip-flops), centering may begin. The receiver may expect a series (e.g., 23bits or UI) of 1's to be sent on at least one designated lane (e.g., the valid) prior to beginning the training patterns on the valid lane and the data lanes (for use in centering these lanes). This series of 1's may follow a previous series of 0's on at least the designated lane. In some implementations, such as shown in FIG. 16, the series of 0s and the series of 1's training patterns may precede both the valid lane and each of the corresponding data lanes to signal the start of training. The receiver may also use the incoming stream of 1's to self-initialize its own LFSR in expectation of the coming PRBSs to be sent on the lanes to be trained. After initial initialization of 23 bits from 1's, the mux control of the LFSR may begin to supply the output of the XOR 1505 in bit 1 of the LFSR such that the LFSR self-initializes and is able to match the PRBSs generated by the transmitter's LFSR (at Rx).In some implementations, a single LFSR (as shown in FIG. 15 ) can be used to generate multiple uncorrelated PRBSs, one for each lane to be trained. For example, Table 1 below illustrates that for a 1UI clock, an XOR operation can be performed for various combinations of two LFSR bits, respectively, to generate the uncorrelated set of PRBSs for training a set of 20 data lanes, a stream lane (S), and a valid lane (V) from a single LSPR. FIG. 17 is a diagram illustrating a portion of the XOR network that may be connected to the LFSR to generate lane-specific uncorrelated PRBSs. For example, FIG. 17 shows XOR gates 1705, 1710 used to generate two uncorrelated PRBSs for data lanes DATA[2] and DATA[3] according to the implementation described in Table 1. It should be appreciated that a full implementation of Table 1 would include 20 additional XOR branches to generate respective PRBSs for the remaining 20 lanes in the example of Table 1. While Table 1 (and FIG. 17 ) describes an implementation of a generated set of uncorrelated PRBS streams for a collection of lanes according to a 1UI clock (e.g., 1 UI per clock cycle), other implementations may include multiple UI clocks, such as a 2 UI clock. Accordingly, two consecutive bits can be generated per clock cycle in the PRBS of a lane. The multiple bits may be provided to a serializer / deserializer (SERDES), which provides the bits as part of a PRBS serial stream on the corresponding lane. In other implementations, an 8UI clock (providing high data rates with potentially constrained clock speeds) may be used. For example, FIG. 18 illustrates a table (analogous to the example of Table 1 and FIG. 17 and above) that shows how 8 instances of uncorrelated PRBS can be generated per cycle from XOR of various combinations of the outputs of a single bit of the LFSR. For example, for valid lane V, eight different PRBS bits (1-8) may be generated (e.g., one for each UI of the clock cycle) of the LFSR according to: where the operator "^" indicates an exclusive OR (XOR) operation on bit outputs at corresponding flip-flops of the LFSR. This set of 8 PRBS bits may then be provided to a SERDES to be serialized across 8 UIs of the MCPL (e.g., for use in training the valid and the other lanes in a set of lanes).A self-sustaining LFSR of a receiver may be synchronized with the LFSR of a remote transmitter and may be used to measure the bit error rate (BER) on the PRBS on all lanes, including the valid lane (e.g., during link training). This approach works even when each lane has a different (uncorrelated) PRBS derived from a linear combination of LFSR bits, as shown and described in some of the examples above. Each of the receiver and transmitter LFSRs may be self-initialized after initialization (e.g., all 1's over 23 UI) so that both LFSRs generate the same sequences for each of the lanes. The receiver can thus check bit by bit on each lane whether the received PRBS matches (after the self-initialization period) that expected to determine the BER for each lane. Similarly, other training activities may exploit synchronization between the receiver and transmitter LFSRs, including during centering.FIG. 19 is a block diagram 1900 illustrating simplified bumpouts of each of the two devices 1905, 1910 connected by a multi-lane MCPL. For example, each device 1905, 1910 may include bumps for transmission on multiple (e.g., 20) data lanes (e.g., TD[0-19]) and bumps for reception on multiple (e.g., 20) data lanes (e.g., RD[0-19]). Similarly, each device 1905, 1910 may include bumps for the corresponding transmit and receive stream lanes (e.g., T_STM, R_STM) and valid transmit and receive lanes (e.g., T_STMD, R_STMD) of the MCPL 1915. Additional bumps may be provided for power (e.g., Vcc, Vss) and clocks (e.g., RCK, TCK), among other examples. The bumps represent physical connection points of the respective devices to corresponding lanes of the MCPL 1915 to the respective device 1905, 1910. As shown, some bumpouts (and corresponding lanes) are in closer proximity to each other than others. Lanes and bumps in close proximity to each other are more likely to produce cross talk between each other than lanes / bumps that are further apart. The physical design of bumpouts and lanes can serve as a basis for developing additional link training tests and activities such as testing for crosstalk.Figure 20 illustrates a table 2000 representing, for each loop or phase, a cross-talk test to be performed on a set of lanes in an MCPL. Different test patterns (e.g., PRBS instances) may be sent on the lanes to test crosstalk. In this example, each loop involves testing two of the lanes for cross-talk stress. On each successive loop, a different pair of lanes is tested. The lanes under test may be spaced sufficiently apart so that they are substantially immune to cross-talk from the other lane under test. For example, in loop 0, lanes DATA[0] and DATA[9] are tested as "victims", lanes DATA[1-3] and DATA[17-19] act as aggressors of lane DATA[0] and lanes DATA[8, 10-12], and stream lane (S) acts as an aggressor of lane DATA[9]. A victim signal (L) is sent from the victim lanes under test and aggressor signals (H) are sent on the respective adjacent aggressor lanes to provide cross-talk stress to the victim lane. In some implementations, these aggressor signals (H) may be complementary to the victim signal (L). For example, in one example, L may be a PRBS pattern while H is the reverse PRBS pattern. During a given test or loop, lanes that are further from the victim lane may carry other signals (e.g., signals "1" -"22") that are different from both victim (L) and aggressor (H) signals (e.g., uncorrelated instances of a PRBS).The example of FIG. 20 illustrates a cross-talk stress test in which the valid lane is not trained along with the other lanes, such as in the example of FIG. 12. To support the training, the valid lane is activated in this example (shown as value "A" in the table of FIG. 20 ) to indicate the arrival of the various (and substantially persistent) cross-talk stress test patterns on the data lanes. However, if the valid lane and its position within the bumpout and link are not taken into account, this may result in suboptimal crosstalk analysis, not only for those lanes adjacent to the valid lane (e.g., the stream lane and the data lane DATA[8]), but also for the valid lane itself. Accordingly, as shown in FIG. 21, in cases where the valid lane (V) is trained along with the data lanes, signals from crosstalk stress tests (including victim and aggressor signals) may be sent on the valid lane as any other data lane. For example, in loop (or test) number 9 of the example shown in Fig. 21, the valid lane is tested by sending a victim signal L on the lane and aggressor signals on the lanes in close proximity to the valid lane. In this case, the PRBS precedes the valid lane the string of 1's to allow self-initialization as described previously. Such a string of previous 1's may be sent for each loop, regardless of the signals L, H or another signal sent on the designated lane (e.g., for signal H, the LFSR may self-initialize with the reverse shift register value, and the PRBS may be generated by reversing the outputs (e.g., DATA(n)) of the implementation shown in FIG. 17).In some cases, the victim and aggressor signals may be based on or include a PRBS, as with other training patterns used to train the valid lane along with other lanes of an MCPL. For example, the signals 2100 provided in the table shown in FIG. 21 may be generated by a self-initialized LFSR, as described above. Indeed, as described in the above centering example (and as shown in FIG. 16, for example), generation of PRBSs for crosstalk from test signals, an LFSR initialization phase for synchronizing the transmitter and receivers, followed by generation of the PRBSs on the various lanes. In some cases, the link may be divided into two groups of lanes (e.g., in cases where there are multiple valid lanes and each group has at least one valid lane). In such cases, there may be multiple designated lanes, but self-initialization for each group occurs after the stream of all 1's (e.g., as described in the previous paragraph) and the PRBS generating phase are common, so that crosstalk may span the group.It should be understood that the specific examples set forth above are provided as non-limiting examples of the application of the more general principles described herein. For example, other systems, link protocols, and technology may also include and use the above functions. Some alternative implementations may take on features different from the specific examples presented above without departing from the scope of the present disclosure. For example, links may be provided with various numbers of lanes, different lane / bump layouts, different clock speeds and data rates, training patterns, link state machines, and link protocols, while still implementing and benefit from the benefit of the various features and principles described herein, among other examples.FIGS. 22A-22B show flowcharts 2200 a-b illustrating example techniques for training a valid lane of a multi-protocol, time-division multiplexed connection link. For example, in FIG. 22A, an offset balanced clock tree may be used to synchronize the sampling of left data lanes with the valid lane because the transmission of data on the data lanes is to be timed based on the beginning of the transmission of valid signals on the valid lane. For example, a strobe or clock signal 2205 may be provided to the clock tree and a phase interpolator may be applied 2210 to the clock signal during centering of the data lanes. The phase interpolator may define an offset of the sampling clock used to sample the data lanes. A delay may be applied 2215 to a branch of the clock tree corresponding to the valid lane to substantially synchronize the sampling 2220 of the valid lane with the sampling of the data lanes. The delay may be defined, for a hardware-based delay circuit, such that the applied delay m remains greater than zero, but equal to or less than 0.5UI during operation (i.e., 0<m<=0.5UI) for each expected process, voltage, or temperature.Referring to FIG. 22B, in an alternative embodiment, the valid lane may be trained along with link data lanes in a single phase or training unit. For example, link training patterns may be received 2230 (e.g., at a receiving device from a transmitting device connected by the link) on each of a set of lanes of the link, the set including the plurality of data lanes and the valid lane. The set of lanes may be trained together 2235 based on the training patterns. Training may include centering the set of lanes and analysis of crosstalk, among other activities. The training may determine operational characteristics to apply to the link during an active, L0, or transmitting link state. Accordingly, the active link state may be entered 2240 after completion of link training, and the valid lane may resume normal or active operation by sending (and receiving 2245) a valid signal on the valid lane to display incoming data to be sent (e.g., in an immediately subsequent signaling window) on the data lanes. This may cause the data lanes (and supporting circuitry at the receiver) to prepare for the arrival 2250 of the data. Concurrent training 2235 of the valid lane and the plurality of data lanes may allow each to be synchronized in scanning, allowing the transmitter to rely on the valid signal to reliably indicate the exact time (e.g., UI) at which the corresponding data will begin to arrive on the data lanes.It should be appreciated that the above-described devices, methods, and systems may be implemented in any electronic device or system as mentioned above. As specific representations, the figures below provide example systems for using the invention as described herein. As the systems are described in more detail below, a number of different connections are disclosed, described, and revised from the discussion above. As can also be readily appreciated, the foregoing advances can be applied to any of these interconnects, fabrics, or architectures.Referring now to FIG. 23, an example implementation of a system 2300 is shown, according to an example embodiment. As shown in FIG. 23, multiprocessor system 2300 is a point-to-point interconnect system and includes a first processor 2370 and a second processor 2380 coupled together via a point-to-point interconnect 2350. Each of the processors 2370 and 2380 may be a version of a processor. In one embodiment, 2352 and 2354 form part of a serial, coherent point-to-point interconnect fabric, such as a high performance architecture. As a result, the invention may be implemented within the QPI architecture.While shown with only two processors 2370, 2380, it should be understood that the scope of the present invention is not so limited. In other embodiments, one or more additional processors may be present in a given processor.Processors 2370 and 2380 are shown including integrated memory controller units 2372 and 2382, respectively. Processor 2370 includes as part of its bus controller units point-to-point (P-P) interfaces 2376 and 2378; similarly, second processor 2380 includes P-P interfaces 2386 and 2388. Processors 2370, 2380 may exchange information via a point-to-point (P-P) interface 2350 using P-P interface circuits 2378, 2388. As in FIG. 23, IMCs 2372 and 2382 couple the processors to respective memories, namely a memory 2332 and a memory 2334, which may be portions of main memory locally attached to the respective processors.Each processor 2370, 2380 exchanges information with a chipset 2390 via individual P-P interfaces 2352, 2354, using point-to-point interface circuits 2376, 2394, 2386, 2398. Chipset 2390 also exchanges information with a high performance graphics circuit 2338 via interface circuit 2392 along a high performance graphics link 2339.A shared cache (not shown) may be included in each processor or external to both processors; however, it is connected to the processors via a P-P connection such that the local cache information from one or both processors may be stored in the shared cache when a processor is placed in a low power mode.Chipset 2390 may be coupled to a first bus 2316 via an interface 2396. In one embodiment, the first bus 2316 may be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, although the scope of the present invention is not so limited.As shown in FIG. 23, different I / O devices 2314 are coupled to a first bus 2316, along with a bus bridge 2318 that couples the first bus 2316 to a second bus 2320. In one embodiment, the second bus 2320 includes a low pin count (LPC) bus. In one embodiment, different devices are coupled to second bus 2320, including, for example, a keyboard and / or mouse 2322, communication devices 2327, and a storage unit 2328, such as a disk drive or other mass storage device, that often includes instructions / code and data 2330. Also shown is audio I / O 2324 coupled to second bus 2320. It should be appreciated that other architectures are possible, with the components and interconnection architectures included varying. For example, instead of the point-to-point architecture of FIG. 23, a system may implement a multipoint bus or other such architecture.While the present invention has been described with reference to a limited number of embodiments, those skilled in the art will appreciate numerous modifications and variations therefrom. It is intended that the appended claims cover all such modifications and variations as fall within the true spirit and scope of this present invention.A design may go through different stages, from creation via simulation to fabrication. Data representing a design may represent the design in a number of ways. First, as is appropriate in simulations, the hardware may be represented using a hardware description language or another function description language. Additionally, a model may be fabricated on circuit levels with logic and / or transistor gates at some stages of the design process. Further, most designs may reach a data plane representing the physical arrangement of various devices in the hardware model at some stage. In the case where conventional semiconductor fabrication techniques are employed, the data representing the hardware model may be the data indicative of the presence or absence of different functions on different mask layers in masks used to fabricate the integrated circuit. In each design representation, the data may be stored in any form of machine readable medium. A memory or magnetic or optical storage such as a disk may be the machine readable medium for storing information transmitted via optical or electrical waves modulated or otherwise generated to transmit such information. When an electric carrier wave indicating or transporting the code or design is transmitted to the extent that copying, buffering, or retransmission of the electric signal is performed, a new copy is made. Thus, a communication service provider or a network service provider may at least temporarily store, on a tangible, machine-readable medium, an article, such as information encoded into a carrier wave, using techniques of embodiments of the present invention.A module as used herein refers to any combination of hardware, software, and / or firmware. As an example, a module includes hardware, such as a microcontroller, connected to a non-transitory medium to store code adapted to be executed by the microcontroller. Therefore, in one embodiment, reference to a module refers to the hardware specifically configured to recognize and / or execute the code stored in a non-transitory medium. Further, in another embodiment, use of a module refers to the non-volatile memory including code specifically adapted to be executed by the microcontroller to perform predetermined operations. And as may be inferred in yet another embodiment, the term module (in this example) may refer to the combination microcontroller and non-volatile medium. Often, module boundaries shown as separate generally vary and potentially overlap. For example, first and second modules may share hardware, software, firmware, or a combination thereof, while potentially retaining some independent hardware, software, or firmware. In one embodiment, the use of the term logic includes hardware such as transistors, registers, or other hardware such as programmable logic packages.The use of the term 'configured to' in one embodiment refers to arranging, assembling, manufacturing, offering, importing, and / or interpreting a device, hardware, logic, or element that is to perform an intended or particular task for sale. In this example, an apparatus or element thereof that is not operating is still 'configured to' perform an intended task when designed, coupled, and / or connected to perform that intended task. As a purely illustrative example, a logic gate may provide 0 or 1 during operation. However, a logic gate 'configured to' apply an enable signal to a clock does not include every potential logic gate that can provide 1 or 0. Instead, the logic gate is one that is coupled in a particular manner so that during operation the 1 or 0 output is to enable the clock. Once more, it should be noted that the use of the term "configured to" does not require operation, but rather focuses on the latent state for which the device, hardware, and / or element is designed to perform a particular task when the device, hardware, and / or element is operating.Further, the use of the terms "to," "capable of," and / or "operable to," in one embodiment, refers to a device, logic, hardware, and / or element designed in such a manner as to enable the use of the device, logic, hardware, and / or element in a manner indicated. As noted above, in one embodiment, the use of 'to', 'capable of' or 'operable' refers to the latent state of a device, logic, hardware, and / or element, where the device, logic, hardware, and / or element are not operating, but are designed in such a manner as to enable the use of a device in a specified manner.A value as used herein includes any known representation of a number, a state, a logical state, or a binary logical state. Often, the use of logic levels, logic values are also referred to as 1s and 0s, which simply represent binary logic states. For example, 1 refers to a high logic level and 0 refers to a low logic level. In one embodiment, a memory cell, such as a transistor or flash cell, may be capable of containing a single logical value or multiple logical values. However, other representations of values are also used in computing systems. The decimal number ten can also be represented, for example, as a binary value of 1010 and as a hexadecimal letter A. Thus, a value includes any representation of information that may be included in a computing system.Moreover, states may be represented by values or parts of values. As an example, a first value, such as a logical one, may represent a default or initial state, while a second value, such as a logical zero, may represent a non-default state. Additionally, in one embodiment, the terms reset and set may refer to a default and an updated value or state. For example, a default value potentially includes a high logical value, i.e., reset, while an updated value potentially includes a low logical value, i.e., set. It should be noted that any combination of values may be used to represent any number of states.The embodiments of methods, hardware, software, firmware, or code set forth above may be implemented via processing element executable instructions or code stored on a machine-accessible, machine-readable, computer-accessible, or computer-readable medium. A non-transitory machine-accessible / readable medium includes any mechanism that provides (i.e., stores and / or transmits) information in a form readable by a machine such as a computer or electronic system. A non-transitory machine-accessible medium includes, for example, random-access memory (RAM), such as static RAM (SRAM) or dynamic RAM (DRAM); ROM; magnetic or optical storage media; flash memory devices; electrical storage devices; optical storage devices; acoustic storage devices; other forms of storage devices for holding information received from transitory (propagated) signals (e.g., carrier waves, infrared signals, digital signals), etc., that are to be distinguished from the non-transitory media that may receive information therefrom.Instructions used to program logic to perform embodiments of the invention may be stored in memory in the system, such as DRAM, cache, flash memory, or other memory. Further, the instructions may be propagated over a network or via other computer readable media. Thus, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), but is not limited to, floppy disks, optical disks, CDs, compact disk read-only memories (CD-ROMs), and magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memories, or a tangible, programmable read-only memories, Machine readable storage used in the transmission of information over the Internet via electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Accordingly, the computer-readable medium includes any type of tangible machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).The following examples are directed to embodiments in accordance with this specification. One or more embodiments may provide a device, system, machine readable storage, machine readable medium, and / or method for receiving one or more link training signals including instances of a link training pattern on multiple lanes from a physical link including at least one valid lane and multiple data lanes that collectively trains the multiple lanes using the link training signals to synchronize scanning of the valid lane with scanning of the multiple data lanes, to enter an active link state, to receive a valid signal on the valid lane during the active link state, the valid signal including and indicating a signal maintained at a value for a defined first duration, receiving data on the plurality of data lanes in a second defined duration after the first duration. The data is to be received during the active link state on the plurality of data lanes during the second defined duration.In one example, the valid signal is received at the beginning of a first one of a series of signal transmission windows defined for the link, each of the series of signal transmission windows including a common defined duration, the first duration including the duration of the first signal transmission window, the second duration including the duration of a second signal transmission window following the first signal transmission window in the series, the data is to be transmitted over a number of consecutive signal transmission windows, and the valid signal is held at the value for the number of consecutive signal transmission windows and is to be deactivated after the end of the number of consecutive signal transmission windows to indicate an end of the data transmitted over the plurality of data lanes.In one example, the link supports multiplexing data from multiple different protocols on the multiple data lanes.In an example, the plurality of lanes further includes a stream lane, and a stream identifier signal is to be transmitted on the stream lane to indicate which of the plurality of protocols applies to the data transmitted on the plurality of data lanes.In one example, the instances of the link training pattern include multiple instances of a pseudo-random binary sequence (PRBS).In one example, the plurality of instances of the PRBS include uncorrelated instances of the PRBS and another of the plurality of instances of the PRBS is to be sent on each lane in the plurality of lanes.In one example, the multiple instances of the PRBS are generated from a common linear feedback shift register (LFSR).In one example, a receiver LFSR is provided that is synchronized to a remote transmitter LFSR.In an example, training the plurality of lanes includes centering the plurality of lanes.In one example, training the plurality of lanes includes performing cross-talk stress tests for each of the plurality of lanes.In one example, detecting a particular signal transmitted on the valid lane to indicate future receipt of the one or more link training signals.In one example, the particular signal transitions to a respective one of the link training signals on the valid lane.In one example, the particular signal is used for allowing an LFSR used to evaluate the link training signals to initialize itself.One or more embodiments may provide for: an apparatus, a system, a machine readable storage, a machine readable medium, and / or a method for transmitting one or more link training signals including instances of a link training pattern on multiple lanes of a physical link including at least one valid lane and multiple data lanes, entering an active link state based on training the multiple lanes using the link training signals, and transmitting a valid signal on the valid lane during the active link state, the valid signal including a signal that is maintained at a value for a defined first duration and indicates that data on the multiple data lanes is in a second defined subsequent to the first duration, It is desired to transmit continuously. The data is to be transmitted on the plurality of data lanes during the active link state during the second defined duration.In one example, a linear feedback shift register (LFSR) generates the instances of the link training pattern.In an example, the instances of the link training pattern include multiple different link training patterns, and each of the multiple link training patterns is to be sent on a respective one of the multiple lanes.In an example, the plurality of link training patterns includes a plurality of instances of a pseudo random binary sequence (PRBS).In one example, the physical layer logic self-initializes the LFSR.In one example, the LFSR causes a first pattern to be generated on the valid lane to indicate that the link training patterns are to be sent based on the self-initialization.In one example, the first pattern is used by a remote receiver device to synchronize an LFSR of the receiver device.In an example, the plurality of lanes further includes a stream lane, and the physical layer logic identifies a type of data to be sent on the plurality of data lanes and sends a stream identifier signal on the stream lane to indicate the type of data.In an example, each duration includes a respective one of a plurality of signal transmission windows, and each signal transmission window has a common duration.In one example, the common duration includes 8 unit interval (UI).One or more embodiments may provide a system including a connection including multiple lanes, where the multiple lanes include multiple dedicated data lanes and at least one dedicated valid signal lane. The system may further include a first device and a second device communicatively coupled to the first device using the connection. The second device may transmit one or more link training signals including instances of a link training pattern on the plurality of lanes to the first device, enter an active link state based on training of the plurality of lanes using the link training signals, and transmit a valid signal to the first device during the active link state on the valid lane in conjunction with data transmitted from the second device to the first device on the data lanes. The valid signal may include a signal maintained at a value across a first signal transmission window and indicating that in a second signal transmission window following the first signal transmission window, data is to be transmitted on the plurality of data lanes, the data on the plurality of data lanes being transmitted to the first device during the active link state during the second signal transmission window.In an example, the first device may receive the one or more link training signals, participate in training the plurality of lanes using the link training signals, enter the active link state, receive the valid signal during the first signal transmission window, and receive the data on the plurality of data lanes during the second signal transmission window.In one example, the first device may determine a phase interpolation applied to the scanning of the plurality of lanes based on the training signals.In an example, the plurality of lanes includes at least one stream signal lane for identifying which of a plurality of supported protocols is to apply to data transmitted during a corresponding signal transmission window on the plurality of data lanes.References throughout this specification to "one embodiment" or "an embodiment" mean that a particular function, structure, or feature described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, the appearances of the phrase "in one embodiment" in various places in this specification are not necessarily all referring to the same embodiment. Further, the particular functions, structures, or features may be combined in any suitable form in one or more embodiments.In the foregoing description, a detailed description has been given with reference to specific exemplary embodiments. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the invention as set forth in the appended claims. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. Further, the foregoing use of embodiment and other exemplary language does not necessarily refer to the same embodiment or example, but may refer to different and different embodiments, as well as potentially the same embodiment.
Claims
An apparatus comprising: physical layer logic to: receive one or more link training signals comprising instances of a link training pattern on a plurality of lanes of a physical link, the plurality of lanes comprising at least one valid lane and a plurality of data lanes; train the plurality of lanes together using the link training signals to synchronize scanning of the valid lane with scanning of the plurality of data lanes; enter an active link state; receiving a valid signal on the valid lane during the active link state, the valid signal comprising a signal maintained at a value for a defined first duration and indicating that data is to be received on the plurality of data lanes in a second defined duration subsequent to the first duration; and receiving the data during the active link state on the plurality of data lanes during the second defined duration.The apparatus of claim 1, wherein the valid signal is received at a beginning of a first one of a series of signal transmission windows defined for the link, each of the series of signal transmission windows comprising a common defined duration, the first duration comprising the duration of the first signal transmission window, the second duration comprising the duration of a second signal transmission window following the first signal transmission window in the series, wherein the data is transmitted over a number of consecutive signal transmission windows, and the valid signal for the number of consecutive signal transmission windows is maintained at the value and is to be disabled after the end of the number of consecutive signal transmission windows to indicate an end of the data transmitted on the plurality of data lanes.The apparatus of claim 1, wherein the link supports multiplexing of data from a plurality of different protocols on the plurality of data lanes.The apparatus of claim 3, wherein the plurality of lanes further comprises a stream lane and a stream identifier signal is to be transmitted on the stream lane to indicate which of the plurality of protocols applies to the data transmitted on the plurality of data lanes.The apparatus of claim 1, wherein the instances of the link training pattern comprise multiple instances of a pseudo-random binary sequence (PRBS).The apparatus of claim 5, wherein the plurality of instances of the PRBS comprise uncorrelated instances of the PRBS, and a different one of the plurality of instances of the PRBS is to be sent on each lane in the plurality of lanes.The apparatus of claim 6, wherein the plurality of instances of the PRBS are generated from a common linear feedback shift register (LFSR).The apparatus of claim 7, further comprising a receiver LFSR synchronized with a remote transmitter LFSR.The apparatus of claim 1, wherein training the plurality of lanes comprises centering the plurality of lanes.The apparatus of claim 1, wherein training the plurality of lanes comprises performing cross-talk stress tests for each of the plurality of lanes.The apparatus of claim 1, wherein the value comprises a logic "1" held for the first duration.The apparatus of claim 1, detecting a particular signal transmitted on the valid lane to indicate future receipt of the one or more link training signals.The apparatus of claim 12, wherein the determined signal transitions to a respective one of the link training signals on the valid lane.The apparatus of claim 12, wherein the determined signal is used to self-initialize an LFSR to evaluate the link training signals.An apparatus comprising: physical layer logic to: transmit one or more link training signals comprising instances of a link training pattern on a plurality of lanes on a physical link, the plurality of lanes comprising at least one valid lane and a plurality of data lanes; enter an active link state based on training the plurality of lanes using the link training signals; transmit a valid signal on the valid lane during the active link state, the valid signal comprising a signal maintained at a value for a defined first duration and indicating that data is to be transmitted on the plurality of data lanes in a second defined duration following the first duration; and transmitting the data, during the active link state, on the plurality of data lanes during the second defined duration.The apparatus of claim 15, further comprising a linear feedback shift register (LFSR) for generating the instances of the link training pattern.The apparatus of claim 16, wherein the instances of the link training pattern comprise a plurality of different link training patterns, and each of the plurality of link training patterns is to be transmitted on a respective one of the plurality of lanes.The apparatus of claim 16, wherein the plurality of link training patterns comprises a plurality of instances of a pseudo random binary sequence (PRBS).The apparatus of claim 16, wherein the physical layer logic is further to cause the LFSR to self-initialize.The apparatus of claim 19, wherein the LFSR causes a first pattern to be generated on the valid lane to indicate that the link training patterns are to be sent based on self-initialization.The apparatus of claim 20, wherein the first pattern is used by a remote receiver device to synchronize an LFSR of the receiver device.The apparatus of claim 15, wherein the plurality of lanes further comprises a stream lane, and the physical layer logic is to: identify a type of data to be sent on the plurality of data lanes; send a stream identification signal on the stream lane to indicate the type of data.The apparatus of claim 15, wherein each duration comprises a respective one of a plurality of signal transmission windows and each signal transmission window has a common duration.The apparatus of claim 23, wherein the common duration comprises 8 unit interval (UI).A method comprising: receiving one or more link training signals comprising instances of a link training pattern on a plurality of lanes of a physical link, the plurality of lanes comprising at least one valid lane and a plurality of data lanes; training the plurality of lanes together using the link training signals to synchronize scanning of the valid lane with scanning of the plurality of data lanes; entering an active link state; receiving a valid signal on the valid lane during the active link state, the valid signal comprising a signal maintained at a value for a defined first duration and indicating that data is to be received on the plurality of data lanes in a second defined duration subsequent to the first duration; and receiving the data during the active link state on the plurality of data lanes during the second defined duration.A system comprising means for carrying out the method of claim 25.A method comprising: transmitting one or more link training signals comprising instances of a link training pattern on the plurality of lanes of a physical link, the plurality of lanes comprising at least one valid lane and a plurality of data lanes; entering an active link state based on the training of the plurality of lanes using the link training signals; transmitting a valid signal on the valid lane during the active link state, the valid signal comprising a signal maintained at a value for a defined first duration and indicating that data is to be transmitted on the plurality of data lanes in a second defined duration subsequent to the first duration; and transmitting the data, during the active link state, on the plurality of data lanes during the second defined duration.A system comprising means for carrying out the method of claim 27.A system comprising: a connection comprising a plurality of lanes, the plurality of lanes including a plurality of own data lanes and at least one own valid signal lane; a first device; and a second device communicatively coupled to the first device using the connection, the second device to: send one or more link training signals comprising instances of a link training pattern to the first device on the plurality of lanes; enter an active link state based on the training of the plurality of lanes using the link training signals; a valid signal to send to the first device on the valid lane during the active link state, the valid signal comprising a signal that is maintained at a value across a first signal transmission window and indicates that data on the plurality of data lanes is to be sent in a second signal transmission window subsequent to the first signal transmission window; and to send the data during the active link state to the first device on the plurality of data lanes during the second signal transmission window.The system of claim 29, wherein the first device: is to receive the one or more link training signals; is to participate in training the plurality of lanes using the link training signals; enter the active link state; is to receive the valid signal during the first signal transmission window; and is to receive the data on the plurality of data lanes during the second signal transmission window.The system of claim 29, wherein the first device is to determine a phase interpolation to be applied to sampling the plurality of lanes based on the training signals.The system of claim 29, wherein the link comprises a multi-protocol, time-multiplexed link.The system of claim 32, wherein the plurality of lanes comprises at least one stream signal lane to identify which of a plurality of supported protocols to apply to data transmitted during a corresponding signal transmission window on the plurality of data lanes.
Citation Information
Patent Citations
Self-synchronizing bit error analyzer and circuit
US20090019326A1
De-correlating training pattern sequences between lanes in high-speed multi-lane links and interconnects
WO2014164004A1