Data rate increase for failed channel recovery in multi-channel data link
By identifying and reconfiguring the continuous working channel set of PCIe links, the data rate drop caused by fault channels is solved, and link bandwidth optimization and performance improvement are achieved.
Patent Information
- Application Number
- CN202380083082.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-14
- Filing Date
- 2023-11-01
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, in the recovery process of fault channels in PCIe links, data rate is difficult to effectively improve, resulting in link bandwidth loss and performance degradation.
By identifying the continuous working channel set and reconfiguring the start address and width of the link, using channel inversion technology, dynamically adjusting the link width and data rate to optimize data transmission after the fault channel recovery.
After the fault channel recovery, the effective utilization of link bandwidth is achieved, the data transmission rate and system performance are improved, and the power consumption is reduced.
Smart Images

Figure CN120303649A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This patent application claims the benefit of priority to pending U.S. Non - Provisional Application No. 18 / 081,396, filed on December 14, 2022, which is assigned to the assignee of the present application and is hereby incorporated by reference in its entirety as if fully set forth herein and for all applicable purposes. Technical Field
[0003] Broadly speaking, aspects of the present disclosure relate to multi - channel data links, and more particularly, to increasing the data rate for failed channel recovery. Background Art
[0004] High - speed interfaces are often used between the circuits and components of mobile wireless devices and other complex systems. For example, certain devices can include processing, communication, storage, and / or display devices that interact with each other through one or more high - speed interfaces. Some of these devices (including synchronous dynamic random access memory (SDRAM)) are capable of providing or consuming data and control information at the processor clock rate. Other devices (e.g., display controllers) can use a variable amount of data at a relatively low video refresh rate.
[0005] The Peripheral Component Interconnect Express (PCIe) interface is a popular high - speed interface that supports high - speed links capable of transmitting data at multiple gigabits per second. The interface also supports multiple speeds and multiple numbers of channels. Compared to parallel buses, PCIe provides lower latency and higher data transfer rates. PCIe is designated for communication between a wide variety of different devices. Generally, one device (e.g., a processor or a hub) acts as a host that communicates with multiple devices called endpoints through a PCIe link. Peripheral devices or components can include graphics adapter cards, network interface cards (NICs), storage accelerator devices, mass storage devices, input / output interfaces, and other high - performance peripheral devices. Summary of the Invention
[0006] The following presents an overview of one or more implementations in order to provide a basic understanding of such implementations. This overview is not an extensive overview of all contemplated implementations, but rather is intended neither to identify key or critical elements of all implementations nor to delineate the scope of any or all implementations. Its sole purpose is to present some concepts of one or more implementations in a simplified form as a prelude to the more detailed description that is presented later.
[0007] In one example, a method includes: detecting a failure of a failed channel of a data link, the data link having a plurality of channels labeled in a continuous sequence; detecting working channels of the data link; and selecting a set of continuous working channels of the data link. The method further includes: defining an operable link as including the selected set of continuous working channels; and transmitting data traffic over the operable link.
[0008] In another example, a non-transitory computer-readable medium has instructions stored therein that, when executed by a processor of an interconnect link, cause the processor to perform the operations of the above method.
[0009] In another example, a device includes interface circuitry configured to provide an interface to a multi-channel data link. A configuration register stores parameters of the data link. A controller is configured to: detect working channels of the data link; select a set of continuous working channels of the data link; define an operable link as including the selected set of continuous working channels; and transmit data traffic over the operable link.
[0010] In another example, a device includes: a unit for detecting a failure of a failed channel of a data link, the data link having a plurality of channels labeled in a continuous sequence; a unit for detecting working channels of the data link; a unit for selecting a set of continuous working channels of the data link; a unit for defining an operable link as including the selected set of continuous working channels; and a unit for transmitting data traffic over the operable link.
[0011] To achieve the foregoing and related purposes, one or more implementations include the features described in detail below and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative aspects of one or more implementations. However, these aspects merely indicate several ways in which the principles of the various implementations may be employed, and the described implementations are intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram of a computing architecture with a PCIe interface applicable to aspects of the present disclosure.
[0013] Figure 2 is a block diagram of a system including a host system and an endpoint device system according to aspects of the present disclosure.
[0014] Figure 3 is a diagram of channels and corresponding drivers in a link according to aspects of the present disclosure.
[0015] Figure 4Is a state diagram showing the operation of a power management state machine according to aspects of the present disclosure.
[0016] Figure 5A Is a diagram of a multi-channel data link having eight channels with channel 2 failed according to aspects of the present disclosure.
[0017] Figure 5B Is a diagram of a multi-channel data link having eight channels with channels 3 - 7 operating according to aspects of the present disclosure.
[0018] Figure 6A Is a diagram of a multi-channel data link having eight channels with channels 0 and 7 failed according to aspects of the present disclosure.
[0019] Figure 6B Is a diagram of a multi-channel data link having eight channels with channels 1 - 4 operating according to aspects of the present disclosure.
[0020] Figure 7A Is a diagram of a multi-channel data link having eight channels with channel 6 failed according to aspects of the present disclosure.
[0021] Figure 7B Is a diagram of a multi-channel data link having eight channels with channels 0 - 5 operating according to aspects of the present disclosure.
[0022] Figure 8A Is a diagram of a multi-channel data link having eight channels operating as x2 with channel 0 failed according to aspects of the present disclosure.
[0023] Figure 8B Is a diagram of a multi-channel data link having eight channels with channels 2 - 3 operating according to aspects of the present disclosure.
[0024] Figure 9 Is a flowchart of an exemplary method for increasing the data rate after the failed channel of a multi-channel data link is restored according to aspects of the present disclosure.
[0025] Figure 10 Is a flowchart of an exemplary method for increasing the data rate after the failed channel of a multi-channel data link is restored according to aspects of the present disclosure. Detailed Description
[0026] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of the various concepts. However, the concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.
[0027] In a PCIe system, for example, a root complex (RC) connected to an endpoint (EP) via a PCIe link, harsh environments and losses can cause faults in the channels of the link. In a controlled and accessible environment, a faulty link can be evaluated and resolved quickly. However, the use of multi-channel links such as PCIe has been well extended beyond such locations and for many different uses, such as vehicle applications and communication infrastructure including cellular radio towers. The reliability of the link can be affected by different types of hardware, firmware, or software faults as well as performance degradation over time. The various components and vendors, fault arrays, and scale challenges in a complete system make it challenging to monitor, collect data, and perform fault isolation on PCIe-based components.
[0028] In many cases, a channel fault will cause a greater restriction in the available link width configuration than may physically be necessary. In various aspects of the present disclosure, a set of continuously working channels can be identified. Additionally, a start address can be identified for the set of continuously working channels and stored in a configuration register. This allows the link to be established starting from any available working channel. As an example, if channel 1 of an x8 data link fails, by using channel 0 as the start address (as is typically assumed), only channel 0 can be used. This gives a maximum link width of x1. However, if a new start address for channel 2 can be identified and stored in the configuration register, then channels 2 through 7 become available and link widths of x1, x2, and x4 can be used.
[0029] In addition, as described herein, the link width can be selected and stored in a configuration register. In an example where there is a channel 1 failure in an x8 data link, no more than 4 channels can be used, even if 7 channels are operational, including 6 contiguous channels. This also applies to using channel inversion, where the same 6 contiguous channels 7 to 2 are available, but the maximum link width is x4, i.e., using channels 7, 6, 5, and 4. By selecting all 6 contiguous channels, i.e., channels 2 to 7, more data can be carried on the link. To allow the RC and EP to communicate using 5 or 6 channels instead of 4 channels, in addition to the start address, additional parameters are identified and stored in the configuration register. The additional parameters can be the link width, where additional link widths, such as x5 and x6, are allowed. Alternatively or additionally, the end address can be identified and stored.
[0030] In the above example of an x5 link starting from channel 2, the start address at channel 2 can be stored together with the link width of x5. The end address will be channel 6. With channel inversion, the start address can be channel 7, where the end address is channel 3 and the link width is x5. By specifically identifying the start and end addresses, a different set of channels can be used than the conventionally allowed set of channels. Instead of using channel 2 as the start address, the RC or EP can identify channel 3 as the start address of an x5 link ending at channel 7. For larger data links, such as x16 and x32, there may be several different possible start addresses available after a channel failure.
[0031] PCIe is a point-to-point interconnect that supports both cross-cable assemblies or internal and external connectivity at the printed circuit board (PCB) level. The connection can be made chip-to-chip without a connector, through an expansion card interface with a board and connector, or on a backplane with multiple boards and connectors. In a complex backplane, there are many reasons why signal integrity can degrade, including crosstalk, reflections, discontinuities, and channel loss. Higher data rates are also more sensitive to loss.
[0032] The connection between any two PCIe devices (e.g., RC and EP) is called a link. A PCIe link is built around a duplex, serial (1-bit), differential, point-to-point connection called a channel. With PCIe, data is transmitted on two signal pairs: two lines (wires, printed circuit board traces, etc.) for transmission and two lines for reception. The transmit and receive pairs are separate differential pairs of a total of four data lines per channel. The link includes a set of channels, and each channel is capable of simultaneously transmitting and receiving data packets between the host and the endpoint.
[0033] In PCIe, multiple consecutive channels are combined to create a link. The first channel is used in an x1 link, referred to herein as Channel 0. The first and second channels are used in an x2 link, referred to herein as Channel 0 and Channel 1. Each link width starts at Channel 0 and additional consecutive channels are added in sequence to achieve the desired link width. A PCIe link can be configured to use fewer than all of its channels to save power or achieve compatibility. As currently defined, the link width can increase in size to a larger width or decrease in size to a smaller width during operation. The current link widths x1, x2, x4, x8, x16, and x32 are permissible for PCIe data links. Using this system, if Channel 3 of an 8-channel data link fails, only Channels 0-1 can be used. The data link can be configured as x1 or x2, but not x4 or x8. Due to the failure of Channel 3, Channels 4-7 may no longer be used. In an extension of PCIe, channel inversion allows the last channel of a sequence to be used as the first channel. With channel inversion, when Channel 3 fails, then Channels 7-4 can be used. This allows an x1 link using Channel 7, an x2 link using Channels 7 and 6, or an x4 link using Channels 7, 6, 5, and 4. If both Channel 0 and Channel 7 fail, a link cannot be formed with or without channel inversion.
[0034] PCIe allows the link width to be changed when the data traffic requirements change. Using a record of which channels have failed, the RC or EP can select a different set of consecutive working channels to meet different link widths. PCIe also allows different data rates, referred to as GEN speeds. The link speed is called GEN speed because newer generations of the PCIe standard provide new higher speeds. For example, PCIe GEN1 allows 2.5 gigatransfers per second (GT / s), PCIe GEN2 allows 5 GT / s, PCIe GEN3 allows 8 GT / s, PCIe GEN4 allows 16 GT / s, PCIe GEN5 allows 32 GT / s, and later generations can provide higher data rates for each channel of the link. Higher speeds provide data transfer benefits but also consume more power and are more error-prone. Therefore, the PCIe link can be managed to operate at a lower GEN speed until there is a high data rate requirement on the link.
[0035] When an RC or EP requests a GEN speed change, the PCIe link is disabled from operating at the first speed and then training, recovery, and configuration are performed at the newly requested GEN speed. After that, the PCIe link becomes active again and is ready to carry data at the newly requested GEN speed. The link width can also be changed simultaneously. The process is the same for increases or decreases in speed and link width. The RC and EP negotiate an appropriate combination of speed and width to carry the expected data traffic.
[0036] Figure 1 is a block diagram of an example computing architecture that uses a PCIe interface. Computing architecture 100 operates using multiple high-speed PCIe interface serial links. The PCIe interface can be characterized as a device that includes a point-to-point topology, where individual serial links connect each device to a host known as a root complex 104 (RC). In computing architecture 100, root complex 104 couples processor 102 to memory devices (e.g., memory subsystem 108) and PCIe switch circuitry 106. In some cases, PCIe switch circuitry 106 includes cascaded switching devices. One or more PCIe endpoint devices 110 (EP) can be directly coupled to root complex 104, while other PCIe endpoint devices 112-1, 112-2, …, 112-N can be coupled to root complex 104 through PCIe switch circuitry 106. Root complex 104 can be coupled to processor 102 using a proprietary local bus interface or a standard-defined local bus interface. Root complex 104 can control configuration and data transactions through the PCIe interface and can generate transaction requests for processor 102. In some examples, root complex 104 is implemented in the same integrated circuit (IC) device that includes processor 102. Root complex 104 supports multiple PCIe ports.
[0037] Root complex 104 can control the communication between processor 102 and memory subsystem 108, which is an example of an endpoint. Root complex 104 also controls the communication between processor 102 and other PCIe endpoint devices 110, 112-1, 112-2, …, 112-N. The PCIe interface can support full-duplex communication between any two endpoints, with no inherent limit on concurrent access across multiple endpoints. Data packets can carry information through any PCIe link. In a multi-channel PCIe link, packet data can be striped across multiple channels. The number of channels in a multi-channel link can be negotiated during device initialization and can be different for different endpoints.
[0038] When one or more lanes of a PCIe link are not fully utilized by a low-bandwidth application that can be adequately served by fewer lanes, the root complex 104 and the endpoint can operate a link with more or fewer lanes. In some examples, one or more lanes can be placed in one or more standby states, where some or all of the lanes operate in a low-power or no-power mode. Changing the number of active lanes for a low-bandwidth application reduces the power consumed to operate the link. Supplying less power reduces current leakage, heat, and power consumption.
[0039] Figure 2 is a block diagram of an exemplary PCIe system in which aspects of the present disclosure may be implemented. System 205 includes a host system 210 and an endpoint device system 250. The host system 210 may be integrated on a first chip (e.g., a system-on-chip or SoC), and the endpoint device system 250 may be integrated on a second chip. Alternatively, the host system (e.g., RC) and / or the endpoint device (EP) system may be integrated in a first package and a second package (e.g., a SiP, a first system board and a second system board having multiple chips) or in other hardware or any combination thereof. In this example, the host system 210 and the endpoint device system 250 are coupled by a PCIe link 285.
[0040] The host system 210 includes one or more host clients 214. Each of the one or more host clients 214 may be implemented on a processor that executes software that performs the functions of the host client 214 discussed herein. For examples with more than one host client, the host clients may be implemented on the same processor or different processors. The host system 210 also includes a host controller 212, which may perform root complex functions. The host controller 212 may be implemented on a processor that executes software that performs the functions of the host controller 212 discussed herein.
[0041] The host system 210 includes a PCIe interface circuit 216, a system bus interface 215, and a host system memory 240. The system bus interface 215 can interface one or more host clients 214 with the host controller 212, and interface each of the one or more host clients 214 and the host controller 212 with the PCIe interface circuit 216 and the host system memory 240. The PCIe interface circuit 216 provides the host system 210 with an interface to a PCIe link 285. In this regard, the PCIe interface circuit 216 is configured to send data (e.g., from a host client 214) to the endpoint device system 250 via the PCIe link 285, and receive data from the endpoint device system 250 via the PCIe link 285. The PCIe interface circuit 216 includes a PCIe controller 218, a physical interface for a PCI Express (PIPE) interface 220, a physical (PHY) transmit (TX) block 222, a clock generator 224, and a PHY receive (RX) block 226. The PIPE interface 220 provides a parallel interface between the PCIe controller 218 and the PHY TX block 222 and the PHY RX block 226. The PCIe controller 218 (which may be implemented in hardware) can be configured to perform the transaction layer, data link layer, and control flow functions specified in the PCIe specification, as further discussed below.
[0042] The host system 210 further includes an oscillator (e.g., a crystal oscillator or “XO”) 230 configured to generate a reference clock signal 232. In one example, the reference clock signal 232 may have a frequency of 19.2 MHz, but is not limited to this frequency. The reference clock signal 232 is input to the clock generator 224, and the clock generator 224 generates a plurality of clock signals based on the reference clock signal 232. In this regard, the clock generator 224 may include a phase-locked loop (PLL) or multiple PLLs, where each PLL generates a corresponding clock signal of the plurality of clock signals by amplifying the frequency of the reference clock signal 232.
[0043] The endpoint device system 250 includes one or more device clients 254. Each device client 254 may be implemented on a processor that executes software that performs the functions of the device client 254 discussed herein. For examples with more than one device client 254, the device clients 254 may be implemented on the same processor or different processors. The endpoint device system 250 further includes a device controller 252. The device controller 252 can be configured to receive bandwidth requests from one or more device clients, and determine whether to change the number of active channels or change the GEN speed based on the bandwidth requests. The device controller 252 may be implemented on a processor that executes software that performs the functions of the device controller.
[0044] The endpoint device system 250 includes a PCIe interface circuit 260, a system bus interface 256, and an endpoint system memory 274. The system bus interface 256 can interface one or more device clients 254 with a device controller 252 and interface each of the one or more device clients 254 and the device controller 252 with the PCIe interface circuit 260 and the endpoint system memory 274. The PCIe interface circuit 260 provides an interface to the PCIe link 285 for the endpoint device system 250. In this regard, the PCIe interface circuit 260 is configured to send data (e.g., from the device client 254) to the host system 210 (also referred to as the host device) via the PCIe link 285 and receive data from the host system 210 via the PCIe link 285. The PCIe interface circuit 260 includes a PCIe controller 262, a PIPE interface 264, a PHY TX block 266, a PHY RX block 270, and a clock generator 268. The PIPE interface 264 provides a parallel interface between the PCIe controller 262 and the PHY TX block 266 and the PHY RX block 270. The PCIe controller 262 (which may be implemented in hardware) can be configured to perform transaction layer, data link layer, and control flow functions.
[0045] The host system memory 240 at the endpoint and the endpoint system memory 274 can be configured to contain registers for the configuration and status of each channel of the PCIe link 285 and the link itself. These registers include group control registers and group status registers. In an example, both the host system memory 240 and the endpoint system memory 274 have link GEN control registers, status registers, capability registers, etc.
[0046] The endpoint device system 250 also includes an oscillator (e.g., a crystal oscillator) 272, which is configured to generate a stable reference clock signal 273 for the endpoint system memory 274. In Figure 2 the example, the clock generator 224 at the host system 210 is configured to generate a stable reference clock signal 273, which is forwarded to the endpoint device system 250 by the PHY RX block 226 via a differential clock line 288. At the endpoint device system 250, the PHY RX block 270 receives the EP reference clock signal on the differential clock line 288 and forwards the EP reference clock signal to the clock generator 268. The EP reference clock signal may have a frequency of 100 MHz, but is not limited to this frequency. The clock generator 268 is configured to generate multiple clock signals based on the EP reference clock signal from the differential clock line 288, as further discussed below. In this regard, the clock generator 268 may include multiple PLLs, where each PLL generates a corresponding clock signal among the multiple clock signals by amplifying the frequency of the EP reference clock signal.
[0047] System 205 also includes a power management integrated circuit (PMIC) 290 coupled to a power supply 292 (e.g., a mains voltage, a battery, or other power supply). The PMIC 290 is configured to convert the voltage of the power supply 292 into multiple supply voltages (e.g., using a switching regulator, a linear regulator, or any combination thereof). In this example, the PMIC 290 generates a voltage 242 for the oscillator 230, a voltage 244 for the PCIe controller 218, and a voltage 246 for the PHY TX block 222, the PHY RX block 226, and the clock generator 224. The voltages 242, 244, and 246 can be programmable, where the PMIC 290 is configured to set the voltage levels (corners) of the voltages 242, 244, and 246 according to instructions (e.g., from the host controller 212).
[0048] The PMIC 290 also generates a voltage 280 for the oscillator 272, a voltage 278 for the PCIe controller 262, and a voltage 276 for the PHY TX block 266, the PHY RX block 270, and the clock generator 268. The voltage 280, the voltage 278, and the voltage 276 can be programmable, where the PMIC 290 is configured to set the voltage levels (corners) of the voltages 280, 278, and 276 according to instructions (e.g., from the device controller 252). The PMIC 290 can be implemented on one or more chips. Although the PMIC 290 is shown as one PMIC in Figure 2 it should be understood that the PMIC 290 can be implemented by two or more PMICs. For example, the PMIC 290 can include a first PMIC for generating the voltages 242, 244, and 246 and a second PMIC for generating the voltages 280, 278, and 276. In this example, both the first and second PMICs can be coupled to the same power supply 292 or different power supplies.
[0049] In operation, the PCIe interface circuit 216 on the host system 210 can send data from one or more host clients 214 to the endpoint device system 250 via the PCIe link 285. When the host controller negotiates the bandwidth for the link, the data from one or more host clients 214 can be directed to the PCIe interface circuit 216 according to the PCIe mapping established by the host controller 212 during an initial configuration (sometimes referred to as link initialization). In the example, the host controller negotiates a first bandwidth for the transmit group of the link and a second bandwidth for the receive group of the link. At the PCIe interface circuit 216, the PCIe controller 218 can perform transaction layer and data link layer functions on the data, such as packing the data, generating error correction codes to be sent with the data, etc.
[0050] The PCIe controller 218 outputs the processed data to the PHY TX block 222 via the PIPE interface 220. The processed data includes data from one or more host clients 214 and overhead data (e.g., packet headers, error correction codes, etc.). In one example, the clock generator 224 may generate a clock 234 for an appropriate data rate or transmission rate based on the reference clock signal 232, and input the clock 234 to the PCIe controller 218 to time the operation of the PCIe controller 218. In this example, the PIPE interface 220 may include a 22-bit parallel bus that transfers 22-bit data to the PHY TX block in parallel for each cycle of the clock 234. At 250 MHz, this translates to a transmission rate of approximately 8 GT / s.
[0051] The PHY TX block 222 serializes the parallel data from the PCIe controller 218 and drives the PCIe link 285 with the serialized data. In this regard, the PHY TX block 222 may include one or more serializers and one or more drivers. The clock generator 224 may generate a high-frequency clock for one or more serializers based on the reference clock signal 232.
[0052] At the endpoint device system 250, the PHY RX block 270 receives the serialized data via the PCIe link 285 and deserializes the received data into parallel data. In this regard, the PHY RX block 270 may include one or more receivers and one or more deserializers. The clock generator 268 may generate a high-frequency clock for one or more deserializers based on the EP reference clock signal. The PHY RX block 270 transfers the deserialized data to the PCIe controller 262 via the PIPE interface 264. The PCIe controller 262 may recover the data from one or more host clients 214 from the deserialized data and forward the recovered data to one or more device clients 254.
[0053] On the endpoint device system 250, the PCIe interface circuit 260 can send data from one or more device clients 254 to the host system memory 240 via the PCIe link 285. In this regard, the PCIe controller 262 at the PCIe interface circuit 260 can perform transaction layer and data link layer functions on the data, such as packing the data, generating error correction codes to be sent with the data, etc. The PCIe controller 262 outputs the processed data to the PHY TX block 266 via the PIPE interface 264. The processed data includes data from one or more device clients 254 and overhead data (e.g., packet headers, error correction codes, etc.). In one example, the clock generator 268 can generate a clock based on the EP reference clock through the differential clock line 288 and input the clock to the PCIe controller 262 to time the operation of the PCIe controller 262.
[0054] The PHY TX block 266 serializes the parallel data from the PCIe controller 262 and uses the serialized data to drive the PCIe link 285. In this regard, the PHY TX block 266 can include one or more serializers and one or more drivers. The clock generator 268 can generate a high-frequency clock for one or more serializers based on the EP reference clock signal.
[0055] At the host system 210, the PHY RX block 226 receives the serialized data via the PCIe link 285 and deserializes the received data into parallel data. In this regard, the PHY RX block 226 can include one or more receivers and one or more deserializers. The clock generator 224 can generate a high-frequency clock for one or more deserializers based on the reference clock signal 232. The PHY RX block 226 transmits the deserialized data to the PCIe controller 218 via the PIPE interface 220. The PCIe controller 218 can recover the data from one or more device clients 254 from the deserialized data and forward the recovered data to one or more host clients 214.
[0056] Figure 3 can be at Figure 1 and Figure 2Diagram of channels in link 385 (e.g., PCIe link 285) used in the system. In this example, link 385 includes multiple channels 310-1 to 310-n, where each channel includes corresponding first differential line pairs 312-1 to 312-n for sending data from host system 210 to endpoint device system 250, and corresponding second differential line pairs 315-1 to 315-n for sending data from endpoint device system to host system 210. From the perspective of the host system, the first channel 310-1 is dual simplex, where the first differential line pair 312-1 serves as the transmission line and the second differential line pair 315-1 serves as the reception line. From the perspective of the endpoint device system, the first channel 310-1 has a reception line and a transmission line. The first differential line pairs 312-1 to 312-n and the second differential line pairs 315-1 to 315-n can be implemented using metal traces on a substrate (e.g., printed circuit board), where the host system can be integrated on a first chip mounted on the substrate, and the endpoint device is integrated on a second chip mounted on the substrate. Alternatively, the link can be implemented through an adapter card slot, a cable, or a combination of different media. The link can also include an optical portion where PCIe packets are encapsulated within different systems. In this example, when sending data from the host system to the endpoint device system across multiple channels, the PHY TX block 222 can include logic for dividing the data among the channels. Similarly, when sending data from the endpoint device system to the host system 210 across multiple channels, the PHY TX block 266 can include logic for dividing the data among the channels.
[0057] Figure 2 The PHY TX block 222 of the host system 210 shown in can be implemented to include transmission drivers 320-1 to 320-n for driving each of the first differential line pairs 312-1 to 312-n to send data, and Figure 2 The PHY RX block 270 of the host shown in can be implemented to include receivers 340-1 to 340-n (e.g., amplifiers) for driving each of the second differential line pairs 312-1 to 312-n to receive data. Each transmission driver 320-1 to 320-n is configured to drive the corresponding differential line pair 312-1 to 312-n with data, and each receiver 340-1 to 340-n is configured to receive data from the corresponding first differential line pairs 312-1 to 312-n. Additionally, in Figure 2In the figure, the PHY TX block 266 of the endpoint device system 250 may include transmit drivers 345-1 to 345-n for each of the second differential line pairs 315-1 to 315-n, and the PHY RX block 226 of the host system 210 may include receivers 325-1 to 325-n (e.g., amplifiers) for each of the second differential line pairs 315-1 to 315-n. Each transmit driver 345-1 to 345-n is configured to drive the corresponding second differential line pair 315-1 to 315-n with data, and each receiver 325-1 to 325-n is configured to receive data from the corresponding second differential line pair 315-1 to 315-n.
[0058] In some aspects, the width of the link 385 may be scaled to match the capabilities of the host system and the endpoint. The link may use one channel (i.e., the first channel 310-1) for an x1 link, two channels (i.e., 310-1, 310-2) for an x2 link, or more channels (up to n channels from 310-1 to 310-n) for a wider link. Although the current link is defined for 1, 2, 4, 8, 16, and 32 channels, different numbers of channels may be used to accommodate a particular implementation.
[0059] In one example, the host system 210 may include a power switch circuit 350 that is configured to individually control the power from the PMIC 290 to the transmit drivers 320-1 to 320-n and the receivers 325-1 to 325-n. Thus, in this example, the number of powered-on drivers and receivers is proportional to the width of the link 385. Similarly, as Figure 2 shown in the endpoint device system 250 may include a power switch circuit 360 that is configured to individually control the power from the PMIC 290 to the transmit drivers 345-1 to 345-n and the receivers 340-1 to 340-n. In this way, the host system sets the number of a plurality of drivers to be selectively powered by the power switch circuit based on the number of powered-on lines to change the number of active transmit or receive lines. With differential signaling, the lines will be set in pairs as active or standby.
[0060] Figure 4 is a state diagram 400 showing the operation of a part of a power management state machine according to certain aspects disclosed herein. The Active State Power Management (ASPM) protocol is a state machine method for reducing power based on link activity detected on a PCIe link between a root complex (RC) and an endpoint (EP) PCIe device. The state diagram is consistent with the Link Training and Status State Machine (LTSSM) defined for PCIe. However, other methods may be used alternatively.
[0061] Power management and bandwidth negotiation can be performed at link initialization, but can also be repeated at a later time. During negotiation, each link partner (e.g., RC and EP) can advertise the number of supported lanes (e.g., link width) and the desired bandwidth at which to operate. For example, the link partners can agree to operate at the highest bandwidth supported by both partners. A number of lanes of the link that the link partners are negotiating are active, and the number of lanes can be changed to a lower rate for link stability reasons. In an example, the link width can be changed autonomously by hardware. As the number of lanes increases, the power used to operate the link also increases. Thus, in some cases, a x16 link can operate at a lower power as a x1 link. This reduces the power consumed by the supporting hardware during low activity periods.
[0062] In the illustrated ASPM method, when data is transmitted over a PCIe link, the link operates in the L0 state 408 (i.e., the link operating state). There can be other operable states, such as the L0p sub-state 418 and the L0s state 410, etc. The L0p sub-state 418 is not an inactive or idle state. When the L0p sub-state 418 is enabled, the link can be configured to the L0p active state, where, based on data rate requirements, the size of the link can be reduced by placing some lanes in electrical idle while still carrying data on other lanes that are in the active L0p sub-state. Similarly, for an increase size request, the electrically idle lanes can be retrained independently to the L0p active sub-state without disturbing ongoing data transmission on other active lanes in the link. The ASPM protocol can also support additional active and standby states and sub-states, such as L1, L2, L1.1, and L1.2, etc. As shown, the L0s state 410 can only be accessed via a connection through the L0 state 408. When the conditions on the link dictate or suggest that a transition between states is desirable or appropriate, an ASPM state change can be initiated. Two communication partners on the link can initiate a power state change request when the conditions are right.
[0063] The PCIe link initiates operations in the Detect state 402. In this state, the host controllers of the RC and EP detect the active connection. The link then moves to the Poll state 404, during which the RC polls the host controller for any active EP connections. Similarly, the EP host controller polls the RC host controller. This allows the available link width and GEN speed to be determined. Any faulty channels will be detected as non-responsive or unreliable during polling. The PCIe link then moves to the Configure state 406, during which the RC and EP ports are configured for a specific link width and GEN speed. Some or all channels are configured as active, and any unused channels are configured as idle, e.g., in the L0s state 410. The active channels are in the L0 state 408. Faulty channels can be powered off.
[0064] When the link is idle (e.g., during a short interval between data bursts), the link can be taken from the L0 state 408 to a standby state, e.g., the L0s state 410 that is accessible only through the connection to the L0 state 408. In this example, the L0s state 410 is a low-power standby of the L0 state 408. The L1 state 416 is a standby state with a lower latency than the L0s state 410. The L0s state 410 acts as a standby state and also acts as an initialization state after power-on, system reset, or after detecting an error condition. In the L0s state 410, device discovery and bus configuration processes can be implemented before the link transitions 428 to the L0 state 408. In the L0 state 408, the PCIe device can be active and responsive to PCIe transactions, and / or can request or initiate PCIe transactions. The L1 state 416 is the primary standby state and allows for a quick return to the L0 state 408 through the Recovery state 412. The L0s state 410 is a lower-power state that allows for an electrical idle state and can be transitioned through the Recovery state 412. When the PCIe device determines that there are no outstanding PCIe requests or pending transactions, it can enter the L0s state 410 through the transition 426. Power consumption can be reduced by disabling or idling the transceivers in the PCIe bus interface, disabling the clock used by the PCI device, and / or disabling the PLL circuit that generates the clock for receiving data. The PCIe device can make the transition 426 to the L0s state 410 through the operation of a hardware controller or some combination of the operating system and hardware control circuits.
[0065] When the PCIe link becomes active while the device (e.g., RC or EP) is in the electrical idle state (e.g., L0s state 410), a return to the L0 state 408 is initiated for the device. A direct transition to the L0 state 408 may not be available. The PCIe link may first transition 422 to the recovery state 412, in which the transceivers in the PCIe bus interface, the clocks used by the PCI device, and / or the PLL circuits are enabled. When the transceivers and other circuits are determined to be operational, a transition 430 from the recovery state 412 to the L0 state 408 can be initiated.
[0066] The ASPM protocol also manages GEN speed changes of the PCIe link. The PCIe link can operate at a lower GEN speed to save power or at a higher GEN speed to provide higher performance. To change the GEN speed, the ASPM protocol takes the active lanes out of the L0 state 408 back to the configuration state 406 to configure the link for the new GEN speed. After configuration, the PCIe link is returned to the L0 state to operate at the newly configured GEN speed. In some aspects, when there are some lanes in the L0 state 408 and other lanes in the L0s state 410, the lanes in the L0 state 408 can continue to operate in the L0 state 408 while the other lanes transition to the configuration state 406. The lanes in the configuration state 406 are configured for the new GEN speed and then transition to the L0 state 408 to operate at the new GEN speed. The lanes in the L0 state can transition to an idle state, e.g., the L0s state 410. These lanes can be reconfigured to be in the configuration state 406 or can remain in the L0s state 410 for later use.
[0067] The ASPM protocol can also manage the link width to reduce or increase the link width, also referred to as sizing up the link width or sizing down the link width. By reducing the link width during low throughput data traffic scenarios, the subsystem of the PCIe link scales down the voltage levels (e.g., to a lower operating level that meets the current throughput on the PCIe link). Scaling down one or more voltage levels reduces power consumption (e.g., reduces leakage current during sustained low throughput traffic or in idle use cases). The number of lanes also affects power consumption. In fact, there is an L0, L0s, L1 transition state diagram for each lane.
[0068] When a faulted channel has been detected through polling or during operation and the faulted channel is active, then a new link width excluding the faulted channel is configured. There may be retry attempts to activate the faulted channel. However, if the channel cannot be used, the link width must be changed to exclude the faulted channel. When a faulted channel is detected with L0p enabled and the number of channels in the electrical idle state is equal to or greater than the current active channels in number, then the link may be reconfigured to maintain the original link width. When an active channel fails, the PCIe data transfer is suspended and the link repeats training by re-entering the Detect State 402. Link training with a faulted channel may take a significant amount of time, e.g., 2 ms, and the subsequently configured link may have a lower link width compared to the link before the channel fault. By changing the start address of the subsequently configured link, the highest available link width can be maintained.
[0069] Figure 5A is a diagram of a multi-channel data link 504, e.g., a PCIe link of a duplex traffic channel between a host 502 and an EP 506. Each channel includes lines in each direction such that the data rate and configuration are the same in both directions. The duplex traffic channel may have the same physical structure as in Figure 3 but is generalized to show the transition from one link width and channel configuration to another. The data link 504 includes eight channels. The duplex traffic channel is the same in structure and capabilities as in the Figure 3 example shown. For the four lines per channel and the thirty-two lines of an x8 link, each channel includes two transmit lines as a differential pair and two receive lines as a differential pair. All channels 510 (channels 0 to 7) identified on the right are active, e.g., in the L0 state, as an x8 link. Data can be sent in one or both directions at a specific data rate or speed (e.g., GEN2 speed) configured. The configuration of the link can be stored in a configuration, control, or capability register 514 and includes a start address 0, an end address 7, and a width x8. In some examples, the register 514 can be maintained in the system memory 240, 274 of the PCIe controller or in another location. An address table with start and end addresses or start address and link width can be maintained at both the RC and the host.
[0070] One of the channels (e.g., channel 2) is a faulty channel 512 and cannot send any traffic. As a result, the entire link is deactivated regardless of the state of any of the channels. The address 2 of the faulty channel will also be stored in the capability register 514. Using the first channel as the start address, the faulty channel (channel 2) results in the loss of 6 channels, even though 5 channels are still operational. This results in a loss of bandwidth through the data link. For x16 or x32 links, the loss will be even more significant, both in absolute terms and as a proportion of the total. With channel inversion, an x4 link can be established starting from channel 7 and including channels 6, 5, and 4. Although eight channels are shown, there can be more or fewer channels to suit a particular implementation. As described above, the PCIe data link operates based on Rule 2 such that x1, x2, x4, x8, x16, and x32 are supported. The data link 504 is configured as an x8 data link, but the link width can be configured as x1, x2, or x4 at any GEN speed to operate more efficiently for the data rate carried by the link.
[0071] Figure 5B A multi-channel data link 504 between a host 502 and an EP 506 after being reconfigured with a new start address and a new link width is shown in accordance with aspects of the present disclosure. There may also be training for a new data rate. In one aspect, when the controller at the host 502 or EP 504 determines that channel 2 has failed, it will determine how to change the link and then send a request to change the start address and link width of the link 504. The data rate of the data link 504 can also change or adapt to changing traffic demands based on the change in link width. The request can be generated by the host controller or the EP controller and then sent to the corresponding controller at the opposite end (e.g., the EP controller or the host controller) respectively.
[0072] In Figure 5B the example, after detecting the faulty channel, a set of consecutive channels with the maximum link width is selected. Polling and configuration status start at the start address or start channel pointer of the selected set of channels. The other working channels of the link can be detected as channels 0, 1, 3, 4, 5, 6, and 7. The consecutive channels are the first set of 0 and 1 and the second set of 3 to 7. The second set is larger and is therefore selected as the operable link 518. This allows the maximum possible link to be used with the remaining channels. Thus, the channels that are not part of the operable link are transitioned to an idle state, e.g., the electrical idle state. This configuration is stored in the register 516 as a start address 3, an end address 6, and a link width of x4. It is also possible to identify the start address as 4 and still maintain the x4 link width. Note that in the normal configuration where channel 0 is the start address, only x1 or x2 link widths are possible.
[0073] Specific locations of the configuration register can be adapted to accommodate different situations. In one aspect, the PCIe specification defines a Device Control 3 register, in which 32 bits are allocated. The fourth bit (i.e., bit 3) is used to indicate whether the L0p state is enabled. The next three bits 4 to 6 ([6:4]) are used to store the target link width. This allows for 8 possible link widths, 6 of which are already defined. The remaining bits [31:7] are reserved. These can be used for start and end addresses. Additionally, one or more than one bit can be used together with the link width bits [6:4] to allow for additional link widths between the link widths allowed by the PCIe data link.
[0074] Once the new operable links with the consecutive working channels 3, 4, 5, and 6 are also brought to the requested data rate, GEN speed, the data traffic of the PCIe link can be carried on these channels. This can be done, for example, by sending an interrupt from the RC controller to the host 502 and the EP 506 to stop the channels that are attempting to use the original x8 link 510 and start using the new operable link 518. All unused channels can be placed directly in electrical idle before or after the new operable link becomes active. Alternatively, one or more of the working channels can be retrained to the new data rate and then placed in electrical idle. If there are other channels in electrical idle that are not being used during this process, those channels can also be trained to the new data rate and then placed in electrical idle. Retraining makes the idle channels available for link width increment requests at the new data rate. In this case, only the link width increment would be adding channel 7 for an x5 link, which can be supported if odd-sized link widths are enabled between the RC and the EP.
[0075] Figure 6A A multi-channel data link 604 is shown, e.g., a PCIe link of a duplex traffic channel between a host 602 and an EP 606. Each channel includes lines in each direction such that the data rate and configuration are the same in both directions. The duplex traffic channel can have the same physical structure as Figure 3 but is generalized to show the transition from one link width and channel configuration to another. The data link 604 includes eight channels. All the channels 510 (channels 0 to 7) identified on the right are active, e.g., in the L0 state, as an x8 link. The configuration of the link can be stored in a configuration, control, or capability register 614 and includes a start address 0, an end address 7, and a width x8.
[0076] Two channels in the channel (channel 0 and channel 7) are faulty channels 612 and cannot send any traffic. As a result, the entire link is disabled regardless of the status of any channel. The addresses 0 and 7 of the faulty channels will also be stored in the capability register 614. Using the first channel as the start address, the link is lost and cannot be damaged. Using channel inversion where the first channel is channel 7, the link is also lost and cannot be restored.
[0077] Figure 6B Illustrated is a multi-channel data link 604 between a host 602 and an EP 606 after being reconfigured with a new start address and a new link width in accordance with aspects of the present disclosure. In one aspect, when the controller at the host 602 or EP 604 determines that channels 0 and 7 have failed, it identifies the start address (1 in this case) and the end address (4 in this case) and establishes the link width (x4 in this case). This leaves channels 5 and 6 unused. As an alternative, an x4 link with a start address of 2 or 3 can be established. Thus, channels not part of the operable link are transitioned to an idle state, e.g., an electrical idle state. This configuration is stored in register 616 as start address 1, end address 4, and link width x4. In this example, the data link 604 is restored even though it may have been lost and non-restorable.
[0078] Figure 7A Illustrated is a multi-channel data link 704, e.g., a PCIe link of a duplex traffic channel between a host 702 and an EP 706. The data link 704 includes eight channels. All channels 710 (channels 0 to 7) identified on the right are active, e.g., in the L0 state, as an x8 link. The configuration of the link can be stored in a configuration, control, or capability register 614 and includes a start address 0, an end address 7, and a width x8. In this example, the faulty channel 712 is channel 6. Thus, there are six consecutive channels, channels 0 to 5 are working channels. Channel 7 is a working channel.
[0079] Using the first channel as the start address, an x4 link can be established using channels 0 to 3. Alternative x4 links can be established using channels 1 to 4 and 2 to 5. Using additional bits for the link width, other link widths can be supported, such as x5 with channels 1 to 4, or 1 to 5, or x6 with all consecutive working channels.
[0080] Figure 7BIllustrated is a multi-channel data link 704 between a host 702 and an EP 706 after being reconfigured with a new start address and a new link width outside the allowable link widths for a PCIe data link. In this example, the link width is x6, which provides 50% more data carrying capacity than an x4 link. This configuration is stored in register 716 as start address 0, end address 5, and link width x6. In this example, the link can be identified by the end address or the link width, or both can be stored in the configuration register 716.
[0081] Figure 8A Illustrated is a multi-channel data link 804, e.g., a PCIe link of a duplex traffic channel between a host 802 and an EP 806. The data link 804 includes eight channels. Only two channels are used as an x2 link 810 in the active state. The remaining channels of the link (channels 2 to 7) are in the idle state. The configuration of the link can be stored in a configuration, control, or capability register 714 and includes start address 0, end address 1, and width x2. In this example, the failed channel 812 is channel 1. Thus, there are six consecutive channels, and channels 2 to 7 are working channels. Channel 1 is also a working channel.
[0082] To obtain two consecutive channels for a new x2 link without a failed channel 810, the controller can select any of the five different channel pairs. In this example, the next two channels (channels 2 and 3) are selected as the set of consecutive working channels of the data link. Figure 8B Illustrated is a multi-channel data link 804 between a host 802 and an EP 806 after being reconfigured with a new start address and the same link width. This configuration is stored in register 816 as start address 2, end address 3, and link width x2. In this example, the link can be identified by the end address or the link width, or both can be stored in the configuration register 816.
[0083] Because the channels (channels 2 and 3) of the new operable link are not part of the original data link (channels 0 and 1), the new operable link can be configured while the original data link is still in operation. The data traffic can be rerouted to these channels by first bringing channels 2 and 3 to the active state by restoring the state. Then, after the traffic has been configured, the traffic can be transmitted to the fully operable link. This additional technique can be used at any time when the new operable link does not include any of the channels of the data link with channel failures (channels 0 and 1).
[0084] Figure 9FIG. 900 is a flow chart of a method for increasing data rate after failure channel recovery of a multi-channel data link (e.g., a PCIe link) in accordance with aspects of the present disclosure. In some aspects, method 900 involves defining an operable link starting from an assignable start address on the link without being restricted to the first or last channel. The link width can also be defined with different variations and is not limited to multiples of two. In some aspects, a new operable link can be prepared before transmitting data so that data is not lost during recovery and configuration states. The method for recovering from a failed channel in a PCIe link can be performed by PCIe controllers 218, 262, which include an RC controller (also referred to as a host controller) and an EP controller. There is negotiation for establishing capabilities and bandwidth requirements. Once the negotiation is complete, the RC controller triggers method 900 to bring a different set of consecutive working channels for the data link. Once the channels are working, the RC controller generates an interrupt to transfer the ongoing data transmission from the first set of active channels to the second newly trained set of channels.
[0085] Method 900 includes, at block 902, detecting a failure of a failed channel of a data link having a plurality of channels labeled in a consecutive sequence. The data link has a first set of channels in an active state and a second set of channels in an idle state. The first set of channels can have a lower or higher number of channels compared to the second set of channels. As described, the link is a PCIe link; however, the method can be adapted to other links having multiple channels. The operation can include detecting only one failed channel or multiple failed channels.
[0086] Method 900 includes, at block 904, detecting the process of working channels of the data link. The working channels correspond to channels that are not failed channels. For a x16 link with one failed channel, there will be 15 working channels. For 3 failed channels, there will be 13 working channels, but in either case, the working channels may not all be consecutive. Method 900 includes, at block 906, the process of selecting a set of consecutive working channels of the data link. Method 900 includes, at block 908, the process of defining an operable link to include the selected set of consecutive working channels.
[0087] Method 900 includes, at block 910, the process of identifying a start address of the operable link. Method 900 includes, at block 912, the process of storing the start address in a configuration register.
[0088] Method 900 includes the process of sending data traffic on an operational link at block 914. An interrupt can be sent to the host controller and the endpoint controller to transfer the data traffic from a first set of channels to the operational link. To facilitate the transfer, the controller can write an indicator to a control register to identify the operational link. For example, the starting channel number and the ending channel number of the operational link as the active set of channels can be written to the control register. Additionally, the controller can write an indicator of the width of the data link to the control register. The controller can also write a channel inversion indicator to the control register to indicate whether the first channel is the highest numbered channel or the lowest numbered channel in the link width. When L0p is enabled, the set of channels can be in the active state L0p, and all other channels can be in electrical idle.
[0089] The first set of channels no longer carries traffic after the transfer. In some examples, method 900 can optionally include changing the first set of channels to an idle state or training the first set of channels to the requested data rate and then changing the first set of channels to an idle state. In some examples, method 900 can also include: training the first set of channels to the requested data rate and then using the first set of channels to increase the link width dimension or decrease the link width dimension.
[0090] Figure 10 A flowchart of another method 1000 for increasing the data rate after a failed channel recovery in a multi-channel data link (e.g., a PCIe link) in accordance with aspects of the present disclosure is shown. In certain aspects, method 1000 involves defining the operational link as starting from an assignable start address on the link and not limited to the first channel or the last channel. The link width can also be defined with different variations and not limited to multiples of two. In some aspects, a new operational link can be prepared before transmitting data so that data is not lost during the recovery and configuration states. The method for recovering from a failed channel in a PCIe link can be performed by PCIe controllers 218, 262, which include an RC controller (also referred to as a host controller) and an EP controller. There is a negotiation for establishing the capabilities and bandwidth requirements. Once the negotiation is complete, the RC controller triggers method 1000 to bring a different set of consecutive working channels for the data link. Once the channels are working, the RC controller generates an interrupt to transfer the ongoing data transmission from the first active set of channels to the second newly trained set of channels.
[0091] Method 1000 includes maintaining an address table at block 1002 to store a set of consecutive working channels of a data link. The set includes corresponding start addresses and link width values. The set in the address table may also include end addresses or any other information. The address table may be stored at the RC and also at the PHY at the EP. The address table may be stored together with a configuration register or be part of the configuration register. The address table may be stored in the system memories 240, 274 or in any other suitable location. The address table may be updated during link negotiation and when a failure of any channel is detected.
[0092] Table 1 is an example of an address table for storing a set of consecutive working channels of an x16 data link. The first row indicates the start address of each set of consecutive working channels. The second row indicates the link width of each start address with the corresponding link width. In this example data set, all channels are in working order. Table 2 is an example of a link width table of an x16 data link where channel 5 has failed. In the PCIe standard, links require consecutive channels numbered as powers of 2, such as x1, x2, x4, x8, x16, x32. Thus, starting from channel 0, 4 consecutive working channels are available. Starting from channel 5, 1 consecutive working channel is available. Starting from channel 6, 8 consecutive working channels are available. Starting from channel 14, 2 consecutive working channels are available. For the maximum number of channels, the controller will select channel 6 as the start channel and a link width of 8. Alternatively, if an x4 link is required, the controller may select channel 0 or channel 6 as the start address. As another alternative, for an x4 link, the controller may select channel 3 or channel 9 for the end address.
[0093] Start address 0 Link width 16
[0094] Table 1
[0095] Start address 0 5 6 14 Link width 4 1 8 2
[0096] Table 2
[0097] Method 1000 includes: at block 1004, detecting a fault in a failed channel of a data link, the data link having a plurality of channels labeled in a consecutive sequence. One or more failed channels can be detected based on the condition of the link. The data link has a first set of channels in an active state and a second set of channels in an idle state. The first set of channels can have a lower or higher number of channels compared to the second set. As described, the link is a PCIe link; however, the method can be adapted to other links having multiple channels. When one or more failed channels are detected, the address table can be updated to change the set of consecutively working channels. For example, if channel 5 fails, the address table will look like Table 2. The failed channels can also be identified in a table stored together with the address table, either as part of the address table or as another table.
[0098] Table 3 is an example of a failed channel table for storing which channels have failed and which channels are working. The information in Table 3 can be combined with the information in Table 2. The first row indicates the addresses of each channel of the x16 link. The second row indicates the channel status as working (W) or failed (F). As in Table 2, the information in Table 3 includes the status information that channel 5 has failed.
[0099] Channel address 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 Status W W W W W F W W W W W W W W W W
[0100] Table 3
[0101] Method 1000 includes a process of detecting the working channels of the data link at block 1006. The address table can be updated to change the set of consecutively working channels to correspond to the detected working channels. Method 1000 includes a process of selecting a set of consecutively working channels of the address table at block 1008. When the address table includes link width values for each set of consecutively working channels, then the operation of selecting a set can be performed by selecting the set with the highest link width value. In another example, there may be a requested link width that is less than the total available link width. As an example, an x16 link can have a request to operate at a link width of 4. Any set with a link width value of 4 or greater can be selected. Method 1000 includes a process of defining the operable link at block 1010 as including the selected set of consecutively working channels.
[0102] Method 1000 includes a process of storing a start address in a configuration register at block 1012. The start address is the start address of the selected set of consecutively working channels. The start address is also stored in the address table. In another example, any other identifier of the selected set can be stored in the configuration register, such as a pointer to an entry in the address table corresponding to the selected set.
[0103] Method 1000 includes the process of sending data traffic on an operable link at block 1014. An interrupt can be sent to the host controller and the endpoint controller to transfer the data traffic from the first set of channels to the operable link. To facilitate the transfer, the controller can write an indicator to a control register to identify the operable link. For example, the start channel number and the end channel number of the operable link as the active set of channels can be written to the control register. Additionally, the controller can write an indicator of the width of the data link to the control register. The controller can also write a channel inversion indicator to the control register to indicate whether the first channel is the highest numbered channel or the lowest numbered channel in the link width. When L0p is enabled, the set of channels can be in the active state L0p, and all other channels can be in electrical idle.
[0104] The first set of channels no longer carries traffic after the transfer. In some examples, method 1000 can optionally include changing the first set of channels to an idle state or training the first set of channels to the requested data rate, and then changing the first set of channels to an idle state. In some examples, method 1000 can also include: training the first set of channels to the requested data rate, and then using the first set of channels to increase the link width dimension or decrease the link width dimension.
[0105] An overview of examples of the present disclosure is provided below.
[0106] Example 1: A method includes: detecting a fault of a failed channel of a data link, the data link having a plurality of channels labeled in a continuous sequence; detecting the working channels of the data link; selecting a set of continuous working channels of the data link; defining an operable link as including the selected set of continuous working channels; and sending data traffic on the operable link.
[0107] Example 2: The method according to example 1, further includes: identifying a start address of the operable link; and storing the start address in a configuration register.
[0108] Example 3: The method according to example 2, wherein the start address identifies a channel other than the end channel at the end of the continuous sequence.
[0109] Example 4: The method according to example 3, further includes: identifying an end address of the last channel of the operable link; and writing the end address to the configuration register.
[0110] Example 5: The method according to example 3 or 4, further includes: identifying a link width of the operable link; and writing the link width to the configuration register.
[0111] Example 6: The method according to Example 5, wherein the link width corresponds to a selected set of consecutive channels.
[0112] Example 7: The method according to any one or more of Examples 1-6, further comprising: sending an interrupt to an endpoint of the data link to transfer data traffic to the operable link.
[0113] Example 8: The method according to Example 7, wherein the interrupt includes a start address of a first channel of the operable link.
[0114] Example 9: The method according to any one or more of Examples 1-8, wherein selecting a set of consecutive working channels comprises: selecting a number of channels equal to an allowable link width for a PCIe data link.
[0115] Example 10: The method according to any one or more of Examples 1-9, wherein selecting a set of consecutive working channels comprises: determining a set of candidate consecutive working channels; determining the number of channels in each candidate set; and selecting the candidate set having the largest number of channels.
[0116] Example 11: The method according to Example 10, wherein selecting a set of consecutive working channels comprises: determining a set of candidate consecutive working channels; determining the number of channels in each candidate set; comparing each number of channels with the allowable link width; and selecting the candidate set having the largest allowable link width.
[0117] Example 12: The method according to any one or more of Examples 1-11, wherein selecting a set of consecutive working channels comprises: determining a set of candidate consecutive working channels; determining the number of channels in each candidate set; and selecting the candidate set having the requested link width.
[0118] Example 13: The method according to any one or more of Examples 1-12, further comprising: writing an enable indicator to a configuration register to indicate that the start address of the operable link is allowed to be an address other than the end address at the end of the consecutive sequence.
[0119] Example 14: The method according to any one or more of Examples 1-13, wherein selecting a set of consecutive working channels comprises: using the configuration register to select a set of consecutive working channels that provides the maximum link width for the data link.
[0120] Example 15: The method according to any one or more of Examples 1-14, further comprising: during training of the working channels of the operable link, sending the data traffic on the channels of the data link having the channel failures.
[0121] Example 16: The method according to any one or more of Examples 1-15 further includes: maintaining an address table to store a set of consecutive working channels of the data link, each set including a corresponding start address and a link width value, wherein selecting the set of consecutive working channels includes: selecting the set of working channels in the address table having the highest link width value, and wherein storing the start address includes: storing the start address in the address table.
[0122] Example 17: An apparatus includes: an interface circuit configured to provide an interface to a multi-channel data link; a configuration register for storing parameters of the data link; and a controller configured to: detect working channels of the data link; select a set of consecutive working channels of the data link; define an operable link as including the selected set of consecutive working channels; and transmit data traffic on the operable link.
[0123] Example 18: The apparatus according to Example 17, wherein the controller is further configured to: identify an end address of the last channel of the operable link; and write the end address to the configuration register.
[0124] Example 19: The apparatus according to Example 17 or 18, wherein the controller is further configured to: identify a link width of the operable link; and write the link width to the configuration register.
[0125] Example 20: The apparatus according to any one or more of Examples 17-19, wherein the controller is further configured to: identify a start address of the operable link; and send an interrupt including the start address to an endpoint of the data link to transfer data traffic to the operable link.
[0126] Example 21: The apparatus according to any one or more of Examples 17-20, wherein selecting the set of consecutive working channels includes: determining candidate sets of consecutive working channels; determining the number of channels in each candidate set; and selecting the candidate set having the largest number of channels.
[0127] Example 22: The apparatus according to any one or more of Examples 17-21, wherein the controller is further configured to: write an enable indicator to the configuration register regarding allowing the start address of the operable link to be an address other than the end address at the end of the consecutive sequence.
[0128] Example 23: An apparatus includes: a unit for detecting a fault in a faulty channel of a data link, the data link having a plurality of channels labeled in a consecutive sequence; a unit for detecting an operating channel of the data link; a unit for selecting a set of consecutive operating channels of the data link; a unit for defining an operable link as including the selected set of consecutive operating channels; and a unit for transmitting data traffic over the operable link.
[0129] Example 24: A non-transitory computer-readable medium having instructions stored therein, the instructions for causing a processor of an interconnect link to perform operations including: detecting a fault in a faulty channel of a data link, the data link having a plurality of channels labeled in a consecutive sequence; detecting an operating channel of the data link; selecting a set of consecutive operating channels of the data link; defining an operable link as including the selected set of consecutive operating channels; and transmitting data traffic over the operable link.
[0130] Example 25: A non-transitory computer-readable medium having instructions stored therein, the instructions for causing a processor of an interconnect link to perform the operations of any one or more of the above methods.
[0131] It should be understood that the present disclosure is not limited to the exemplary terms used above to describe aspects of the present disclosure. For example, bandwidth may also be referred to as throughput, data rate, or another term.
[0132] Although aspects of the present disclosure have been discussed above using examples of the PCIe standard, it should be understood that the present disclosure is not limited to that example and may be used with other standards.
[0133] The host client 214, host controller 212, device controller 252, and device client 254 discussed above may each be implemented using a controller or processor configured to execute these functions by software including code for performing the functions described herein. The software may be stored on a non-transitory computer-readable storage medium (e.g., RAM, ROM, EEPROM, optical disk, and / or magnetic disk), shown as host system memory 240, endpoint system memory 274, or another memory.
[0134] Any reference to elements herein using names (e.g., "first", "second", etc.) generally does not limit the number or order of those elements. Instead, these names are used herein as a convenient way to distinguish between two or more elements or instances of an element. Thus, a reference to a first and a second element does not mean that only two elements may be employed, or that the first element must precede the second element.
[0135] Within this disclosure, the word "exemplary" is used to mean "serving as an example, instance, or illustration". Any implementation or aspect described herein as "exemplary" is not necessarily to be construed as more preferred or advantageous than other aspects of the disclosure. Similarly, the term "aspect" does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation. The term "coupled" is used herein to refer to a direct or indirect electrical or other communication coupling between two structures. Additionally, the term "approximate" means within ten percent of the stated value.
[0136] The foregoing description of the disclosure has been provided to enable a person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method, comprising: Detecting a fault in a fault channel of a data link, the data link having a plurality of channels labeled in a continuous sequence; Detecting the working channels of the data link; Selecting a set of continuous working channels of the data link; Defining an operable link as including the selected set of continuous working channels; And Transmitting data traffic on the operable link.
2. The method according to claim 1, further comprising: Identifying a start address of the operable link; And Storing the start address in a configuration register.
3. The method according to claim 2, wherein, The start address identifies a channel other than an end channel at the end of the continuous sequence.
4. The method according to claim 3, further comprising: Identifying an end address of the last channel of the operable link; And writing the end address to the configuration register.
5. The method according to claim 3 further comprises: Identifying a link width of the operable link; and writing the link width to the configuration register.
6. The method according to claim 5, wherein The link width corresponds to the selected set of continuous channels.
7. The method according to claim 1, further comprising: Sending an interrupt to an endpoint of the data link to transfer data traffic to the operable link.
8. The method according to claim 7, wherein The interrupt includes the start address of the first channel of the operable link.
9. The method according to claim 1, wherein Selecting a set of continuous working channels includes: selecting a number of channels equal to an allowable link width for a PCIe data link.
10. The method according to claim 1, wherein, Selecting a set of continuous working channels includes: determining a set of candidate continuous working channels; determining the number of channels in each candidate set; and selecting the candidate set having the largest number of channels.
11. The method according to claim 10, wherein, Selecting a set of continuous working channels includes: determining a set of candidate continuous working channels; determining the number of channels in each candidate set; comparing each number of channels with the allowable link width; and selecting the candidate set having the largest allowable link width.
12. The method according to claim 1, wherein Selecting a set of continuous working channels includes: determining a set of candidate continuous working channels; determining the number of channels in each candidate set; and selecting the candidate set having a requested link width.
13. The method according to claim 1 further comprises: Writing an enable indicator to the configuration register to indicate that the start address of the operable link is allowed to be an address other than the end address at the end of the continuous sequence.
14. The method according to claim 1, wherein Selecting a set of continuous working channels includes: using the configuration register to select a set of continuous working channels that provides the maximum link width for the data link.
15. The method according to claim 1 further comprises: During training the working channels of the operable link, transmitting the data traffic on the channels of the data link having the channel fault.
16. The method according to claim 1 further comprises: Maintaining an address table to store sets of continuous working channels of the data link, each set including a corresponding start address and link width value, wherein selecting a set of continuous working channels includes: selecting the set of working channels in the address table having the highest link width value, and wherein storing the start address includes: storing the start address in the address table.
17. An apparatus, comprising: An interface circuit configured to provide an interface with a multi-channel data link; A configuration register for storing parameters of the data link; And A controller configured to: Detect the working channels of the data link; Select a set of continuous working channels of the data link; Define an operable link as including the selected set of continuous working channels; And Transmit data service on the operable link.
18. The device according to claim 17, wherein, The controller is further configured to: identify an end address of a last channel of the operable link; and write the end address into the configuration register.
19. The device according to claim 17, wherein The controller is further configured to: identify a link width of the operable link; and write the link width into the configuration register.
20. The apparatus according to claim 17, wherein The controller is further configured to: identify a start address of the operable link; and send an interrupt including the start address to an endpoint of the data link to transfer data service to the operable link.
21. The apparatus according to claim 17, wherein Selecting a set of consecutive working channels includes: determining a set of candidate consecutive working channels; determining the number of channels in each candidate set; and selecting the candidate set having the largest number of channels.
22. The device according to claim 17, wherein The controller is further configured to: write an enable indicator regarding allowing the start address of the operable link to be an address other than the end address at the end of the consecutive sequence into the configuration register.
23. An apparatus, comprising: a unit for detecting a fault of a faulty channel of a data link, the data link having a plurality of channels labeled in a consecutive sequence; a unit for detecting working channels of the data link; a unit for selecting a set of consecutive working channels of the data link; a unit for defining an operable link as including the selected set of consecutive working channels; and a unit for transmitting data service on the operable link.
24. A non-transitory computer-readable medium having instructions stored therein, the instructions for causing a processor of an interconnect link to perform operations including the following: detect a fault of a faulty channel of a data link, the data link having a plurality of channels labeled in a consecutive sequence; detect working channels of the data link; select a set of consecutive working channels of the data link; define an operable link as including the selected set of consecutive working channels; and transmit data service on the operable link.