Latency reduction for link speed switching in multi-channel data link
By receiving data rate change requests at the PCIe link controller, changing and training channel states, the problem of PCIe link prolonging during data rate switching is solved, and continuous transmission of data services and system performance is improved.
Patent Information
- Application Number
- CN202380069248.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-04
- Filing Date
- 2023-07-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-07-28
AI Technical Summary
The existing PCIe links delay delays during data rate switching, affecting system performance.
By receiving a data rate change request at the controller, changing the state of the idle channel to active, training the channel to a new data rate, and transmitting the data traffic from the active channel to the newly trained channel after training.
It realizes that while maintaining data service continuity, reduce link speed switching delays and improve system performance.
Smart Images

Figure CN119948466A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This patent application claims priority to pending U.S. non-provisional application No. 17 / 960,050, filed on October 4, 2022, which is assigned to the assignee of the present application and is hereby expressly incorporated herein by reference as if fully set forth below and for all applicable purposes. Technical Field
[0003] Aspects of the present disclosure relate generally to multi-lane data links, and more particularly to reducing latency during link speed switching. Background Art
[0004] High-speed interfaces are often used between circuits and components in mobile wireless devices and other complex systems. For example, some devices may include processing, communication, storage, and / or display devices that interact with each other through one or more high-speed interfaces. Some of these devices, including synchronous dynamic random access memory (SDRAM), may be able to provide or use data and control information at the processor clock rate. Other devices, such as display controllers, may use variable amounts of data at relatively low video refresh rates.
[0005] The Peripheral Component Interconnect Express (PCIe) interface is a popular high-speed interface that supports high-speed links capable of sending data at speeds of several gigabits per second. The interface also supports multiple speeds and multiple lane counts. PCIe provides lower latency and higher data transfer rates than parallel buses. PCIe is specified for communication between a wide variety of different devices. Typically, one device (e.g., a processor or a hub) acts as a host, which communicates with multiple devices (called endpoints) over a PCIe link. Peripheral devices or components can include graphics adapter cards, network interface cards (NICs), storage accelerator devices, mass storage devices, input / output interfaces, and other high-performance peripherals. Summary of the invention
[0006] The following content presents a summary of one or more implementations in order to provide a basic understanding of such implementations. This summary is not an exhaustive overview of all contemplated implementations, and is not intended to identify key or important elements of all implementations, nor is it intended to delineate the scope of any or all implementations. Its sole purpose is to present some concepts of one or more implementations in a simplified form as a preface to the more detailed description that is subsequently presented.
[0007] In one example, a device having an interface circuit and a controller for a peripheral component interconnect express (PCIe) link is disclosed. The device includes an interface circuit and a controller configured to provide an interface with a peripheral component interconnect express (PCIe) link. The controller is configured to: receive a request at the controller to change a data rate of a data link to a requested data rate; change a second channel set from an idle state to an active state; train the second channel set to the requested data rate; transfer data traffic from a first channel set to a second channel set after training; and send data traffic on the second channel set.
[0008] In another example, a method includes receiving, at a controller, a request to change a data rate of a multi-channel data link to a requested data rate, the data link having a first channel set in an active state and a second channel set in an idle state. The second channel set is changed from the idle state to the active state. The second channel set is trained to the requested data rate. After training, data traffic is transmitted from the first channel set to the second channel set, and the data traffic is sent on the second channel set.
[0009] In another example, a non-transitory computer readable medium has instructions stored therein for causing a processor of an interconnect link to perform the operations of the above-described method.
[0010] In another example, an apparatus includes: a component for providing an interface with a multi-channel data link; and a component for receiving a request to change the data rate of the data link to a requested data rate, the data link having a first set of channels in an active state and a second set of channels in an idle state. The apparatus also includes: a component for changing the second set of channels to an active state; a component for training the second set of channels to the requested data rate; a component for transferring data traffic from the first set of channels to the second set of channels after training; and a component for sending data traffic on the second set of channels.
[0011] To achieve the foregoing and related purposes, one or more implementations include the features fully described below and specifically pointed out in the claims. The following description and the accompanying figures set forth in detail certain illustrative aspects of one or more implementations. However, these aspects are merely indicative of several of the various ways in which the principles of each implementation may be employed, and the described implementations are intended to cover all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a block diagram of a computing architecture with a PCIe interface suitable for use with various aspects of the present disclosure.
[0013] Figure 2 is a block diagram of a system including a host system and an endpoint device system according to aspects of the present disclosure.
[0014] Figure 3 is a diagram of channels and corresponding drivers in a link according to aspects of the present disclosure.
[0015] Figure 4 is a state diagram illustrating the operation of a power management state machine in accordance with aspects of the present disclosure.
[0016] Figure 5A is a diagram of a multi-channel data link in accordance with aspects of the present disclosure.
[0017] Figure 5B is a diagram of a multi-channel data link with a second set of channels in recovery and configuration in accordance with aspects of the present disclosure.
[0018] Figure 5C is an illustration of a multi-channel data link after a configuration state speed change in accordance with aspects of the present disclosure.
[0019] Figure 5D is a diagram of a multi-lane data link with data traffic transmitted to a second set of lanes in accordance with aspects of the present disclosure.
[0020] Fig. 6A is a diagram of a multi-lane data link operating in lane reversal in accordance with aspects of the present disclosure.
[0021] Figure 6B is a diagram of a multi-channel data link with a first set of channels in recovery and configuration in accordance with aspects of the present disclosure.
[0022] Figure 6C is an illustration of a multi-lane data link with an enlarged data link width in accordance with aspects of the present disclosure.
[0023] Fig. 7A is a diagram of a multi-lane data link operating in lane flipping in accordance with aspects of the present disclosure.
[0024] Figure 7B is a diagram of a multi-channel data link with a first set of channels in recovery and configuration in accordance with aspects of the present disclosure.
[0025] Figure 7C is an illustration of a multi-channel data link after a configuration state speed change in accordance with aspects of the present disclosure.
[0026] Fig.7D is a diagram of a multi-lane data link with data traffic transmitted to a first set of lanes in accordance with aspects of the present disclosure.
[0027] Fig. 8A is a diagram of a multi-lane data link operating in a x1 width in accordance with aspects of the present disclosure.
[0028] Figure 8B is a diagram of a multi-channel data link with a second set of channels in recovery and configuration in accordance with aspects of the present disclosure.
[0029] Figure 8C is an illustration of a multi-channel data link after a configuration state speed change in accordance with aspects of the present disclosure.
[0030] Fig.8D is a diagram of a multi-lane data link with data traffic transmitted to a first set of lanes in accordance with aspects of the present disclosure.
[0031] Fig. 9 is a flow chart of an exemplary method for bandwidth-based power management according to aspects of the present disclosure.
[0032] Fig.10 is a block diagram of a PCIe link interface processing circuit according to aspects of the present disclosure. DETAILED DESCRIPTION
[0033] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. In order to provide a comprehensive understanding of the various concepts, the detailed description includes specific details. However, these concepts may be practiced without these specific details. In some instances, in order to avoid obscuring such concepts, well-known structures and components are shown in block diagram form.
[0034] In a PCIe system (e.g., a root complex (RC) connected to an endpoint (EP) via a PCIe link), the bandwidth or data rate requirements for the PCIe link change over time. To accommodate this, a dynamic GEN speed switch can be used. When the RC or EP requests a GEN speed switch, the PCIe link is disabled from operating at the first speed, and then training, recovery, and configuration are performed at the newly requested GEN speed. Thereafter, the PCIe link is active again and ready to carry data at the newly requested GEN speed. The process is the same for an increase or decrease in speed. During training, recovery, and configuration, no data is transmitted or received over the PCIe link. This stops communication between the RC and the PC, and also requires buffering of data, which has the risk that some data may be lost.
[0035] As described herein, when an idle lane is available on a PCIe link, the idle lane can perform training, recovery, and configuration at the newly requested GEN speed while the active lane continues to operate at the current GEN speed. After the previously active lane is configured and active, data traffic can be transferred to the newly active lane at the newly requested GEN speed. This allows data communications to continue as previously described while the new GEN speed lane is being prepared. No interruption of data traffic through the link is required.
[0036] The connection between any two PCIe devices (e.g., RC and EP) is called a link. PCIe links are built around duplex, serial (1-bit), differential, point-to-point connections called channels. With PCIe, data is transmitted over two signal pairs: two lines (wires, circuit board traces, etc.) for sending and two lines for receiving. The send and receive pairs are separate differential pairs for a total of four data lines per channel. The link includes a collection of channels, and each channel is capable of simultaneously transmitting and receiving data packets between the host and the endpoint.
[0037] PCIe links as currently defined are configured with a specific link width. Link widths can scale from one lane to 32 individual lanes. Lanes are defined as a number that is a power of 2. Common deployments have 1, 2, 4, 8, 12, 16, or 32 lanes, which may be labeled x1, x2, x4, x8, x12, x16, or x32, respectively, where the number is actually the number of lanes. In one example, a PCIe x1 implementation has four lanes to connect a lane pair in each direction, while a PCIex16 implementation has 16 times that number, 16 lanes, or 64 lanes.
[0038] A PCIe link can be configured to use less than all of its lanes in order to save power or for compatibility. As an example, a x4 PCIe add-on board can be installed into a x4, x8, x12, x16, or x32 slot or connector. After configuration and training, the link through the connector can operate at a width of x1, x2, or x4, but not greater widths because the add-on card does not support more than 4 lanes. As currently defined, during operation, the link width can be scaled up to a larger width or scaled down to a smaller width.
[0039] PCIe links can also be configured for different link speeds, also known as data rates, which depend on the specified speed capabilities of the link. During configuration and training, the fastest PCIe link speed is determined, which is usually the slowest maximum speed between the host and the endpoint. PCIe links can be configured at this maximum common speed or a slower speed to change power. Link speeds can change during operation. Link speeds are called GEN speeds because newer generations of PCIe standards provide new higher speeds. For example, PCIe GEN1 allows 2.5 gigabit transfers per second (GT / s), PCIe GEN2 allows 5GT / s, PCIe GEN3 allows 8GT / s, PCIe GEN4 allows 16GT / s, PCIe GEN5 allows 32GT / s, and subsequent versions may provide higher data rates for each channel of the link. Higher speeds provide data transmission benefits, but also consume more power and are more prone to errors. Therefore, PCIe links can be managed to operate at lower GEN speeds until there is a high data rate demand on the link.
[0040] Power management and bandwidth negotiation can be performed when the link is initialized, but can also be repeated at a later time. During the negotiation, each link partner (e.g., RC and EP) can announce the number of channels supported (e.g., link width) and the bandwidth desired for operation. For example, the link partners can agree to operate at the highest bandwidth supported by both partners. The link partners negotiate that a certain number of channels of the link are active, and that number of channels can be changed to a lower rate for reasons of link stability. In one example, the link width can be changed autonomously by hardware. As the number of channels increases, the power to operate the link also increases. Therefore, in some cases, a x16 link can operate as a x1 link at lower power. This reduces the power consumed by the supporting hardware during periods of low activity.
[0041] Figure 11 is a block diagram of an example computing architecture using a PCIe interface. The computing architecture 100 operates using multiple high-speed PCIe interface serial links. The PCIe interface can be characterized as an apparatus including a point-to-point topology, in which a separate serial link connects each device to a host, which is referred to as a root complex 104 (RC). In the computing architecture 100, the root complex 104 couples the processor 102 to a memory device (e.g., a memory subsystem 108) and a PCIe switching circuit 106. In some instances, the PCIe switching circuit 106 includes a cascaded switching device. One or more PCIe endpoint devices 110 (EP) can be directly coupled to the root complex 104, while other PCIe endpoint devices 112-1, 112-2, ..., 112-N can be coupled to the root complex 104 through the PCIe switching circuit 106. The root complex 104 can be coupled to the processor 102 using a proprietary local bus interface or a standard defined local bus interface. Root complex 104 may control configuration and data transactions over a PCIe interface and may generate transaction requests for processor 102. In some examples, root complex 104 is implemented in the same integrated circuit (IC) device that includes processor 102. Root complex 104 supports multiple PCIe ports.
[0042] The root complex 104 may control communication between the processor 102 and a memory subsystem 108, which is one example of an endpoint. The root complex 104 also controls communication between the processor 102 and other PCIe endpoint devices 110, 112-1, 112-2, ..., 112-N. The PCIe interface may support full-duplex communication between any two endpoints, with no inherent restrictions on concurrent access across multiple endpoints. Data packets may carry information over any PCIe link. In a multi-lane PCIe link, packet data may be striped across multiple lanes. The number of lanes in a multi-lane link may be negotiated during device initialization and may be different for different endpoints.
[0043] When one or more lanes of a PCIe link are not fully utilized by low-bandwidth applications (which can be adequately served by fewer lanes), the root complex 104 and the endpoints can operate the link with more or fewer lanes. In some examples, one or more lanes can be placed in one or more standby states in which some or all of the lanes operate in a low-power or no-power mode. Changing the number of active lanes for low-bandwidth applications reduces the power to operate the link. Providing less power reduces current leakage, heat, and power consumption.
[0044] Figure 2205 is a block diagram of an exemplary PCIe system in which aspects of the present disclosure may be implemented. System 205 includes a host system 210 and an endpoint device system 250. Host system 210 may be integrated on a first chip (e.g., a system on a chip or SoC), and endpoint device system 250 may be integrated on a second chip. Alternatively, the host system and / or the endpoint device system may be integrated in a first package and a second package (e.g., a SiP), a first system board and a second system board having multiple chips, or in other hardware or any combination. In this example, host system 210 and endpoint device system 250 are coupled via PCIe link 285.
[0045] The host system 210 includes one or more host clients 214. Each of the one or more host clients 214 can be implemented on a processor executing software that performs the functions of the host client 214 discussed herein. For examples of more than one host client, the host clients can be implemented on the same processor or on different processors. The host system 210 also includes a host controller 212 that can perform root complex functions. The host controller 212 can be implemented on a processor executing software that performs the functions of the host controller 212 discussed herein.
[0046] The host system 210 includes a PCIe interface circuit 216, a system bus interface 215, and a host system memory 240. The system bus interface 215 can interface one or more host clients 214 with the host controller 212, and each of the one or more host clients 214 and the host controller 212 with the PCIe interface circuit 216 and the host system memory 240. The PCIe interface circuit 216 provides an interface to the PCIe link 285 to the host system 210. In this regard, the PCIe interface circuit 216 is configured to send data (e.g., from the host client 214) to the endpoint device system 250 via the PCIe link 285 and receive data from the endpoint device system 250 via the PCIe link 285. The PCIe interface circuit 216 includes a PCIe controller 218, a physical interface 220 for a PCI Express (PIPE) interface, a physical (PHY) transmit (TX) block 222, a clock generator 224, and a PHY receive (RX) block 226. PIPE interface 220 provides a parallel interface between PCIe controller 218 and PHY TX block 222 and PHY RX block 226. PCIe controller 218 (which may be implemented in hardware) may be configured to perform transaction layer, data link layer, and control flow functions specified in the PCIe specification, as discussed further below.
[0047] The host system 210 also includes an oscillator (e.g., a crystal oscillator or "XO") 230 configured to generate a reference clock signal 232. In one example, the reference clock signal 232 may have a frequency of 19.2 MHz, but is not limited to this frequency. The reference clock signal 232 is input to the clock generator 224, which generates a plurality of clock signals based on the reference clock signal 232. In this regard, the clock generator 224 may include one or more phase-locked loops (PLLs), each of which generates a respective one of the plurality of clock signals by multiplying the frequency of the reference clock signal 232.
[0048] The endpoint device system 250 includes one or more device clients 254. Each device client 254 may be implemented on a processor executing software that performs the functions of the device client 254 discussed herein. For examples of more than one device client 254, the device clients 254 may be implemented on the same processor or on different processors. The endpoint device system 250 also includes a device controller 252. The device controller 252 may be configured to receive bandwidth requests from one or more device clients and determine whether to change the number of active channels or change the GEN speed based on the bandwidth requests. The device controller 252 may be implemented on a processor executing software that performs the functions of the device controller.
[0049] The endpoint device system 250 includes a PCIe interface circuit 260, a system bus interface 256, and an endpoint system memory 274. The system bus interface 256 may interface one or more device clients 254 with the device controller 252, and interface each of the one or more device clients 254 and the device controller 252 with the PCIe interface circuit 260 and the endpoint system memory 274. The PCIe interface circuit 260 provides an interface to a PCIe link 285 to the endpoint device system 250. In this regard, the PCIe interface circuit 260 is configured to send data (e.g., from the device client 254) to a host system 210 (also referred to as a host device) through the PCIe link 285 and receive data from the host system 210 via the PCIe link 285. The PCIe interface circuit 260 includes a PCIe controller 262, a PIPE interface 264, a PHY TX block 266, a PHY RX block 270, and a clock generator 268. PIPE interface 264 provides a parallel interface between PCIe controller 262 and PHY TX block 266 and PHY RX block 270. PCIe controller 262 (which may be implemented in hardware) may be configured to perform transaction layer, data link layer, and control flow functions.
[0050] The host system memory 240 and the endpoint system memory 274 at the endpoint can be configured to contain registers for the status of each lane of the PCIe link 285. These registers include group control registers and group status registers. In an example, both the host system memory 240 and the endpoint system memory 274 have link GEN control registers, status registers, and capability registers, among others.
[0051] The endpoint device system 250 also includes an oscillator (eg, a crystal oscillator) 272 configured to generate a stable reference clock signal 273 for an endpoint system memory 274. Figure 2 In the example of , the clock generator 224 at the host system 210 is configured to generate a stable reference clock signal 273, which is forwarded by the PHY RX block 226 to the endpoint device system 250 via the differential clock line 288. At the endpoint device system 250, the PHY RX block 270 receives the EP reference clock signal on the differential clock line 288 and forwards the EP reference clock signal to the clock generator 268. The EP reference clock signal may have a frequency of 100 MHz, but is not limited to this frequency. The clock generator 268 is configured to generate a plurality of clock signals based on the EP reference clock signal from the differential clock line 288, as discussed further below. In this regard, the clock generator 268 may include a plurality of PLLs, each of which generates a respective one of the plurality of clock signals by multiplying the frequency of the EP reference clock signal.
[0052] The system 205 also includes a power management integrated circuit (PMIC) 290 coupled to a power source 292 (e.g., a mains voltage, a battery, or other power source). The PMIC 290 is configured to convert the voltage of the power source 292 to a plurality of supply voltages (e.g., using a switching regulator, a linear regulator, or any combination thereof). In this example, the PMIC 290 generates a voltage 242 for the oscillator 230, a voltage 244 for the PCIe controller 218, and a voltage 246 for the PHY TX block 222, the PHY RX block 226, and the clock generator 224. The voltages 242, 244, and 246 may be programmable, wherein the PMIC 290 is configured to set the voltage levels (angles) of the voltages 242, 244, and 246 according to instructions (e.g., from the host controller 212).
[0053] The PMIC 290 also generates a voltage 280 for the oscillator 272, a voltage 278 for the PCIe controller 262, and a voltage 276 for the PHY TX block 266, the PHY RX block 270, and the clock generator 268. The voltages 280, 278, and 276 may be programmable, wherein the PMIC 290 is configured to set the voltage levels (angles) of the voltages 280, 278, and 276 according to instructions (e.g., from the device controller 252). The PMIC 290 may be implemented on one or more chips. Although the PMIC 290 may be implemented on a single chip, the PMIC 290 may be implemented on a single chip. Figure 2 290 is shown as one PMIC, but it should be understood that PMIC 290 may be implemented by two or more PMICs. For example, PMIC 290 may include a first PMIC for generating voltages 242, 244, and 246 and a second PMIC for generating voltages 280, 278, and 276. In this example, both the first PMIC and the second PMIC may be coupled to the same power supply 292 or different power supplies.
[0054] In operation, the PCIe interface circuit 216 on the host system 210 can send data from one or more host clients 214 to the endpoint device system 250 via the PCIe link 285. When the host controller negotiates bandwidth for the link, data from the one or more host clients 214 can be directed to the PCIe interface circuit 216 according to the PCIe mapping established by the host controller 212 during initial configuration (sometimes referred to as link initialization). In an example, the host controller negotiates a first bandwidth for a transmit group of the link and a second bandwidth for a receive group of the link. At the PCIe interface circuit 216, the PCIe controller 218 can perform transaction layer and data link layer functions on the data, such as packetizing the data, generating error correction codes to be sent with the data, etc.
[0055] The PCIe controller 218 outputs the processed data to the PHY TX block 222 via the PIPE interface 220. The processed data includes data from one or more host clients 214 and overhead data (e.g., packet headers, error correction codes, etc.). In one example, the clock generator 224 may generate a clock 234 for an appropriate data rate or transfer rate based on the reference clock signal 232, and input the clock 234 to the PCIe controller 218 to time the operation of the PCIe controller 218. In this example, the PIPE interface 220 may include a 22-bit parallel bus that transmits 22 bits of data in parallel to the PHY TX block for each cycle of the clock 234. At a frequency of 250 MHz, the transfer rate is approximately 8 GT / s.
[0056] The PHY TX block 222 serializes the parallel data from the PCIe controller 218 and drives the PCIe link 285 with the serialized data. In this regard, the PHY TX block 222 may include one or more serializers and one or more drivers. The clock generator 224 may generate a high frequency clock for the one or more serializers based on the reference clock signal 232.
[0057] At the endpoint device system 250, the PHY RX block 270 receives the serialized data via the PCIe link 285 and deserializes the received data into parallel data. In this regard, the PHY RX block 270 may include one or more receivers and one or more deserializers. The clock generator 268 may generate a high frequency clock for the one or more deserializers based on the EP reference clock signal. The PHY RX block 270 transmits the deserialized data to the PCIe controller 262 via the PIPE interface 264. The PCIe controller 262 may recover data from the one or more host clients 214 from the deserialized data and forward the recovered data to the one or more device clients 254.
[0058] On the endpoint device system 250, the PCIe interface circuit 260 may send data from one or more device clients 254 to the host system memory 240 via a PCIe link 285. In this regard, the PCIe controller 262 at the PCIe interface circuit 260 may perform transaction layer and data link layer functions on the data, such as packetizing the data, generating error correction codes to be sent with the data, etc. The PCIe controller 262 outputs the processed data to the PHY TX block 266 via the PIPE interface 264. The processed data includes data from the one or more device clients 254 and overhead data (e.g., packet headers, error correction codes, etc.). In one example, the clock generator 268 may generate a clock based on the EP reference clock via a differential clock line 288 and input the clock to the PCIe controller 262 to time the operation of the PCIe controller 262.
[0059] The PHY TX block 266 serializes the parallel data from the PCIe controller 262 and drives the PCIe link 285 with the serialized data. In this regard, the PHY TX block 266 may include one or more serializers and one or more drivers. The clock generator 268 may generate a high frequency clock for the one or more serializers based on the EP reference clock signal.
[0060] At the host system 210, the PHY RX block 226 receives the serialized data via the PCIe link 285 and deserializes the received data into parallel data. In this regard, the PHY RX block 226 may include one or more receivers and one or more deserializers. The clock generator 224 may generate a high frequency clock for the one or more deserializers based on the reference clock signal 232. The PHY RX block 226 transmits the deserialized data to the PCIe controller 218 via the PIPE interface 220. The PCIe controller 218 may recover data from the one or more device clients 254 from the deserialized data and forward the recovered data to the one or more host clients 214.
[0061] Figure 3 is available in Figure 1 and Figure 2 2 is a diagram of channels in a link 385 (e.g., PCIe link 285) used in a system of FIG. 210. In this example, link 385 includes a plurality of channels 310-1 to 310-n, where each channel includes a respective first differential line pair 312-1 to 312-n for transmitting data from a host system 210 to an endpoint device system 250, and a respective second differential line pair 315-1 to 315-n for transmitting data from the endpoint device system to the host system 210. From the perspective of the host system, the first channel 310-1 is dual simplex, where the first differential line pair 312-1 acts as a transmit line and the second differential line pair 315-1 acts as a receive line. From the perspective of the endpoint device system, the first channel 310-1 has a receive line and a transmit line. The first differential line pair 312-1 to 312-n and the second differential line pair 315-1 to 315-n may be implemented with metal trace lines on a substrate (e.g., a printed circuit board), wherein the host system may be integrated on a first chip mounted on the substrate, and the endpoint device is integrated on a second chip mounted on the substrate. Alternatively, the link may be implemented through an adapter card slot, a cable, or a combination of different media. The link may also include an optical portion in which the PCIe packets are encapsulated within different systems. In this example, when data is transmitted from the host system to the endpoint device system across multiple channels, the PHYTX block 222 may include logic for dividing data between channels. Similarly, when data is transmitted from the endpoint device system to the host system 210 across multiple channels, the PHY TX block 266 may include logic for dividing data between channels.
[0062] Figure 2 The PHY TX block 222 of the host system 210 shown in FIG. 1 may be implemented to include a transmission driver 320-1 to 320-n for driving each first differential line pair 312-1 to 312-n to transmit data, and Figure 2The PHYRX block 270 of the host shown in the figure may be implemented to include receivers 340-1 to 340-n (e.g., amplifiers) for driving each second differential line pair 312-1 to 312-n to receive data. Each transmit driver 320-1 to 320-n is configured to drive the corresponding differential line pair 312-1 to 312-n with data, and each receiver 340-1 to 340-n is configured to receive data from the corresponding first differential line pair 312-1 to 312-n. In addition, in Figure 2 In the embodiment of the present invention, the PHY TX block 266 of the endpoint device system 250 may include a transmit driver 345-1 to 345-n for each second differential line pair 315-1 to 315-n, and the PHY RX block 226 of the host system 210 may include a receiver 325-1 to 325-n (e.g., an amplifier) for each second differential line pair 315-1 to 315-n. Each transmit driver 345-1 to 345-n is configured to drive the corresponding second differential line pair 315-1 to 315-n with data, and each receiver 325-1 to 325-n is configured to receive data from the corresponding second differential line pair 315-1 to 315-n.
[0063] In some aspects, the width of the link 385 can be scaled to match the capabilities of the host system and the endpoints. The link can use one lane 310-1 for a x1 link, two lanes 310-1, 310-2 for a x2 link, or multiple lanes for wider links, up to n lanes from 310-1 to 310-n. Current links are defined as 1, 2, 4, 8, 16, and 32 lanes, but different numbers of lanes can be used to suit a particular implementation.
[0064] In one example, the host system 210 may include a power switch circuit 350 configured to individually control power to the transmit drivers 320-1 to 320-n and the receivers 325-1 to 325-n from the PMIC 290. Thus, in this example, the number of drivers and receivers powered is proportional to the width of the link 385. Similarly, Figure 2 The endpoint device system 250 shown in FIG. 2 may include a power switch circuit 360 configured to individually control power supply from the PMIC 290 to the transmit drivers 345-1 to 345-n and the receivers 340-1 to 340-n. In this manner, the host system sets the number of multiple drivers to be selectively powered by the power switch circuit to change the number of active transmit or receive lines based on the number of lines powered. Using differential signaling, the lines will be set to active or standby states in pairs.
[0065] Figure 44 is a state diagram 400 illustrating the operation of a portion of a power management state machine according to certain aspects disclosed herein. The active state power management (ASPM) protocol is a state machine method for reducing power based on link activity detected on a PCIe link between a root complex (RC) and an endpoint (EP) PCIe device. The state diagram is consistent with the link training and state state machine (LTSSM) defined for PCIe. However, other methods may be used alternatively. In this method, when data is being transmitted over a PCIe link, the link operates in an L0 state 408 (i.e., a link operational state). There may be other operational states, such as an L0p substate 418 and an L0s state 410, etc. The L0p substate 418 is not an inactive or idle state. When the L0p substate 418 is enabled, the link can be configured in an L0p active state, where based on data rate requirements, the link can be reduced by placing several of the channels in electrical idle while still carrying data on other channels in an active L0p substate. Similarly, for an amplification request, the electrically idle channel can be independently retrained to the L0p active substate without interfering with ongoing data transmissions on other active channels in the link. The ASPM protocol can also support additional active and standby states and substates, such as L1, L2, L1.1, and L1.2, etc. As shown, the L0s state 410 can only be accessed via a connection via the L0 state 408. When the status on the link indicates or suggests that a transition is required or appropriate between states, an ASPM state change can be initiated. When the status is normal, both communication partners on the link can initiate a power state change request.
[0066] The PCIe link initiates operation in a detect state 402. In this state, the host controllers of the RC and EP detect active connections. The link then moves to a polling state 404, during which the host controller of the RC polls any active EP connections. Similarly, the EP host controller polls the RC host controller. This allows the available link width and GEN speed to be determined. The PCIe link then moves to a configuration state 406, during which the RC and EP ports are configured for a specific link width and GEN speed. Some or all of the channels are configured as active, and any unused channels are configured as idle, such as in a L0s state 410. Active channels are in a L0 state 408.
[0067] When the link is idle (e.g., in a short time interval between data bursts), the link can be placed in a standby state (e.g., L0s state 410) from L0 state 408, which can only be accessed through a connection to L0 state 408. In this example, L0s state 410 is a low-power standby for L0 state 408. L1 state 416 is a standby state with lower latency than L0s state 410. L0s state 410 is used as a standby state and is also used as an initialization state after power-on, system reset, or after an error condition is detected. In L0s state 410, device discovery and bus configuration processes can be implemented before the link transitions 428 to L0 state 408. In L0 state 408, PCIe devices can be active and respond to PCIe transactions, and / or can request or initiate PCIe transactions. L1 state 416 is the main standby state and allows a quick return to L0 state 408 through recovery state 412. The L0s state 410 is a lower power state that allows an electrical idle state and can be transitioned through a resume state 412. When the PCIe device determines that there are no outstanding PCIe requests or pending transactions, the L0s state 410 can be entered through a transition 426. Power consumption can be reduced by disabling or idling transceivers in the PCIe bus interface, disabling, gating clocks used by PCI devices, and disabling PLL circuits used to generate clocks for receiving data. The PCIe device can make a transition 426 to the L0s state 410 through the operation of a hardware controller or some combination of an operating system and hardware control circuits.
[0068] When a PCIe link becomes active while a device (e.g., RC or EP) is in an electrical idle state (e.g., L0s state 410), a return to L0 state 408 is initiated for the device. A direct transition to L0 state 408 may not be available. The PCIe link may first transition 422 to a recovery state 412, where the transceiver in the PCIe bus interface, the clock used by the PCI device, and / or PLL circuits are enabled. When it is determined that the transceiver and other circuits are functional, a transition 430 from the recovery state 412 to the L0 state 408 may be initiated.
[0069] The ASPM protocol also manages the GEN speed changes of the PCIe link. The PCIe link can operate at a lower GEN speed to save power, or operate at a higher GEN speed to provide higher performance. In order to change the GEN speed, the ASPM protocol returns the active channel from the L0 state 408 to the configuration state 406 to configure the link to the new GEN speed. After configuration, the PCIe link is returned to the L0 state to operate at the GEN speed of the new configuration. In some aspects, when there are some channels in the L0 state 408 and other channels in the L0s state 410, the channel in the L0 state 408 can continue to operate in the L0 state 408, and the other channels are converted to the configuration state 406. The channel in the configuration state 406 is configured to the new GEN speed, and then converted to the L0 state 408 to operate at the new GEN speed. The channel in the L0 state can be converted to an idle state, for example, the L0s state 410. These channels may be reconfigured in the Configure state 406 or may remain in the L0s state 410 for later use.
[0070] The ASPM protocol can also manage link width to reduce or increase link width, also known as scaling up and down link width. By reducing link width during low throughput data traffic scenarios, the subsystem of the PCIe link reduces voltage levels (e.g., to a lower operating level that satisfies the current throughput on the PCIe link). The reduced voltage level or levels reduces power consumption (e.g., reduces leakage current during sustained low throughput traffic or in idle use cases). The number of channels also affects power consumption. In practice, there is a L0, L0s, L1 transition state diagram for each channel.
[0071] Figure 5A A multi-lane data link 540 is shown, such as a PCIe link of duplex traffic channels between host 502 and EP 506. The channels each include a line in each direction so that the data rate and configuration are the same in both directions. A duplex traffic channel may have Figure 3 The same physical structure as in , but is generalized to show the transition from one data rate to another. Data link 540 includes four channels. The duplex traffic channel has Figure 3. Each channel includes two transmit lines as a differential line pair and two receive lines as a differential line pair, for a total of four lines for each channel, and a total of sixteen lines for a x4 link. The first channel set 511 (channel 0 and channel 1) is in an active state as a x2 link, such as an L0 state. Data can be sent in one direction or both directions at a specific data rate or speed configuration (e.g., GEN2 speed). The second channel set 512 (channel 2 and channel 3) is in an idle state (e.g., when L0p is enabled) and does not send any business. Although four channels are shown, there may be more or fewer channels to suit a specific implementation. As described above, the PCIe data link operates according to the rule of 2, thereby supporting x1, x2, x4, x8, x16, and x32. The data link 540 is configured as a x4 data link, but the link width is configured as x2 GEN2 to operate more efficiently for the data rate carried by the link.
[0072] Figure 5B The multi-channel data link 542 between the host 502 and the EP 506 is shown training to a new data rate. In one aspect, when a controller at the host 502 or EP 504 receives a request to change the data rate of the data link 542 (e.g., from GEN2 to GEN3), the data link 542 is changed. The request can be generated by the host controller or the EP controller and then transmitted to the corresponding controller at the opposite end, such as the EP controller or the host controller, respectively. As previously described, the first channel set 511 continues to send data traffic. The second channel set 520 (channels 2 and 3) that was in the idle state changes to the active state, then transitions to the recovery state, and then transitions to the configuration state. This changes the second channel set from the idle state to the active state and trains the second channel set to the requested data rate. The second channel set 520 is not in electrical idle as before, but no data is sent on the second channel set 520 during this time. While the second set of channels 520 is undergoing recovery and configuration, the multi-channel data link 542 between the host 502 and the EP 504 continues to send data traffic on the original first set of channels 511 (channel 0 and channel 1) at the original data rate.
[0073] Figure 5C The multi-channel data link 544 is shown after the configuration state speed change of the second channel set 522. The first channel set 511 is still in the same active state, such as the L0 state, and is running at the original speed (e.g., Figure 5A and Figure 5BThe second set of lanes 522 (lane 2 and lane 3) has been reconfigured to be active at the new requested data rate (e.g., GEN3 speed). Traffic is still being carried on the first set of lanes 511 at the original speed. To transition the data traffic of data link 544 to the new speed, Figure 5B After the training in , data traffic is transferred from the first set of lanes 512 to the second set of lanes 522. The specific data rates GEN2 and GEN3 are provided as examples. A similar process is performed for any data rate change, and the link width is also scaled up or down.
[0074] Figure 5D A multi-lane data link 546 between the host 502 and the EP 506 is shown after data traffic has been transferred to the second lane set 522. The configuration of the second lane set has been changed to the new data rate. Data traffic is sent on the second lane set 522 of the data link. The second lane set 522 of the data link 546 is carrying data traffic at the requested data rate GEN3 instead of the first lane set 514. The first lane set 514 of the data link 546 does not need to support the x2 link width, and therefore it is changed to the electrical idle state.
[0075] The x2 link width is using lanes 2 and 3 instead of lanes 0 and 1. In some aspects, lane flipping can be enabled to support the use of the highest numbered lanes (e.g., lanes 3 and 2) instead of the lowest numbered lanes (e.g., lanes 0 and 1). Figure 5B During the recovery and configuration shown, data traffic on the first channel set 511 is not interrupted, while the second channel set 520 is being trained to the requested data rate. After reconfiguring the link 546, the first channel set is no longer needed to support the requested data rate and link width. Therefore, the first channel set 514 transitions to an idle state. The specific data rates GEN2 and GEN3 are provided as examples. A similar process can be performed for any data rate change, and the link width can also be enlarged or reduced.
[0076] like Figure 5B As shown in , the data rate can be configured differently for different channels. A first set of channels has a different GEN speed than a second set of channels. This allows Figure 5BAs shown, one or more channels are retrained with a new GEN speed configuration without disturbing other active channels operating at a different GEN speed. A second set of channels 520 are retrained to the GEN3 speed without disturbing the first set of channels 511 operating at the GEN2 speed. A second set of channels 512 in electrical idle is used to meet the GEN speed change requirements. During a data rate switch from the GEN2 speed to the GEN3 speed, the same number of electrically idle channels are brought back from idle and configured to the new data rate (GEN3 speed).
[0077] once Figure 5C If the new channel set shown in FIG. 5 , namely the second channel set 522, is also raised to the requested data rate, namely the GEN3 speed, the data traffic of the PCIe link can be transferred to the new channel set operating at the new data rate. This can be accomplished by, for example, transmitting an interrupt from the RC controller to the host 502 and the EP 506 to stop using the first channel set operating at the old data rate and start using the second channel set operating at the new data rate. Figure 5D As shown, the first set of channels 514 can be placed directly in electrical idle, or retrained to the new data rate and then placed in electrical idle. If there are other channels in electrical idle that are not used in the process, these channels can also be trained to the new data rate and then placed in electrical idle. This retraining makes the idle channels available for link width enlargement requests at the new data rate.
[0078] Both the host 502 and the EP 506 include a system memory, and the system memory includes control registers, such as a link width control register, a first channel control register, and an enable control register. The control registers may use an enable indication flag to enable or disable data rate switching, such as Figure 5B As shown in , the idle channels are activated and trained to the new speed. Figure 5C As shown, the default method is that active channels always start with channel 0. For channel rollover, active channels always start with the highest numbered channel, for example, channel 1, channel 3, channel 7, channel 15, or channel 31. The indicator flags in the link width control register then determine how many additional channels are active. Figure 5C In the example shown in , the link width control register indicates a width of 2 with no lane inversion. Figure 5D In the example shown in Figure 1, the active lane can be indicated by lane flipping and a link width of 2.
[0079] Fig. 6A A multi-channel data link 642 is shown between the host 602 and the EP 606, where the first channel set 612 is in an idle state and the second channel set 614 is in an active state. The data link 642 has Figure 5DThe second lane set 614 of the data link 642 sends data traffic at the configured data rate (eg, GEN3). The first lane set 612 is in an electrical idle state. This provides a x2 link width with lane flipping at GEN3 speed.
[0080] Figure 6B A multi-channel data link 644 is shown between the host 602 and the EP 606 during recovery and reconfiguration. Figure 6B In some aspects, the width of link 644 can be scalable by controlling the number of channels that are active. The host 602 or EP 606 can initiate a specific link width or a change in link width to determine the number of business channels that are powered to send and receive data over link 644. In the case of high transmit or receive traffic, all business channels are activated for the x4 link. If both transmit and receive traffic are low or if there is business inactivity, one or more duplex business channels can be placed in a low power state. The x1 configuration is operated using only channel 0. The x2 configuration is operated using channel 0 and channel 1, or as shown, using channel 2 and channel 3 and channel flipping. The x4 configuration uses all four channels.
[0081] The data link 644 operates at x2 using only the second set of data lanes 614. The second set of data lanes 614 continues to send data traffic between the host 602 and the EP 604 at the requested data rate, such as Fig. 6A As shown. In response to the link width request, the first lane set 616 changes from the idle state to the configuration and recovery state. In this example, the first lane set is trained to the requested data rate GEN3.
[0082] Figure 6C 6 shows a multi-lane data link 646 between host 602 and EP 606 after the link width is enlarged to x4. Figure 6C , the first channel set 616 is restored and configured to be able to operate at the requested data rate GEN3, which together with the second channel set forms a single channel set 618 that is all active and operates at the requested data rate. Fig. 6A The previous x2 data link has been converted to a x4 link using two free lanes.
[0083] In some aspects, the transition to the new data rate and to the increased data link width can be performed simultaneously. As an example, Figure 5BThe data link 544 shown in FIG. 5 has a first active channel set 511 under GEN2 and a second active channel set 522 under GEN3. Then, by transitioning the first channel set 616 to recovery and configuration instead of transitioning to idle state, the data link can be directly from Figure 5B The configuration changes to Figure 6B Then, the first channel set 616 can be configured with the new data rate GEN3 and used to enlarge the link width, such as Figure 6C shown.
[0084] Fig. 7A is a diagram of a multi-lane data link 740, such as a PCIe link with a duplex traffic channel between a host 702 and an EP 706. The duplex traffic channel has Figure 5D . The first channel set 711 (channel 0 and channel 1) is in an idle state. The second channel set 712 (channel 2 and channel 3) is in an active state. Data traffic can be sent in one direction or both directions at a specific data rate or speed configuration (e.g., GEN2 speed). Data link 740 operates in channel flipping. This configuration may be due to FIG. 5A to FIG. 5D sequence or for any other suitable reason. Although four channels are shown, there may be more or fewer channels to suit a particular implementation. Depending on the configuration, the data link may transition back to normal channel sequencing with or without a change in speed.
[0085] Figure 7B A multi-channel data link 742 between the host 702 and the EP 706 during recovery and configuration of the first channel set 714 is shown. In one aspect, when the controller at the host 702 or EP 704 receives a request to change the data rate of the data link 742 to the requested data rate (e.g., from GEN2 to GEN3), the state of the first channel set 714 is changed to an active state and placed in a recovery state and then placed in a configuration state. This can also occur in response to a request to flip the channel flip. The second channel set 712 remains in an active state and sends data traffic over the data link 742 between the host 702 and the EP 704. While the first channel set 714 is undergoing recovery and configuration, the multi-channel data link 742 between the host 702 and the EP 704 continues to send data traffic at the original data rate on the original second channel set 712 (channel 0 and channel 1).
[0086] Figure 7C7 shows a multi-channel data link 744 between the host 702 and the EP 706 after the configuration state speed of the first channel set 716 is changed. The second channel set 712 is still in the same active state, such as the L0 state. Data can be sent in both directions at the original speed (e.g., GEN2), such as Fig. 7A and Figure 7B The first channel set 716 (channel 0 and channel 1) has been reconfigured to be active at the new requested data rate (eg, GEN3 speed). Figure 7B After the training in , the data traffic is transferred from the second channel set 712 to the first channel set 716. The data traffic is then sent on the first channel set.
[0087] Fig.7D A multi-channel data link 746 between a host 702 and an EP 706 is shown after data traffic is transferred to a first channel set 718 and a second channel set 716 is changed to an idle state. The first channel set 716 of the data link 746, rather than the second channel set 714, is now carrying data traffic at the requested data rate GEN3. The second channel set 718 of the data link 746 is not required to support a x2 link width. Therefore, the second channel set 718 can be changed to an electrical idle state. In some aspects, the second channel set is trained and configured to the new data rate before the state of the second channel set is changed to the idle state. Fig.7D In the example shown, the data rate of the data link has been changed and lane flipping is no longer required. The x2 link is carried on the first set of lanes (lane 0 and lane 1). In some examples, lane flipping may be disabled. In this example, as in other examples, the data rate has been increased from GEN2 to GEN3. However, the same method can be used to reduce the data rate from GEN3 to GEN2. The data rate can also be changed from or to other rates, including non-adjacent rates, such as from GEN5 to GEN2. This method can be used for any data rate change and can also be used to increase or decrease the link width.
[0088] Fig. 8A A multi-channel data link 840 is shown, such as a PCIe link with a duplex traffic channel between the host 802 and the EP 806. The duplex traffic channel has Figure 5A. The first channel set 810 has a single channel (channel 0) and is in an active state, such as L0 state, as a x1 link and transmits data in one or both directions at a specific data rate or speed configuration (e.g., GEN4 speed). The second channel set 814 (channel 1, channel 2, and channel 3) is in an idle state and does not transmit any traffic. Although four channels are shown, there may be more or fewer channels to suit a particular implementation.
[0089] Figure 8B A multi-channel data link 842 between a host 802 and an EP 806 is shown being trained to a new data rate. In one aspect, a controller at the host 802 or EP 804 receives a request to change the data rate of the data link 842 from GEN4 to GEN3 and to change the link width from x1 to x2. This is an example, and other combinations are possible. The data link 842 is changed by changing an idle state channel of a second channel set to an active state via a recovery state and a configuration state. As previously described, the first channel set 810 continues to send data traffic. This is to change the second channel set from an idle state to an active state and to train the second channel set to the requested data rate. While the second channel set 820 is undergoing recovery and configuration, the data link 842 continues to send data traffic at the original data rate on the original first channel set 811 including only channel 0.
[0090] Figure 8C 8 shows a multi-channel data link 844 after the configuration state speed of the second channel set 818 is changed. The first channel set 811 is still in the same active state, such as the L0 state. Data can be sent in both directions at the original speed (e.g., GEN4), such as Fig. 8A and Figure 8B as shown in . Figure 8B The second channel set 816 is configured into two separate sets. Channels 2 and 3 have been reconfigured into a new second channel set 818 and changed to an active state at a new requested data rate (e.g., GEN3 speed). Traffic is still carried at the original speed on the first channel set 811. Channel 1 is the only channel in the third channel set 820 and is changed to an idle state. Figure 8B After the training in, data traffic is transmitted from the first channel set 810 to the second channel set 818.
[0091] Fig.8DA multi-channel data link 846 between the host 802 and the EP 806 is shown after data traffic is transferred to the second channel set 818. The configuration of the second channel set has been changed to the new data rate. The data traffic is sent on the second channel set 822 of the data link. The first channel set 812 of the data link 846 does not need to support the x2 link width, and therefore it is changed to the electrical idle state. The first channel set 810 and the third channel set 820 are combined to form a new first channel set 812, and both the first channel set and the third channel set are changed to the idle state. The request for the new data rate and the new link width is completed.
[0092] The x2 link width is using lane 2 and lane 3 instead of lane 0 and lane 1. In some aspects, lane flipping can be enabled to support the use of lane 3 and lane 2 instead of lane 0 and lane 0. Lane 1 is not used in the new configuration in a x4 link because the active lanes are adjacent and include the first lane (lane 0) or, in the case of lane flipping, the last lane (lane 3). Similar principles can be applied to other link widths and other transitions. The first lane set 812 can be trained directly to the new data rate or remain in an idle state.
[0093] Fig. 9 A flow chart of a method 900 for reducing the latency of link speed switching in a multi-channel data link (e.g., a PCIe link) according to various aspects of the present disclosure is illustrated. In some aspects, method 900 involves switching the speed of an idle service channel of a link while an active channel continues to send data traffic. When a previously idle service channel is configured to a new speed, data traffic is transferred to the configured channel. This allows data traffic to continue without interruption during the speed switch. Some modifications may be made to the control register to indicate the change in operation. The method of switching the speed of a service channel in a PCIe link may be performed by a PCIe controller including an RC controller 218 (also referred to as a host controller) and an EP controller 262. There will be a negotiation to establish capabilities and bandwidth requirements. Once the negotiation is complete, the RC controller will trigger method 900 to propose a second channel set with a requested GEN speed. Once the channel is turned on, the RC controller will trigger an interrupt to transfer the ongoing data transmission from the first active channel set to the second new training channel set.
[0094] Method 900 includes receiving at a controller at block 902 a request for a multi-channel data link to change the data rate of the data link to a requested data rate. The data link has a first set of channels in an active state and a second set of channels in an idle state. The first set of channels may have a lower or higher channel number than the second set of channels. As described, the link is a PCIe link, however the method may be adapted to accommodate other links having multiple channels.
[0095] The method 900 includes a process of changing the second channel set to an active state at block 904. When L0p is enabled, the channel set may be in the active state L0p, and all other channels may be in electrical idle. The method 900 includes a process of training the second channel set to a requested data rate at block 906. The training may be performed by restoring a state and configuring a state or using a different process to accommodate a specific data link. In some aspects, data traffic is sent on the first channel set while the second channel set is being trained.
[0096] Method 900 includes a process of transmitting data traffic from a first channel set to a second channel set after training at box 908. An interrupt may be transmitted to a host controller and an endpoint controller to transmit data traffic from the first channel set to the second channel set. To facilitate transmission, the controller may write an indication to a control register to identify the second channel set. For example, a starting channel number for the second channel set (which is an active channel set) may be written to the control register. In addition, the controller may write an indication of the width of the data link to the control register. The controller may also write a channel flip indication to the control register to indicate whether the first channel is the highest numbered channel or the lowest numbered channel in the link width. The method includes a process of sending data traffic on the second channel set at box 910.
[0097] After the transmission, the first channel set no longer carries traffic. In some examples, the method 900 may optionally include changing the first channel set to an idle state or training the first channel set to a requested data rate and then changing the first channel set to an idle state. In some examples, the method 900 may also include training the first channel set to a requested data rate and then using the first channel set to enlarge or reduce the link width.
[0098] Fig.101004 is a block diagram of a link interface processing circuit. Processing circuit 1004 is a device that can be part of a host or an endpoint. The processing circuit is coupled to a link 1002 having multiple duplex channels, such as a PCIe link. Link 1002 is coupled to another PCIe device, such as an endpoint or a host, at the opposite end. Data and control information transmitted as packets through link 1002 are coupled to PCIe interface 1020, which provides a PHY level interface to the link and converts baseband signals into packets. Data and control packets are transmitted to other components of processing circuit 1004 through bus 1010 via interface 1020.
[0099] The processing circuit 1004 includes a memory 1008. In addition, the processing circuit 1004 includes a data rate / link width change block 1012 that is coupled to the bus 1010 to receive and transmit requests to change the data rate and / or link width of the link 1002. The data rate / link width change block 1012 can access the computer-readable storage medium 1008 to access code 1032 for receiving data rate / link width requests. In various aspects, the storage medium 1008 is a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium 1008 has instructions stored therein for causing the processor 1006 of the interconnect link to perform operations such as Fig. 9 and Fig.10 The data rate / link width change block 1012 may also access an enable and channel configuration register 1040 containing enable bits and configuration bits in a storage medium 1008, such as a control register in a system memory of the processing circuit 1004.
[0100] In addition, within the processing circuit 1004, a channel state change block 1014 can be configured to change the state of a channel in the link 1002 from active to inactive. The channel state change block can operate, for example, using the active state power management (ASPM) protocol of the PCIe system or using another protocol or method. The channel state change block 1014 can access code for managing the channel state 1034 and access enable and channel configuration registers 1040 and other control registers of the processing circuit through the bus 1010. These registers can be used to store changes in the channel state for use by other blocks.
[0101] A channel training block 1016 within the processing circuit 1004 may access the bus 1010 to train the channel set to a higher or lower data rate, for example, to train the first channel set or the second channel set to a requested data rate. The channel training block 1016 may access code in the storage medium 1008 for training the channel set 1036, and may also access the enable and channel configuration registers 1040 to store the results.
[0102] The data traffic transmission / sending block 1018 within the processing circuit 1004 can access the bus 1010 to transmit the data traffic from the first set of channels to the second set of channels after training, and to send the data traffic on the second set of channels. The data traffic transmission / sending block can access the code for transmitting and sending the data traffic 1038 in the storage medium 1008, and can also access the enable and channel configuration register 1040 to determine the active channel. Each of the aforementioned blocks can be coupled to the bus 1010 to enable the blocks to communicate with each other, with the storage medium 1008, and to the processor 1006. The processor 1006 controls the operation of the other blocks and starts an instance of each block according to the operation of the processing circuit 1004.
[0103] The interface configuration block 1022 can access code for configuring the PCIe interface 1042. When executing the code, the interface configuration block 1022 reads and writes values from various configuration registers. These registers include enable and channel configuration registers 1040, which can be part of the control, status and capability registers. These registers can be accessed and read at the beginning of link initialization, and then updated as the state changes and the activity changes. The registers can also be modified in response to power management and bandwidth negotiation or in response to changing the state of one or more channels of the link 1002.
[0104] The following provides an overview of various embodiments of the disclosure.
[0105] Embodiment 1: A device, comprising: an interface circuit, the interface circuit being configured to provide an interface with a multi-channel data link, the data link having a first channel set in an active state and a second channel set in an idle state; and a controller, the controller being configured to: receive a request at the controller to change the data rate of the data link to a requested data rate; change the second channel set from an idle state to an active state; train the second channel set to the requested data rate; transmit data services from the first channel set to the second channel set after the training; and send the data services on the second channel set.
[0106] Embodiment 2: The apparatus of Embodiment 1, wherein the controller is further configured to change the first set of channels to an idle state.
[0107] Embodiment 3: According to the device of embodiment 1 or 2, the device further comprises a link width control register, and the controller is configured to write an indication of the width of the data link including the second channel set into the link width control register.
[0108] Embodiment 4: According to the device described in Embodiment 3, the device further comprises a first channel control register, and the controller is configured to write an indication identifier of the first channel in the second channel set as an active channel set into the first channel control register.
[0109] Embodiment 5: According to the device described in any one or more of Embodiments 1 to 4, the device also includes an enable control register, and the controller is configured to write an enable indication flag into the enable control register to indicate that data services can be transmitted from the first channel set to the second channel set.
[0110] Embodiment 6: A method, comprising: receiving at a controller a request for a multi-channel data link to change the data rate of the data link to a requested data rate, the data link having a first channel set in an active state and a second channel set in an idle state; changing the second channel set to an active state; training the second channel set to the requested data rate; transmitting data traffic from the first channel set to the second channel set after the training; and sending the data traffic on the second channel set.
[0111] Embodiment 7: According to the method described in Embodiment 6, the method further comprises changing the first channel set to an idle state after transmitting the data service.
[0112] Embodiment 8: The method according to embodiment 6 or 7 further comprises training the first channel set to the requested data rate during sending the data service on the second channel set.
[0113] Embodiment 9: The method according to any one or more of embodiments 6 to 8 further comprises enlarging the data link to include the first set of channels.
[0114] Embodiment 10: According to the method described in any one or more of Embodiments 6 to 9, the method further comprises writing an indication of the width of the data link including the second channel set into a control register of the data link.
[0115] Embodiment 11: The method according to embodiment 10, wherein writing the indication mark of the width further comprises writing the indication mark of the first channel in the second channel set as the active channel set into the control register of the data link.
[0116] Embodiment 12: According to the method described in any one or more of Embodiments 6 to 11, the method further comprises writing an enable indication flag into a control register of the data link to indicate that the data service can be transmitted from the first channel set to the second channel set.
[0117] Embodiment 13: The method according to any one or more of Embodiments 6 to 12, wherein training the second channel set includes training the second channel set during sending the data service on the first channel set.
[0118] Embodiment 14: A method according to any one or more of Embodiments 6 to 13, wherein transmitting the data service includes, after training the second channel set, transmitting an interrupt to a host controller and an endpoint controller to transmit the data service from the first channel set to the second channel set.
[0119] Embodiment 15: A method according to any one or more of Embodiments 6 to 14, wherein the data link is a Peripheral Component Interconnect Express (PCIe) link, and wherein the training includes a PCIe recovery state and a PCIe configuration state.
[0120] Embodiment 16: The method of Embodiment 15, wherein the first set of lanes is in a PCIe L0p state in the active state.
[0121] Embodiment 17: The method of Embodiment 15 or 16, wherein the second set of channels is in an electrically idle state.
[0122] Embodiment 18: The method according to any one or more of Embodiments 15 to 17, wherein changing the second set of lanes to the active state includes changing the second set of lanes to a PCIe L0p state.
[0123] Embodiment 19: A non-transitory computer-readable medium having instructions stored therein, the instructions being used to cause a processor of an interconnection link to perform operations, the operations comprising: receiving at a controller a request for a multi-channel data link to change the data rate of the data link to a requested data rate, the data link having a first channel set in an active state and a second channel set in an idle state; changing the second channel set to an active state; training the second channel set to the requested data rate; transmitting data traffic from the first channel set to the second channel set after the training; and sending the data traffic on the second channel set.
[0124] Embodiment 20: The non-transitory computer readable medium of embodiment 19, wherein training the second set of channels comprises training the second set of channels during transmission of the data traffic on the first set of channels.
[0125] Embodiment 21: A non-transitory computer-readable medium according to embodiment 19 or 20, wherein the instructions for transmitting the data service include instructions for transmitting an interrupt to a host controller and an endpoint controller to transmit the data service from the first channel set to the second channel set after training the second channel set.
[0126] Embodiment 22: An apparatus, the apparatus comprising: means for providing an interface with a multi-channel data link; means for receiving a request to change a data rate of the data link to a requested data rate, the data link having a first set of channels in an active state and a second set of channels in an idle state; means for changing the second set of channels to an active state; means for training the second set of channels to the requested data rate; means for transferring data traffic from the first set of channels to the second set of channels after the training; and means for sending the data traffic on the second set of channels. [This section will be completed in final draft.]
[0127] It should be understood that the present disclosure is not limited to the exemplary terms used above to describe aspects of the present disclosure. For example, bandwidth may also be referred to as throughput, data rate, or another term.
[0128] Although aspects of the present disclosure are discussed above using the example of the PCIe standard, it should be understood that the present disclosure is not limited to this example and may be used with other standards.
[0129] Each of the host client 214, host controller 212, device controller 252, and device client 254 discussed above may be implemented using a controller or processor configured to perform the functions described herein by executing software including code for performing the functions described herein. The software may be stored on a non-transitory computer-readable storage medium, such as RAM, ROM, EEPROM, optical disk, and / or magnetic disk, shown as host system memory 240, endpoint system memory 274, or another memory.
[0130] Any reference to an element using designations such as "first," "second," etc. herein does not generally limit the number or order of those elements. Rather, these designations are used herein as a convenient method of distinguishing two or more elements or instances of elements. Thus, reference to a first element and a second element does not mean that only two elements can be used or that the first element must be located before the second element.
[0131] Within this disclosure, the word "exemplary" is used to mean "serving as an example, instance, or illustration." Any specific implementation or aspect described herein as "exemplary" is not necessarily to be construed as preferred or superior to other aspects of the disclosure. Likewise, the term "aspect" does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation. The term "coupled" is used herein to refer to a direct or indirect electrical or other communicative coupling between two structures. Additionally, the term "approximately" means within ten percent of a stated value.
[0132] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples described herein, but should be granted the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A device, comprising: an interface circuit configured to provide an interface with a multi-lane data link having a first set of lanes in an active state and a second set of lanes in an idle state; as well as A controller, the controller being configured to: receiving, at the controller, a request to change a data rate of the data link to a requested data rate; changing the second set of channels from an idle state to an active state; training the second set of channels to the requested data rate; transmitting data traffic from the first set of channels to the second set of channels after the training; as well as The data service is sent on the second set of channels. 2 . The apparatus of claim 1 , wherein the controller is further configured to change the first channel set to an idle state. 3 . The apparatus according to claim 1 , further comprising a link width control register, wherein the controller is configured to write an indication of a width of the data link including the second lane set into the link width control register. 4 . The device according to claim 3 , further comprising a first channel control register, wherein the controller is configured to write an indication flag of the first channel in the second channel set as an active channel set into the first channel control register. 5 . The device according to claim 1 , further comprising an enable control register, wherein the controller is configured to write an enable indication flag into the enable control register to indicate that data services can be transmitted from the first channel set to the second channel set.
6. A method comprising: receiving, at a controller, a request for a multi-lane data link to change a data rate of the data link to a requested data rate, the data link having a first set of lanes in an active state and a second set of lanes in an idle state; changing the second channel set to an active state; training the second set of channels to the requested data rate; transmitting data traffic from the first set of channels to the second set of channels after the training; as well as The data service is sent on the second set of channels. 7 . The method according to claim 6 , further comprising changing the first channel set to an idle state after transmitting the data service.
8. The method of claim 6, further comprising training the first set of channels to the requested data rate during transmission of the data traffic on the second set of channels.
9. The method of claim 6, further comprising enlarging the data link to include the first set of channels.
10. The method according to claim 6, further comprising writing an indication of a width of the data link including the second set of lanes into a control register of the data link. 11 . The method according to claim 10 , wherein writing the indication of the width further comprises writing an indication of the first channel in the second channel set as an active channel set into a control register of the data link.
12. The method according to claim 6, further comprising writing an enable indication flag into a control register of the data link to indicate that the data service can be transmitted from the first channel set to the second channel set.
13. The method of claim 6, wherein training the second set of channels comprises training the second set of channels during transmission of the data traffic on the first set of channels.
14. The method of claim 6, wherein transferring the data traffic comprises, after training the second set of channels, transmitting an interrupt to a host controller and an endpoint controller to transfer the data traffic from the first set of channels to the second set of channels.
15. The method of claim 6, wherein the data link is a Peripheral Component Interconnect Express (PCIe) link, and wherein the training includes a PCIe recovery state and a PCIe configuration state.
16. The method of claim 15, wherein the first set of lanes is in a PCIe L0p state in the active state. The method of claim 15 , wherein the second set of channels is in an electrically idle state.
18. The method of claim 15, wherein changing the second set of lanes to the active state comprises changing the second set of lanes to a PCIe L0p state.
19. A device, comprising: Components for providing an interface to a multi-channel data link; means for receiving a request to change a data rate of the data link to a requested data rate, the data link having a first set of lanes in an active state and a second set of lanes in an idle state; means for changing the second set of channels to an active state; means for training said second set of channels to said requested data rate; means for transmitting data traffic from said first set of channels to said second set of channels after said training; as well as Means for sending the data service on the second set of lanes.
20. The apparatus of claim 19, wherein means for training the second set of channels comprises means for training the second set of channels during transmission of the data traffic on the first set of channels.
Citation Information
Patent Citations
Seamless addition of high bandwidth lanes
CN107924378A
Chip to chip interface with scalable bandwidth
US10521391B1
Controlling A Physical Link Of A First Protocol Using An Extended Capability Structure Of A Second Protocol
US20140108697A1
Link speed control systems for power optimization
US20170280385A1
Partial link width states for multilane links
US20190196991A1
Cited By
Link rate switching method, device, equipment, medium and product
CN120768712A