Peripheral Component Interconnect Power Management

By introducing the L0ps state, the host and endpoint can independently switch to a low-power state in FLIT mode, solving the problem that the PCIe link is difficult to save power in a low-activity state and achieving more efficient energy consumption management.

CN119948469BActive Publication Date: 2025-09-16QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380069266.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-07-31
Publication Date
2025-09-16
Estimated Expiration
2043-07-31

AI Technical Summary

Technical Problem

Existing PCIe links have difficulty saving power efficiently in low-activity states, especially in FLIT mode, where the L0s state is unavailable, making it impossible to independently transition to a low-power state, affecting device energy consumption management.

Method used

A new link state, L0ps, is introduced, allowing the host and endpoint to enter the L0ps state independently of L0p in FLIT mode, trigger a fast transition through an inactivity timer, and recover without going through the recovery state, combining electrical idle ordered sets and fast training sequences for channel de-skew.

Benefits of technology

It achieves more efficient power saving in FLIT mode, reduces the energy consumption of devices in low activity states, and improves the power management efficiency of the link.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948469B_ABST
    Figure CN119948469B_ABST
Patent Text Reader

Abstract

A new Peripheral Component Interconnect Express (PCIe) link state can enhance the power saving capabilities of a PCIe link operating in flow control unit (FLIT) mode. A device can operate a data link with a host in FLIT mode using fixed-size packets, the data link being in a partial-width link state (PLS) in which a first set of lanes of the data link are in an electrically idle state and a second set of lanes of the data link are in an active state capable of being used for data traffic with the host. The device can transition one or more lanes of the second set of lanes from the PLS to a partial-width standby link state (PSLS) in which the one or more lanes of the second set of lanes are in a standby state having lower power consumption than the active state.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims priority to pending U.S. non-provisional application No. 17 / 959,996, filed on October 4, 2022, which is assigned to the assignee of this application and is hereby expressly incorporated herein by reference as if fully set forth below and for all applicable purposes. Technical Field

[0003] The techniques discussed below relate generally to Peripheral Component Interconnect Express (PCIe) devices, and more particularly to techniques for managing link power consumption of PCIe devices. Background Art

[0004] High-speed interfaces are often used between circuits and components in mobile wireless devices and other complex systems. For example, some devices may include processing devices, communication devices, storage devices, and / or display devices that interact with each other through one or more high-speed interfaces. Some of these devices, including synchronous dynamic random access memory (SDRAM), may be able to provide or use data and control information at the processor clock rate. Other devices, such as display controllers, may use variable amounts of data at relatively low video refresh rates.

[0005] The Peripheral Component Interconnect Express (PCIe) standard is a high-speed interface that supports high-speed data links capable of sending data at speeds of several gigabits per second. The PCIe interface also has multiple standby modes for when the link is inactive. Compared to parallel buses, PCIe offers lower latency and higher data transfer rates. PCIe can be used for communication between a wide variety of different devices. Typically, one device (e.g., a processor or hub) acts as a host, which communicates with multiple devices (called endpoints) over PCIe links (data links). Peripheral devices or components can include graphics adapter cards, network interface cards (NICs), storage accelerator devices, mass storage devices, input / output (I / O) interfaces, and other high-performance peripherals.

[0006] The connection between any two PCIe devices is called a link. PCIe links are built around duplex, serial (1-bit), differential, point-to-point connections called lanes. With PCIe, data is transmitted via two signal pairs: two lines (wires, circuit board traces, etc.) for transmit and two lines for receive. The transmit and receive pairs are separate differential pairs for a total of four data lines per lane. The link consists of a set of lanes, and each lane is capable of simultaneously transmitting and receiving data packets between the host and endpoints. As currently defined, a PCIe link can scale from one lane to 32 individual lanes. Common deployments have 1, 2, 4, 8, 12, 16, or 32 lanes, labeled x1, x2, x4, x8, x12, x16, or x32, respectively, where the number refers to the number of lanes. In one example, a PCIe x1 implementation has four lanes to connect a pair of lanes in each direction, while a PCIe x16 implementation has 16 times that number—16 lanes, or 64 lanes.

[0007] There are various link power management states that a PCIe physical link can enter and exit in response to stateful power management activities, such as L0, multiple L0s, and L1. These link power management states allow PCIe devices to use power more efficiently depending on the traffic conditions or state of the PCIe link. Summary of the Invention

[0008] The following content presents a summary of one or more implementations in order to provide a basic understanding of such implementations. This summary is not an exhaustive overview of all contemplated implementations and is not intended to identify key or important elements of all implementations, nor is it intended to delineate the scope of any or all implementations. Its sole purpose is to present some concepts of one or more implementations in a simplified form as a prelude to the more detailed description that will be presented later.

[0009] In one example, a method of operating an endpoint for data communication is disclosed. The method includes operating a data link with a host in a flow control unit (FLIT) mode using fixed-size packets, the data link being in a partial width link state (PLS), wherein a first set of lanes of the data link is in an electrically idle state, and a second set of lanes of the data link is in an active state available for data traffic with the host. The method also includes transitioning one or more lanes of the second set of lanes of the data link from the PLS to a partial width standby link state (PSLS), wherein the one or more lanes of the second set of lanes are in a standby state having lower power consumption than the active state.

[0010] In one example, an endpoint for a Peripheral Component Interconnect Express (PCIe) link is provided. The endpoint includes an interface circuit configured to provide an interface with the PCIe link connected to a host. The endpoint also includes a controller configured to operate the PCIe link in a flow control unit (FLIT) mode using fixed-size packets. The PCIe link is in a partial width link state (PLS), in which a first group of lanes of the PCIe link is in an electrically idle state and a second group of lanes of the PCIe link is in an active state available for data traffic with the host. The controller is further configured to transition one or more lanes of the second group of lanes PCIe link from the PLS to a partial width standby link state (PSLS), in which the one or more lanes of the second group of lanes are in a standby state having lower power consumption than the active state.

[0011] In one example, a host for a Peripheral Component Interconnect Express (PCIe) link is provided. The host includes an interface circuit configured to provide an interface with the PCIe link connected to an endpoint. The host also includes a controller configured to operate the PCIe link in a flow control unit (FLIT) mode using fixed-size packets. The PCIe link is in a partial width link state (PLS), in which a first group of channels of the PCIe link is in an electrically idle state and a second group of channels of the PCIe link is in an active state that can be used for data traffic with the endpoint. The controller is further configured to transition one or more lines of the second group of channels PCIe link from the PLS to a partial width standby link state (PSLS), in which the one or more lines of the second group of channels are in a standby state having lower power consumption than the active state.

[0012] To accomplish the foregoing and related objectives, one or more implementations include the features fully described below and particularly pointed out in the claims. The following description and the accompanying figures set forth in detail certain illustrative aspects of one or more implementations. However, these aspects are merely indicative of a few of the various ways in which the principles of various implementations may be employed, and the described implementations are intended to encompass all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a block diagram of a computing architecture with a Peripheral Component Interconnect Express (PCIe) interface suitable for use with aspects of the present disclosure.

[0014] Figure 2 is a block diagram of a system including a host system and an endpoint device system according to aspects of the present disclosure.

[0015] Figure 3 is a diagram of lanes and corresponding drivers in a PCIe link according to aspects of the present disclosure.

[0016] Figure 4 is a state diagram illustrating the operation of a power management state machine according to aspects of the present disclosure.

[0017] Figure 5 is a diagram illustrating a multi-lane PCIe link between a host and an endpoint according to aspects of the present disclosure.

[0018] Figure 6 is a diagram illustrating a first example of a PCIe link operating in flow control unit (FLIT) mode between a host and an endpoint, according to some aspects of the present disclosure.

[0019] Figure 7 is a diagram illustrating a second example of a PCIe link operating in FLIT mode between a host and an endpoint, according to aspects of the present disclosure.

[0020] Figure 8 is a diagram illustrating a PCIe configuration space structure according to some aspects of the present disclosure.

[0021] Figure 9 is a diagram illustrating exemplary PCIe registers according to aspects of the present disclosure.

[0022] Figure 10 is a block diagram of a PCIe link interface processing circuit according to aspects of the present disclosure.

[0023] Figure 11 is a flow chart of an exemplary method for link state management of a PCIe link according to aspects of the present disclosure. DETAILED DESCRIPTION

[0024] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein may be practiced. The detailed description includes specific details to provide a thorough understanding of the various concepts. However, these concepts may be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.

[0025] Recent Peripheral Component Interconnect Express (PCIe) specifications (e.g., PCIe 6.0) may use Flow Control Module (FLIT) encoding to improve the latency and efficiency of PCIe links. When a PCIe link is in FLIT mode, error correction is performed on fixed-size packets (flits). At the physical layer of the PCIe link, the data transmission unit is a flit. In addition, PCIe 6.0 introduces a partial-width link state (L0p) available in FLIT mode. In L0p, some channels (i.e., partial widths) may be in an electrically idle mode. When the channels are electrically idle, the corresponding line drivers may be set to a static or high-impedance state, and the differential voltage of the channels may be fixed (e.g., 0 volts). In some aspects, the standby link state L0s is not available in FLIT mode. Therefore, for L0p, the L1 state becomes a power-saving state with minimal recovery latency (e.g., approximately 64 μs). In FLIT mode, the transition from L0p to the power-saving state L1 is triggered by an L1 state inactivity timer. However, for unidirectional transmission (i.e., only the receive (RX) line or only the transmit (TX) line is occupied), both the TX line and the RX line remain in the active L0p state because both devices connected by the PCIe link (e.g., host and endpoint) need to transition to L1. Since L0s is not available in FLIT mode, power cannot be saved by transitioning the transmitter (e.g., host or endpoint) alone to a low power state without any negotiation or handshake sequence required to transition the device to other available low power states (e.g., L1).

[0026] In an exemplary handshake sequence, the physical layer (PHY) of a PCIe device (e.g., an endpoint) may detect a certain idle period on the PCIe link (e.g., based on a PCIe inactivity timer). This idle period may be implementation-specific (e.g., 7 microseconds (μs) to 10 μs). The device then blocks new outbound PCIe transactions (e.g., PCIe traffic) to the system (e.g., a host). The PCIe device may continue transmitting PM_Active_State_Request_L1 Data Link Layer Packets (DLLPs) to the system until the device receives a PM_Request_ACK from the system. When the system receives the PM_Active_State_Request_L1 DLLP, the system blocks new transactions to the device and continues transmitting the PM_Request_ACK until the system receives an Electrically Idle Ordered Set. When the device receives the PM_Request_ACK, it transmits an Electrically Idle Ordered Set and places the device's transmitter in an Electrically Idle state. When the system receives the Electrically Idle Ordered Set, the system places its transmitter in Electrically Idle. At this point, the PCIe link is in the L1 state. The system or device may initiate an exit from the L1 state.

[0027] Various aspects of the present disclosure provide technologies for implementing new PCIe link states to enhance the power saving capabilities of PCIe links in FLIT mode. These technologies enable hosts and endpoints to enter new link states (referred to as L0ps in this disclosure) independently of the L0p in FLIT mode. In some aspects, a PCIe link in the L0p state can (e.g., based on an L0ps inactivity timer) quickly enter the L0ps state and recover from the L0ps state without experiencing a recovery state. In some aspects, a channel can enter the L0ps state after receiving an electrical idle ordered set (EIOS) for a PCIe channel. In some aspects, when the channel returns to the L0p state from the L0ps state, the channel reestablishes bit lock, symbol lock, or block alignment and performs channel-to-channel de-skew. Although the channels of a multi-channel PCIe link can send data symbols simultaneously, when the data symbols of different channels arrive at the receiver at different times, channel-to-channel skew occurs. The arrival time difference is referred to as channel-to-channel skew. For example, a device exiting L0ps can de-skew the channels by transmitting an exit pattern on the idle channels to train and de-skew them. As an example, the exit pattern can include an Electrical Idle Exit Ordered Set (EIEOS) and a Fast Training Sequence (FTS).

[0028] In some aspects, the TX and RX lanes on a PCIe lane can be switched to the L0ps state independently (e.g., not simultaneously). In some aspects, the TX lanes and RX lanes can be independently in the L0ps state based on separate inactivity timeouts without any handshaking between the host and the endpoint.

[0029] Figure 11 is a block diagram of an exemplary computing architecture using a PCIe interface. Computing architecture 100 operates using multiple high-speed PCIe interface serial links. The PCIe interface can be characterized as an arrangement comprising a point-to-point topology, wherein a separate serial link connects each device to a host, which can be referred to as a root complex 104. In computing architecture 100, root complex 104 couples processor 102 to memory devices (e.g., memory subsystem 108) and PCIe switch circuitry 106. In some examples, PCIe switch circuitry 106 includes cascaded switch devices. One or more PCIe endpoint devices 110 can be directly coupled to root complex 104, while other PCIe endpoint devices 112-1, 112-2, ..., 112-N can be coupled to root complex 104 via PCIe switch circuitry 106. Root complex 104 can be coupled to processor 102 using a proprietary local bus interface or a standard-defined local bus interface. Root complex 104 may control configuration and data transactions over a PCIe interface and may generate transaction requests for processor 102. In some examples, root complex 104 is implemented in the same integrated circuit (IC) device that includes processor 102. Root complex 104 may support multiple PCIe ports.

[0030] The root complex 104 can control communication between the processor 102 and a memory subsystem 108, which is one example of an endpoint. The root complex 104 (host) also controls communication between the processor 102 and other PCIe endpoint devices 110, 112-1, 112-2, ..., 112-N. The PCIe interface can support full-duplex communication between any two endpoints, with no inherent restrictions on concurrent access across multiple endpoints. Data packets can carry information across any PCIe link. In a multi-lane PCIe link, packet data can be striped across multiple lanes. The number of lanes in a multi-lane link can be negotiated during device initialization and can be different for different endpoints.

[0031] When one or both traffic directions of the lanes of a PCIe link are underutilized by low-bandwidth applications (which can be adequately served by fewer lanes), the root complex 104 and endpoints can operate the link with more or fewer transmit and receive lanes in one or both directions. In some aspects, the host (e.g., root complex 104) and endpoints can operate in FLIT mode and change between partial-width link states (e.g., L0p and L0ps) based on the traffic conditions of the link.

[0032] In some aspects, the computing architecture 100 can be implemented based on the PCIe M.2 specification. The M.2 form factor can be used for mobile adapters. M.2 enables the expansion, contraction, and higher integration of functionality into a single form factor module solution. For example, Figure 1 Any of the described PCIe endpoints may be implemented as an M.2 adapter, and the root complex 104 may be implemented as an M.2 platform.

[0033] Figure 2 is a block diagram of an exemplary PCIe system in which aspects of the present disclosure may be implemented. System 205 includes a host system 210 and an endpoint device system 250, which may communicate with Figure 1 The host and endpoint are the same. For example, the host system 210 can be a PCIe M.2 platform and the endpoint device system 250 can be an M.2 adapter. The host system 210 can be integrated on a first chip (e.g., a system on a chip or SoC), and the endpoint device system 250 can be integrated on a second chip. Alternatively, the host system and / or the endpoint device system can be integrated in a first package and a second package (e.g., a SiP), a first system board and a second system board having multiple chips, or integrated in other hardware or any combination. In this example, the host system 210 and the endpoint device system 250 are coupled via a PCIe link 285.

[0034] Host system 210 includes one or more host clients 214. Each of the one or more host clients 214 can be implemented on a processor that executes software that performs the functions of host client 214 discussed herein. For examples with more than one host client, the host clients can be implemented on the same processor or on different processors. Host system 210 also includes a host controller 212 that can perform root complex functions. Host controller 212 can be implemented on a processor that executes software that performs the functions of host controller 212 discussed herein.

[0035] The host system 210 includes PCIe interface circuitry 216, a system bus interface 215, and a host system memory 240. The system bus interface 215 can interface one or more host clients 214 with the host controller 212, and each of the one or more host clients 214 and the host controller 212 with the PCIe interface circuitry 216 and the host system memory 240. The PCIe interface circuitry 216 provides the host system 210 with an interface to a PCIe link 285. In this regard, the PCIe interface circuitry 216 is configured to transmit data (e.g., from the host client 214) to the endpoint device system 250 via the PCIe link 285 and to receive data from the endpoint device system 250 via the PCIe link 285. The PCIe interface circuitry 216 includes a PCIe controller 218, a physical interface 220 for a PCI Express (PIPE) interface, a physical (PHY) transmit (TX) block 222, a clock generator 224, and a PHY receive (RX) block 226. PIPE interface 220 provides a parallel interface between PCIe controller 218 and PHY TX block 222 and PHY RX block 226. PCIe controller 218 (which may be implemented in hardware) may be configured to perform transaction layer, data link layer, and control flow functions specified in the PCIe specification, as described further below.

[0036] The host system 210 also includes an oscillator (e.g., a crystal oscillator or "XO") 230 configured to generate a reference clock signal 232. In one example, the reference clock signal 232 may have a frequency of 19.2 MHz, but is not limited to this frequency. The reference clock signal 232 is input to a clock generator 224, which generates a plurality of clock signals based on the reference clock signal 232. In this regard, the clock generator 224 may include one or more phase-locked loops (PLLs), each of which generates a respective one of the plurality of clock signals by multiplying the frequency of the reference clock signal 232.

[0037] Endpoint device system 250 includes one or more device clients 254. Each device client 254 may be implemented on a processor executing software that performs the functionality of the device client 254 discussed herein. In examples where there are more than one device client 254, the device clients 254 may be implemented on the same processor or on different processors. Endpoint device system 250 also includes a device controller 252. The device controller 252 may be configured to receive bandwidth requests from one or more device clients and, based on the bandwidth requests, determine whether to change the number of transmit lines or the number of receive lines. The device controller 252 may be implemented on a processor executing software that performs the functionality of the device controller.

[0038] The endpoint device system 250 includes a PCIe interface circuit 260, a system bus interface 256, and an endpoint system memory 274. The system bus interface 256 can interface one or more device clients 254 with the device controller 252, and each of the one or more device clients 254 and the device controller 252 with the PCIe interface circuit 260 and the endpoint system memory 274. The PCIe interface circuit 260 provides the endpoint device system 250 with an interface to a PCIe link 285. In this regard, the PCIe interface circuit 260 is configured to send data (e.g., from the device client 254) to the host system 210 (also referred to as a host device) via the PCIe link 285 and to receive data from the host system 210 via the PCIe link 285. The PCIe interface circuit 260 includes a PCIe controller 262, a PIPE interface 264, a PHY TX block 266, a PHY RX block 270, and a clock generator 268. PIPE interface 264 provides a parallel interface between PCIe controller 262 and PHY TX block 266 and PHY RX block 270. PCIe controller 262 (which may be implemented in hardware) may be configured to perform transaction layer, data link layer, and control flow functions.

[0039] Host system memory 240 and endpoint system memory 274 at the endpoint can be configured to contain registers containing the status of each transmit and receive lane of PCIe link 285. The transmit lanes can be configured as a differential transmit lane pair, and the receive lanes can be configured as a differential receive lane pair.

[0040] The endpoint device system 250 also includes an oscillator (eg, a crystal oscillator) 272 configured to generate a stable reference clock signal 273 for the endpoint system memory 274. Figure 2 In the example of FIG, the clock generator 224 at the host system 210 is configured to generate a stable reference clock signal 273, which is forwarded by the PHY RX block 226 to the endpoint device system 250 via the differential clock line 288. At the endpoint device system 250, the PHY RX block 270 receives the endpoint (EP) reference clock signal on the differential clock line 288 and forwards the EP reference clock signal to the clock generator 268. The EP reference clock signal may have a frequency of 100 MHz, but is not limited to such a frequency. The clock generator 268 may be configured to generate multiple clock signals based on the EP reference clock signal from the differential clock line 288, as discussed further below. In this regard, the clock generator 268 may include multiple phase-locked loops (PLLs), each of which generates a respective one of the multiple clock signals by multiplying the frequency of the EP reference clock signal.

[0041] The system 205 also includes a power management integrated circuit (PMIC) 290 coupled to a power source 292 (e.g., mains voltage, a battery, or other power source). The PMIC 290 is configured to convert the voltage of the power source 292 to a plurality of supply voltages (e.g., using a switching regulator, a linear regulator, or any combination thereof). In this example, the PMIC 290 generates a voltage 242 for the oscillator 230, a voltage 244 for the PCIe controller 218, and a voltage 246 for the PHY TX block 222, the PHY RX block 226, and the clock generator 224. The voltages 242, 244, and 246 can be programmable, wherein the PMIC 290 is configured to set the voltage levels (angles) of the voltages 242, 244, and 246 according to instructions (e.g., from the host controller 212).

[0042] The PMIC 290 also generates a voltage 280 for the oscillator 272, a voltage 278 for the PCIe controller 262, and a voltage 276 for the PHY TX block 266, the PHY RX block 270, and the clock generator 268. The voltages 280, 278, and 276 may be programmable, wherein the PMIC 290 is configured to set the voltage levels (angles) of the voltages 280, 278, and 276 according to instructions (e.g., from the device controller 252). The PMIC 290 may be implemented on one or more chips. Although the PMIC 290 may be implemented on a single chip, the PMIC 290 may be implemented on a single chip. Figure 2 290 is shown as one PMIC, but it should be understood that PMIC 290 can be implemented by two or more PMICs. For example, PMIC 290 can include a first PMIC for generating voltages 242, 244, and 246 and a second PMIC for generating voltages 280, 278, and 276. In this example, both the first PMIC and the second PMIC can be coupled to the same power supply 292 or different power supplies.

[0043] In operation, the PCIe interface circuitry 216 on the host system 210 can send data from one or more host clients 214 to the endpoint device system 250 via the PCIe link 285. When the host controller negotiates bandwidth for the link, data from the one or more host clients 214 can be directed to the PCIe interface circuitry 216 according to the PCIe mapping established by the host controller 212 during initial configuration (sometimes referred to as link initialization). At the PCIe interface circuitry 216, the PCIe controller 218 can perform transaction layer and data link layer functions on the data, such as packetizing the data, generating error correction codes to be sent with the data, etc.

[0044] The PCIe controller 218 outputs the processed data to the PHY TX block 222 via the PIPE interface 220. The processed data includes data from one or more host clients 214 as well as overhead data (e.g., packet headers, error correction codes, etc.). In one example, the clock generator 224 can generate a clock 234 for an appropriate data rate or transfer rate based on the reference clock signal 232 and input the clock 234 to the PCIe controller 218 to time the operation of the PCIe controller 218. In this example, the PIPE interface 220 can include a 22-bit parallel bus that transmits 22 bits of data in parallel to the PHY TX block for each cycle of the clock 234. At a frequency of 250 MHz, the transfer rate is approximately 8 GT / s.

[0045] The PHY TX block 222 serializes the parallel data from the PCIe controller 218 and drives the PCIe link 285 with the serialized data. In this regard, the PHY TX block 222 may include one or more serializers and one or more drivers. The clock generator 224 may generate a high-frequency clock for the one or more serializers based on the reference clock signal 232.

[0046] At the endpoint device system 250, the PHY RX block 270 receives the serialized data via the PCIe link 285 and deserializes the received data into parallel data. In this regard, the PHY RX block 270 may include one or more receivers and one or more deserializers. The clock generator 268 may generate a high-frequency clock for the one or more deserializers based on the EP reference clock signal. The PHY RX block 270 transmits the deserialized data to the PCIe controller 262 via the PIPE interface 264. The PCIe controller 262 may recover data from the one or more host clients 214 from the deserialized data and forward the recovered data to the one or more device clients 254.

[0047] On the endpoint device system 250, the PCIe interface circuit 260 can transmit data from one or more device clients 254 to the host system memory 240 via a PCIe link 285. In this regard, the PCIe controller 262 at the PCIe interface circuit 260 can perform transaction layer and data link layer functions on the data, such as packetizing the data, generating error correction codes to be sent with the data, etc. The PCIe controller 262 outputs the processed data to the PHY TX block 266 via the PIPE interface 264. The processed data includes data from the one or more device clients 254 and overhead data (e.g., packet headers, error correction codes, etc.). In one example, the clock generator 268 can generate a clock based on the EP reference clock via the differential clock line 288 and input the clock to the PCIe controller 262 to control the timing operation of the PCIe controller 262.

[0048] The PHY TX block 266 serializes the parallel data from the PCIe controller 262 and drives the PCIe link 285 with the serialized data. In this regard, the PHY TX block 266 may include one or more serializers and one or more drivers. The clock generator 268 may generate a high-frequency clock for the one or more serializers based on the EP reference clock signal.

[0049] At the host system 210, the PHY RX block 226 receives the serialized data via the PCIe link 285 and deserializes the received data into parallel data. In this regard, the PHY RX block 226 may include one or more receivers and one or more deserializers. The clock generator 224 may generate a high-frequency clock for the one or more deserializers based on the reference clock signal 232. The PHY RX block 226 transmits the deserialized data to the PCIe controller 218 via the PIPE interface 220. The PCIe controller 218 may recover data from the one or more device clients 254 from the deserialized data and forward the recovered data to the one or more host clients 214.

[0050] In some aspects, the host system 210 and the endpoint system 250 can operate the PCIe link 285 in FLIT mode and switch the link 285 between a partial-width link state (e.g., L0p) and a standby state (e.g., L0ps) in FLIT mode without going through a recovery state.

[0051] Figure 3 It is available in Figure 1 and Figure 2 FIG. 3 is a diagram of exemplary channels in a link 385 used in a system of FIG. For example, the link 385 may be implemented as Figure 2PCIe link 285. In this example, link 385 includes multiple lanes 310-1 to 310-n, where each lane includes a corresponding first differential line pair 312-1 to 312-n for transmitting data from host system 210 to endpoint device system 250, and a corresponding second differential line pair 315-1 to 315-n for transmitting data from the endpoint device system to host system 210. From the perspective of the host system, the first lane 310-1 is dual simplex, where the first differential line pair 312-1 serves as a transmit line and the second differential line pair 315-1 serves as a receive line. From the perspective of the endpoint device system, the first lane 310-1 has a receive line and a transmit line. The first differential line pairs 312-1 to 312-n and the second differential line pairs 315-1 to 315-n can be implemented using metal traces on a substrate (e.g., a printed circuit board), wherein the host system can be integrated on a first chip mounted on the substrate, and the endpoint device is integrated on a second chip mounted on the substrate. Alternatively, the link can be implemented via an adapter card slot (e.g., a PCIe M.2 slot), a cable, or a combination of different media. The link can also include an optical section in which PCIe packets are encapsulated within different systems. In this example, when data is transmitted from the host system to the endpoint device system across multiple channels, the PHY TX block 222 may include logic for dividing the data between the channels. Similarly, when data is transmitted from the endpoint device system to the host system 210 across multiple channels, the PHY TX block 266 may include logic for dividing the data between the channels.

[0052] Figure 2 The PHY TX block 222 of the host system 210 shown in FIG. 1 may be implemented to include transmit drivers 320 - 1 to 320 - n to drive each first differential line pair 312 - 1 to 312 - n to transmit data, and Figure 2 The PHY RX block 270 of the endpoint device system 250 shown in FIG can be implemented to include receivers 340-1 to 340-n (e.g., amplifiers) to receive data from each second differential line pair 312-1 to 312-n. Each transmit driver 320-1 to 320-n is configured to drive the corresponding differential line pair 312-1 to 312-n with data, and each receiver 340-1 to 340-n is configured to receive data from the corresponding first differential line pair 312-1 to 312-n. In addition, in FIG Figure 2In the embodiment, the PHY TX block 266 of the endpoint device system 250 may include a transmit driver 345-1 to 345-n for each second differential line pair 315-1 to 315-n, and the PHY RX block 226 of the host system 210 may include a receiver 325-1 to 325-n (e.g., an amplifier) ​​for each second differential line pair 315-1 to 315-n. Each transmit driver 345-1 to 345-n is configured to drive the corresponding second differential line pair 315-1 to 315-n with data, and each receiver 325-1 to 325-n is configured to receive data from the corresponding second differential line pair 315-1 to 315-n.

[0053] In some aspects, the width of the link 385 can be scalable to match the capabilities of the host system and endpoints. The link can use one lane 310-1 for a x1 link, two lanes 310-1, 310-2 for a x2 link, or more lanes, up to n lanes from 310-1 to 310-n, for wider links. Currently, links (x1, x2, x4, x8, x16, and x32) are defined for 1, 2, 4, 8, 16, and 32 lanes, but a different number of lanes may be used to suit a particular implementation.

[0054] In one example, the host system 210 may include a power switch circuit 350 configured to individually control the power supplied from the PMIC 290 to the transmit drivers 320-1 to 320-n and the receivers 325-1 to 325-n. Thus, in this example, the number of drivers and receivers powered is proportional to the width of the link 385. Similarly, Figure 2 The endpoint device system 250 shown in FIG may include a power switch circuit 360 that is configured to individually control power supply from the PMIC 290 to the transmit drivers 345-1 to 345-n and the receivers 340-1 to 340-n. In this manner, the host system can set the number of multiple drivers to be selectively powered by the power switch circuit to change the number of active transmit and / or receive lines based on the number of lines that are powered (active) or electrically idle. Using differential signaling, lines can be set to active or standby in pairs. In some aspects, the transmit and receive lines of a differential pair can be independently set to active or standby.

[0055] ASPM status

[0056] Figure 44 is a diagram illustrating the operation of a power management state machine according to some aspects disclosed herein. In some aspects, a PCIe system (e.g., system 205) may use the Active State Power Management (ASPM) protocol to manage power. The ASPM protocol is a power management mechanism for PCIe devices to reduce power usage based on link activity detected on a PCIe link between a host (e.g., a root complex) and an endpoint PCIe device. State diagram 400 illustrates some PCIe link states consistent with the Link Training and Status State Machine (LTSSM) defined for PCIe, and other link states may be omitted for brevity. In this example, the link may operate in L0 state 404 (i.e., an active link operating state), in which data may be transmitted in both directions over the PCIe link. In L0, a PCIe device (e.g., a host or endpoint) may be active and responding to PCIe transactions and / or may request or initiate PCIe transactions. Figure 4 Also shown is the standby state L1 406 defined in the PCIe specification. As shown, the L1 state 406 is accessible through a connection to the L0 state 404. When conditions on the link indicate or suggest that a transition between states is appropriate, an ASPM state change can be initiated. Both communication partners of the link (e.g., a host and an endpoint) can initiate a power state change request when conditions are correct (e.g., idle or low data traffic).

[0057] When the link is idle (e.g., with no data traffic in the short time intervals between data bursts or in a time interval greater than a predetermined threshold), the link can enter the standby state L0s 408 from the L0 state, which is only accessible through the L0 state. In PCIe, L0s 408 is a power-saving state accessible from L0. The link can also change to L1 state 406, which is a standby state with a higher exit latency than L0s state 402, so that it takes longer to return from L1 to L0 than from L0s to L0. However, L1 can provide greater power savings than L0s. In some aspects, the link can return from L1 to L0 via recovery state 410. In the recovery state, devices using the link (e.g., a host and endpoint) can exchange training sequences to negotiate various link parameters, including, for example, lane polarity, link / lane number, equalization parameters, data rate, etc. The exit latency of L0s and L1 refers to the time it takes for a device to return to the L0 state. When a device (e.g., host or endpoint) enters L0s, the transmitting device can transmit an Electrical Idle Ordered Set (EIOS) to the receiving device and then turn off the power to its transmitter. When the device returns from L0s to L0, it can send a specific number of small ordered sets called Fast Training Sequences (FTS) in the PCIe specification so that the receiver can regain receiver lock and be able to receive traffic on the link.

[0058] In L0s, data can be transmitted in both directions or in only one direction, so that two devices connected by the link (e.g., a host and an endpoint) can each independently set their transmitters to idle. In some aspects, the L0s state can serve as a low-latency standby state. Power saving techniques available during L0s may include, but are not limited to, powering down at least a portion of the transceiver circuitry and clock gating at least the link layer logic. In L0s, device discovery and bus configuration processes can be implemented prior to the link transition 422 from L0s to L0.

[0059] In L1, data is not transmitted over the link, so that portions of the PCIe transceiver logic and / or PHY circuitry can be shut down or disabled to achieve higher power savings than those achievable in L0s. For example, a PCIe device can shut down most link transceiver circuitry and / or PLLs. A PCIe device can also use application clock gating (i.e., reduce the clock rate) for most PCIe architecture logic. The L1 state is a primary standby state with higher latency and power savings than L0s. When a PCIe device determines that there are no outstanding PCIe requests or pending transactions or services, it can enter the L1 state by transitioning 424 from L0. In some examples, power consumption in L1 can be reduced by disabling or idling transceivers in the PCIe bus interface, disabling, gating, or slowing down the clocks used by PCI devices, and disabling PLL circuits used to generate clocks for receiving data. A PCIe device can make the transition 424 to L1 through operation of a hardware controller or some combination of an operating system and hardware control circuitry.

[0060] In some aspects, the ASPM protocol may determine whether to transition to L0s or L1 based on a finite time interval or a threshold defined as L0s / L1 entry latency. For example, whenever a PCIe link is inactive for a given L0s or L1 entry latency duration, the PCIe controller may request the link partner to enter a low-power or standby link state (L0s or L1) to save power. In some instances, the L0s entry latency duration and the L1 entry latency duration may be selected based on overall system parameters, activity, and / or pending operations. The ASPM state machine may initiate a transition to a low-power state (e.g., L0s or L1) after an observed period of link inactivity. The specific entry latency duration may be adjustable to accommodate different system architectures and device characteristics. In some aspects, in some implementations, the packet latency between read / write requests in the PCIe interface may vary, for example, between 1 μs and 40 μs. Generally, L0s entry latency is shorter than L1 entry latency.

[0061] FLIT mode

[0062] In some aspects, a PCIe link can operate in FLIT mode, which can improve PCIe link latency and efficiency by performing error correction on fixed-size packets (flits). A partial-width link state, L0p 430, is available in FLIT mode. In L0p, some lanes of the link can be placed in an electrically idle (EI) state while other lanes remain available for transmitting PCIe traffic. By idling some lanes during low-throughput data traffic scenarios, the L0p state can reduce power consumption while active data communications can continue over the link.

[0063] In L0p, the link can have a partial width. In some cases, each direction of the link can have a different width. Therefore, flits can be transmitted on the link with different widths. The link can exit to other link states, such as a low-power link state (e.g., L1), based on certain received and transmitted messages or other events. However, in FLIT mode, according to the current PCIe specification, the standby state L0s is not available. In this case, the transition from L0p to the power-saving state (e.g., L1) must wait until the L1 state inactivity timer is triggered.

[0064] L0ps status in FLIT mode

[0065] In some aspects, a new partial width standby link state 432 (L0ps) is available when the link is in FLIT mode. When the link is in FLIT mode, L0ps provides a low-power standby state with lower latency than L1. In L0p, when any RX line or TX line of the link becomes idle or inactive (e.g., within a short time interval or predetermined threshold between data bursts), the idle RX / TX line can change from L0p to a new standby state, L0ps, which is only accessible from L0p when FLIT mode is enabled. The RX line and TX line of a channel can enter the L0ps state independently. In some aspects, a separate L0p / L0ps transition state diagram exists for each line, so that the RX line and TX line of a channel can switch independently between the L0p state and the L0ps state.

[0066] In some aspects, L0ps is a low-power state that allows a PCIe link to quickly enter and recover from it without undergoing recovery. The lanes of a PCIe link can independently enter and exit the L0ps state. For example, when the RX / TX lanes are in L0ps, the lane transmitters (e.g., drivers 320-1 to 320-n) and lane receivers (e.g., receivers 340-1 to 340-n) can stop or reduce their clock rates (e.g., using dynamic clock gating) to reduce power consumption. In some aspects, a PCIe device controls its transmitter to enter L0ps and send an ordered set (e.g., EIOS), and the receiver enters L0ps after receiving the ordered set from the transmitter. In some aspects, a device (e.g., an endpoint) can transition a data link from L0p to L0ps without obtaining permission from the PCIe host. In contrast, a transition to L1 involves a device (e.g., an endpoint) first requesting permission from an upstream device (e.g., a host) to enter the deeper, power-saving L1 state. After confirmation, both devices may turn off their transmitters and enter electrical idle in the L1 state.

[0067] Figure 5 is a diagram of a PCIe link between a host 502 and an endpoint 504 according to some aspects. Link 506 includes multiple duplex traffic channels that may have Figure 3 The physical structure described is the same physical structure, but is generalized to show four lanes in an exemplary x4 configuration. For example, link 506 may include four lanes 511, 512, 513, 514, but more or fewer lanes may be used. Each lane includes two transmit (TX) lanes (TX signal lanes) as a differential lane pair and two receive (RX) lanes (RX signal lanes) as a differential lane pair. In this example, link 506 has four lanes per lane. Figure 5 In the example, the TX line carries traffic in the direction from the host 502 to the endpoint 504, and the RX line carries traffic in the direction from the endpoint 504 to the host 502. Figure 3 In the example, Figure 2 The PHY TX block 222 shown in FIG may be implemented to include a transmit driver for each differential pair of TX lines, and Figure 2 The PHY RX block 270 shown in FIG. 2 may be implemented to include a receiver for each differential pair of RX lines.

[0068] In L0p, the width of link 506 can be changed by controlling the number of active and / or idle lanes 511, 512, 513, and 514 without disrupting data flow (i.e., always keeping at least one lane active when changing the link width). The host 502 or endpoint 504 can change the link width by configuring the number of traffic lanes powered to send and receive data over the link. L0p is a partial-width state in which some lanes (e.g., lanes 511 and 512) can be active and some lanes (e.g., lanes 513 and 514) can be electrically idle (EI). In L0p, active lanes 511 and 512 can be used to send and / or receive traffic, and the EI lanes are unused (idle). If both TX traffic and / or RX traffic are low, or if there is no traffic activity on the active lanes, one or more RX lanes and / or TX lanes can be placed in the L0ps state (standby state) to reduce link power consumption.

[0069] Figure 6 is a diagram illustrating an example of a PCIe link operating in FLIT mode between a host 602 and an endpoint 604 according to some aspects. The host 602 and the endpoint 604 may communicate with Figure 5 The host 502 and endpoint 504 in are the same. Link 606 may have a similar Figure 5 6 . In one example, lanes 611 and 612 can be active lanes in the L0p state, and lanes 613 and 614 can be electrically idle (EI). Due to low or no RX activity, RX lanes 620 of lane 611 and RX lanes 622 of lane 612 can change to the L0ps state to reduce link power consumption. In this case, TX lanes 624 of lane 611 and TX lanes 626 of lane 612 remain in the L0p state for PCIe traffic. In some aspects, a transmitter (e.g., host 602 / endpoint 604) can use an inactivity timer or threshold to determine the timing of changing the lanes to the L0ps state due to low traffic or inactivity. In one aspect, when an inactivity timer expires or is triggered, a transmitter (e.g., endpoint 604) may send an EIOS on the corresponding channel and enter L0ps, and a receiver (e.g., host 602) may enter L0ps after receiving the EIOS on the corresponding channel. The RX and TX lanes of the same channel may independently enter L0ps based on different inactivity timers (e.g., RX inactivity timer and TX inactivity timer). Similarly, the RX and TX lanes may independently return to the L0p state.

[0070] Figure 7is a diagram illustrating another example of a PCIe link operating in FLIT mode between a host 702 and an endpoint 704 according to some aspects. The host 702 and the endpoint 704 may communicate with Figure 5 and Figure 6 The host and endpoint in the same. Link 706 between host 702 and endpoint 704 may have four lanes 711, 712, 713, and 714. In one example, lanes 711 and 712 may be active lanes in the L0p state, and lanes 713 and 714 may be electrically idle (EI). Due to low TX traffic or no TX traffic, TX line 720 of lane 711 and TX line 722 of lane 712 may be placed in the L0ps state to reduce link power consumption. RX line 724 of lane 711 and RX line 726 of lane 712 remain in the L0p state for PCIe traffic. In some aspects, a transmitter (e.g., host 702 / endpoint 704) may use an inactivity timer or threshold to determine the timing for changing a lane to the L0ps state due to low traffic or inactivity. When the inactivity timer expires or is triggered, the transmitter (e.g., host 702 or endpoint 704) may send an EIOS on the corresponding channel and enter L0ps, and the receiver (e.g., endpoint 704) may enter L0ps after receiving the EIOS on the corresponding channel. The RX and TX lanes of the same channel may enter L0ps independently. The RX and TX lanes of the same channel may enter L0ps independently based on different inactivity timers (e.g., RX inactivity timer and TX inactivity timer).

[0071] The new L0ps state described above can also provide power savings for PCIe links in FLIT mode (e.g., in the L0p state) even when individual TX or RX lines are idle. Because a line can enter L0ps without a handshake between the host and the endpoint (unlike the transition to L1), the overhead of a handshake sequence can be avoided. Therefore, transitions between the L0ps state and the L0p state can provide significant overhead power savings over a period of time. Furthermore, the L0ps state power savings can scale with increasing link width. Compared to L1, the L0ps state can provide lower latency (e.g., exit latency) because the transition from L0ps to L0p does not require a recovery state involving the exchange of training sequences to negotiate various link parameters (including, for example, lane polarity, link / lane number, equalization parameters, data rate, etc.).

[0072] PCIe Capability Structure

[0073] In some aspects, configuration and control of PCI devices (eg, hosts and endpoints) may be performed using a set of registers referred to as a configuration space in the PCIe specification. PCIe devices may have an extended configuration space that provides additional registers. Figure 8 is a diagram illustrating an exemplary PCIe configuration space (eg, a device 3 extended capabilities structure) according to some aspects. The device 3 extended capabilities structure 800 may be configured to support the above-described Figures 4 to 7 The device 3 extended capability structure 800 may include a PCIe extended capability header 802 , a device capability 3 register 804 , a device control 3 register 806 , and a device status 3 register 808 .

[0074] Figure 9 3 is a diagram illustrating a device capability 3 register 900 and a device control 3 register 901 according to some aspects. The device capability 3 register 900 and the device control 3 register 901 may be associated with Figure 8 The device capability 3 register included in the device 3 extended capability structure 800 is the same as the device control 3 register. The device capability 3 register may have 32 bits, some of which bits 902 (e.g., bits 0-9) are configured for various functions according to the current PCIe specification. For example, bit 3 may indicate whether L0p is supported by the receiver, bits 4-6 may indicate the port L0p exit latency, and bits 7-9 may indicate the retimer L0p exit latency. The device capability 3 register also has reserved bits 904 (e.g., bits 10-31) that may be used to implement new features, such as those described above with respect to Figures 4 to 7 In one example, bit 10 may be used to indicate whether L0ps is supported, and bits 11-12 may provide a L0ps exit delay value. The L0ps exit delay specifies the delay to exit the L0ps state (eg, return to L0p).

[0075] The device control 3 register 901 may have 32 bits, some of which bits 906 (e.g., bits 0-6) are configured for various functions in the current PCIe specification. For example, bit 3 may indicate whether the L0p state is enabled. The device control 3 register also has reserved bits 908 (e.g., bits 7-31) that may be used to implement new features, such as those described above with respect to Figures 4 to 7 In one example, bit 7 may be used to indicate whether L0ps is enabled (eg, 1 for enabled, 0 for disabled).

[0076] Figure 10 is a block diagram of a link interface processing circuit. Processing circuit 1004 is a device that can be part of a host or an endpoint. The processing circuit is coupled to a Figures 5 to 7The link 1002 is similar to the duplex channels described, for example, a PCIe link. The link 1002 can be coupled to another PCIe device (e.g., an endpoint or a host) at the opposite end. The data and control information conveyed as packets by the link 1002 are coupled to a link interface 1020 (e.g., a PCIe interface), which provides a PHY-level interface to the link 1002 and converts the baseband signals into packets. The data and control packets are transmitted to other components of the processing circuit 1004 via the link interface 1020 through the bus 1010. The link interface 1020 has a direct connection to the interface configuration circuit 1018 for configuration and control settings for the operation of the link 1002.

[0077] Processing circuitry 1004 also includes memory 1021, which can be used to store data and information used by the processor during various operations. Processing circuitry 1004 also includes timer circuitry 1012 coupled to bus 1010. Timer circuitry 1012 can be configured for various timing-related functions, such as timing for latency, inactivity, acknowledgments, and transitions between PCIe states (e.g., L0, L0p, L0ps, L1, L2, and L3). Timer circuitry 1012 can access computer-readable storage medium 1008 to access code 1032 for managing timers. In some aspects, the storage medium is a non-transitory computer-readable medium. Timer circuitry 1012 can also access registers stored in storage medium 1008 (and / or memory 1021) that contain receive (RX) traffic timing thresholds 1034 and transmit (TX) traffic timing thresholds 1036 that can be used to determine the timing of transitions between PCIe link states.

[0078] The processing circuit 1004 may also include a power management circuit 1014 that manages power to each line / lane of the link 1002 and to other components of the processing circuit 1004. The power management circuit 1014 accesses code 1040 for managing PCIe power, as well as transmit line status registers 1042 and receive line status registers 1044 via the bus 1010. These registers may be used to store the status of each transmit line and each receive line, or the transmit side of the link and the receive side of the link. This status may be determined using code 1032 for managing timers, code 1040 for managing PCIe power, or in another manner.

[0079] Processing circuitry 1004 may also include link traffic monitoring circuitry 1016 that monitors transmit and receive traffic activity on link 1002. For example, link traffic monitoring circuitry 1016 may monitor traffic activity and inactivity to determine the current link state and transitions between link states. Link traffic monitoring circuitry 1016 may access code 1050 in storage medium 1008 for monitoring link traffic and may also access registers to store results and obtain traffic activity thresholds. Transmit traffic activity threshold 1052 and receive traffic activity threshold 1054 may be used to monitor transmit traffic activity and receive traffic activity, respectively.

[0080] The power management circuit 1014 can manage the power of the transmit line and the power of the receive line based on the transmit traffic activity and the receive traffic activity. The interface configuration circuit 1018 can modify the configuration in response to the power management circuit 1014. For example, the interface configuration circuit 1018 can change the link state (e.g., L0, L0p, L0ps) of the link 1002.

[0081] Interface configuration circuitry 1018, like link traffic monitoring circuitry 1016, power management circuitry 1014, and timer circuitry 1012, is coupled to bus 1010, enabling each of these blocks to communicate with one another, with storage medium 1008, and with processor 1006. Processor 1006 can control the operation of the other components and appropriately initiate instances of each component or its functionality based on the operation of processing circuitry 1004. Interface configuration circuitry 1018 also has access to code 1060 for configuring the PCIe interface. When executing this code, interface configuration circuitry 1018 can read and write values ​​from various configuration registers. For example, these registers include TX control, status, and capability register 1062 and RX control, status, and capability register 1064. These registers can be accessed and read at the beginning of link initialization and then updated with the results of initialization. Registers can also be modified in response to power management and bandwidth negotiation, or changes in the state of one or more transmit or receive lines of link 1002.

[0082] In some aspects, the interface configuration circuitry 1018 and the link interface 1020 may configure the link 1002 to use both L0p and L0ps link states in FLIT mode, as described above. In some aspects, the link 1002 in the L0p state may quickly enter the L0ps state (e.g., based on an L0ps inactivity timer maintained by, for example, timer 1012), and recover from the L0ps state without going through a recovery state. In some aspects, the interface configuration circuitry 1018 and the link interface 1020 may send or receive an Electrically Idle Ordered Set (EIOS) via the link 1002. The EIOS may cause the link to enter L0ps. In some aspects, when the lanes of the link 1002 switch from L0ps back to L0p, the interface configuration circuitry 1018 and the link interface 1020 may reestablish bit lock, symbol lock, or block alignment, and perform lane-to-lane de-skew. In one example, the interface configuration circuitry 1018 and the link interface 1020 can de-skew the lanes by transmitting exit patterns on the idle lanes to train and de-skew them. As an example, the exit patterns can include an Electrical Idle Exit Ordered Set (EIEOS) and a Fast Training Sequence (FTS).

[0083] Processing circuitry 1004 can initialize link 1002, manage power, link status, and change the number of active lines in link 1002. During operation, bandwidth requests may also be received from a host or endpoint. These bandwidth requests may trigger bandwidth negotiation, which may subsequently change the values ​​set in control, status, and capability registers. The number of active lines may then be changed in response to transmit and receive traffic activity. Link traffic monitoring circuitry 1016 may also monitor TX traffic activity on the transmit lines of link 1002 and RX traffic activity on the receive lines of link 1002. TX and RX traffic activity are evaluated to determine changes in the number of active lines. Power management circuitry 1014 may change the link status of one or more TX or RX lines. The status changes may then be recorded in TX line status register 1042 and RX line status register 1044. This evaluation may be performed in various ways. In some examples, link traffic monitoring circuitry 1016 compares TX traffic activity to one or more thresholds in TX traffic threshold register 1052, and RX traffic activity to one or more thresholds in RX traffic threshold register 1054. The message may then be transmitted over link 1002 to a connected device (eg, a host or endpoint).

[0084] After changing the number of active lines or the link state, the power management circuit 1014 may change the voltage level of one or more of the voltages 276, 278, and 280 by commanding the PMIC 290 to set the voltage level of one or more of the voltages supplied by the PMIC 290, such as Figure 2 As shown. The power management circuit 1014 can also connect or disconnect power to the drivers and receivers of the affected lines based on the new number of active lines. For example, if the number of active lines decreases, the power management circuit 1014 can power down the drivers in the PHY TX block 222 and / or the receivers in the PHY RX block 226 corresponding to the lines in link 1002 that are deactivated due to the change. The power management circuit 1014 can power down selected drivers and / or receivers by transmitting instructions to the power switch circuit to shut down the selected drivers and / or receivers. Thus, power is managed according to the negotiated bandwidth by providing one or more voltages to the interface circuitry of the link and by setting the levels of the one or more voltages.

[0085] Figure 11 A flow chart illustrating a method 1100 for link state management of a link (e.g., a PCIe link) according to aspects of the present disclosure is provided. In certain aspects, the method 1100 provides techniques for link state management of a PCIe link operating in FLIT mode. As described herein, the link may be a PCIe link, however, the method may be adapted to accommodate other data links having transmit and receive lines operating in various link states. In some aspects, the method 1100 may be implemented at either the described host or endpoint.

[0086] At 1102, the method includes a process of operating a PCIe link (data link with a host or endpoint) in a flow control unit (FLIT) mode using a fixed-size packet. The PCIe link is in a partial width link state (PLS), in which a first group of channels of the PCIe link are in an electrical idle state, and a second group of channels of the PCIe link are in an active state that can be used to carry PCIe services to / from a host or endpoint. In some aspects, PLS may correspond to the L0p state of the PCIe link in the FLIT mode, as described herein. For example, the PCIe link may be identical to link 506, 606, or 706, wherein some channels (e.g., channels 513, 514, 613, 614, 713, or 714) are in an electrical idle state such as in the L0p. In one aspect, interface configuration circuitry 1018 and PCIe interface 1020 may provide means for operating a PCIe link in the PLS using the FLIT mode.

[0087] At 1104, the method includes a process of transitioning one or more lanes (e.g., Rx lanes and / or Tx lanes) of the second set of lanes from the PLS to a partial width standby link state (PSLS), wherein the one or more lanes of the second set of lanes are in a standby state having lower power consumption than an active state. In some aspects, the PSLS corresponds to the L0ps state of the PCIe link in FLIT mode. As described above in Figures 4 to 9 As described in , when the link is in FLIT mode, L0ps provides a lower power standby state from L0p, in which FLIT mode, the other standby state L0s is not available. In one aspect, the interface configuration circuit 1018 and the PCIe interface 1020 may provide a means for transitioning a PCIe link between PLS and PSLS.

[0088] The following provides an overview of various embodiments of the disclosure.

[0089] Embodiment 1: A method of operating an endpoint for data communication, the method comprising: operating a data link with a host in a flow control unit (FLIT) mode using fixed-size packets, the data link being in a partial width link state (PLS), in which a first group of channels of the data link is in an electrically idle state, and a second group of channels of the data link is in an active state capable of being used for data services with the host; and transitioning one or more lines of the second group of channels from the PLS to a partial width standby link state (PSLS), in which the one or more lines of the second group of channels are in a standby state, the standby state having lower power consumption than the active state.

[0090] Embodiment 2: The method of embodiment 1, further comprising transitioning the one or more lanes of the second group of channels from the PSLS back to the PLS without undergoing a recovery state.

[0091] Embodiment 3: The method according to embodiment 1 further comprising: receiving an electrically idle ordered set from the host via the data link; and transitioning the one or more lanes of the second group of channels from the PLS to the PSLS in response to the electrically idle ordered set.

[0092] Embodiment 4: The method of embodiment 1 further comprising transitioning the one or more lanes of the second set of lanes from the PSLS to the PLS, the transition comprising at least one of: establishing at least one of bit lock, symbol lock, or block alignment of the data link with the host; or de-skewing the lane-to-lane skew of the data link.

[0093] Embodiment 5: A method according to embodiment 1, 2, 3 or 4, wherein the one or more lines of the second group of channels include a transmit signal line for transmitting signals to the host and a receive signal line for receiving signals from the host, and wherein converting the one or more lines of the second group of channels from the PLS to the PSLS includes at least one of the following: converting the transmit signal line to the PSLS independently of the receive signal line; or converting the receive signal line to the PSLS independently of the transmit signal line.

[0094] Example 6: A method according to Example 5, wherein transitioning the one or more lines of the second group of channels from the PLS to the PSLS includes: transitioning the transmit signal line to the PSLS based on a first inactivity timer; and transitioning the receive signal line to the PSLS based on a second inactivity timer independent of the first inactivity timer.

[0095] Embodiment 7: The method of embodiment 5, further comprising: transitioning the one or more lanes of the second group of channels from the PLS to the PSLS without obtaining permission from the host.

[0096] Embodiment 8: The method of embodiment 1, 2, 3, or 4, wherein the PLS includes a Peripheral Component Interconnect Express (PCIe) L0p state configured to operate the data link in the FLIT mode using fixed-size packets, and the PSLS includes a PCIe L0ps state configured to operate the data link in the FLIT mode.

[0097] Embodiment 9: An endpoint for a Peripheral Component Interconnect Express (PCIe) link, the endpoint comprising: an interface circuit configured to provide an interface with the PCIe link connected to a host; and a controller configured to: operate the PCIe link in a flow control unit (FLIT) mode using fixed-size packets, the PCIe link being in a partial-width link state (PLS), in which a first group of channels of the PCIe link are in an electrically idle state and a second group of channels of the PCIe link are in an active state capable of being used for data services with the host; and transition one or more lines of the second group of channels from the PLS to a partial-width standby link state (PSLS), in which the one or more channels of the second group of channels are in a standby state, the standby state having lower power consumption than the active state.

[0098] Embodiment 10: The endpoint of Embodiment 9, wherein the controller is further configured to transition the one or more lanes of the second group of lanes from the PSLS back to the PLS without undergoing a recovery state.

[0099] Embodiment 11: The endpoint of embodiment 9, wherein the controller is further configured to: receive an electrically idle ordered set from the host via the PCIe link; and in response to the electrically idle ordered set, transition the one or more lanes of the second group of channels from the PLS to the PSLS.

[0100] Embodiment 12: The endpoint of embodiment 9, wherein, to transition the one or more lanes of the second set of lanes from the PSLS to the PLS, the controller is further configured to: establish at least one of bit lock, symbol lock, or block alignment of the PCIe link with the host; or de-skew lane-to-lane skew of the PCIe link.

[0101] Embodiment 13: An endpoint according to embodiment 9, 10, 11 or 12, wherein the one or more lines of the second group of channels include a transmit signal line for transmitting signals to the host and a receive signal line for receiving signals from the host, and wherein, in order to transition the one or more lines of the second group of channel PCIe links from the PLS to the PSLS, the controller is further configured to perform at least one of the following: transition the transmit signal line to the PSLS independently of the receive signal line; or transition the receive signal line to the PSLS independently of the transmit signal line.

[0102] Embodiment 14: An endpoint according to embodiment 13, wherein, in order to transition the one or more lines of the second group of channels from the PLS to the PSLS, the controller is further configured to: transition the transmit signal line to the PSLS based on a first inactivity timer; and transition the receive signal line to the PSLS based on a second inactivity timer independent of the first inactivity timer.

[0103] Embodiment 15: The endpoint of Embodiment 13, wherein the controller is further configured to transition the one or more lanes of the second set of channels from the PLS to the PSLS without obtaining permission from the host.

[0104] Embodiment 16: An endpoint according to embodiment 9, 10, 11 or 12, wherein the PLS includes a PCIe L0p state configured to operate the PCIe link in the FLIT mode using fixed-size packets, and the PSLS includes a PCIe L0ps state configured to operate the PCIe link in the FLIT mode.

[0105] Embodiment 17: A host for a Peripheral Component Interconnect Express (PCIe) link, the host comprising: an interface circuit configured to provide an interface with the PCIe link connected to an endpoint; and a controller configured to: operate the PCIe link in a flow control unit (FLIT) mode using fixed-size packets, the PCIe link being in a partial-width link state (PLS), in which a first group of channels of the PCIe link are in an electrically idle state and a second group of channels of the PCIe link are in an active state capable of being used for data services with the endpoint; and transition one or more lines of the second group of channels from the PLS to a partial-width standby link state (PSLS), in which the one or more channels of the second group of channels are in a standby state, the standby state having lower power consumption than the active state.

[0106] Embodiment 18: The host of embodiment 17, wherein, in order to transition the one or more lanes of the second set of lanes from the PSLS to the PLS, the controller is further configured to perform at least one of: establishing at least one of bit lock, symbol lock, or block alignment of the PCIe link with the endpoint; or de-skew lane-to-lane skew of the PCIe link.

[0107] Embodiment 19: A host according to embodiment 17 or 18, wherein the one or more lines of the second group of channels include a transmit signal line for transmitting signals to the endpoint and a receive signal line for receiving signals from the endpoint, and wherein, in order to convert the one or more lines of the second group of channels from the PLS to the PSLS, the controller is further configured to perform at least one of the following: converting the transmit signal line to the PSLS independently of the receive signal line; or converting the receive signal line to the PSLS independently of the transmit signal line.

[0108] Example 20: A host according to Example 19, wherein, in order to transition the one or more lines of the second group of channels from the PLS to the PSLS, the controller is further configured to: transition the transmit signal line to the PSLS based on a first inactivity timer; and transition the receive signal line to the PSLS based on a second inactivity timer independent of the first inactivity timer.

[0109] It should be understood that the present disclosure is not limited to the exemplary terms used above to describe aspects of the present disclosure. For example, bandwidth may also be referred to as throughput, data rate, or another term.

[0110] Although aspects of the present disclosure are discussed above using the example of the PCIe standard, it should be understood that the present disclosure is not limited to this example and may be used with other standards.

[0111] Each of the host client 214, host controller 212, device controller 252, and device client 254 discussed above may be implemented using a controller or processor configured to perform the functions described herein by executing software including code for performing the functions described herein. The software may be stored on a non-transitory computer-readable storage medium (e.g., RAM, ROM, EEPROM, optical and / or magnetic disks, shown as host system memory 240, endpoint system memory 274, or another memory).

[0112] Any reference to an element herein using designations such as "first," "second," etc. does not generally limit the number or order of those elements. Rather, these designations are used herein as a convenient method of distinguishing between two or more elements or instances of an element. Thus, a reference to a first element and a second element does not mean that only two elements can be used or that the first element must precede the second element.

[0113] Throughout this disclosure, the word "exemplary" is used to mean "serving as an example, instance, or illustration." Any specific implementation or aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects of the disclosure. Likewise, the term "aspect" does not require that all aspects of the disclosure include the discussed feature, advantage, or mode of operation. The term "coupled" is used herein to refer to a direct or indirect electrical or other communicative coupling between two structures. Additionally, the term "approximately" means within ten percent of a stated value.

[0114] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of operating an endpoint for data communication, the method comprising: operating a data link with a host in a flow control unit FLIT mode using fixed-size packets, the data link being in a partial width link state PLS, wherein a first group of lanes of the data link is in an electrically idle state and a second group of lanes of the data link is in an active state capable of carrying Peripheral Component Interconnect Express (PCIe) traffic with the host; using an inactivity timer to detect low traffic or inactivity on one or more lines of the second set of channels; as well as In response to detecting the low traffic or inactivity, transitioning the one or more lanes of the second group of lanes from the PLS to a partial width standby link state (PSLS) without handshaking between the host and the endpoint, wherein the one or more lanes of the second group of lanes are in a standby state having lower power consumption than the active state.

2. The method according to claim 1, further comprising: The one or more lanes of the second group of channels are transitioned from the PSLS back to the PLS without going through a recovery state.

3. The method according to claim 1, further comprising: receiving an electrically idle ordered set from the host via the data link; as well as Responsive to the electrically idle ordered set, the one or more lanes of the second group of channels are transitioned from the PLS to the PSLS.

4. The method of claim 1 , further comprising transitioning the one or more lanes of the second set of lanes from the PSLS to the PLS, the transition comprising at least one of: establishing at least one of bit lock, symbol lock, or block alignment of the data link with the host; or A lane-to-lane skew of the data link is de-skewed.

5. The method of claim 1 , wherein the one or more lines of the second set of channels include a transmit signal line for transmitting signals to the host and a receive signal line for receiving signals from the host, and wherein transitioning the one or more lanes of the second set of channels from the PLS to the PSLS comprises at least one of: switching the transmit signal line to the PSLS independently of the receive signal line; or The receive signal line is switched to the PSLS independently of the transmit signal line.

6. The method of claim 5 , wherein transitioning the one or more lanes of the second set of lanes from the PLS to the PSLS comprises: transitioning the transmit signaling line to the PSLS based on a first inactivity timer; as well as The receive signal line is transitioned to the PSLS based on a second inactivity timer that is independent of the first inactivity timer.

7. The method according to claim 5, further comprising: The one or more lanes of the second set of channels are transitioned from the PLS to the PSLS without obtaining permission from the host.

8. The method of claim 1 , wherein the PLS comprises a Peripheral Component Interconnect Express (PCIe) LOp state configured to operate the data link in the FLIT mode using fixed-size packets, and the PSLS comprises a PCIe LOps state configured to operate the data link in the FLIT mode.

9. An endpoint for a Peripheral Component Interconnect Express (PCIe) link, the endpoint comprising: an interface circuit configured to provide an interface with the PCIe link connected to a host; and A controller configured to: operating the PCIe link in a flow control unit FLIT mode using fixed-size packets, the PCIe link being in a partial width link state PLS, wherein a first group of lanes of the PCIe link is in an electrically idle state and a second group of lanes of the PCIe link is in an active state capable of carrying Peripheral Component Interconnect Express (PCIe) traffic with the host; as well as using an inactivity timer to detect low traffic or inactivity on one or more lines of the second set of channels; as well as responsive to detecting the low traffic or inactivity, transitioning the one or more lanes of the second set of lanes from the PLS to a partial width standby link state (PSLS) without handshaking between the host and the endpoint, in which the one or more lanes of the second set of lanes are in a standby state, The standby state has lower power consumption than the active state.

10. The endpoint of claim 9, wherein the controller is further configured to: The one or more lanes of the second group of channels are transitioned from the PSLS back to the PLS without going through a recovery state.

11. The endpoint of claim 9, wherein the controller is further configured to: receiving an electrical idle ordered set from the host via the PCIe link; and Responsive to the electrically idle ordered set, the one or more lanes of the second group of channels are transitioned from the PLS to the PSLS.

12. The endpoint of claim 9, wherein: To transition the one or more lanes of the second set of channels from the PSLS to the PLS, the controller is further configured to: establishing at least one of bit locking, symbol locking, or block alignment of the PCIe link with the host; or A lane-to-lane skew of the PCIe link is de-skewed.

13. The endpoint of claim 9 , wherein the one or more lines of the second set of channels include a transmit signal line for transmitting signals to the host and a receive signal line for receiving signals from the host, and in, To transition the one or more lanes of the second set of channels from the PLS to the PSLS, the controller is further configured to do at least one of the following: switching the transmit signal line to the PSLS independently of the receive signal line; or The receive signal line is switched to the PSLS independently of the transmit signal line.

14. The endpoint of claim 13, wherein: To transition the one or more lanes of the second set of channels from the PLS to the PSLS, the controller is further configured to: transitioning the transmit signaling line to the PSLS based on a first inactivity timer; as well as The receive signal line is transitioned to the PSLS based on a second inactivity timer that is independent of the first inactivity timer.

15. The endpoint of claim 13, wherein the controller is further configured to: The one or more lanes of the second set of channels are transitioned from the PLS to the PSLS without obtaining permission from the host.

16. The endpoint of claim 9, wherein the PLS includes a PCIe LOp state configured to operate the PCIe link in the FLIT mode using fixed-size packets, and the PSLS includes a PCIe LOps state configured to operate the PCIe link in the FLIT mode.

17. A host for a Peripheral Component Interconnect Express (PCIe) link, the host comprising: an interface circuit configured to provide an interface with the PCIe link connected to an endpoint; and A controller configured to: operating the PCIe link in a flow control unit FLIT mode using fixed-size packets, the PCIe link being in a partial width link state PLS, wherein a first group of lanes of the PCIe link is in an electrically idle state and a second group of lanes of the PCIe link is in an active state capable of carrying Peripheral Component Interconnect Express (PCIe) traffic with the endpoint; using an inactivity timer to detect low traffic or inactivity on one or more lines of the second set of channels; as well as The one or more lanes of the second group of lanes are transitioned from the PLS to a partial width standby link state (PSLS) without handshaking between the host and the endpoint, wherein the one or more lanes of the second group of lanes are in a standby state having lower power consumption than the active state.

18. The host according to claim 17, wherein: To transition the one or more lanes of the second set of channels from the PSLS to the PLS, the controller is further configured to do at least one of the following: establishing at least one of bit lock, symbol lock, or block alignment of the PCIe link with the endpoint; or A lane-to-lane skew of the PCIe link is de-skewed.

19. The host of claim 17, wherein the one or more lines of the second set of channels include a transmit signal line for transmitting signals to the endpoint and a receive signal line for receiving signals from the endpoint, and in, To transition the one or more lanes of the second set of channels from the PLS to the PSLS, the controller is further configured to do at least one of the following: switching the transmit signal line to the PSLS independently of the receive signal line; or The receive signal line is switched to the PSLS independently of the transmit signal line.

20. The host according to claim 19, wherein To transition the one or more lanes of the second set of channels from the PLS to the PSLS, the controller is further configured to: transitioning the transmit signaling line to the PSLS based on a first inactivity timer; and The receive signal line is transitioned to the PSLS based on a second inactivity timer that is independent of the first inactivity timer.

Citation Information

Patent Citations

  • Partial link width states for bidirectional multilane links

    US20200226084A1

  • Dynamic network controller power management

    US20210041929A1