Tool box for transferring data between root complex and endpoint

By bridging hardware components with different bandwidths through the toolbox, the problem that switches and retimers in PCIe topologies cannot simultaneously meet the requirements of high bandwidth and low latency is solved, achieving efficient and low-latency data transmission and reducing system costs.

CN121501720APending Publication Date: 2026-02-10MARVELL ASIA PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511112300.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-08-07
Filing Date
2025-08-08
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing PCIe topologies, switches and retimers cannot simultaneously meet the requirements of high bandwidth and low latency when bridging hardware components with different bandwidths, resulting in performance loss and increased costs.

Method used

The toolbox is used, which includes independent physical and data link layers to form independent links, bridge hardware components with different bandwidths, support the latest version of the PCIe standard, and manage link bandwidth and error conditions through a control module, avoiding transaction layer buffers and reducing latency.

Benefits of technology

It enables efficient data transmission between hardware components with different bandwidths, reduces latency and cost, provides an alternative to switches and retimers, and is suitable for PCIe and CXL communication standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121501720A_ABST
    Figure CN121501720A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a toolbox for communicating data between a root complex and an endpoint. An example kit for connecting between a root complex and an endpoint in a computing device includes a first port configured to connect to the root complex, a second port configured to connect to the endpoint, a first physical layer connected to the first port, and a second physical layer connected to the second port, and a first data link layer and a second data link layer, the first data link layer being connected between the second data link layer and the first physical layer, and the second data link layer being connected between the first data link layer and the second physical layer. The first physical layer, the first data link layer, the second physical layer, and the second data link layer are configured to form one or more channels for communicating data between the root complex and the endpoint.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefits of U.S. Provisional Application No. 63 / 681,031, filed August 8, 2024, and U.S. Non-Provisional Application No. 19 / 293,529, filed August 7, 2025. The entire disclosure of the above-cited applications is incorporated herein by reference. Technical Field

[0003] This disclosure relates to a toolbox for transferring data between root complexes and endpoints in a computing system. Background Technology

[0004] The background description provided herein is intended to generally present the context of this disclosure. The work of the currently named inventors, to the extent that it is described in this background section, and in aspects that may not otherwise be described as prior art at the time of filing, is neither expressly nor impliedly acknowledged as prior art to this disclosure.

[0005] Computing systems typically transfer data between hardware components. Using this data communication, computing systems can adhere to specific communication standards, such as the High-Speed ​​Peripheral Component Interconnect (PCIe) used to establish point-to-point connections between different hardware components. In a PCIe topology, data can be transferred via a separate serial link between a root complex (e.g., the host) and one or more endpoints. The root complex is the device that connects the central processing unit (CPU) and memory in the computing system to one or more endpoints. The root complex controls other PCIe components in the hierarchy. Endpoints are peripheral devices in the computing system that provide specific functions, such as non-volatile memory (NVM) fast solid-state drives (SSDs), network interface controllers (NICs), graphics processing units (GPUs), and additional memory cards.

[0006] A root complex can be directly connected to endpoints or connected to one or more endpoints via another PCIe component. Specifically, a switch can be used to connect the root complex to multiple endpoints via a single PCIe link on the root complex side and multiple PCIe links on the endpoint side. With this configuration, each endpoint is associated with its own PCIe link, and traffic flows between the PCIe links on the root complex side and the endpoint side via the switch. Alternatively, a retimer can be used to connect the root complex to one or more endpoints via one-to-one link connectivity. Specifically, the retimer implements an independent PCIe link for each endpoint. Thus, if two endpoints are used, the retimer connects the root complex to the first endpoint via a single PCIe link on the root complex side and a single PCIe link on the endpoint side, and connects the root complex to the second endpoint via another single PCIe link on the root complex side and another single PCIe link on the endpoint side. With this configuration, traffic between the root complex side and each endpoint flows through the retimer and each independent PCIe link, without crossing between independent PCIe links. Summary of the Invention

[0007] An example toolkit for connecting a root complex and an endpoint in a computing device includes: a first port configured to connect to the root complex; a second port configured to connect to the endpoint; a first physical layer connected to the first port and a second physical layer connected to the second port; a first data link layer and a second data link layer, the first data link layer being connected between the second data link layer and the first physical layer, and the second data link layer being connected between the first data link layer and the second physical layer. The first physical layer, the first data link layer, the second physical layer, and the second data link layer are configured to form one or more channels for transmitting data between the root complex and the endpoint.

[0008] In some examples, the first physical layer and the first data link layer form a first link, and the second physical layer and the second data link layer form a second link independent of the first link.

[0009] In some examples, the first physical layer includes a physical coding sublayer (PCS) connected to a first port and a media access control (MAC) sublayer connected to a first data link layer, and the second physical layer includes a PCS connected to a second port and a MAC sublayer connected to a second data link layer.

[0010] In some examples, the PCS of the first physical layer is configured to provide a data interface between the first port and the MAC sublayer, and the PCS of the second physical layer is configured to provide a data interface between the second port and the MAC sublayer of the second physical layer.

[0011] In some examples, the MAC sublayer of the first physical layer includes a data path module between the PCS of the first physical layer and the first data link layer, and a first link training and state machine (LTSSM) module, and the MAC sublayer of the second physical layer includes a data path module between the PCS of the second physical layer and the second data link layer, and a second LTSSM module.

[0012] In some examples, the first data link layer includes a first finite state machine (FSM) module configured to connect to a second link, and the second data link layer includes a second FSM module configured to connect to the first link.

[0013] In some examples, the second FSM module is configured to detect a link down condition on the second link, and in response to the link down condition, the first LTSSM module is configured to create a link down condition on the first link.

[0014] In some examples, the first FSM module is configured to detect a link outage on the first link, and in response to the link outage, the second LTSSM module is configured to enter a disabled state.

[0015] In some examples, the toolbox includes a control module configured to determine whether the bandwidth on the first link and the bandwidth on the second link are the same, and in response to the bandwidth on the first link and the bandwidth on the second link being the same, allow data link layer packets (DLLPs) to be passed between the root complex and the endpoints via the toolbox.

[0016] In some examples, the control module is configured to prevent DLLP from being passed between the root complex and endpoints via the toolbox in response to a difference in bandwidth on the first link and bandwidth on the second link.

[0017] In some examples, the toolkit includes a control module configured to detect transient error conditions on one of the first or second links and, in response to such transient error conditions, send a negative acknowledgment (NAK) signal on the other of the first or second links.

[0018] In some examples, the toolbox includes a control module that is configured to enable a low-power state for one of the first or second links when no Transaction Layer Packet (TLP) is present in one of the links of the first or second link.

[0019] In some examples, the control module is configured to initiate a low-power state request to a root complex or endpoint that can connect to another link in the first or second link, and in response to the low-power state request being rejected, to switch one of the first or second links back to an active state.

[0020] In some examples, the control module is configured to exit the low-power state in response to a request to the root complex or endpoint.

[0021] In some examples, the bandwidth on the first link is the same as the bandwidth on the second link.

[0022] In some examples, the data rate at the first port is different from the data rate at the second port.

[0023] In some examples, the data rate at the first port is the same as the data rate at the second port.

[0024] In some examples, the toolkit does not include the transaction layer.

[0025] In some examples, the toolbox is configured to pass data conforming to the High Speed ​​Peripheral Component Interconnect (PCIe) standard between the root complex and endpoints.

[0026] An example computing system for transmitting data conforming to the High-Speed ​​Peripheral Component Interconnect (PCIe) standard includes: a root complex configured to connect to a processor and memory, an endpoint, and a toolbox connecting the root complex and the endpoint. The toolbox includes a first port connected to the root complex and a second port connected to the endpoint, and the toolbox is configured to transmit data conforming to the High-Speed ​​Peripheral Component Interconnect (PCIe) standard between the root complex and the endpoint.

[0027] In some examples, the toolbox includes a first physical layer connected to a first port, a second physical layer connected to a second port, a first data link layer and a second data link layer, the first data link layer being connected between the second data link layer and the first physical layer, the second data link layer being connected between the first data link layer and the second physical layer, and the first physical layer, the first data link layer, the second physical layer and the second data link layer being configured to form one or more channels for transmitting data between the root complex and the endpoints.

[0028] In some examples, the first physical layer and the first data link layer form a first link, and the second physical layer and the second data link layer form a second link independent of the first link.

[0029] In some examples, the first physical layer includes a physical coding sublayer (PCS) connected to a first port and a media access control (MAC) sublayer connected to a first data link layer, and the second physical layer includes a PCS connected to a second port and a MAC sublayer connected to a second data link layer. The PCS of the first physical layer is configured to provide a data interface between the first port and the MAC sublayer, and the PCS of the second physical layer is configured to provide a data interface between the second port and the MAC sublayer of the second physical layer.

[0030] In some examples, the MAC sublayer of the first physical layer includes a data path module between the PCS of the first physical layer and the first data link layer, and a first link training and state machine (LTSSM) module; the MAC sublayer of the second physical layer includes a data path module between the PCS of the second physical layer and the second data link layer, and a second LTSSM module; the first data link layer includes a first finite state machine (FSM) module configured to be connected to the second link; and the second data link layer includes a second FSM module configured to be connected to the first link.

[0031] In some examples, the second FSM module is configured to detect a link outage condition on the second link, and in response to the link outage condition, the first LTSSM module is configured to create a link outage condition on the first link.

[0032] In some examples, the first FSM module is configured to detect a link outage on the first link, and in response to the link outage, the second LTSSM module is configured to enter a disabled state.

[0033] In some examples, the toolbox also includes a control module configured to determine whether the bandwidth on the first link and the bandwidth on the second link are the same, and in response to the bandwidth on the first link and the bandwidth on the second link being the same, allow data link layer packets (DLLPs) to be passed between the root complex and the endpoints via the toolbox.

[0034] In some examples, the control module is configured to prevent DLLP from being passed between the root complex and endpoints via the toolbox in response to a difference in bandwidth on the first link and bandwidth on the second link.

[0035] In some examples, the toolbox also includes a control module configured to detect transient error conditions on one of the first or second links and, in response to such transient error conditions, send a negative acknowledgment (NAK) signal on the other of the first or second links.

[0036] In some examples, the toolbox also includes a control module configured to: enable a low-power state for one of the first or second links if no Transaction Layer Packet (TLP) exists in one of the links, initiate a low-power state request for the root complex or endpoint connected to the other link in the first or second link, and, in response to the low-power state request being rejected, transition one of the first or second links back to an active state.

[0037] In some examples, the control module is configured to exit the low-power state in response to a request to the root complex or endpoint.

[0038] In some examples, the bandwidth on the first link is the same as the bandwidth on the second link.

[0039] In some examples, the data rate at the first port is different from the data rate at the second port.

[0040] In some examples, the data rate at the first port is the same as the data rate at the second port.

[0041] In some examples, the toolkit does not include the transaction layer.

[0042] An example computing system for transmitting data includes a root complex configured to connect to a processor and memory, endpoints, and a toolbox connecting the root complex and endpoints. The toolbox includes a first port connected to the root complex and a second port connected to the endpoints, a first physical layer connected to the first port and a second physical layer connected to the second port, and a first data link layer and a second data link layer, the first data link layer connecting the second data link layer and the first physical layer, and the second data link layer connecting the first data link layer and the second physical layer. The first physical layer, the first data link layer, the second physical layer, and the second data link layer are configured to form one or more channels for transmitting data between the root complex and the endpoints.

[0043] In some examples, the data rate at the first port is different from the data rate at the second port.

[0044] In some examples, the data rate at the first port is the same as the data rate at the second port.

[0045] Other applications of this disclosure will become apparent from the description, claims, and drawings. The detailed description and specific embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description

[0046] Figure 1 This is a functional block diagram of an example computing system according to embodiments of the present disclosure, including a root complex, endpoints, and a toolbox connecting the root complex and the endpoints.

[0047] Figure 2 This is a functional block diagram of another example computing system according to an embodiment of the present disclosure, the computing system including a toolbox with two independent links, each of which is implemented by a physical layer and a data link layer.

[0048] Figure 3This is a functional block diagram of an example toolbox including physical layers according to embodiments of the present disclosure, each physical layer having a physical coding sublayer and a media access control sublayer.

[0049] Figures 4A-4B This is a functional block diagram of another example toolbox according to an embodiment of the present disclosure.

[0050] Figure 5 This is a flowchart of a control process for resolving link interruption conditions on downstream link segments in a toolbox, according to embodiments of the present disclosure.

[0051] Figure 6 This is a flowchart of a control process for resolving link interruption conditions on an upstream link segment in a toolkit, according to embodiments of the present disclosure.

[0052] Figure 7 This is a flowchart of a control process for resolving transient error conditions on an upstream link segment in a toolkit, according to embodiments of the present disclosure.

[0053] Figure 8 This is a flowchart of a control process for resolving transient error conditions on a downstream link segment in a toolkit, according to embodiments of the present disclosure.

[0054] Figure 9 This is a flowchart of a control process for resolving bandwidth mismatch situations in a toolbox, according to embodiments of the present disclosure.

[0055] Figure 10 This is a flowchart of a control process for resolving dynamic state changes in a toolbox, according to an embodiment of the present disclosure.

[0056] In the accompanying drawings, reference numerals may be used repeatedly to identify similar and / or identical elements. Detailed Implementation

[0057] In computing systems, data transfer typically occurs between hardware components. The speed of data transfer is a key metric for performance; higher speeds indicate faster data exchange. In some examples, computing systems may follow specific communication standards such as High-Speed ​​Peripheral Component Interconnect (PCIe) to establish point-to-point connections for data transfer between different hardware components. PCIe is a high-speed standard used to connect the central processing unit (CPU) and memory to endpoints such as graphics cards, sound cards, solid-state drives, memory cards, network interface controllers, etc.

[0058] In a PCIe topology, data can be transferred via a separate serial link between a root complex (e.g., a host) and one or more endpoints. The root complex connects the CPU and memory to one or more endpoints. In many cases, switches or retimers are used to connect the root complex to multiple endpoints. Components in a PCIe topology can include different implementation layers (e.g., in the protocol stack) depending on their functionality. Such implementation layers are defined in the PCIe specification, so that the evolution of data rates does not require a complete redesign of PCIe components.

[0059] Specifically, the PCIe implementation layers are defined as the Physical Layer, Data Link Layer, and Transaction Layer. The Physical Layer typically manages low-level telecommunication and data transmission over the links between PCIe components (e.g., encoding, decoding, framing, clock recovery, etc.). The Data Link Layer typically manages flow control, error detection (CRC), link management, etc., to ensure reliable transmission of data packets between PCIe components. The Transaction Layer typically manages the actual data transfer between PCIe components and the routing of data within the stack. For example, the Transaction Layer can break data down into Transaction Layer Packets (TLPs), which are then passed down to the Data Link Layer for processing. Typically, the root complex and endpoints have a single set of Physical Layer, Data Link Layer, and Transaction Layer layers. A switch has multiple sets of Physical Layer, Data Link Layer, and Transaction Layer layers, with one set on the root complex side, two or more sets on the endpoint side (one set for each endpoint device connected to the switch), and the switch interconnects between them. A retimer has at least two sets of Physical Layer layers.

[0060] When hardware components in a computing system move to higher bandwidth capabilities, switches and retimers in a PCIe topology may not be able to adapt to these higher data transfer rates while maintaining low latency and maximum available bandwidth. For example, retimers may limit operation to the highest common data rate connecting the hardware components. As an example, CPUs typically migrate to higher bandwidths as the PCIe specification evolves, while endpoint devices tend to migrate to higher bandwidths more slowly. For instance, if a CPU with two PCIe links moves to a newer version of the PCIe specification that supports higher bandwidth (e.g., 64 GT / s data rate per lane), but endpoint devices, each with a single PCIe link, remain on an older version of the PCIe specification with lower bandwidth (e.g., 32 GT / s data rate per lane), the retimers between the CPU and endpoint devices will operate at the lower bandwidth of the endpoint devices, even if the CPU's PCIe link is capable of operating at higher bandwidths. In this example, a CPU with 8 lanes on each PCIe link has 1024G (2×(2×4×64G)) bandwidth, while two endpoint devices with 4 lanes on each PCIe link have 512G (2×(2×4×32G)) bandwidth. Thus, due to the retimer limitation, the CPU bandwidth is limited to the lower bandwidth of the endpoint devices, failing to fully utilize the CPU's increased performance. However, due to its simple design, the retimer provides low latency for data transfer between the CPU and the endpoint devices.

[0061] A similar situation arises when using switches in a PCIe topology. Specifically, when a CPU migrates to a newer version of the PCIe specification that supports higher bandwidth, while endpoint devices remain with an older version of the PCIe specification that has lower bandwidth, the CPU bandwidth is again limited to the lower bandwidth of the endpoint devices when using a two-port switch. For example, a CPU with one PCIe link with 8 lanes has, for example, 1024G (8×2×64G) bandwidth, while two endpoint devices with 4 lanes each on the PCIe link have, for example, 512G (2×(2×4×32G)) bandwidth. To address this, one solution could be to move to a 5-port switch, where the CPU with one PCIe link with 8 lanes is connected to 4 endpoint devices, each with 4 lanes on each PCIe link. In this case, the CPU bandwidth remains at 1024G, and the endpoint device bandwidth increases to 1024G (4×(2×4×32G)). However, using this approach, the switch experiences latency penalties because the added ports cause the switch to perform more functions at the transaction layer. Therefore, although switches can be used to bridge bandwidth differences between the CPU and endpoint devices, the time it takes for data to propagate between the CPU and endpoint devices increases. This results in a performance penalty in the system.

[0062] Furthermore, the performance impact may increase over time. For example, as the data rate of hardware components (e.g., CPU) increases, additional latency is introduced due to the greater performance impact of switches. For instance, as the data rate increases, the impact of the increased interconnect latency from switches is more significant, and larger buffers (e.g., increased area and increased cost) are needed at endpoints and root complexes to hide the interconnect latency and achieve full bandwidth capability.

[0063] The example computing system disclosed herein utilizes a unique gearbox for connecting the root complex and one or more endpoint devices. Such a gearbox is capable of bridging bandwidth differences between a faster CPU and slower endpoint devices (e.g., different data rates on each side of the gearbox) while experiencing minimal latency during data transfer. Thus, the gearbox provides the benefits of switches (e.g., bridging bandwidth differences) and retimers (e.g., low latency) without the drawbacks associated with switches (e.g., higher latency) and retimers (bandwidth limitations). Therefore, the gearbox provides a superior alternative to switches and retimers in computing systems conforming to specific communication standards (e.g., PCIe).

[0064] The examples disclosed herein provide a toolbox implemented with a set of multiple PLs and DLs, as well as multiple independent links. In this way, each link implements its own set of PLs and DLs(s) independently of the other links. Using this approach, independent links can support new versions of the PCIe specification (e.g., Gen7 line rates of 128 GT / s per lane) or other communication standards. For example, since the Compute Express Link (CXL) uses both the PCIe physical and data link layers, the toolbox here can be implemented using the CXL communication standard. For example, CXL is typically used for coherent system memory that is highly sensitive to latency (e.g., introduced when data is transmitted through the employed switch). Therefore, the toolbox provides an improved solution to replace both the CXL / PCIe switch and the CXL / PCIe retimer.

[0065] Furthermore, the example toolkit in this paper offers additional benefits beyond switches in data transmission. For example, due to the simple design of the collection of PLs and DLs (e.g., two collections, etc.) and independent links (e.g., two links), the toolkit requires a significantly smaller area than a switch. Because of this reduced area, the implementation cost of the toolkit will be lower than that of a switch.

[0066] The following example illustrates how to use a toolkit to establish a topology for point-to-point connections for data transfer between different hardware components in a computing system that conforms to communication standards such as PCIe and CXL.

[0067] For example, Figure 1 A computing system 100 for transferring data between hardware components in a computing device is shown. For example, Figure 1 The computing system 100 typically includes a CPU 102, a memory 104, and an endpoint 106, wherein data can be transferred between the CPU 102 and / or the memory 104 and the endpoint 106. Figure 1 In the example, data communication follows and is described relative to the PCIe standard. However, it should be understood that data communication with respect to computing system 100 (or any other system or component herein) may conform to another suitable standard, such as the CXL standard.

[0068] exist Figure 1 In the example, CPU 102 and memory 104 can be any suitable processor and memory circuitry in a computing device (e.g., a server, personal computer, laptop computer, etc.). For example, memory 104 may include one or more volatile memory circuitry, such as static random access memory (SRAM) circuitry or dynamic random access memory (DRAM) circuitry. Figure 1In this system, CPU 102 includes processor 112. In such an example, processor 112 may include a single processor circuit or multiple processor circuits for executing executable instructions stored, for example, in memory 104. Endpoint 106 may be any suitable device associated with a computing device. For example, endpoint 106 may include a graphics card (e.g., a graphics processing unit in a card), a sound card, a solid-state drive, an interposer memory card, a network interface controller, and / or any other peripheral devices in computing system 100 that provide specific functionality. Although Figure 1 The computing system 100 is shown as including one endpoint, but it should be understood that the computing system 100 or other computing systems herein may include multiple endpoints (e.g., multiple devices) communicating with the CPU 102.

[0069] The computing system 100 also includes a root complex 108 and a toolbox 110. The root complex 108 can be used as follows: Figure 1 The components are either integrated into or external to CPU 102. Root complex 108 acts as a bridge between CPU 102 (more specifically, processor 112) and the PCIe framework, thereby connecting CPU 102 and memory 104 to endpoint 106 and / or other PCIe devices in computing system 100. Root complex 108 performs various functions, including control over endpoint 106 and / or other PCIe devices in the hierarchy. For example, root complex 108 can detect and identify endpoint 106 connected to a communication bus (e.g., PCI bus) in the computing device, and route traffic between CPU 102, memory 104, and endpoint 106, etc.

[0070] As shown in the figure, toolbox 110 is connected between root complex 108 and endpoint 106. For example, although Figure 1 Not shown, but toolbox 110 includes a port for connecting to root complex 108 and another port for connecting to endpoint 106. In such an example, each port may include a Physical Media Attachment (PMA) transmitter (Tx) and a PMA receiver (Rx). Although Figure 1 The computing system 100 is shown as including a toolbox, but it should be understood that in some exemplary embodiments, the computing system 100 or other computing systems herein may include multiple toolboxes.

[0071] Toolbox 110 transfers PCIe-compliant data between root compound 108 and endpoint 106. For example, root compound 108 may initiate a data request (e.g., a memory read request (MRd)) which is passed to endpoint 106 via toolbox 110. Upon receiving the data request, endpoint 106 may then return a response containing data to root compound 108 (via toolbox 110). In other examples, endpoint 106 may initiate a data request that is passed to root compound 108 via toolbox 110. Upon receiving the data request, root compound 108 may then return a response containing data to endpoint 106 (e.g., from memory 104). In some cases, data may be transferred between endpoint 106 and another endpoint. In such an example, endpoint 106 may initiate a data request that is passed to root compound 108 via toolbox 110. Then, upon receiving a data request, root complex 108 forwards the request to another endpoint (e.g., a completer endpoint) via toolbox 110 or another toolbox (or generates another data request for that other endpoint). Next, the completer endpoint may return a response with data to root complex 108 via toolbox 110 or another toolbox. Root complex 108 then forwards the response to endpoint 106 via toolbox 110 (or generates another response for endpoint 106).

[0072] Toolbox 110 may include various layers (not shown) for facilitating data transmission between root complex 108 and endpoint 106. For example, toolbox 110 may include the physical layer and data link layer from a protocol stack (e.g., a PCIe protocol stack). In these examples, the physical layer typically manages low-level telecommunications and data transmission (e.g., encoding, decoding, framing, clock recovery, etc.) on the link connecting root complex 108 and endpoint 106. The data link layer typically manages flow control, error detection (CRC), link management, etc., to ensure reliable transmission of data packets between root complex 108 and endpoint 106.

[0073] For example, Figure 2 A computing system 200 is shown that is similar to computing system 100 but has a toolbox 210 that includes a physical layer and a data link layer. Figure 2 In the middle, toolbox 210 is connected to Figure 1 Root complex 108 and Figure 1 Between the endpoints 106. Figure 2 Toolbox 210 works in a similar manner to the other toolboxes in this article.

[0074] exist Figure 2In this example, toolbox 210 includes a first set of physical layer 212 and data link layer 214, and a second set of physical layer 216 and data link layer 218. With this configuration, physical layer 212 is connected to or includes a port (not shown) for connection to root complex 108, and data link layer 214 is connected between data link layer 218 and physical layer 212. Additionally, physical layer 216 is connected to or includes another port (not shown) for connection to endpoint 106, and data link layer 218 is connected between data link layer 214 and physical layer 216. In such an example, each port may include a PMA Tx and a PMA receiver Rx, as described above.

[0075] exist Figure 2 In this example, toolbox 210 implements two independent links 220 and 222. Link 220 (e.g., a PCIe link) connects the root complex 108 to toolbox 210, and link 222 (e.g., a PCIe link) connects the endpoint 106 to toolbox 210. Each link 220 and 222 operates independently, allowing each link to negotiate its own parameters (e.g., data rate, etc.). In this example, link 220 is formed with physical layer 212 and data link layer 214 for communication with root complex 108, and link 222 is formed with physical layer 216 and data link layer 218 for communication with endpoint 106.

[0076] In various examples, links 220 and 222 may include one or more channels for transmitting data between root complex 108 and endpoint 106. In such examples, each channel includes two lines for sending and receiving data. For example, in link 220, physical layer 212 and data link layer 214 may form one or more channels. Additionally, in link 222, physical layer 216 and data link layer 218 may form one or more channels. The number of channels associated with link 220 and the number of channels associated with link 222 may be the same or different.

[0077] exist Figure 2In the example, toolbox 210 is implemented without a transaction layer. Specifically, links 220 and 222 are not associated with a transaction layer. In other words, the protocol stacks associated with links 220 and 222 consist only of physical layers 212 and 216 and data link layers 214 and 218, respectively. The protocol stacks do not include a transaction layer. Using this implementation, toolbox 210 can experience reduced latency compared to other PCIe devices (e.g., switches). For example, because there is no transaction layer, there is no need to use transaction layer buffers such as store-and-forward buffers or virtual channel (VC) buffers to implement toolbox 210. Such transaction layer buffers typically introduce latency because the entire data packet in a request or response must be received and processed before the forwarding of the request or response can occur.

[0078] exist Figure 2 In the example, the bandwidth on link 220 and the bandwidth on link 222 are the same, such as 512G, 1024G, 2048G, etc. For example, because toolbox 210 does not include a transaction layer buffer, toolbox 210 cannot throttle transaction layer packets received from root complex 108 or endpoint 106. Thus, the bandwidth matching on each of links 220 and 222 allows transaction layer packets that pass link layer checks (as explained further below) on the RX side of data link layers 214 and 216 to be sent out on the TX side of data link layers 214 and 216 without throttling.

[0079] Furthermore, the data rates on links 220 and 222 can be the same or different, as long as the bandwidth on each link is the same. For example, the data rate at a port connected to root complex 108 can be the same or different from the data rate at a port connected to endpoint 106. For example, and as described above, a CPU connected to root complex 108 can be migrated to a higher bandwidth configuration with a higher data rate (e.g., 64 GT / s per channel, 128 GT / s per channel, etc.) than endpoint 106, which has a lower data rate (e.g., 32 GT / s per channel). Using this method, toolbox 210 can bridge bandwidth differences between the root complex side (e.g., a faster host device) and the endpoint side (e.g., one or more endpoints with slower speeds) by, for example, connecting additional endpoints, adding channels on the endpoint side, etc. In other examples, CPUs connected to root complex 108 and endpoint 106 can have the same bandwidth and the same data rate (e.g., 32 GT / s per channel, 64 GT / s per channel, 128 GT / s per channel, etc.).

[0080] In some examples, the physical layers 212 and 216 of toolbox 210 can be decomposed into multiple sub-layers. For example, each physical layer 212 and 216 may include a physical coding (PCS) sub-layer, a media access control (MAC) sub-layer, etc.

[0081] As an example, Figure 3 It shows the relationship with Figure 2 Toolbox 210 is similar to toolbox 310, except that physical layers 212 and 216 have multiple sub-layers. Although Figure 3 Not shown, but toolbox 310 can be connected to the root complex (e.g., Figure 1 The root complex 108) and one or more endpoints (e.g., Figure 1 Between the endpoints 106. Figure 3 Toolbox 310 works in a similar way to the other toolboxes here.

[0082] exist Figure 3 In the example, toolbox 310 includes physical layers 212 and 216 and data link layers 214 and 218. Physical layer 212 and data link layer 214 correspond to link 220 (e.g., a PCIe link) connecting the root complex and toolbox 310, and physical layer 216 and data link layer 218 correspond to link 222 (e.g., a PCIe link) connecting one or more endpoints and toolbox 310, as described above. Links 220 and 222 are independent of each other.

[0083] like Figure 3 As shown, physical layers 212 and 216 each include a port, a PCS sublayer, and a MAC sublayer. Specifically, in Figure 3 In this configuration, physical layer 212 includes port 330, a PCS sublayer 332 connected to port 330, and a MAC sublayer 334 connected to data link layer 214. Similarly, physical layer 216 includes port 336, a PCS sublayer 338 connected to port 336, and a MAC sublayer 340 connected to data link layer 218. With this configuration, PCS sublayer 332 provides a data interface between port 330 and MAC sublayer 334, and PCS sublayer 338 provides a data interface between port 336 and MAC sublayer 340.

[0084] exist Figure 3In this protocol stack, each port 330, 336 can be a sublayer comprising PMA Tx and PMA Rx. In such an example, port 330 is connected to the root complex interface, and port 336 is connected to the endpoint interface. Furthermore, each PCS sublayer 332, 338 facilitates the conversion of data into a format suitable for transmission over its respective PMA and MAC sublayers. For example, and as further described herein, each PCS sublayer 332, 338 typically performs data encoding and decoding, alignment mark insertion and removal, and channel block synchronization. Each MAC sublayer 334, 340 manages access to its respective link and typically performs framing, data encoding and decoding, flow control, error detection, etc. In such an example, data link layer 214 and physical layer 212 (with port 330, PCS sublayer 332, and MAC sublayer 334) are part of a protocol stack, and data link layer 218 and physical layer 216 (with port 336, PCS sublayer 338, and MAC sublayer 340) are part of another protocol stack.

[0085] Additionally, toolbox 310 includes control module 342. Figure 3 In the example, control module 342 may include a control and status register (CSR) for storing information about the state of control module 342 and controlling rate matching of data passing through data link layers 214, 218. For example, the CSR (or more generally, control module 342) may store interrupt status, operating mode, flags, etc., and control interrupt enabling / disabling, changing operating modes, setting flags, etc. Furthermore, in some embodiments, control module 342 may detect errors on data link layers 214, 218 and / or physical layers 212, 216, as further explained below.

[0086] Figures 4A-4B Another example toolbox 400 for connecting a root complex to one or more endpoints is shown. Toolbox 400 is similar to... Figure 3 The toolbox 310 differs in that it includes various modules for implementing data transfer between the root complex and endpoints. Figures 4A-4B Toolbox 400 works in a similar way to the other toolboxes here.

[0087] exist Figures 4A-4B The toolbox 400 includes... Figure 3 The physical layers 212 and 216 and the data link layers 214 and 218. Specifically, in Figure 4A In this context, the physical layer 212 includes port 330, PCS sublayer 332, and MAC sublayer 334. Additionally, in... Figure 4B In this layer, physical layer 216 includes port 336, PCS sublayer 338, and MAC sublayer 340. Physical layer 212 and data link layer 214 are connected to the root complex (…). Figure 4A (not shown in the diagram) forms an independent link, and the physical layer 216 and data link layer 218 connect with the endpoint ( Figure 4B (Not shown in the image) to form another independent link.

[0088] Physical layers 212, 216 and data link layers 214, 218 include similar modules along the data receive (RX) path and data transmit (TX) path for independent links. Figures 4A-4B As shown, modules along the data receive (RX) path of one link are positioned in a mirror configuration relative to modules along the data receive (RX) path of another link. Similarly, modules along the data transmit (TX) path of one link are positioned in a mirror configuration relative to modules along the data transmit (TX) path of another link.

[0089] exist Figure 4A In this configuration, port 330 includes PMA Rx 402 and PMA Tx 404, and PCS sublayer 332 includes block alignment module 406, elastic buffer 408, and encoding / loopback module 410. As shown, block alignment module 406 and elastic buffer 408 are on the data receive (RX) path of PCS sublayer 332. Encoding / loopback module 410 is on the data transmit (TX) path of PCS sublayer 332.

[0090] in addition, Figure 4A The MAC sublayer 334 includes a collection of two data path functional modules and a Link Training and Status State Machine (LTSSM) module 438. The LTSSM module 438 typically manages the initialization and configuration of its associated links, such as link width negotiation, data rate negotiation, and equalization, to ensure stable and reliable connections between devices. One data path corresponds to the data reception (RX) path, which includes a precoding module 412, a Gray codec module 414, a descrambler module 416, a deskew module 418, a Non-Flit Mode (NFM) TLP (Transaction Layer Packet) digest (TD) & marker module 420, a Flit Mode (FM) CRC (Cyclic Redundancy Check) and FEC (Forward Error Correction) check module 422, and a multiplexer 424. The other data path corresponds to the data transmission (TX) path, which includes a precoding module 426, a Gray coding module 428, a scrambler module 430, a lane striping module 432, a multiplexer 434, and an FM TX (frequency modulation transmission) retry and buffer module 436.

[0091] Data link layer 214 includes a set of two data path functional modules and a link finite state machine (FSM) module 444, which manages and controls the behavior of its associated links. One data path corresponds to a data receive (RX) path with a TLP check and drop module 440 and a data link layer packet (DLLP) check module 442. The other data path corresponds to a data transmit (TX) path with a DLLP generator 446 and an NFM TX retry and buffer module 448.

[0092] exist Figure 4B In this configuration, port 336 includes PMA Tx 450 and PMA Rx 452, and PCS sublayer 338 includes an encoding / return module 454 in the data transmission (TX) path, and a block alignment module 456 and a flexible buffer 458 in the data reception (RX) path. Additionally, Figure 4B The MAC sublayer 340 includes a set of two data path functional modules and an LTSSM module 486. Similar to the LTSSM module 438, the LTSSM module 486 typically manages the initialization and configuration of its associated interconnected links, such as link width negotiation, data rate negotiation, and equalization, to ensure stable and reliable connections between devices. One data path corresponds to the data transmission (TX) path, which includes a precoding module 460, a Gray coding module 462, a scrambler module 464, a channel striping module 466, a multiplexer 468, and an FM TX retry and buffer module 470. The other data path corresponds to the data reception (RX) path, which includes a precoding module 472, a Gray coding module 474, a descrambler module 476, an anti-skew module 478, an FM CRC & FEC check module 480, an NFM TD & marker module 482, and a multiplexer 484. Data link layer 218 includes a set of two data path functional modules and a link FSM module 492, which manages and controls the behavior of its associated links. One data path corresponds to a data transmission (TX) path with an NFM TX retry and buffering module 488 and a DLLP generator 490. The other data path corresponds to a data reception (RX) path with a DLLP checking module 494 and a TLP checking and discarding module 496.

[0093] exist Figures 4A-4BIn this configuration, data link layers 214 and 218 communicate with each other and partially with MAC sublayers 334 and 340. For example, the TLP check and drop module 440 in data link layer 214 passes data (e.g., TLP) to the NFM TX retry and buffer module 488 in data link layer 218 and the FM TX retry and buffer module 470 in MAC sublayer 340. Additionally, the TLP check and drop module 496 in data link layer 218 passes data (e.g., TLP) to the NFM TX retry and buffer module 448 in data link layer 214 and the FM TX retry and buffer module 436 in MAC sublayer 334. Furthermore, the DLLP checking module 442 in data link layer 214 transmits data (e.g., DLLP) to the DLLP generator 490 in data link layer 218 via a remote fiber optic channel (FC), and the DLLP checking module 494 in data link layer 218 transmits data (e.g., DLLP) to the DLLP generator 446 in data link layer 214 via another remote FC. DLLP checking module 442 and DLLP generator 446 communicate via a local FC, and DLLP checking module 494 and DLLP generator 490 communicate via another local FC.

[0094] During operation, data (e.g., TLP) is received via PMA Rx 402, 452 and passed to block alignment modules 406, 456 and elastic buffers 408, 458 in PCS sublayers 332, 338. Block alignment modules 406, 456 typically synchronize data transmissions, allowing for the interpretation of appropriate data blocks. Elastic buffers 408, 458 are used to compensate for time differences (e.g., to recover the time difference between the clock and the local clock). The data then passes through the data path functional modules associated with MAC sublayers 334, 340.

[0095] For example, precoding modules 412, 472 receive and precode the scrambled data bits from elastic buffers 408, 458. Then, Gray code decoding modules 414, 474 typically convert the data bits in Gray code format (e.g., a binary number system where two consecutive values ​​differ by only one bit) back to their equivalent binary representation. Descrambler modules 416, 476 then descramble the data (e.g., restoring the data stream to its original form). Next, the data is passed to correction modules 418, 478, which align the data across multiple channels to compensate for inter-channel bias. Then, before passing through multiplexers 424, 484, the data is passed to the NFM TD and tagger modules 420, 482 and the FM CRC and FEC check modules 422, 480 in the MAC sublayers 334, 340.

[0096] Each FM CRC and FEC check module 422, 480 implements a retry buffer, forward error correction (FEC), and cyclic redundancy check (CRC). For example, FEC adds redundant data to the data stream, enabling error detection and correction without requiring retransmissions. CRC is another error detection tool that calculates a value for the data (e.g., a checksum) and then sends that value along with the data and compares it to a similarly calculated value downstream. The retry buffer can store retransmitted data affected by errors (e.g., data with detected errors). Each NFM TD and tagger module 420, 482 implements TLP and DLLP tags that define the boundaries of the passed TLP and DLLP.

[0097] Then, data is passed from multiplexers 424 and 484 to TLP checking and discarding modules 440 and 496 and DLLP checking modules 442 and 494 in data link layers 214 and 218. DLLP checking modules 442 and 494 perform a verification process to verify the integrity of the received DLLP. TLP checking and discarding modules 440 and 496 perform a verification process to verify the integrity of the received TLP. If a fatal error exists in the TLP, the packet can be discarded.

[0098] Then, data (e.g., TLP) is transferred from one data link layer 214, 218 to another data link layer 214, 218 via a remote FC. Specifically, data is transferred from DLLP checking modules 442, 494 in one data link layer 214, 218 to DLLP generators 446, 490 in another data link layer 214, 218. Additionally, data is transferred from TLP checking and dropping modules 440, 496 in one data link layer 214, 218 to NFM TX retry and buffering modules 448, 488 in another data link layer 214, 218. Through this transfer, data is transferred from the receive (Rx) path of one link to the transmit (TX) path of another link.

[0099] In their respective transport (TX) paths, DLLP generators 446 and 490 create DLLPs. The created DLLPs are then passed from DLLP generators 446 and 490 to FM TX retry and buffer modules 436 and 470 in MAC sublayers 334 and 340, and to NFM TX retry and buffer modules 448 and 488 in data link layer 214. Additionally, each FM TX retry and buffer module 436 and 470 and each NFM TX retry and buffer module 448 and 488 implements a TLP sequence numbering process, CRC, and a retry buffer. The TLP sequence numbering process adds a sequence number to a TLP with non-fatal errors. The CRC calculates the value of the data (e.g., a checksum), then sends that value along with the data and compares it to a similarly calculated value downstream. As described above, the retry buffer can store retransmitted data affected by errors (e.g., data with detected errors). Then, data is transmitted from the FM TX retry and buffer modules 436, 470 and the NFM TX retry and buffer modules 448, 488 through multiplexers 434, 468 in the MAC sublayers 334, 340.

[0100] The data then passes through the channel striping modules 432 and 466, scrambler modules 430 and 464, Gray coding modules 428 and 462, and precoding modules 426 and 460 in the MAC sublayers 334 and 340. For example, each channel striping module 432 and 466 distributes data packets (e.g., TLP and DLLP) across multiple channels of its link. The scrambler modules 430 and 464 randomize (or scramble) the data stream before transmission. The Gray coding modules 428 and 462 convert the data bits into Gray code format (e.g., a binary number system where two consecutive values ​​differ by only one bit), and the precoding modules 426 and 460 precode the scrambled data bits. The data is then passed through the transmission (TX) path in the PCS sublayers 332 and 338 to the encoding / return modules 410 and 454. The encoding / return modules 410 and 454 test the transmitted data by sending encoded data and receiving the same data returned for verification. Then, the data is passed to PMA Tx 404, 450 before flowing to the root complex and endpoints.

[0101] Additionally, such as Figure 4A As shown, the toolbox 400 includes a control module 498, which can communicate with... Figure 3The control module 342 operates in a similar manner. For example, the control module 498 may include a CSR for storing information about the control state and controlling the rate matching of data passing through data link layers 214, 218. Furthermore, in some embodiments, the control module 342 may detect errors on data link layers 214, 218 and / or physical layers 212, 216.

[0102] For example, the PCIe specification has defined components (e.g., root complexes, retimers, switches, endpoints, etc.) and defined functions for each of those components. Toolboxes such as any of the toolboxes disclosed herein are not components defined in the PCIe specification. However, if components are compatible with the PCIe specification, they can interoperate with other PCIe-compatible components in a computing system. Accordingly, toolbox 400 (or any other toolbox disclosed herein) can function to ensure that they do not cause interoperability problems when used. In other words, toolbox 400 can function as if it did not exist in a computing system with PCIe-compatible components (e.g., invisible). Therefore, and as further explained below, toolbox 400 can be designed to handle different PCIe-related functions without causing any problems.

[0103] Toolbox 400 can be designed to handle challenges associated with bandwidth mismatch and errors (e.g., physical layer errors, data link layer errors, etc.). For example, because Toolbox 400 does not have a transaction layer with a transaction layer buffer, it cannot throttle received TLPs. In order for Toolbox 400 to function as if it were not present, the bandwidth on each link must be matched so that TLPs checked by the link layer on the data receive (RX) path (e.g., CRC, FEC, and FM CRC and FEC checking modules 422, 480 in NFM TD and marker modules 420, 482) can be transmitted on the transmit (TX) path without any throttling.

[0104] For example, toolbox 400 can prevent link segments between upstream PCIe components (e.g., root complex, etc.) and toolbox 400, and between toolbox 400 and downstream PCIe components (e.g., endpoints, etc.), from entering the DL_Active state unless the bandwidth on the link segments matches. In the DL_Active state, data link layers 214, 218 are active and ready for packet transmission. Using this method, control module 498 can determine whether the bandwidth on the link between the upstream PCIe component and toolbox 400 is the same as the bandwidth on another link between toolbox 400 and downstream PCIe components. If the bandwidths are different, control module 498 or another suitable module (e.g., link FSM modules 444, 492) can prevent DLLP from being transmitted via toolbox 400 between the root complex and endpoints. This can be achieved, for example, by preventing toolbox 400 from entering the DL_Active state.

[0105] In response to equal (or matched) bandwidth, control module 498 can allow toolbox 400 to enter the DL_Active state, and packets (e.g., DLLPs) can be passed between the root complex and endpoints via toolbox 400. In such instances, control module 498 may switch one or both links into recovery / configuration mode to attempt to match the bandwidth on the two link segments. Once the bandwidth on the link segments matches, toolbox 400 will allow the DLLP to pass, which causes the link segments between the upstream and downstream components and toolbox 400 to enter the DL_Active state.

[0106] Additionally, once a link is active (e.g., in DL_Active state), errors (e.g., bit error rate (BER) etc.) may occur on one or two link segments. Such errors are typically transient error conditions corrected by a recovery cycle. During this period, the transmit (TX) path is blocked. Thus, traffic via toolbox 400 can be temporarily suspended during this recovery. Therefore, control module 498 or another suitable module in toolbox 400 can detect transient error conditions (e.g., BER etc.) on one of these links. In response to the detection of a transient error condition, control module 498 (or one of link FSM modules 444, 492) can send a negative acknowledgment (NAK) signal on the other (active) link for the received packets. In some examples, a prolonged recovery may cause replay number flips in the active link. Thus, toolbox 400 can set replay number flips as a non-fatal error. This allows the recovery cycle to continue operating.

[0107] Furthermore, once the link is active, one of the link FSM modules 444 and 492 can detect link outage conditions associated with its link. For example, the downstream link FSM module can detect link outage conditions caused by, for example, endpoint removal (e.g., accidental hot-plugging), or potential link degradation that prevents the downstream LTSSM module from connecting. In such an example, if the downstream link FSM module (e.g., link FSM module 444) detects a link outage condition on the downstream link segment, the upstream LTSSM module (e.g., LTSSM module 486) can create an outage condition on the upstream link segment.

[0108] In other instances, the upstream link FSM module (e.g., link FSM module 492) can detect link interruption conditions. This may occur, for example, if the upstream LTSSM module fails to link or fails to detect a status condition. In response to the upstream link FSM module detecting a link interruption condition, the downstream LTSSM module (e.g., LTSSM module 438) can enter a disabled state.

[0109] Toolbox 400 can also be designed to address challenges associated with low-power support. For example, PCIe links can exist in different power states to manage power consumption. Power states can include active (or normal power) state (L0), idle (or low-latency standby) state (L0s), low power (or standby) state (L1), low power sub-state (L1 SS), dynamically adjustable state (L0P), sleep state (L2), and link off state (L3). In such an example, power state entries can be initiated by downstream and upstream components (e.g., root complex, endpoints, etc.).

[0110] For example, a downstream component can initiate a lower power state (L1). In such an example, if there is no TLP to forward in the downstream transmission (TX) path, toolbox 400 can accept a lower power request. Therefore, if no TLP exists in the link associated with the downstream component, control module 498 can enable the low power state (L1) of that link.

[0111] Then, once the downstream link is in a low-power state (L1), a request can be made to the upstream component. For example, after the downstream link is in a low-power state (L1), the control module 498 can initiate a low-power state (L1) request to the upstream component (e.g., the root complex, etc.) connected to the upstream link. The upstream component can accept or reject the low-power state (L1) request. If rejected, the control module 498 can transition the downstream link back to an active state (L0).

[0112] Additionally, power state exit can be initiated by downstream and upstream components (e.g., root complex, endpoints, etc.). For example, if a link is in a low-power state (L1), an upstream or downstream component can initiate a request to exit the low-power state (L1), thereby causing its associated link to begin transitioning out of the low-power state (L1). In response to this request, control module 498 can control the link to also begin transitioning out of the low-power state (L1).

[0113] Other power state transitions can be handled in a similar manner to entering and exiting a low-power state (L1). For example, low-power sub-states (L1 SS) can be managed in a similar way to low-power states (L1). Additionally, active states (L0) can be switched at the link segment level (e.g., similar to the behavior of a switch).

[0114] Figures 5 to 10 The control processes 500, 600, 700, 800, 900, and 1000, which enable the toolbox to handle different PCIe-related functions, are illustrated. (Relative to...) Figures 4A to 4B Toolbox 400 (e.g., control module 498, link FSM modules 444, 492, etc.) explains the operation of control processes 500, 600, 700, 800, 900, and 1000. However, it should be understood that the operation of control processes 500, 600, 700, 800, 900, and 1000 can be implemented using any other toolbox disclosed herein.

[0115] Figure 5 The control process 500 illustrates an example of addressing a link outage condition on a downstream link segment. Figure 5 In the process, control procedure 500 begins at 502, where control module 498 determines whether the downstream and upstream links are active (e.g., DL_Active state). If not, control procedure 500 returns to 502. If yes, control procedure 500 proceeds to 504.

[0116] At 504, toolbox 400 determines whether a link outage condition has been detected on the downstream link segment. This detection can be performed by the downstream link FSM module (e.g., link FSM module 444) in toolbox 400. If no, control procedure 500 returns to 502. Otherwise, if yes at 504, control procedure 500 proceeds to 506. At 510, the upstream LTSSM module (e.g., LTSSM module 486) can create an unexpected outage condition on the upstream link segment.

[0117] Figure 6 The control process 600 illustrates an example of addressing a link outage condition on an upstream link segment. Figure 6In the process, control procedure 600 begins at 602 by determining, via control module 498, whether the downstream and upstream links are in an active state (e.g., DL_Active state), as explained above. If not, control procedure 600 returns to 602. If yes, control procedure 600 proceeds to 604.

[0118] At 604, toolbox 400 determines whether a link interruption condition has been detected on the upstream link segment. This detection can be performed by the upstream link FSM module (e.g., link FSM module 492) in toolbox 400. If no, control procedure 600 returns to 602. Otherwise, if yes at 604, control procedure 600 proceeds to 606. At 606, the downstream LTSSM module (e.g., LTSSM module 438) enters a disabled state on the downstream link.

[0119] Figure 7 The control process 700 illustrates an example of transient error conditions on the addressing upstream link segment. Figure 7 In this process, control procedure 700 begins at 702, where control module 498 determines whether the downstream and upstream links are active (e.g., DL_Active state), as described above. If not, control procedure 700 returns to 702. If yes, control procedure 700 proceeds to 704.

[0120] At 704, toolbox 400 detects whether there is an error condition on the upstream link segment (e.g., a transient error condition caused by a high BER). If not, control procedure 700 returns to 702. If yes at 704, control procedure 700 proceeds to 706, where recovery mode is entered in the case of an upstream link segment blockage. Then control procedure 700 proceeds to 710.

[0121] At 710, a NAK is sent for the received packet on the downstream link segment (e.g., another active link segment). Control procedure 700 then proceeds to 712, where toolbox 400 determines whether recovery is complete. If yes, control procedure 700 returns to 702. Otherwise, if no at 712, control procedure 700 returns to 710.

[0122] Figure 8 The control process 800 illustrates an example of transient error conditions on the downstream link segment. Figure 8 In the process described above, control process 800 begins at 802 by determining, through control module 498, whether the downstream and upstream links are in an active state (e.g., DL_Active state). If not, control process 800 returns to 802. If yes, control process 800 proceeds to 804.

[0123] At 804, toolbox 400 detects whether there is an error condition on the downstream link segment (e.g., a transient error condition caused by a high BER). If not, control procedure 800 returns to 802. If yes at 804, control procedure 800 proceeds to 806, where it enters recovery mode, where the downstream link segment is blocked. Then control procedure 800 proceeds to 810.

[0124] At 810, a NAK is sent for the received packet on the upstream link segment (e.g., another active link segment). Then, control procedure 800 proceeds to 812, where toolbox 400 determines whether recovery is complete. If yes, control procedure 800 returns to 802. Otherwise, if no at 812, control procedure 800 returns to 810.

[0125] exist Figure 9 In the process, control process 900 begins at 902 by determining, via control module 498, whether the bandwidth on the link segments connecting the upstream component and toolbox 400, and connecting toolbox 400 and downstream components, matches. If not, control process 900 proceeds to 904 and 906. At 904, toolbox 400 enters a recovery state. At 906, the link segment is prevented from entering an active state (e.g., DL_Active state). Control process 900 can then return to 902.

[0126] However, if yes is true at 902, then control procedure 900 proceeds to 908. At 908, the link segment is allowed to enter its DL_active state. This allows packets (e.g., DLLPs) to be transmitted between upstream and downstream components.

[0127] exist Figure 10 In this process, control process 1000 begins at 1002 by determining, via control module 498, whether a request to enter a low-power state has been received from a downstream component. If not, control process 1000 may return to 1002. If yes, control process 1000 proceeds to 1004.

[0128] At 1004, it is determined whether any TLP exists in the downstream transmission (TX) path. If yes, control process 1000 proceeds to 1006, where the downstream link is prevented from entering a low-power state. Control process 1000 can then return to 1002. If no at 1004, control process 1000 proceeds to 1008, where the downstream link enters a low-power state. Control process 1000 then proceeds to 1010.

[0129] At 1010, a request to enter a low-power state is sent to the upstream component. Control process 1000 then proceeds to 1012, where control module 498 determines whether the request is accepted by the upstream component. If not, control process 1000 proceeds to 1014, where the downstream link exits the low-power state. Control process 1000 can then return to 1002.

[0130] However, if the condition is yes at 1012, control process 1000 proceeds to 1016. At 1016, the upstream link enters a low-power state. Control process 1000 then proceeds to 1018, where at 1080 control module 498 determines whether a request to exit the low-power state has been received from either the upstream or downstream component. If no, control process 1000 returns to 1018. Otherwise, if the upstream or downstream component requests to exit the low-power state, both the upstream and downstream links exit the low-power state at 1020. Control process 1000 can then either terminate or return to 1002.

[0131] As described above, this toolbox offers the advantages of switches and retimers without their associated disadvantages. Thus, the toolbox provides hybrid components that can replace switches and retimers in a computing system according to PCIe or other relevant communication standards. For example, as shown in Table 1 below, the toolbox offers low-latency options similar to retimers. For instance, the latency associated with toolbox 400 could be (a) RX NFM latency – PMA RX + PCS + MAC + DLL; (b) RX FM latency – PMA RX + PCS + MAC; (c) RX to TX interconnect latency; and (d) TX FM / NFM latency – DLL + MAC + PMA TX. Additionally, similar to switches, the toolbox can operate with the same or different bandwidths, the same or different data rates, and the same or different numbers of lanes. Furthermore, similar to retimers, the toolbox offers reduced area requirements and associated costs.

[0132]

[0133] The foregoing description is merely illustrative in nature and is in no way intended to limit this disclosure, its application, or use. The broad teachings of this disclosure can be implemented in various forms. Therefore, although this disclosure includes specific examples, its true scope should not be limited thereto, as other modifications will become apparent upon examination of the accompanying drawings, specification, and appended claims. It should be understood that one or more steps within the method may be performed in a different order (or simultaneously) without altering the principles of this disclosure. Furthermore, although each embodiment has been described above as having certain features, any one or more of those features described with respect to any embodiment of this disclosure may be implemented in and / or combined with features of any other embodiment, even if such combination is not explicitly described. In other words, the described embodiments are not mutually exclusive, and substitution of one or more embodiments for each other remains within the scope of the invention.

[0134] Various terms are used to describe spatial and functional relationships between elements (e.g., between modules, circuit elements, semiconductor layers, etc.), including “connected,” “joined,” “coupled,” “adjacent,” “closely adjacent,” “on top of,” “above,” “below,” and “set.” Unless explicitly described as “direct,” when describing a relationship between first and second elements in the foregoing disclosure, the relationship can be a direct relationship in which no other intervening element exists between the first and second elements, or an indirect relationship (spatially or functionally) in which one or more intervening elements exist between the first and second elements. As used herein, at least one of the phrases A, B, and C should be interpreted as indicating a logic of non-exclusive OR (A or B or C) and should not be interpreted as indicating “at least one of A, at least one of B, and at least one of C.”

[0135] In the accompanying drawings, the direction of the arrows, as indicated by the arrows, typically represents the flow of information of interest (e.g., data or instructions). For example, when unit A and unit B exchange various types of information, but the information sent from unit A to unit B is relevant to the illustration, the arrow may point from unit A to unit B. This unidirectional arrow does not imply that no other information is sent from unit B to unit A. Furthermore, for information sent from unit A to unit B, unit B may send a request for that information to unit A or receive confirmation of that information.

[0136] In this application, including the following definitions, the term "module" or "controller" may be replaced by the term "circuit". The term "module" may refer to or include some of the following: application-specific integrated circuit (ASIC); digital, analog, or mixed analog / digital discrete circuit; digital, analog, or mixed analog / digital integrated circuit; combinational logic circuit; field-programmable gate array (FPGA); processor circuitry (shared, dedicated, or grouped) that executes code; memory circuitry (shared, dedicated, or grouped) that stores code executed by the processor circuitry; other suitable hardware components that provide the functions described; or some or all of the above, such as in a system-on-a-chip.

[0137] This module may include one or more interface circuits. In some examples, the interface circuits may include wired or wireless interfaces connected to a local area network (LAN), the Internet, a wide area network (WAN), or a combination thereof. The functionality of any given module in this disclosure can be distributed among multiple modules connected via the interface circuits. For example, multiple modules can allow for load balancing. In another example, a server (also referred to as a remote or cloud) module may perform certain functions on behalf of a client module.

[0138] As described above, the term "code" can include software, firmware, and / or microcode, and can refer to programs, routines, functions, classes, data structures, and / or objects. The term "shared processor circuitry" includes a single processor circuitry that executes some or all of the code from multiple modules. The term "group processor circuitry" covers processor circuitry that, in combination with additional processor circuitry, executes some or all of the code from one or more modules. References to multiple processor circuitry include multiple processor circuitry on a discrete die, multiple processor circuitry on a single die, multiple cores of a single processor circuitry, multiple threads of a single processor circuitry, or combinations thereof. The term "shared memory circuitry" includes a single memory circuitry that stores some or all of the code from multiple modules. The term "group memory circuitry" includes memory circuitry that, in combination with additional memory, stores some or all of the code from one or more modules.

[0139] The term "storage circuit" is a subset of the term "computer-readable medium." As used herein, the term "computer-readable medium" does not include transient electrical or electromagnetic signals propagating through a medium (such as on a carrier wave); therefore, the term "computer-readable medium" can be considered tangible and non-transient. Non-limiting examples of non-transient tangible computer-readable media include non-volatile memory circuits (e.g., flash memory circuits, erasable programmable read-only memory circuits, or mask read-only memory circuits), volatile memory circuits (e.g., static random access memory circuits or dynamic random access memory circuits), magnetic storage media (e.g., analog or digital magnetic tape or hard disk drives), and optical storage media (e.g., CDs, DVDs, or Blu-ray discs).

[0140] In this application, a device unit described as having specific attributes or performing specific operations is specifically configured to have those specific attributes and perform those specific operations. Specifically, the description of a unit performing an action means that the element is configured to perform that action. The configuration of the unit may include programming the unit, for example by encoding instructions on a non-transitory, tangible computer-readable medium associated with the unit.

[0141] The apparatus and methods described in this application can be implemented, in part or in whole, by a special-purpose computer created by configuring a general-purpose computer to perform one or more specific functions contained in a computer program. The aforementioned function blocks, flowchart components, and other elements serve as software specifications that can be translated into a computer program through the routine work of a skilled technician or programmer.

[0142] A computer program includes processor-executable instructions stored on at least one non-transitory tangible computer-readable medium. A computer program may also include or depend on stored data. A computer program may include a basic input / output system (BIOS) for interacting with the hardware of a special-purpose computer, device drivers for interacting with specific devices of the special-purpose computer, one or more operating systems, user applications, background services, background applications, etc.

[0143] Computer programs may include: (i) descriptive text to be parsed, such as HTML (Hypertext Markup Language), XML (Extensible Markup Language), or JSON (JavaScript Object Notation); (ii) assembly code; (iii) object code generated from source code by a compiler; (iv) source code for execution by an interpreter; and (v) source code for compilation and execution by a just-in-time (JIT) compiler, etc. As an example only, source code may come from languages ​​including C, C++, C#, Objective-C, Swift, Haskell, Go, SQL, R, Lisp, etc. Fortran, Perl, Pascal, Curl, OCaml, HTML5 (Hypertext Markup Language, Fifth Revision), Ada, ASP (Active Server Pages), PHP (PHP: Hypertext Preprocessor), Scala, Eiffel, Smalltalk, Erlang, Ruby, Visual Lua, MATLAB, SIMULINK and

Claims

1. A toolbox for connecting root complexes and endpoints in a computing device, the toolbox comprising: The first port is configured to connect to the root complex; The second port is configured to connect to the endpoint; A first physical layer and a second physical layer, wherein the first physical layer is connected to the first port and the second physical layer is connected to the second port; as well as A first data link layer and a second data link layer, wherein the first data link layer is connected between the second data link layer and the first physical layer, and the second data link layer is connected between the first data link layer and the second physical layer. The first physical layer, the first data link layer, the second physical layer, and the second data link layer are configured to form one or more channels for transmitting data between the root complex and the endpoints.

2. The toolbox according to claim 1, wherein: The first physical layer and the first data link layer form a first link; and The second physical layer and the second data link layer form a second link independent of the first link.

3. The toolbox according to claim 2, wherein: The first physical layer includes a Physical Coding Sublayer (PCS) connected to the first port and a Media Access Control (MAC) sublayer connected to the first data link layer; and The second physical layer includes a PCS connected to the second port and a MAC sublayer connected to the second data link layer.

4. The toolbox according to claim 3, wherein: The PCS of the first physical layer is configured to provide a data interface between the first port and the MAC sublayer; and The PCS of the second physical layer is configured to provide a data interface between the second port and the MAC sublayer of the second physical layer.

5. The toolbox according to claim 3, wherein: The MAC sublayer of the first physical layer includes a data path module between the PCS of the first physical layer and the first data link layer, and a first link training and state machine LTSSM module; and The MAC sublayer of the second physical layer includes a data path module between the PCS of the second physical layer and the second data link layer, and a second LTSSM module.

6. The toolbox according to claim 5, wherein: The first data link layer includes a first finite state machine (FSM) module configured to connect to the second link; and The second data link layer includes a second FSM module configured to connect to the first link.

7. The toolbox according to claim 6, wherein: The second FSM module is configured to detect link interruption status on the second link; as well as In response to the link interruption condition, the first LTSSM module is configured to create a link interruption condition on the first link.

8. The toolbox according to claim 6, wherein: The first FSM module is configured to detect link interruption status on the first link; as well as In response to the link interruption, the second LTSSM module is configured to enter a disabled state.

9. The toolbox according to claim 6 further includes a control module, the control module being configured to: Determine whether the bandwidth on the first link and the bandwidth on the second link are the same; and In response to the fact that the bandwidth on the first link and the bandwidth on the second link are the same, data link layer packets (DLLPs) are allowed to be transmitted between the root complex and the endpoints via the toolbox.

10. The toolbox of claim 9, wherein the control module is configured to prevent the DLLP from being transmitted between the root complex and the endpoint via the toolbox in response to a difference in bandwidth between the first link and the second link.

11. The toolbox according to claim 6, further comprising a control module, the control module being configured to: Detect transient error conditions on one of the first or second links; and In response to the transient error condition, a negative acknowledgment (NAK) signal is sent on the other link of the first link or the second link.

12. The toolbox of claim 6 further includes a control module configured to enable a low-power state of one of the first or second links in the event that no Transaction Layer Packet (TLP) is present in either the first or second link.

13. The toolbox of claim 12, wherein the control module is configured to: Initiate a low-power state request for the root complex or the endpoint that can connect to the first link or the second link; and In response to the rejection of the low power state request, one of the first or second links is switched back to the active state.

14. The toolbox of claim 12, wherein the control module is configured to exit the low-power state in response to a request to the root complex or the endpoint.

15. The toolbox of claim 2, wherein the bandwidth on the first link and the bandwidth on the second link are the same.

16. The toolbox of claim 2, wherein the data rate at the first port is different from the data rate at the second port.

17. The toolbox of claim 2, wherein the data rate at the first port is the same as the data rate at the second port.

18. The toolbox of claim 1, wherein the toolbox does not include a transaction layer.

19. The toolbox of claim 1, wherein the toolbox is configured to transmit data conforming to the PCIe standard for high-speed peripheral component interconnect between the root complex and the endpoint.

20. A computing system for transmitting data conforming to the PCIe (Peripheral Component Interconnect) standard, the computing system comprising: The root complex is configured to connect to the processor and memory; Endpoint; as well as A toolbox, connected between the root complex and the endpoint, the toolbox including a first port connected to the root complex and a second port connected to the endpoint, the toolbox being configured to transmit data conforming to the High Speed ​​Peripheral Component Interconnect (PCIe) standard between the root complex and the endpoint.

21. The computing system of claim 20, wherein the toolbox comprises: The first physical layer is connected to the first port; The second physical layer is connected to the second port; First data link layer and second data link layer; The first data link layer is connected between the second data link layer and the first physical layer; The second data link layer is connected between the first data link layer and the second physical layer; as well as The first physical layer, the first data link layer, the second physical layer, and the second data link layer are configured to form one or more channels for transmitting data between the root complex and the endpoint.

22. The computing system of claim 21, wherein the first physical layer and the first data link layer form a first link, and the second physical layer and the second data link layer form a second link independent of the first link.

23. The computing system according to claim 21, wherein: The first physical layer includes a physical coding sublayer (PCS) connected to the first port and a media access control (MAC) sublayer connected to the first data link layer; The second physical layer includes a PCS connected to the second port and a MAC sublayer connected to the second data link layer; The PCS of the first physical layer is configured to provide a data interface between the first port and the MAC sublayer; as well as The PCS of the second physical layer is configured to provide a data interface between the second port and the MAC sublayer of the second physical layer.

24. The computing system according to claim 23, wherein: The MAC sublayer of the first physical layer includes a data path module between the PCS of the first physical layer and the first data link layer, and a first link training and state machine LTSSM module. The MAC sublayer of the second physical layer includes a data path module between the PCS of the second physical layer and the second data link layer, and a second LTSSM module; The first data link layer includes a first finite state machine (FSM) module configured to connect to the second link; and The second data link layer includes a second FSM module configured to connect to the first link.

25. The computing system according to claim 24, wherein: The second FSM module is configured to detect link interruption status on the second link; as well as In response to the link interruption condition, the first LTSSM module is configured to create a link interruption condition on the first link.

26. The computing system according to claim 24, wherein: The first FSM module is configured to detect link interruption status on the first link; as well as In response to the link interruption, the second LTSSM module is configured to enter a disabled state.

27. The computing system of claim 24, wherein the toolbox further comprises a control module, the control module being configured to: Determine whether the bandwidth on the first link and the bandwidth on the second link are the same; and In response to the fact that the bandwidth on the first link and the bandwidth on the second link are the same, data link layer packets (DLLPs) are allowed to be transmitted between the root complex and the endpoint via the toolbox.

28. The computing system of claim 27, wherein the control module is configured to prevent the DLLP from being transmitted between the root complex and the endpoint via the toolbox in response to a difference in bandwidth between the first link and the second link.

29. The computing system of claim 24, wherein the toolbox further comprises a control module, the control module being configured to: Detect transient error conditions on one of the first or second links; and In response to the transient error condition, a negative acknowledgment (NAK) signal is sent on the other link of the first link or the second link.

30. The computing system of claim 24, wherein the toolbox further comprises a control module, the control module being configured to: If there is no Transaction Layer Packet (TLP) in either the first link or the second link, enable the low-power state of either the first link or the second link. Initiate a low-power state request for the root complex or the endpoint connected to another link in the first link or the second link; as well as In response to the rejection of the low power state request, one of the first or second links is switched back to the active state.

31. The computing system of claim 30, wherein the control module is configured to exit the low-power state in response to a request for the root complex or the endpoint.

32. The computing system of claim 22, wherein the bandwidth on the first link and the bandwidth on the second link are the same.

33. The computing system of claim 22, wherein the data rate at the first port is different from the data rate at the second port.

34. The computing system of claim 22, wherein the data rate at the first port is the same as the data rate at the second port.

35. The computing system of claim 20, wherein the toolbox does not include a transaction layer.

36. A computing system for transmitting data, the computing system comprising: The root complex is configured to connect to the processor and memory; Endpoint; as well as A toolbox, connected between the root complex and the endpoint, includes: a first port connected to the root complex and a second port connected to the endpoint; a first physical layer connected to the first port and a second physical layer connected to the second port; and a first data link layer and a second data link layer, wherein the first data link layer is connected between the second data link layer and the first physical layer, and the second data link layer is connected between the first data link layer and the second physical layer. The first physical layer, the first data link layer, the second physical layer, and the second data link layer are configured to form one or more channels for transmitting data between the root complex and the endpoint.

37. The computing system of claim 36, wherein the data rate at the first port is different from the data rate at the second port.

38. The computing system of claim 36, wherein the data rate at the first port is the same as the data rate at the second port.