A PCIe switching structure and switching chip

By designing a PCIe switching structure including PCIe controller, Router routing, Crossbar switch, DMA engine and NOC, the problem of waste of resources and insufficient flexibility in the existing DMA implementation solutions is solved, and flexible mounting of the DMA engine and data migration between multiple ports are realized.

CN119484435BActive Publication Date: 2025-05-09WELL CORE MICROELECTRONICS TECH (TIANJIN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510065444.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-09
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

In existing PCIe-related DMA implementation solutions, DMA engines can only be mounted on specific ports, resulting in waste of resources and increased costs, and it is difficult to achieve the flexibility of mounting any port.

Method used

A PCIe switching structure is designed, including PCIe controller, Router routing, Crossbar switch, DMA engine and NOC. The configuration space of P2P Function and DMA Function realizes flexible mounting of the DMA engine, and supports data transfer within any switching port or between any two switching ports.

Benefits of technology

In the PCIe switching implementation that supports multiple virtual switches, there is no need to implement multiple DMA modules, and DMA is flexible mountable as a Function in the PCIe structure, maximizes the use of resources, and supports data migration between arbitrary ports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119484435B_ABST
    Figure CN119484435B_ABST
Patent Text Reader

Abstract

The present invention provides a PCIe switching structure and a switching chip, which includes: a PCIe controller, multiple Routers, a Crossbar switch, a DMA engine and a NOC. The Crossbar switch is connected to multiple Routers; the PCIe controller is connected to the Crossbar switch through the Router; a DMA Function and / or a P2P Function are provided in the PCIe controller; and the DMA engine is connected to the Crossbar switch through a Router. In the PCIe switching supporting multiple virtual switches, this solution does not require multiple DMA modules, and DMA can be used as a Function in the PCIe structure to be attached to any virtual switch upstream port as needed, supporting data movement within any switch port or between any two switch ports.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of integrated circuit design, data transmission and data communication, and in particular to a PCIe switching structure and a switching chip, the implementation of which includes but is not limited to ASIC implementation or FPGA implementation. Background Art

[0002] A PCIe switch chip is an integrated circuit specifically designed to enhance the interconnection capabilities of the PCIe bus architecture in a computer system. It acts as a communication bridge between PCIe devices, allowing multiple PCIe devices to communicate in a system with higher bandwidth and lower latency.

[0003] The typical PCIe switch application topology is a tree structure. The device at the top (Root Port) is usually the CPU, which runs the PCIe management software and endpoint drivers. The port connected to it in the PCIe switch is called the Upstream Port. The devices at the end of the tree structure are usually various endpoint devices, and the ports connected to the endpoints are called Downstream Ports. In a switch, there is only one UP port and multiple DP ports.

[0004] DMA is a data transfer method that allows data to be exchanged directly between certain hardware subsystems and main memory without the direct involvement of the CPU. This technology greatly improves the speed and efficiency of data transfer because the CPU can continue to perform other tasks without having to wait for the data transfer to complete.

[0005] As a multi-port device, a PCIe switch can connect multiple PCIe links to enable communication between devices. Inside a PCIe switch, DMA can be used to transfer data between different PCIe devices. For example, GPUs, network interface cards (NICs), or storage controllers can read or write memory directly through DMA.

[0006] Existing PCIe-related DMA implementation solutions are all based on PCIe endpoints. Among the solutions mainly adopted in the prior art, some DMA engines are mounted as endpoints under the DP port of the PCIe exchange. In this case, the mounted DP port can only be used for DMA function, which wastes the expansion port; some DMA engines can only be mounted on specific UP ports. If the flexibility of mounting any port is to be achieved, it needs to be implemented through multiple DMA modules, which wastes resources and increases costs. Summary of the invention

[0007] In order to solve at least some of the problems existing in the above-mentioned prior art, this solution proposes a PCIe switch chip DMA structure and a switch chip, so as to realize that in the PCIe switch implementation supporting multiple virtual switches (Virtual Switch), there is no need to implement multiple DMA modules. Specifically, the present invention provides the following technical solutions:

[0008] On the one hand, the present invention provides a PCIe switching structure, which includes: a PCIe controller, multiple Routers, a Crossbar switch, a DMA engine and a NOC;

[0009] The Crossbar switch connects multiple Routers and exchanges data with each Router. The PCIe controller connects to the Crossbar switch through the Router, and the Router exchanges data with the PCIe controller. The PCIe controller is provided with a P2P Function, or a DMA Function and a P2P Function. The DMA engine is connected to the Crossbar switch through a Router.

[0010] The PCIe controller is connected to the NOC and exchanges data;

[0011] The P2P Function has an independent configuration space; the DMA Function is mapped to the configuration space of the DMA engine through the NOC.

[0012] Preferably, the PCIe controller includes: a physical layer, a data link layer and a transaction layer;

[0013] The physical layer is connected to the data link layer and exchanges data; the data link layer is connected to the transaction layer and exchanges data; the transaction layer is connected to the NOC and Router.

[0014] The transaction layer includes a TLP processing unit, an RX buffer, a TX buffer, a Conf Mux unit, and a PCIe Config0 unit; the TLP processing unit is connected to the data link layer and performs data interaction, the TLP processing unit is connected to the RX buffer and the TX buffer, the RX buffer and the TX buffer are connected to the Conf Mux unit and the Router, the Conf Mux unit and the PCIe Config0 unit are connected to the NOC respectively; the PCIe Config0 unit is the configuration space of the P2P Function;

[0015] The Conf Mux unit receives the configuration message in the RX buffer and sends the Type0 configuration message to the NOC or PCIe Config0 unit based on the Function ID in the Type0 configuration message.

[0016] The Type 0 configuration message is used to configure a DMA engine or a P2P Function; the Function ID is used to determine whether it is a configuration message of a DMA Function or a configuration message of a P2P Function.

[0017] Preferably, the TLP processing unit receives and parses the TLP message, and sends the parsed TLP message to the RX buffer; after the TLP message is read out from the RX buffer, if it is determined to be a configuration message of configuration information, it is sent to the Conf Mux unit.

[0018] Preferably, in the Conf Mux unit, before initiating access to the configuration space of the P2P Function or the configuration space of the DMA engine, the relevant information in the configuration message is first converted into a type that is convenient for switching and forwarding on the NOC, and the relevant information includes read / write information, the address in the configuration message, and the data in the configuration message.

[0019] Preferably, the Router routing is used to implement address routing and ID routing of the PCIe protocol;

[0020] The address routing is used to request the routing of the TLP message, and the value of the Address field in the Header of the TLP message data packet is compared with the Base and Limit fields in the configuration space of each P2P Function and each DMA Function to determine the port to which the message is sent;

[0021] The ID routing is used to configure the access request, using the Bus Number field of the Requester ID in the Header of the CPL / CPLD message data packet to compare with the Secondary Bus Number field and the Subordinate Bus Number field in each port to determine the port to which the request is sent;

[0022] When the Bus Number field of the Requester ID is equal to the Bus Number field of the port where the DMA Function is located, the Function Number (ie, function serial number) field of the DMA Function is further checked to determine whether to send it to the port where the DMA engine is located.

[0023] Preferably, the Crossbar switch is used to implement a point-to-point network. After the Router determines the destination port to be sent to, the TLP message is forwarded between each port and the DMA engine through the Crossbar switch.

[0024] Preferably, the configuration space register of the DMA engine and all the port configuration registers of the PCIe switch structure are uniformly encoded; each PCIe controller and the NOC are connected via an access master port and an access slave port;

[0025] The access to the master port is initiated by the Conf MUX unit in the PCIe controller and is used to access the configuration space register of the DMA engine or the configuration space register of the P2P Function of other ports; the access to the slave port is used by other master ports to access the P2P Function configuration space register inside the port; the other master port refers to the UP port that can access all registers of the chip where the PCIe switching structure is located.

[0026] Preferably, when the upstream port (i.e., Upstream UP) receives a Cfg1 type message, and the BusNumber thereof is equal to the Secondary Bus Number of the UP port, it will be sent to the Conf Mux module, converted into a NOC format (such as an AHB interface format) through a unified encoding, and sent to the NOC.

[0027] For each Function with a 4KB space, it is uniformly encoded in units of 4KB according to the Device Number in the Cfg1 message; the space for accessing the DMA Function is also encoded into the unified NOC network with a 4KB offset; so that the NOC can forward to the corresponding port according to the address above 4KB, including the port where the DMA engine is located.

[0028] Preferably, the controller register of the DMA engine is implemented in the configuration space of the DMA Function; the main control device controls the DMA data movement by accessing the controller register of the DMA engine.

[0029] Preferably, by configuring the Multi-Function Device field in the Header Type register in the general configuration space of the PCIe Type 0 / 1 of the UP port to 1, the multi-Function function of the UP port is enabled, so as to discover the DMA Function of the DMA engine during the operation of the PCIe switch structure;

[0030] When the DMA Function is discovered, the DMA driver starts to be loaded to configure the DMA engine to implement the data moving function.

[0031] Preferably, the RC (root complex, i.e., the root complex device in the PCIe system) device accesses the DMA Function through a Cfg0 message with a FunctionNumber field of 1; the Cfg0 message with a Function Number field of 1 is transmitted to the configuration space of the DMA Function through the NOC; before the Cfg0 message enters the NOC, it is converted into a type that is convenient for transmission on the NOC through a Conf Mux unit.

[0032] Preferably, the process of data movement by the DMA engine is:

[0033] S1, configure the descriptor information into the configuration space of the DMA Function and complete the configuration, then start the DMA engine;

[0034] S2. The DMA engine initiates an MRd read request to the source register based on the address of the data to be moved in the descriptor. The MRd read request carries the BDF information of the DMA Function. The Router routes the message to the corresponding port based on the address information in the MRd.

[0035] S3, the device attached to the source register responds to the MRd read request and returns the CPLD data. The Requester ID in the CPLD data will be consistent with the MRd read request. The Router forwards the CPLD data to the DMA engine based on the Requester ID.

[0036] S4. The DMA engine takes out the Data field in the CPLD data and constructs an MWr write request. The Router routes the message to the corresponding port according to the address information in the MWr write request, and the device connected to the port writes the data to the destination memory.

[0037] On the other hand, the present invention further provides a PCIe switching chip, wherein the switching chip includes the PCIe switching structure as described above.

[0038] Compared with the prior art, this solution has at least the following beneficial effects:

[0039] In the PCIe switching implementation that supports multiple virtual switches, there is no need to implement multiple DMA modules. Instead, DMA can be used as a Function in the PCIe structure to attach to any virtual switch upstream port as needed. This implementation can achieve maximum flexibility under limited resource consumption (the implementation of each DMA function consumes hardware resources). This solution supports data movement within any switch port or between any two switch ports. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 is a structural block diagram of an embodiment of the present invention;

[0042] Figure 2 A schematic diagram of a PCIe controller implementation structure according to an embodiment of the present invention;

[0043] Figure 3a This is an example of the request header format for 64-bit addressing;

[0044] Figure 3b This is an example of the request header format for 32-bit addressing;

[0045] Figure 4 A schematic diagram of a PCIe completion header format specified by the PCIe protocol according to an embodiment of the present invention;

[0046] Figure 5a Example of the ID routing field of the 4DW header specified for the PCIe protocol;

[0047] Figure 5b Example of ID routing field of 3DW header specified for PCIe protocol;

[0048] Figure 6 A Type 1 configuration space header specified in the PCIe protocol of an embodiment of the present invention;

[0049] Figure 7 A schematic diagram of NOC network connection according to an embodiment of the present invention;

[0050] Figure 8 A schematic diagram of discovering multiple Functions according to an embodiment of the present invention;

[0051] Fig. 9 The figure is a schematic diagram of the DMA engine data moving process according to an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. It should be clear that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] Those skilled in the art should know that the following specific embodiments or specific implementations are a series of optimized settings listed by the present invention to further explain the specific content of the invention, and these settings can be combined or used in association with each other, unless the present invention clearly states that some or a specific embodiment or implementation cannot be associated or used together with other embodiments or implementations. At the same time, the following specific embodiments or implementations are only used as the most optimized settings, and are not to be understood as limiting the scope of protection of the present invention.

[0054] In view of the problems existing in the prior art, this solution does not need to implement multiple DMA modules in the PCIe switching implementation that supports multiple virtual switches (i.e., Virtual Switch). Instead, DMA can be used as a Function in the PCIe structure to be attached to any virtual switch upstream port as needed. This implementation can achieve maximum flexibility under limited resource consumption (the implementation of each DMA function consumes hardware resources). It supports data movement within any switch port or between any two switch ports.

[0055] The so-called virtual switch refers to a function implemented inside a PCIe switch, which allows multiple physical ports (such as upstream and downstream ports) of a switch chip to be organized into multiple independent virtual PCIe domains or networks, thereby providing more fine-grained traffic management and resource isolation. When a switch chip is configured as multiple virtual switches, it is equivalent to having multiple logically independent switch chips.

[0056] In a specific embodiment, the overall structure of this solution is as follows Figure 1 As shown, the overall structure includes one or more PCIe controllers (i.e., PCIe Controller), one or more Routers, Crossbar switches, and a NOC (network on chip) on-chip configuration network, wherein the PCIe controller is provided with a P2P Function or a DMA Function and a P2P Function are provided at the same time. Figure 1 The light-colored part represents the Function attached to the PCIe controller, Figure 1 The DMA part and the P2P part in the figure, where the P2P Function is fixed and marked with a solid line; the DM AFunction is selected according to the configuration and is optional, marked with a dotted line.

[0057] Further explanation is that P2P function and DMA function can actually be configuration spaces accessible to RC devices. If they can be accessed, the corresponding function is considered to exist. The configuration space of DMA function can be based on the configuration of NOC. The port to which it is mounted can make the master port of that port access the DMA function.

[0058] In this embodiment, the structure including multiple PCIe controllers, multiple Router routes, multiple DMA Functions and multiple P2P Functions is used as an example for explanation. In the preferred implementation structure, the Crossbar switch is connected to multiple Router routes, and data is exchanged between the Crossbar switch and the Router routes; the Router route is connected to the PCIe controller, and each PCIe controller is provided with a P2P Function and a DMA Function, or only one P2PFunction is provided. The DMA Function is connected to the NOC and exchanges configuration information data. Data is exchanged between the Router route and the PCIe controller, and more specifically, data is exchanged between the Router route and the TLP processing unit in the PCIe controller.

[0059] In this embodiment, in the PCIe protocol, "Function" refers to one or more logical units within a PCIe device. Each Function can be regarded as an independent device with its own set of configuration space and functions. A physical PCIe device can contain one or more Functions, depending on the design and functional complexity of the device.

[0060] 1. The following describes the main functional modules in the solution in combination with this embodiment.

[0061] 1. PCIe Controller: The PCIe controller is mainly used to understand and execute the PCIe protocol, implement the functions of the physical layer, data link layer and transport layer specified in the PCIe protocol, and ensure that the communication between all devices follows the specified rules and standards. The implementation structure of the PCIe controller in this embodiment is as follows: Figure 2As shown in the figure, in addition to implementing the PCI-to-PCI bridge type function, the transport layer can also optionally implement the DMA function, which is the DMA function mentioned above. The P2P type function is a Type 1 function, which, as the name implies, is a function with a Type 1 configuration space header; the DMA function is a Type 0 function with a Type 0 configuration space header, which is commonly known as an endpoint. Figure 2 The PCIe Confg0 in the figure is the configuration space of P2PFunction; the configuration space corresponding to the DMA Function is accessed through NOC.

[0062] The PCIe controller structure mainly includes the physical layer (i.e., PHYMac Layer), the data link layer (i.e., Data Link Layer), and the transaction layer (i.e., Transaction Layer), wherein the physical layer is connected to the data link layer and performs data exchange, the data link layer is connected to the transaction layer and performs data exchange, and the transaction layer is connected to the NOC. The transaction layer includes the TLP processing unit (i.e., TLP Process), RX / TX buffer, Conf Mux unit, and PCIe Config0 unit. The TLP processing unit is connected to the data link layer and performs data exchange, the TLP processing unit is connected to the RX / TX buffer, the RX / TX buffer is connected to the Conf Mux unit and the PCIe Config0 unit, the Conf Mux unit and the PCIe Config0 unit are connected to each other and perform data exchange, and the Conf Mux unit and the PCIe Config0 unit are respectively connected to the NOC. It is explained here that the Conf Mux unit is a configuration access split unit, which splits the access to each configuration space register and combines the response; PCIe Config0 is the Function0 configuration space of PCIe, which corresponds to the P2P Function.

[0063] In the transaction layer logic, after being decoded by the TLP processing unit, the correct TLP message (i.e., transport layer message) is parsed and enters the RX buffer. After the TLP message is read out from the RX buffer, it is determined that the Bus field in the Cfg1 (i.e., Type1 Configuration Request) message is equal to the Cfg1 in the Secondary Bus Number register value of the UP port, and it is forwarded to the Conf Mux unit. For other messages, they are transparently transmitted, that is, forwarded through Router routing and Crossbar; the configuration access message will be distributed in the Conf Mux unit. If it is to access Function 0, it will be sent to the PCIe Config0 unit, which contains the Type1 configuration access space of the P2P Function; if it is to access Function 1, it will be converted into a bus format and sent to the external NOC bus, and forwarded to the DMA configuration space through the NOC bus. If it is determined that it is not a Cfg0 configuration message after reading from the RX buffer, it will be transparently transmitted, that is, forwarded through the Router and Crossbar switch; in addition, it is further explained that the non-transparently transmitted messages, in addition to the Cfg0 configuration message, also include the Cfg1 configuration message whose Bus field is equal to the Secondary Bus Number register value of the UP port. It is further explained here that Function 0 refers to the FunctionNumber field in the Config0 TLP message is 0, which is used to access the P2P configuration space register of the PCIe Config0 module, which is a Type 1 configuration space; Function 1 refers to the Function Number field in Config0 is 1, which corresponds to the DMA configuration space register, which is a Type 0 configuration space.

[0064] In the Conf Mux unit, before initiating access to the P2P configuration space or DMA configuration space, the access message for accessing the P2P configuration space or DMA configuration space will first be converted into APB / AHB / AXI and other bus types that are convenient for switching and forwarding on the NOC. The most commonly used is the AHB bus. It is further explained that this access message is not forwarded and terminated inside the chip. For example, the specific conversion access is to piece together the {Function Number, Device Number, Ext Reg Number, RegisterNumber, 2'00} fields in the Config message as the address of the AHB bus, and Payload as the access data.

[0065] Furthermore, the judgment and processing method of the Cfg1 message received by the UP to access the P2P Function configuration space of other DP ports is as follows:

[0066] When the Upstream UP receives a Cfg1 type message, and the Bus Number field is equal to the Secondary Bus Number field of the UP port, it will also be sent to the Conf Mux module, converted into the NOC format (such as the AHB interface format) through a unified encoding and sent to the NOC. For each Function with a 4KB space, it is uniformly encoded in units of 4KB according to the Device Number in the Cfg1 message; the space for accessing the DMA Function is also encoded with 4KB as an offset to enter the unified NOC network, so that the NOC can forward it to the corresponding port according to the address above 4KB, including the port where the DMA engine is located.

[0067] In this embodiment, Function mainly includes three types:

[0068] Type 0 Function: Represents a standard PCIe device, such as a graphics card, network adapter, etc.

[0069] Type 1 Function: This function is usually used for bridge devices, such as PCIe to PCI bridge.

[0070] Type 2 Function: Mainly represents a specific type of device, such as certain types of bridges or special-purpose devices.

[0071] It is further explained here that Figure 1 The P2P (i.e. P2P Function) is a Type 1 Function; the DMA Function is a Type 0 Function.

[0072] 2. Router routing: implement various routing methods specified by the PCIe protocol. In this embodiment, address routing and ID routing are mainly implemented.

[0073] Address routing, as the name implies, is address-based routing, which is used to request the routing of TLP messages, including MRd and MWr messages used in DMA data movement. Figure 3a and Figure 3bAs shown, the Header of these two data packets contains an Address field. In the Router routing of this embodiment, the value of this field is compared with the Base and Limit fields in each P2P configuration space to ultimately determine the port to which the packet is sent. Here, an example is used to illustrate address routing:

[0074] Assume a basic PCIe topology, which includes a Root Complex (RC), a PCIe switch, and two end points (EPs): EP1 is a network adapter and EP2 is a graphics card.

[0075] System initialization:

[0076] When the operating system starts, it will traverse all connected PCIe devices and allocate address space to them. For EP1 (network adapter), the operating system allocates memory mapping address space 0x4000_0000 - 0x40FF_FFFF. For EP2 (graphics card), the operating system allocates memory mapping address space 0x8000_0000 - 0x8FFF_FFFF.

[0077] Configure BAR:

[0078] The system writes the above address range into the BAR of the corresponding device and the PCIe switch port connected to it. For example, for EP1, the DP connected to it is set to: Base: 0x4000_0000; Limit: 0x40FF_FFFF. For EP2, the DP connected to it is set to: Base: 0x8000_0000; Limit: 0x8FFF_FFFF.

[0079] Address routing process:

[0080] Assume that the CPU needs to send some control information to EP1, and it constructs a Memory Write TLP message with the destination address 0x4000_1234. When the TLP message arrives at the PCIe switch, the Router checks whether the address belongs to the known device address range. According to the configuration, the Router knows that 0x4000_1234 falls within the address range of EP1 (0x4000_0000 - 0x40FF_FFFF), so the Router forwards this TLP message to the switch port connected to EP1. After receiving the TLP message, EP1 recognizes that this is a request for itself. After processing the data, if necessary, it can generate an appropriate Completion TLP message to return to the initiator.

[0081] ID routing, as the name implies, is routing based on ID. It is used to configure access requests and complete the routing of TLP and other messages, including CPL / CPLD messages encountered in DMA data movement. Figure 3a , 3b and Figure 4 As shown in the figure, the header of the completed data packet contains the Requester ID field, which is used during the routing process. Figure 5a , Figure 5b The Bus Number field is compared to determine the port to which the packet is sent. Figure 5a , 5b The Bus Number field in Figure 6 The secondary bus number (Secondary Bus) and subordinate bus number (Subordinate Bus) registers in the P2P configuration space of each port are compared.

[0082] 3. Crossbar switching: After the Router determines the destination port to be sent, the TLP packet is forwarded between the various ports of the Crossbar switch and the DMA engine during the switching process. The Crossbar (cross switch) switching structure is used to implement a point-to-point network architecture. The Crossbar switching structure can provide non-blocking direct connection, which allows any input end to be directly and independently connected to any output end. It should be noted here that the internal structure of the Crossbar switch can be implemented through existing technologies and is not understood as the inventive point of the present invention.

[0083] 4. NOC: The full name is Network-on-Chip, which is a communication architecture used in integrated circuit design to solve the data communication problems between multi-core processors, SoC (System-on-a-Chip) and large-scale integrated circuits. Figure 7As shown, the NOC in this embodiment is a simplified implementation, which only needs to support 4-byte read and write access, so a simple transmission protocol such as APB or AHB can be used. In the application of this embodiment, the port configuration registers of the whole chip, including the configuration space registers of DMA, are uniformly encoded. For example, in the PCIe exchange of N ports, the P2P configuration space of Port0 occupies [0-4KB), the P2P configuration space of PortM occupies [(M)*4KB, (M+1)*4KB), the P2P configuration space of PortN-1 occupies [(N-1)*4KB, N*4KB), the first DMA configuration space occupies [(N)*4KB, (N+1)*4KB), and the second DMA configuration space occupies [(N+1)*4KB, (N+2)*4KB). In this embodiment, since access in NOC must be routed based on the address, and in the original Cfg message, access is required based on Device Num, Function Num and Register Number, and this traditional access method cannot be well applied to the solution of the present invention, this solution adopts a unified coding method for conversion.

[0084] Each PCIe controller and NOC are connected through an access master port and an access slave port, and the interface is in the form of APB / AHB bus. The access to the master port is initiated by the Conf MUX unit in the PCIe controller, which is used to access the DMA configuration space and the P2P space of the DP port. The initiator of the access to the master port is usually the Upstream port in the PCIe network, which can access the configuration space of the Downstream port. The access slave port is used by other master ports to access the configuration space inside the port (the port and the DP port). The other master port is the UP port (i.e., Upstream port) that can access the registers of the entire chip, such as IIC, JTAG, etc., or it can be an upstream port with full chip access rights. In actual applications, the Bar0 space of a certain upstream port (Upstream Port) can also be opened to manage the registers of the entire chip. It is further explained here that accessing the configuration space inside the port by accessing the slave port mainly includes: RC accesses the configuration space of UP and DP in the PCIe switch through Cfg messages, among which Cfg0 (Type0 Configuration Request) messages access the UP port and Cfg1 messages access the Dp port; Cfg1 can also access downstream devices, and can determine whether to access the DP port of this chip through the Bus number. If the DP port of this chip is not accessed, it is transparently transmitted to the downstream device through the Router and Crossbar switch.

[0085] 5. DMA engine: implements DMA-related functional logic. In this embodiment, the controller register of the DMA engine is implemented in the configuration space of the DMA Function; further, the DMA Function is a set of configuration spaces, including two parts, one is the PCIe protocol part presented to the RC as specified by the protocol, and the other is the configuration registers used for DMA work; the DMA engine works under the configuration guidance of the configuration registers used for DMA work, such as the source and destination ports, length, etc. when moving data, all of which need to be configured in the DMA Function. The DMA engine is the functional implementation logic of the DMA, and works under the configuration of the internal registers of the DMA Function; the DMA Function and the DMA engine together constitute and implement the DMA function.

[0086] The overall function of DMA is Figure 1 The DMA engine (i.e. DMA Engine) and DMA (i.e. DMA Function) in the lower right corner are composed. As mentioned above, DMA Function is a set of configuration spaces, which contains the configuration information and control information required for DMA work; DMA engine implements the functional logic of DMA work. DMA Function is mounted on the UP port as Function 1, and the upstream RC device accesses DMA Function through Cfg0 message with Function Number 1; the above Cfg0 message will be converted into AHB access in the Cfg Mux unit and reach the DMA Function configuration space through the NOC network. After configuring the relevant information of DMA work, the DMA engine starts to work and generates access requests to be forwarded through Router and Crossbar to perform DMA data movement function.

[0087] The master device controls DMA data movement by accessing the registers of these configuration spaces. The DMA engine can support immediate-value-based movement, descriptor-based data movement, and descriptor-chain-based data movement. Immediate-value-based movement is to put the moved data and destination address into the register, and then start DMA to initiate the movement. This application is usually suitable for a small amount of data movement; descriptor-based data movement is to configure the DMA descriptor (including the source address, destination address, length, interrupt mode and other information of the data to be moved) into the register, and then start DMA to initiate the movement. This application is usually suitable for data movement in a data set; descriptor-chain-based data movement is to store the descriptor information in the memory space of the switching connection device. After starting the DMA movement, DMA first moves the descriptor from the memory to the inside of the chip, and then identifies the information therein for data movement. In the data movement of the descriptor chain, in addition to the conventional descriptor information, it can also contain the address of the next descriptor. When the data indicated by the descriptor is moved, the next description is automatically read for processing until the address of the next descriptor contained in the descriptor is invalid.

[0088] 2. The working process of the DMA structure of this embodiment is explained below.

[0089] 1. Enable the multi-function function of a port by configuring the Multi-Function Device field of the Header Type register in the common configuration space (i.e. Common ConfigurationSpace) of the PCIe Type 0 / 1 of a UP port to 1. In this way, the DMA endpoint Function will be discovered during the software enumeration process. The software (i.e., the software running on the CPU and other devices connected to the UP port) enumeration process is not described here. For the implementation structure in this embodiment, the configuration message sent by the software pointing to Function 1 is forwarded to the DMA configuration space through the NOC, so that the software can discover the DMAFunction and thus have the ability to manage the DMA engine.

[0090] When DMA is configured on a UP port, from the perspective of software enumeration, it will be found that there are multiple functions on the UP port, such as Figure 8 The P2P bridge function (i.e., P2P bridge function) and DMA function (i.e., DMA function) in the diagram are as follows Figure 8 shown.

[0091] 2. After the software discovers the DMA Function, it will load the DMA driver, thereby having the function of configuring the DMA engine to realize data movement.

[0092] 3. The process of DMA data transfer is as follows Fig. 9 As shown:

[0093] Step 1: The software configures the descriptor information into the controller registers related to the DMA engine in the configuration space (ConfigurationSpace) of the DMA Function, and starts the DMA engine. The format of the descriptor and the content of the register are not the focus of this article, so they will not be described here;

[0094] Step 2: The DMA engine initiates an MRd read request to SourceMemory (i.e., source register) based on the Source Address (address of the data to be moved) in the descriptor. The Requester ID field in the MRd request will carry the BDF information of the DMAFunction (i.e., Bus Number (i.e., bus number, assigned during enumeration), Device Number (i.e., device number, usually fixed by hardware and consistent with the port number), Function Number (i.e., function number, value is 1)); the Router routing module will route the message to the corresponding PCIe controller based on the Address information in MRd;

[0095] Step 3: The device attached to the Source Memory responds to the MRd read request and returns the CPLD data. The Requester ID in the CPLD will be consistent with the content in the MRd. The Router module forwards the CPLD message data to the DMA engine according to the Requester ID. It should be further explained that the complete CPLD message is forwarded here, and the payload data is included in the CPLD message.

[0096] Step 4: The DMA engine takes the Data field in the CPLD data and constructs a MWr write request based on the DestinationAddress field in the descriptor; similarly, the Router will route the message to the corresponding PCIe port (such as Figure 1 ), and finally the PCIe controller, then enters Figure 2 The data is then sent to the connected device after being processed by the transaction layer, data link layer, and physical layer. Finally, the device connected to the port writes the data to the destination memory.

[0097] The DMA implementation structure of the present solution is described in detail above. In addition, in another embodiment, the present solution further provides a PCIe switch chip, in which the DMA structure described above is provided to support the functions of the switch chip.

[0098] In addition, the above-mentioned switching chip is also provided with other conventional structures that cooperate with the DMA structure to support the normal operation of the chip, such as multiple types of external data interfaces, other necessary registers, etc.

[0099] In another implementation mode, the present solution can be implemented by a device, which may include one or more modules in the system modules in the above embodiments. The module may be one or more hardware modules specially configured to perform the corresponding steps, or implemented by a processor configured to perform the corresponding steps, or stored in a computer-readable medium for implementation by a processor, or implemented by some combination.

[0100] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A PCIe switching structure, characterized in that: The switching structure includes: a high-speed peripheral component interconnect PCIe controller, multiple routing routers, a crossbar switch, a direct memory access DMA engine and an on-chip network NOC; The crossbar switch connects multiple routing routers and exchanges data with each routing router; the high-speed peripheral component interconnection PCIe controller connects to the crossbar switch through the routing router, and the routing router exchanges data with the high-speed peripheral component interconnection PCIe controller; the high-speed peripheral component interconnection PCIe controller is provided with a point-to-point function component P2P Function, or the high-speed peripheral component interconnection PCIe controller is provided with a direct memory access function component DMA Function and a point-to-point function component P2P Function; the direct memory access DMA engine is connected to the crossbar switch through a routing router; The high-speed peripheral component interconnection PCIe controller connects and exchanges data with the on-chip network NOC; The point-to-point function component P2P Function has an independent configuration space; the direct memory access function component DMAFunction is mapped to the configuration space of the direct memory access DMA engine through the on-chip network NOC; The high-speed peripheral component interconnect (PCIe) controller includes: a physical layer, a data link layer and a transaction layer; The physical layer is connected to the data link layer and exchanges data; the data link layer is connected to the transaction layer and exchanges data; the transaction layer is connected to the on-chip network NOC and the router; The transaction layer includes a transaction layer data packet TLP processing unit, a receiving RX buffer, a sending TX buffer, a configuration multiplexing Conf Mux unit and a high-speed peripheral component interconnection configuration PCIe Config0 unit; the transaction layer data packet TLP processing unit is connected to the data link layer and performs data interaction, the transaction layer data packet TLP processing unit is connected to the receiving RX buffer and the sending TX buffer, the receiving RX buffer and the sending TX buffer are connected to the configuration multiplexing Conf Mux unit and the routing Router, the configuration multiplexing Conf Mux unit and the high-speed peripheral component interconnection configuration PCIe Config0 unit are respectively connected to the on-chip network NOC; the high-speed peripheral component interconnection configuration PCIe Config0 unit is the configuration space of the point-to-point functional component P2P Function; The configuration multiplexing Conf Mux unit receives the configuration message in the RX buffer, and sends the Type0 type configuration message to the on-chip network NOC or the high-speed peripheral component interconnect configuration PCIe Config0 unit based on the function number Function ID in the Type0 type configuration message; The Type0 configuration message is used to configure a direct memory access DMA engine or a point-to-point function component P2PFunction; the function number Function ID is used to determine whether it is a configuration message of a direct memory access function component DMA Function or a configuration message of a point-to-point function component P2P Function.

2. The PCIe switching structure according to claim 1, characterized in that: The transaction layer data packet TLP processing unit receives and parses the transaction layer data packet TLP message, and sends the parsed transaction layer data packet TLP message to the receiving RX buffer; After the transaction layer data packet TLP message is read out from the receiving RX buffer, if it is determined to be a configuration message of configuration information, it is sent to the configuration multiplexing Conf Mux unit.

3. The PCIe switching structure according to claim 2, characterized in that: In the configuration multiplexing Conf Mux unit, before initiating access to the configuration space of the point-to-point functional component P2P Function or the configuration space of the direct memory access DMA engine, the relevant information in the configuration message is first converted into a type that is convenient for switching and forwarding on the on-chip network NOC, and the relevant information includes read / write information, the address in the configuration message, and the data in the configuration message.

4. The PCIe switching structure according to claim 1, characterized in that: The routing Router is used to implement address routing and ID routing of the high-speed peripheral component interconnection PCIe protocol; The address routing is used to request the routing of the transaction layer data packet TLP message, and the value of the address field in the transaction layer data packet header TLP Header is compared with the base Base and limit Limit fields in the configuration space of each point-to-point function component P2P Function and each direct memory access function component DMA Function to determine the port to be sent; The ID routing is used to configure the access request, using the Bus Number field of the requester ID in the completion / data completion message header CPL / CPLDHeader to compare with the SecondaryBus Number field and the SubordinateBus Number field in each port to determine the port to which the request is sent; When the Bus Number field of the requester ID is equal to the Bus Number field of the port where the direct memory access function component DMAFunction is located, the Function Number field of the direct memory access function component DMA Function is further determined to determine whether to send to the port where the direct memory access DMA engine is located.

5. The PCIe switching structure according to claim 4, characterized in that: The crossbar switch is used to realize a point-to-point network. When the router determines the destination port to be sent, the transaction layer data packet TLP message is forwarded between each port and the direct memory access DMA engine through the crossbar switch.

6. The PCIe switching structure according to claim 5, characterized in that: The root component RC device accesses the direct memory access functional component DMA Function through a Class 0 configuration Cfg0 message with a function number Function Number field of 1; the Class 0 configuration Cfg0 message with a function number Function Number field of 1 is transmitted to the configuration space of the direct memory access functional component DMA Function through the on-chip bus NOC; before the Class 0 configuration Cfg0 message enters the on-chip bus NOC, it is converted into a type that is convenient for transmission on the on-chip bus NOC through the configuration multiplexing Conf Mux unit.

7. The PCIe switching structure according to claim 1, characterized in that: The configuration space register of the direct memory access DMA engine and all the port configuration registers of the PCIe switch structure are uniformly encoded; each high-speed peripheral component interconnection PCIe controller and the on-chip bus NOC are connected through an access master port and an access slave port; The access to the master port is initiated by the configuration multiplexing Conf MUX unit in the high-speed peripheral component interconnect PCIe controller, and is used to access the configuration space register of the direct memory access DMA engine, or access the configuration space register of the point-to-point functional component P2P Function of other ports; the access slave port is used by other master ports to access the configuration space register of the point-to-point functional component P2P Function inside the port; the other master port refers to the upstream UP port that can access all registers of the chip where the PCIe switching structure is located.

8. The PCIe switching structure according to claim 1, characterized in that: The controller register of the direct memory access DMA engine is implemented in the configuration space of the direct memory access functional component DMA Function; the main control device implements the control of direct memory access DMA data movement by accessing the controller register of the direct memory access DMA engine.

9. The PCIe switching structure according to claim 8, characterized in that: By configuring the Multi-Function Device field in the Header Type register in the general configuration space of the high-speed peripheral component interconnect 0 / 1 type PCIe Type 0 / 1 of the upstream UP port to 1, the multi-function Function of the upstream UP port is enabled to discover the direct memory access function component DMAFunction of the direct memory access DMA engine in the operation of the PCIe switch structure; When the direct memory access functional component DMA Function is discovered, the direct memory access DMA driver starts to be loaded to configure the direct memory access DMA engine to implement the data moving function.

10. A PCIe switching chip, characterized in that: The switching chip includes the PCIe switching structure as described in any one of claims 1-9.

Citation Information

Patent Citations

  • PCIE exchange chip port configuration system and method supporting virtual exchange

    CN111092773A

  • KR20240032680A