A PCIe system and a memory access method thereof

CN122527054BActive Publication Date: 2026-09-04SHENZHEN NANFEI MICROELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611015094.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-04
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

1、单点性能瓶颈:如图1所示,当设备EndPoint0访问设备EndPoint1时,数据流如箭头所示,所有设备的DMA流量都汇聚到RC侧的IOMMU进行转换,在高并发I/O场景下,IOMMU成为系统瓶颈,增加延迟并限制吞吐量

Benefits of technology

[0012] According to the PCIe system and its memory access method described in the above embodiments, the IOMMU function is distributed and offloaded to the PCIe switching chip. During initialization, the PCIe system's management module obtains the identity of the physical endpoint device connected to the PCIe switching chip, the physical address of the peer device, and the virtual address of the corresponding virtual endpoint device. This information is then sent as routing and authentication information to the IOMMU of the PCIe switching chip. When the PCIe switching chip receives a packet from a virtual device, the IOMMU can match the device identifier and virtual address in the packet with multiple routing and authentication information, thereby converting the virtual address of the destination device in the packet to its physical address, and then routing based on the physical address. It is evident that this method of accessing memory eliminates the need for packets to pass through the host (root node), requiring only the PCIe switching chip, thus reducing DMA latency and eliminating the performance bottleneck of the root complex.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122527054B_ABST
    Figure CN122527054B_ABST
Patent Text Reader

Abstract

The PCIe system and the memory access method thereof disclosed by the application distribute the IOMMU function to the PCIe switching chip; the management module of the PCIe system obtains the identity of the physical endpoint device connected to the PCIe switching chip, the physical address of the peer device and the virtual address of the virtual endpoint device corresponding to the peer device during initialization, and delivers the information as routing and verification information to the input / output memory management unit of the PCIe switching chip. Thus, the input / output memory management unit in the PCIe switching chip can match the device identity and the virtual address in the message in the plurality of routing and verification information after receiving the message of a device, so as to convert the virtual address of the destination device in the message to obtain the physical address, and then route based on the physical address. It can be seen that the message does not need to pass through the host (root node) in the memory access mode, but only needs to pass through the PCIe switching chip, so as to reduce the DMA delay and eliminate the performance bottleneck of the root complex.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of PCIe systems, and more specifically to a PCIe system and its memory access method. Background Technology

[0002] In computer systems, the efficiency and security of the interaction between I / O devices and main memory directly affect the overall system performance and reliability. Early systems used direct physical address access for I / O devices, meaning devices directly read and write to main memory via physical memory addresses. With the development of data centers, cloud computing, and high-performance computing, the performance and security requirements for I / O virtualization and device pass-through technologies have increased, leading to the emergence of IOMMU technology. In other words, in PCIe systems, physical endpoint devices need to be isolated. This is typically achieved by virtualizing virtual machines and virtual endpoint devices within the system. Physical endpoint devices only know about their own corresponding virtual endpoint devices and are unaware of the virtual endpoint devices corresponding to other physical endpoint devices. This ensures that security issues arising from interactions between virtual endpoint devices will not affect physical endpoint devices, thus improving system security. Physical endpoint devices can be devices that support the PCIe protocol and typically provide at least one of the functions required by a PCIe system (such as computing or storage). Operations between virtual endpoint devices (such as reading and writing data in each other's memory) ultimately need to be implemented on the physical endpoint devices (reading and writing data in their memory). The IOMMU (Input / Output Memory Management Unit) is a key component in a computer system responsible for managing direct memory access (DMA) of I / O devices. Its core function is to implement address translation (converting virtual addresses initiated by devices into system physical addresses) and DMA access permission control for PCIe (Peripheral Component Interconnect Express) virtual devices, preventing devices from illegally accessing memory and ensuring system security.

[0003] In a PCIe system, the IOMMU is located in the RC (Root Complex) or the host CPU. All DMA requests for PCIe virtual devices must be processed by the IOMMU on the RC or CPU side. Specifically, for example... Figure 1As shown, in the prior art, a packet sent from device EndPoint0 to device EndPoint1 first reaches the PCIe switch chip. The PCIe switch chip recognizes this as a packet with an untranslated address and redirects it to the RC (Redirecting Controller). Upon arrival at the RC, the packet first queries the page table based on the identifier and virtual address. If a match is found, the virtual address is translated to a physical address, and the packet address type is modified to the translated address. Then, the packet is forwarded from the uplink port to the PCIe switch chip, which routes it to device EndPoint1 based on the physical address. The PCIe switch chip primarily addresses the problem of insufficient CPU interfaces. The host CPU has a limited number of PCIe interfaces, making it impossible to directly connect all endpoint devices. The PCIe switch chip acts like an overpass, with one uplink port (upstream port) connecting to the CPU or the next-level PCIe switch chip, and multiple downlink ports (downstream ports) connecting to various endpoint devices such as graphics cards, NVMe hard drives, and network cards, thus greatly expanding the system's device connectivity. PCIe switching chips can also parse the destination of data packets, accurately forwarding data from the uplink port to the target downlink port, or directly forwarding it between different downlink ports (P2P). For data involving memory access, since virtual address to physical address translation is required, the data must be sent uplink to the host's IOMMU for address translation before the PCIe switching chip can forward it to the target downlink port.

[0004] However, with the development of cloud computing, edge computing, and other scenarios, the number of devices in virtualized environments has increased dramatically, and this centralized architecture has the following inherent drawbacks: 1. Single-point performance bottleneck: such as Figure 1 As shown, when device EndPoint0 accesses device EndPoint1, the data flow is as indicated by the arrow. All DMA traffic from all devices converges to the IOMMU on the RC side for conversion. In high-concurrency I / O scenarios, the IOMMU becomes a system bottleneck, increasing latency and limiting throughput.

[0005] 2. High latency: For devices connected downstream of remote switches, their DMA requests need to be transmitted through multiple switches to the root complex for conversion before returning, resulting in unnecessary round-trip latency.

[0006] Therefore, the memory access methods in the existing PCIe system still need to be improved and enhanced. Summary of the Invention

[0007] This invention provides a PCIe system and its memory access method, aiming to improve the efficiency of memory access in the PCIe system.

[0008] One embodiment provides a memory access method for a PCIe system. The PCIe system includes a host and one or more PCIe switching chips. The host is connected to the PCIe switching chips via a PCIe bus. The PCIe switching chips are equipped with an input / output memory management unit and a routing unit. The host is equipped with a management module. The method includes: The management module initializes the PCIe system to obtain the identity and physical address of the physical endpoint devices connected to the one or more PCIe switching chips, the virtual address of the virtual endpoint devices corresponding to the physical endpoint devices, and the topology of the PCIe system. The management module obtains the peer devices that the physical endpoint device can reach based on the topology; The management module sends the identity identifier of the physical endpoint device, the physical address of the peer device, and the virtual address of the virtual endpoint device corresponding to the peer device as routing and authentication information to the input / output memory management unit of the PCIe switching chip. After receiving a message sent by the first virtual endpoint device to operate on the second virtual endpoint device, the input / output memory management unit of the PCIe switching chip obtains the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device from the message. The input / output memory management unit uses the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device to match multiple routing and authentication information. If a match is found, the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched routing and authentication information. The routing unit of the PCIe switching chip routes the packet to the physical address in the packet.

[0009] One embodiment provides a PCIe system, including a host and one or more PCIe switching chips. The host is connected to the PCIe switching chips via a PCIe bus. The PCIe switching chips are provided with an input / output memory management unit and a routing unit. The host is provided with a management module. The host's management module is used to initialize the PCIe system, obtain the identity and physical address of the physical endpoint devices connected to the one or more PCIe switching chips, the virtual address of the virtual endpoint devices corresponding to the physical endpoint devices, and the topology of the PCIe system; obtain the peer devices that the physical endpoint devices can reach based on the topology; and send the identity of the physical endpoint devices, the physical addresses of the peer devices, and the virtual addresses of the virtual endpoint devices corresponding to the peer devices as routing and authentication information to the input / output memory management unit of the PCIe switching chip. The input / output memory management unit of the PCIe switching chip is used to, upon receiving a message sent by the first virtual endpoint device for operation on the second virtual endpoint device, obtain the identity identifier of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device from the message; match the identity identifier of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device in multiple routing and authentication information; if a match is found, modify the virtual address of the second virtual endpoint device in the message to the physical address in the matched routing and authentication information; The routing unit of the PCIe switching chip is used to route the packet to the physical address in the packet.

[0010] One embodiment provides a host for a PCIe system, the host being used to connect to one or more PCIe switching chips via a PCIe bus, including: a management module and a virtual machine; The management module is used to enumerate the physical endpoint devices connected to the one or more PCIe switching chips, obtain the identity, physical address, and topology of the physical endpoint devices, and the PCIe system topology; obtain the peer devices that the physical endpoint devices can reach based on the topology; pass the physical endpoint devices through to the virtual machine; simulate the virtual endpoint devices corresponding to the physical endpoint devices for the virtual machine; obtain the virtual address of the virtual endpoint devices corresponding to the physical endpoint devices from the virtual machine; and send the identity of the physical endpoint devices, the physical addresses of the peer devices, and the virtual addresses of the virtual endpoint devices corresponding to the peer devices as routing and authentication information to the input / output memory management unit of the PCIe switching chip, so that the input / output memory management unit can verify and route the packets received by the PCIe switching chip according to the routing and authentication information. The virtual machine is used to enumerate the virtual endpoint devices simulated by the management module and assign identity identifiers and virtual addresses to the virtual endpoint devices.

[0011] One embodiment provides a PCIe switching chip, comprising: Uplink port, used to connect to the host via the PCIe bus; Multiple downlink ports, which are used to connect to physical endpoint devices; An input / output memory management unit is configured to receive routing and authentication information sent by the host, the routing and authentication information including: the identity identifier of the physical endpoint device connected to the PCIe switching chip, the physical address of the peer device, and the virtual address of the virtual endpoint device corresponding to the peer device; the peer device is other physical endpoint devices that the physical endpoint device can be routed to; after receiving a message sent by the first virtual endpoint device to operate on the second virtual endpoint device, the unit obtains the identity identifier of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device from the message; the unit matches the identity identifier of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device in multiple routing and authentication information; if a match is found, the unit modifies the virtual address of the second virtual endpoint device in the message to the physical address in the matched routing and authentication information; A routing unit is used to route the message to the physical address in the message.

[0012] According to the PCIe system and its memory access method described in the above embodiments, the IOMMU function is distributed and offloaded to the PCIe switching chip. During initialization, the PCIe system's management module obtains the identity of the physical endpoint device connected to the PCIe switching chip, the physical address of the peer device, and the virtual address of the corresponding virtual endpoint device. This information is then sent as routing and authentication information to the IOMMU of the PCIe switching chip. When the PCIe switching chip receives a packet from a virtual device, the IOMMU can match the device identifier and virtual address in the packet with multiple routing and authentication information, thereby converting the virtual address of the destination device in the packet to its physical address, and then routing based on the physical address. It is evident that this method of accessing memory eliminates the need for packets to pass through the host (root node), requiring only the PCIe switching chip, thus reducing DMA latency and eliminating the performance bottleneck of the root complex. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the IOMMU data flow in an existing PCIe system; Figure 2 This is a structural block diagram of an embodiment of the PCIe system provided by the present invention; Figure 3 This is a flowchart of an embodiment of the memory access method for a PCIe system provided by the present invention; Figure 4 yes Figure 3 A flowchart of an embodiment of step 1; Figure 5 This is a diagram illustrating the enumeration of the management module; Figure 6 This is a diagram illustrating virtual machine enumeration; Figure 7 This is an IOMMU entry in one embodiment of the input / output memory management module; Figure 8 This is a schematic diagram of the IOMMU data flow in the PCIe system provided by this invention. Detailed Implementation

[0014] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of this application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to this application are not shown or described in the specification. This is to avoid obscuring the core parts of this application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0015] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.

[0016] The serial numbers assigned to components in this document, such as "first" and "second," are used only to distinguish the described objects and have no sequential or technical meaning. The terms "connection" and "linkage" used in this application, unless otherwise specified, include both direct and indirect connections (linkages).

[0017] The PCIe system provided by this invention involves the host's hypervisor (management module) monitoring operations on the IOMMU and identifying the IOMMU entries of endpoint devices (PCIe devices) requiring this function. The hypervisor then distributes the critical information needed for address translation to the IOMMU entries in the PCIe switch chip via the PCIe switch chip driver interface. When a virtual device initiates a packet with an untranslated address (GPA) and passes through the PCIe switch chip, the IOMMU entries are queried to identify the packet. The PCIe switch chip performs address translation on the packet and routes it directly to its final destination based on the translated address (HPA), eliminating the need for forwarding to the host and thus improving memory access efficiency in the PCIe system. The following embodiments illustrate this in detail.

[0018] like Figure 2 As shown, the PCIe system provided by this invention includes a host 10 (root complex) and one or more PCIe switch chips 20. For simplicity, only one PCIe switch chip 20 is shown in the figure; typically, a PCIe system includes multiple PCIe switch chips 20. The host 10 is connected to the PCIe switch chip 20 via a PCIe bus (peripheral component interconnect express), specifically to the uplink port of the PCIe switch chip 20. The PCIe switch chip 20 has multiple downlink ports, which are used to connect to physical endpoint devices (EPs) 30 or the uplink port of another PCIe switch chip 20. The PCIe system may also include physical endpoint devices connected to each PCIe switch chip 20. In the PCIe system, the host 10, all the PCIe switch chips 20, and the physical endpoint devices connected to the PCIe switch chips 20 form a tree-like topology, with the host 10 as the root node.

[0019] In the PCIe system of this invention, the host 10 is equipped with a management module 120 (Hypervisor) and a virtual machine 110. The virtual machine 110 can be virtualized by the management module 120. The management module 120 is used to manage the PCIe system, such as being responsible for initiating and managing all bus operations. Typically, the hardware of the host 10 includes a CPU and memory. The data of each physical endpoint device in the PCIe system is stored in the memory of the host 10 and indexed by the host physical address (HPA). Of course, the data of the physical endpoint devices can also be stored inside the physical endpoint devices themselves.

[0020] The PCIe switching chip 20 is equipped with an Input / Output Memory Management Unit (IOMMU) 210 and a routing unit 220.

[0021] The host 10 and the PCIe switch chip 20 work together to improve the efficiency of device access to memory in the PCIe system, as follows: Figure 3 As shown, it includes the following steps: Step 1: The management module initializes the PCIe system, obtaining the identity identifier, physical address, virtual address of the corresponding virtual endpoint device, and the topology of the PCIe system for the physical endpoint device 30 connected to the PCIe switching chip. The identity identifier (BDF) of the physical endpoint device 30 identifies the physical endpoint device. This embodiment uses the unique identifier of the physical endpoint device 30 as an example. The identity identifier can include: Bus, Device, and Function, etc. In this embodiment, the physical address of the physical endpoint device can be the address of its data in memory, such as the Host Physical Address (HPA) or the address of the Base Address Register (BAR). The memory can be the memory of the host 10 or the memory of the physical endpoint device 30. The virtual endpoint device is virtualized based on the physical endpoint device. The address (virtual address) of the virtual endpoint device is its own "Guest Physical Address (GPA)," but it is a virtual address for the host 10.

[0022] The physical endpoint device 30 can be of various types depending on the needs or functions of the system, such as GPU (graphics processing unit), various types of storage devices, and various types of network devices.

[0023] like Figure 4 As shown, this step may specifically include the following steps: Step 11: Management module 120 enumerates the physical endpoint devices 30 connected to PCIe switch chip 20 (e.g., ...). Figure 5 The enumeration of physical endpoint devices 30 (GPU0 and GPU1 in the system) includes assigning identity identifiers (such as BDF) and physical addresses (such as BAR addresses) to each physical endpoint device 30, thus obtaining the identity identifier and physical address of each physical endpoint device in the system. Enumeration also reveals the connection relationships between each PCIe switch chip 20 and each physical endpoint device 30, which is the topology of the PCIe system.

[0024] Step 12: Management module 120 starts virtual machine 110 and passes physical endpoint device 30 through to virtual machine 110. Passing through physical endpoint device means that virtual machine 110 directly "exclusively" owns and controls a physical endpoint device, as if the device were physically plugged into the virtual machine itself, requiring almost no intervention from management module 120. Management module 120 also simulates virtual endpoint devices corresponding to physical endpoint device 30 for virtual machine 110; for example, the virtual endpoint device corresponding to physical endpoint device GPU0 is vGPU0, and the virtual endpoint device corresponding to physical endpoint device GPU1 is vGPU1.

[0025] Step 13: Virtual machine 110 enumerates the management module 120 for its simulated virtual endpoint devices (such as...). Figure 6 The virtual endpoint devices in the dataset are vGPU0 and vGPU1. Assign identity identifiers (such as virtual BDFs) and virtual address GPAs (such as virtual BAR addresses) to these virtual endpoint devices. Figure 6 As shown.

[0026] Step 2: The management module 120 obtains the peer devices that physical endpoint device 30 can route to based on the topology. Once the topology of the PCIe system is obtained, it can be determined which physical endpoint devices each physical endpoint device 30 can route to. The peer devices of physical endpoint device 30 are the other physical endpoint devices that physical endpoint device 30 can route to in the PCIe system.

[0027] Step 3: The management module 120 sends the identity of the physical endpoint device 30, the physical address of the peer device, and the virtual address of the corresponding virtual endpoint device as routing and authentication information to the input / output memory management unit 210 of the PCIe switching chip 20. The virtual address of the virtual endpoint device is allocated by the virtual machine 110. The management module 120 can obtain the virtual address of the virtual endpoint device corresponding to each physical endpoint device by intercepting the enumeration request of the virtual machine 110.

[0028] The input / output memory management unit 210 of the PCIe switching chip 20 receives routing and authentication information from each physical endpoint device 30 and records it in the IOMMU table entry. The IOMMU table entry is as follows: Figure 7 As shown, “BUS”, “DEVICE”, and “FUNCTION” correspond to the identity identifier BDF of the physical endpoint device 30, GPA is the virtual address of the virtual endpoint device corresponding to the physical endpoint device 30, and HPA is the physical address of the physical endpoint device 30.

[0029] Step 4: The input / output memory management unit 210 of the PCIe switch chip 20 receives a message sent by the first virtual endpoint device (e.g., vGPU0) to operate on the second virtual endpoint device (e.g., vGPU1). Essentially, the PCIe switch chip 20 receives a message sent from one virtual endpoint device to another. The message typically includes information such as Request ID, Process Address Space ID (PASID), destination address, AT field, and message type. The Request ID is the ID of the message initiator, i.e., the identifier of the physical endpoint device corresponding to the first virtual endpoint device. For messages requiring address translation, the destination address is the virtual address of the second virtual endpoint device, which is the virtual address of the message's destination. The first virtual endpoint device can be any of the virtual endpoint devices created by the system, while the second virtual endpoint device is the destination virtual endpoint device of the message sent by the first virtual endpoint device. That is, "first" and "second" in "first virtual endpoint device" and "second virtual endpoint device" are not specific (limited), but are used for ease of distinction and description.

[0030] Step 5: For a message that requires address translation, the input / output memory management unit 210 obtains from the message the identity of the physical endpoint device 30 (e.g., GPU0) corresponding to the first virtual endpoint device (e.g., vGPU0) and the virtual address of the second virtual endpoint device (e.g., vGPU1).

[0031] Step 6: The input / output memory management unit 210 matches the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device against multiple routing and authentication information. If a match is found, the virtual address of the second virtual endpoint device in the packet is modified to the physical address in the matched routing and authentication information. From the above, it can be seen that the routing and authentication information of physical endpoint device GPU0 includes the identity BDF of GPU0, as well as the physical address HPA and the corresponding virtual address GPA of other physical endpoint devices GPU1 that GPU0 can route to. For the first virtual endpoint device vGPU0, its sent packet contains the identity of GPU0 and the virtual address of the second virtual endpoint device vGPU1. Therefore, it can match the routing and authentication information of physical endpoint device GPU0, indicating that the packet can be transmitted to the destination. If no match is found, it means that there is no routing relationship between the physical endpoint device to which the identity in the packet belongs and the physical endpoint device corresponding to the virtual endpoint device to which the destination address in the packet belongs, and the packet cannot be transmitted.

[0032] Step 7: The routing unit 220 of the PCIe switching chip 20 routes the packet to the physical address in the packet, thereby completing the corresponding operation (such as read and write operation) of the packet sent by the first virtual endpoint device GPU0 to the second virtual endpoint device GPU1.

[0033] PCIe system topologies are typically very large. For example, large AI models require numerous GPUs (physical endpoint devices). The virtual GPUs corresponding to these GPUs often need to perform read and write operations on memory. If each packet requires root node verification and address translation, it can lead to memory access latency and create a performance bottleneck in the PCIe system. However, the PCIe system and memory access method provided in this invention creatively utilize the "intelligent transportation hub" characteristics of the PCIe switching chip 20, offloading some functions of the IOMMU (verification, address translation, etc.) to the PCIe switching chip 20. Furthermore, considering the computing power and processing speed limitations of the PCIe switching chip 20, another part of the IOMMU's functions—the determination and maintenance of IOMMU entries (routing and verification information)—is retained in the host 10. This achieves collaborative management of input and output between the host and the PCIe switching chip, such as... Figure 8 As shown by the arrow, the message only needs to pass through the PCIe switching chip 20 to be routed to the destination, making data transmission (memory access) efficient and fast.

[0034] The address of the destination device in the message may be untranslated or translated, so this can be determined. For example... Figure 3 As shown, step 5 may specifically include: The input / output memory management unit 210 identifies the message and obtains the identity, address, and type of the physical endpoint device corresponding to the first virtual endpoint device. For example, the Request ID (BUS+DEVICE+FUNCTION) in the message is the identity of the physical endpoint device corresponding to the first virtual endpoint device. The AT field in the message indicates the address type. AT is a 2-bit field defined in the PCIe Specification, representing Address Type: 00 indicates the message address is not translated, 01 indicates the message is an address translation request, 10 indicates the message address has been translated, and 11 is a reserved field. The input / output memory management unit 210 determines whether the address has been translated based on its type. If the address type is "untranslated," the input / output memory management unit 210 uses this address as the virtual address of the second virtual endpoint device, and then proceeds to step 6. If the address type is "translated," it means the address in the message is the physical address corresponding to the destination device. Therefore, the input / output memory management unit 210 uses this address as the physical address of the physical endpoint device corresponding to the second virtual endpoint device, and then proceeds to step 7.

[0035] The message is sent by the first virtual endpoint device (e.g., vGPU0) to perform an operation on the second virtual endpoint device (e.g., vGPU1). Not every virtual endpoint device necessarily has permission to perform specific operations on another virtual endpoint device. Therefore, permission checks can be performed in the routing and authentication information, which in this embodiment refers to the routing and authentication information of the physical endpoint device (e.g., ...). Figure 7 The IOMMU entries shown also include: the physical endpoint device's access permissions to the peer device. Access permissions can be divided into three types: read permissions, write permissions, and permissions that allow both read and write access.

[0036] Correspondingly, there are also various types of messages. This embodiment will use three types of messages—read request messages, write request messages, and atomic operation messages—as examples for illustration.

[0037] In some embodiments, the routing and authentication information of the physical endpoint device may also include: Process Address Space ID (PASID). For example... Figure 7 As shown, in the routing and authentication information, the device identifier (such as BUS, DEVICE, FUNCTION) of the physical endpoint device, the process address space ID (in some embodiments, PASID may not be included), and the virtual address (GPA) of the virtual endpoint device corresponding to the peer device can be used as KEY. The operation permissions of the physical endpoint device to the peer device and the physical address (HPA) of the peer device can be used as DATA.

[0038] In step 6, the input / output memory management unit 210 can specifically use the Request ID + PASID (or not including PASID) + GPA from the message as the KEY (index value) to query the IOMMU entries of each physical endpoint device recorded in the input / output memory management unit 210 (e.g., ...). Figure 7When an IOMMU table entry is matched and the AT field of the message is 00, it indicates that the message needs to undergo IOMMU address translation. The input / output memory management unit 210 also obtains the message type. For read request messages, i.e., when the message type is a read request message, it determines whether the operation permissions (WRITE PERMISSION and READ PERMISSION) in the matched routing and authentication information include read permission. If so (READ PERMISSION in the DATA section of the matched IOMMU table is 1), the permission check passes, and the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched routing and authentication information (the ADDR (GPA) of the message is modified to HPA in the IOMMU table entry DATA). In some embodiments, whether to modify the address also involves the ER (Execute-Requested) field in the message. When the ER field is 1, it indicates that the read request is to read an executable instruction, requiring execution permission; when it is 0 or not present, it indicates that the read request is a normal memory read, without execution intent, and does not depend on execution permission. Therefore, it can be modified in the READ field. If PERMISSION is 1 and the ER field of the message is 0 or does not exist, the permission check is successful. Otherwise, if the operation permissions in the matched route and authentication information do not include read permission, the input / output memory management unit 210 redirects the message to the management module 120 for processing. For write request messages, that is, when the received message type is a write request message, the input / output memory management unit 210 determines whether the operation permissions in the matched route and authentication information include write permission. If so (WRTIE PERMISSION in the DATA section of the matched IOMMU table is 1), the permission check is successful, and the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched route and authentication information. Otherwise, if the operation permissions in the matched route and authentication information do not include write permission, the input / output memory management unit 210 redirects the message to the management module 120 for processing. For atomic operation messages, that is, when the received message type is atomic operation (uninterruptible operation) message, the input / output memory management unit 210 determines whether the operation permissions in the matched routing and authentication information include read permission and write permission. If so (the WRTIE PERMISSION and READ PERMISSION in the DATA part of the matched IOMMU table are both 1), it means that the permission check is passed, and the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched routing and authentication information; otherwise, if the operation permissions in the matched routing and authentication information do not include read permission or write permission, the input / output memory management unit 210 redirects the message to the management module 120 for processing.

[0039] In existing PCIe systems, after the RC or host CPU receives a packet uploaded by the PCIe switch chip, the RC or host CPU's IOMMU queries the BDF (Browser Definition Table) in the packet using page tables. Specifically, it first queries using the BUS number, then the DEVICE and FUNCTION numbers, and finally the virtual address. Each query narrows the query range, allowing resources to be created based on the actual table entries used (because the IOMMU page tables in the RC use system memory DRAM, which can be dynamically allocated).

[0040] For PCIe switching chips, the resources required are reserved according to the peak budget (the maximum number of entries that can be supported), and the resources used are generally SRAM. Therefore, the required resources are roughly the same regardless of the lookup method used. Thus, in this embodiment, the IOMMU of the PCIe switching chip uses direct indexing or hashing to look up existing routes and authentication information using the Business Found Function (BDF) and virtual address as the key. This requires only one lookup, resulting in lower latency and higher efficiency compared to existing technologies.

[0041] In some embodiments, after address translation in step 6, the input / output memory management unit 210 also modifies the AT field of the message to 10, and then the routing unit 220 performs routing.

[0042] As can be seen, when a virtual endpoint device wants to operate on the data in memory of another virtual endpoint device (such as read and write operations), the memory access method provided by the present invention can achieve this very efficiently and quickly. The packet does not need to pass through the root node. The larger the PCIe system, the more obvious the improvement in memory access efficiency.

[0043] like Figure 2 As shown, the present invention also provides a host for a PCIe system, the host being used to connect to one or more PCIe switch chips 20 via a PCIe bus. The host includes a management module 120 and a virtual machine 110.

[0044] The management module 120 is used to enumerate the physical endpoint devices 30 connected to the one or more PCIe switching chips, obtain the identity, physical address, and topology of the physical endpoint devices 30, and the topology of the PCIe system; obtain the peer devices that can be routed to by the physical endpoint devices 30 according to the topology; pass the physical endpoint devices 30 through the virtual machine 110; simulate the virtual endpoint device corresponding to the physical endpoint devices 30 for the virtual machine 110; obtain the virtual address of the virtual endpoint device corresponding to the physical endpoint devices from the virtual machine 110; and send the identity of the physical endpoint devices 30, the physical address of the peer devices, and the virtual address of the virtual endpoint devices corresponding to the peer devices as routing and authentication information to the input / output memory management unit 210 of the PCIe switching chip 20, so that the input / output memory management unit 210 can verify and route the packets received by the PCIe switching chip 20 according to the routing and authentication information. The specific process of the management module 120 performing these functions has been described in detail in the foregoing embodiments and will not be repeated here.

[0045] Virtual machine 110 is used to enumerate the virtual endpoint devices simulated by management module 120 and assign identity identifiers and virtual addresses to the virtual endpoint devices. The specific process of virtual machine 110 performing these functions has been described in detail in the foregoing embodiments and will not be repeated here.

[0046] like Figure 2 and Figure 8 As shown, the present invention also provides a PCIe switching chip, including: an uplink port 230, a memory management unit 210, a routing unit 220 and multiple downlink ports 240.

[0047] Uplink port 230 is used to connect to the host via the PCIe bus.

[0048] Downlink port 240 is used to connect to physical endpoint device 30, and can also be connected to uplink port 230 of the next-level PCIe switching chip.

[0049] The memory management unit 210 receives routing and authentication information from the host. This information includes: the identity of the physical endpoint device 30 connected to the PCIe switch chip, the physical address of the peer device, and the virtual address of the virtual endpoint device corresponding to the peer device; other physical endpoint devices that the peer device can route to from the physical endpoint device 30; upon receiving a message sent by the first virtual endpoint device to operate on the second virtual endpoint device, the unit retrieves the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device from the message; it matches the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device against multiple routing and authentication messages; if a match is found, the unit modifies the virtual address of the second virtual endpoint device in the message to the physical address in the matched routing and authentication message. The specific process by which the memory management unit 210 performs these functions has been described in detail in the preceding embodiments and will not be repeated here.

[0050] The routing unit 220 is used to route packets to the physical addresses in the packets. The specific process by which the routing unit 220 performs these functions has been described in detail in the foregoing embodiments and will not be repeated here.

[0051] In the above embodiments, implementation can be achieved, in whole or in part, by software, hardware, firmware, or any combination thereof. Furthermore, as those skilled in the art will understand, the principles herein can be reflected in a computer program product on a computer-readable storage medium pre-loaded with computer-readable program code. Any tangible, non-transitory computer-readable storage medium may be used, including magnetic storage devices (hard disks, floppy disks, etc.), optical storage devices (CDs, DVDs, Blu-ray discs, etc.), flash memory, and / or the like. These computer program instructions can be loaded onto a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to form a machine, such that instructions executing on the computer or other programmable data processing apparatus can generate means for performing a specified function. These computer program instructions can also be stored in a computer-readable storage medium that can instruct the computer or other programmable data processing apparatus to operate in a particular manner, such that instructions stored in the computer-readable storage medium can form an article of manufacture, including means for implementing the specified function. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to perform a series of operational steps on the computer or other programmable apparatus to produce a computer-implemented process, such that instructions executing on the computer or other programmable apparatus can provide steps for implementing the specified function.

[0052] This document describes various exemplary embodiments with reference to them. However, those skilled in the art will recognize that changes and modifications can be made to the exemplary embodiments without departing from the scope of this document. For example, various operational steps and components for performing operational steps can be implemented in different ways depending on the specific application or considering any number of cost functions associated with the operation of the system (e.g., one or more steps can be deleted, modified, or combined with other steps).

[0053] While the principles herein have been illustrated in various embodiments, numerous modifications to the structures, arrangements, proportions, elements, materials, and components, particularly suited to specific environments and operational requirements, may be used without departing from the principles and scope of this disclosure. These modifications and other alterations or alterations will be included within the scope of this document. Those skilled in the art will recognize that many changes can be made to the details of the above embodiments without departing from the fundamental principles of the invention.

Claims

1. A memory access method for a PCIe system, the PCIe system comprising a host and one or more PCIe switching chips, the host being connected to the PCIe switching chips via a PCIe bus, characterized in that, The PCIe switching chip is equipped with an input / output memory management unit and a routing unit; The host is equipped with a management module; the method includes: The management module initializes the PCIe system to obtain the identity and physical address of the physical endpoint devices connected to the one or more PCIe switching chips, the virtual address of the virtual endpoint devices corresponding to the physical endpoint devices, and the topology of the PCIe system. The management module obtains the peer devices that the physical endpoint device can reach based on the topology; The management module sends the identity identifier of the physical endpoint device, the physical address of the peer device, and the virtual address of the virtual endpoint device corresponding to the peer device as routing and authentication information to the input / output memory management unit of the PCIe switching chip; the routing and authentication information of the physical endpoint device also includes the operation permissions of the physical endpoint device to the peer device; After receiving a message sent by the first virtual endpoint device to operate on the second virtual endpoint device, the input / output memory management unit of the PCIe switching chip obtains the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device from the message. The input / output memory management unit uses the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device to match multiple routing and authentication information; If a match is found, the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched route and verification information, including: obtaining the type of the message; when the message type is a read request message, determining whether the operation permissions in the matched route and verification information include read permission, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information; when the message type is a write request message, determining whether the operation permissions in the matched route and verification information include write permission, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information; when the message type is an atomic operation message, determining whether the operation permissions in the matched route and verification information include both read and write permissions, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information. The routing unit of the PCIe switching chip routes the packet to the physical address in the packet.

2. The method as described in claim 1, characterized in that, Also includes: When the message type is a read request message, if the operation permissions in the matched route and verification information do not include permission to read, the input / output memory management unit will redirect the message to the management module for processing. When the message type is a write request message, if the operation permissions in the matched route and verification information do not include write permission, the input / output memory management unit will redirect the message to the management module for processing. When the message type is an atomic operation message, if the operation permissions in the matched routing and verification information do not include read permissions or write permissions, the input / output memory management unit will redirect the message to the management module for processing.

3. The method as described in claim 1, characterized in that, The step of obtaining the identity identifier of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device from the message includes: The message is identified to obtain the identity, address, and type of the physical endpoint device corresponding to the first virtual endpoint device; when the type of the address is an untranslated address, the address is used as the virtual address of the second virtual endpoint device.

4. The method as described in claim 3, characterized in that, Also includes: When the address type is a converted address, the address is used as the physical address of the physical endpoint device corresponding to the second virtual endpoint device.

5. The method as described in claim 1, characterized in that, The host also has a virtual machine; the management module initializes the PCIe system, obtaining the identity and physical address of the physical endpoint devices connected to the one or more PCIe switching chips, the virtual address of the virtual endpoint device corresponding to the physical endpoint device, and the topology of the PCIe system, including: The management module enumerates the physical endpoint devices connected to the one or more PCIe switching chips to obtain the identity, physical address and topology of the PCIe system of the physical endpoint devices. The physical endpoint device is transparently transmitted to the virtual machine; a virtual endpoint device corresponding to the physical endpoint device is simulated for the virtual machine; The virtual machine enumerates the virtual endpoint devices simulated by the management module and assigns identity identifiers and virtual addresses to the virtual endpoint devices.

6. The method as described in claim 1, characterized in that, The routing and authentication information also includes a process address space ID; the input / output memory management unit matches the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device in multiple routing and authentication information, including: Using the identity identifier of the physical endpoint device corresponding to the first virtual endpoint device in the message, the process address space ID of the physical endpoint device, and the virtual address of the virtual endpoint device corresponding to the peer device as index values, the routing and verification information of each physical endpoint device recorded by the input / output memory management unit is queried.

7. A PCIe system, comprising a host and one or more PCIe switching chips, wherein the host is connected to the PCIe switching chips via a PCIe bus, characterized in that, The PCIe switching chip is equipped with an input / output memory management unit and a routing unit; the host is equipped with a management module; The host's management module is used to initialize the PCIe system, obtain the identity and physical address of the physical endpoint devices connected to the one or more PCIe switching chips, the virtual address of the virtual endpoint devices corresponding to the physical endpoint devices, and the topology of the PCIe system; and obtain the peer devices that the physical endpoint devices can be routed to based on the topology. The identity identifier of the physical endpoint device, the physical address of the peer device, and the virtual address of the virtual endpoint device corresponding to the peer device are sent as routing and authentication information to the input / output memory management unit of the PCIe switching chip; the routing and authentication information of the physical endpoint device also includes the operation permissions of the physical endpoint device to the peer device; The input / output memory management unit of the PCIe switching chip is used for: Upon receiving a message from the first virtual endpoint device that is performing an operation on the second virtual endpoint device, the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device are obtained from the message; the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device are then used for matching in multiple routing and authentication information. If a match is found, the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched routing and verification information, including: obtaining the type of the message; When the message type is a read request message, determine whether the operation permissions in the matched route and verification information include read permission. If so, modify the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information. When the message type is a write request message, determine whether the operation permissions in the matched route and verification information include write permission. If so, modify the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information. When the message type is an atomic operation message, determine whether the operation permissions in the matched route and verification information include both read and write permissions. If so, modify the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information. The routing unit of the PCIe switching chip is used to route the packet to the physical address in the packet.

8. A host for a PCIe system, the host being configured to connect to one or more PCIe switching chips via a PCIe bus, characterized in that, include: Management modules and virtual machines; The management module is used to enumerate the physical endpoint devices connected to the one or more PCIe switching chips, and obtain the identity, physical address and topology of the physical endpoint devices and the PCIe system. Based on the topology, obtain the peer devices that the physical endpoint device can reach by routing; The physical endpoint device is transparently transmitted to the virtual machine; a virtual endpoint device corresponding to the physical endpoint device is simulated for the virtual machine; the virtual address of the virtual endpoint device corresponding to the physical endpoint device is obtained from the virtual machine; the identity of the physical endpoint device, the physical address of the peer device, and the virtual address of the virtual endpoint device corresponding to the peer device are sent as routing and authentication information to the input / output memory management unit of the PCIe switching chip, so that the input / output memory management unit can verify and route the packets received by the PCIe switching chip according to the routing and authentication information; the routing and authentication information of the physical endpoint device also includes the physical endpoint device's operation permissions for the peer device; The virtual machine is used to enumerate the virtual endpoint devices simulated by the management module and assign identity identifiers and virtual addresses to the virtual endpoint devices. The PCIe switching chip's input / output memory management unit verifies and routes the packets received by the PCIe switching chip based on the routing and verification information, including: Upon receiving a message sent by the first virtual endpoint device to perform an operation on the second virtual endpoint device, the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device are obtained from the message. The identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device are used to match multiple routing and authentication information; If a match is found, the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched routing and verification information, including: obtaining the type of the message; when the message type is a read request message, determining whether the operation permissions in the matched routing and verification information include read permission, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched routing and verification information; when the message type is a write request message, determining whether the operation permissions in the matched routing and verification information include write permission, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched routing and verification information; when the message type is an atomic operation message, determining whether the operation permissions in the matched routing and verification information include both read and write permissions, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched routing and verification information.

9. A PCIe switching chip, characterized in that, include: Uplink port, used to connect to the host via the PCIe bus; Multiple downlink ports, which are used to connect to physical endpoint devices; Input / output memory management unit, used for: The system receives routing and authentication information from the host, which includes: the identity of the physical endpoint device connected to the PCIe switching chip, the operation permissions of the physical endpoint device to the peer device, the physical address of the peer device, and the virtual address of the virtual endpoint device corresponding to the peer device; the peer device is other physical endpoint devices that the physical endpoint device can route to. Upon receiving a message sent by the first virtual endpoint device to perform an operation on the second virtual endpoint device, the identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device are obtained from the message. The identity of the physical endpoint device corresponding to the first virtual endpoint device and the virtual address of the second virtual endpoint device are used to match multiple routing and authentication information; If a match is found, the virtual address of the second virtual endpoint device in the message is modified to the physical address in the matched route and verification information, including: obtaining the type of the message; when the message type is a read request message, determining whether the operation permissions in the matched route and verification information include read permission, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information; when the message type is a write request message, determining whether the operation permissions in the matched route and verification information include write permission, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information; when the message type is an atomic operation message, determining whether the operation permissions in the matched route and verification information include both read and write permissions, and if so, modifying the virtual address of the second virtual endpoint device in the message to the physical address in the matched route and verification information. A routing unit is used to route the message to the physical address in the message.

Citation Information

Patent Citations

  • Message forwarding method and device applied to end-to-end communication

    CN108471384A

  • PCIe switch chip with RDMA acceleration function and PCIe switch

    CN118093468A