Method, device and cloud server for extending a PCIe system and storage medium
Patent Information
- Application Number
- CN202310664708.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-06
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2043-06-06
AI Technical Summary
目前,通常采用基于高速串行计算机扩展总线标准(peripheral component interconnectexpress,PCIe)机制的热插拔技术,来实现弹性设备的灵活选用,参见图1所示,为基于PCIe机制的热插拔拓扑结构示意图,其中,PCIe交换器(switch)包含一个上游端口(uplinkport,UP)和多个下游端口(downlink port,DP),一个DP下只能挂载一个可热插拔设备,而通常情况下一个计算机系统只支持一个PCIe域(domin),一个PCIe域最多提供256个总线号,且PCIe协议为点对点传输协议,从而受限于DP数量和总线数量,使得整个拓扑中可提供的物理功能(physical function,PF)设备是有限的
[0028] In this embodiment of the application, a first number of first type DPs and a second number of second type DPs in the target PCIe system are allocated according to the target number of required PF devices to generate a target system topology corresponding to the target PCIe system, and then the host configures the target PCIe system according to the topology. In the target system topology, the second type of DP connects multiple PF devices through its corresponding bridge group. Each bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect PF devices. Each PCI bridge can also connect multiple PF devices. Thus, a single second type DP can connect multiple PF devices. It is evident that by using multi-level bridging, the number of PF devices that a single DP can connect is expanded, freeing the PCIe system from the limitations of the number of DPs and buses. This increases the number of PF devices available in the PCIe system, thereby enabling hot-swappable configuration and management of large-scale flexible devices, expanding the number of hot-swappable devices visible to the user, and significantly reducing system management complexity. Furthermore, since a second type DP can connect multiple PF devices, these connected PF devices can share the bus resources of the second type DP. In scenarios with a large number of PF devices, this greatly reduces bus resource overhead and improves bus resource utilization.
Smart Images

Figure CN119088740B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly to the field of cloud server technology, providing a method, device, cloud server, and storage medium for expanding a PCIe system. Background Technology
[0002] A bare metal cloud server (CBM) is a hardware device that combines the characteristics of a traditional physical server with the virtualization service capabilities of cloud computing technology. Simply put, bare metal refers to computer hardware without an operating system, or it can be understood as not completely devoid of operating systems and software, but rather allowing tenants to select and configure the operating systems and software they need. Furthermore, bare metal servers are physically isolated from other tenants, thus possessing the characteristic of secure physical isolation, making them particularly suitable for scenarios with high security requirements.
[0003] When bare-metal cloud services are needed, tenants can select elastic devices according to their own requirements. Correspondingly, bare-metal cloud service providers can hot-swap or hot-plug these elastic devices to meet the tenant's needs. Currently, hot-plugging technology based on the high-speed serial computer extended bus standard (Peripheral Component Interconnect Express, PCIe) is commonly used to achieve flexible selection of elastic devices. See [link to relevant documentation]. Figure 1 The diagram shows a hot-swappable topology based on the PCIe mechanism. A PCIe switch contains one upstream port (UP) and multiple downstream ports (DP). Only one hot-swappable device can be connected to a DP. However, a computer system typically supports only one PCIe domain, which provides a maximum of 256 bus numbers. Furthermore, the PCIe protocol is a point-to-point transmission protocol. Therefore, the number of DPs and buses limits the number of physical function (PF) devices that can be provided in the entire topology.
[0004] Therefore, how to achieve hot-swapping of large-scale flexible devices is an urgent problem to be solved. Summary of the Invention
[0005] This application provides a method, device, cloud server, and storage medium for expanding a PCIe system, which uses multi-level bridging technology to expand user-visible hot-swappable devices, thereby achieving hot-swappable management of large-scale flexible devices.
[0006] On the one hand, a method for extending a PCIe system, applied to a smart network interface card (NIC), is provided, the method comprising:
[0007] Obtain the target number of Physical Function PF devices required for the target PCIe system to be configured;
[0008] Based on the target number and the pre-configured upper limit of the number of downstream port DPs, the first number of first type DPs with a single PF device mounted in the target PCIe system and the second number of second type DPs with multiple PF devices mounted are obtained.
[0009] Based on the first quantity and the second quantity, a target system topology corresponding to the target PCIe system is generated; wherein, in the target system topology, the second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect PF devices.
[0010] The target system topology is notified to the host to instruct the host to configure the target PCIe system according to the target system topology.
[0011] On the one hand, a method for extending a PCIe system is provided, applied to a host, the method comprising:
[0012] Receive the target system topology sent by the smart network card; wherein, in the target system topology, each second type DP is connected to multiple PF devices through its corresponding bridge group, the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect the PF devices.
[0013] From the PF devices included in itself, select PF devices that conform to the target system topology, and call the standard hot-swap controller (SHPC) of each PF device to initialize the state of each PF device.
[0014] Each PF device is assigned a system identifier, and the assigned system identifier is returned to the smart network interface card; wherein, the system identifier is generated based on the target system topology and is a unique identifier for the target PF device in the target PCIe system.
[0015] On the one hand, a smart network interface card (NIC) is provided, including:
[0016] The receiving unit is used to obtain the target number of Physical Function PF devices required by the target PCIe system to be configured;
[0017] The DP configuration unit is used to obtain, based on the target number and the pre-configured upper limit of the number of downstream port DPs, the first number of first type DPs with a single PF device mounted in the target PCIe system and the second number of second type DPs with multiple PF devices mounted.
[0018] A topology generation unit is used to generate a target system topology corresponding to the target PCIe system based on the first quantity and the second quantity; wherein, in the target system topology, the second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect the PF devices.
[0019] The sending unit is used to notify the host of the target system topology, so as to instruct the host to configure the target PCIe system according to the target system topology.
[0020] On the one hand, a host is provided, including:
[0021] A receiving unit is used to receive a target system topology sent by a smart network card; wherein, in the target system topology, each second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect the PF devices.
[0022] The PCIe system configuration unit is used to select PF devices that conform to the target system topology from the PF devices it includes, and call the standard hot-swap controller (SHPC) of each PF device to initialize the state of each PF device.
[0023] The resource configuration unit is used to assign system identifiers to each of the PF devices and return the assigned system identifiers to the smart network interface card; wherein, the system identifier is generated based on the target system topology and is a unique identifier for the target PF device in the target PCIe system.
[0024] On the one hand, a cloud server is provided, including the aforementioned smart network card and the aforementioned host.
[0025] On one hand, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above methods.
[0026] On the one hand, a computer storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the above methods.
[0027] On one hand, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and executes the computer program, causing the computer device to perform the steps of any of the methods described above.
[0028] In this embodiment of the application, a first number of first type DPs and a second number of second type DPs in the target PCIe system are allocated according to the target number of required PF devices to generate a target system topology corresponding to the target PCIe system, and then the host configures the target PCIe system according to the topology. In the target system topology, the second type of DP connects multiple PF devices through its corresponding bridge group. Each bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect PF devices. Each PCI bridge can also connect multiple PF devices. Thus, a single second type DP can connect multiple PF devices. It is evident that by using multi-level bridging, the number of PF devices that a single DP can connect is expanded, freeing the PCIe system from the limitations of the number of DPs and buses. This increases the number of PF devices available in the PCIe system, thereby enabling hot-swappable configuration and management of large-scale flexible devices, expanding the number of hot-swappable devices visible to the user, and significantly reducing system management complexity. Furthermore, since a second type DP can connect multiple PF devices, these connected PF devices can share the bus resources of the second type DP. In scenarios with a large number of PF devices, this greatly reduces bus resource overhead and improves bus resource utilization. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0030] Figure 1 A schematic diagram of a hot-plug topology based on the PCIe mechanism provided in this application embodiment;
[0031] Figure 2 This is a schematic diagram illustrating an application scenario provided in the embodiments of this application;
[0032] Figure 3 This is an architecture diagram of the bare metal cloud server service system provided in the embodiments of this application;
[0033] Figure 4A flowchart illustrating the method for extending the PCIe system provided in this application embodiment;
[0034] Figure 5 A schematic diagram of a configuration page provided in an embodiment of this application;
[0035] Figures 6a-6c A schematic diagram of a scalable hot-swappable topology based on the SHPC mechanism provided in an embodiment of this application;
[0036] Figure 7 Another flowchart illustrating the method for extending the PCIe system provided in this application embodiment;
[0037] Figure 8 This is a schematic diagram of the process for storing device information of a PF device provided in an embodiment of this application;
[0038] Figure 9 This is a schematic diagram illustrating the principle of a multi-level table entry lookup technique based on CRC encoding provided in an embodiment of this application.
[0039] Figure 10 A schematic diagram illustrating the device information query process provided in this application embodiment;
[0040] Figure 11 A flowchart illustrating the hot-swap management process provided in an embodiment of this application;
[0041] Figure 12 A schematic diagram illustrating the state management principle of a high-reliability hot-swappable state machine provided in this application embodiment;
[0042] Figure 13 This is a schematic diagram of the structure of a smart network card provided in an embodiment of this application;
[0043] Figure 14 A schematic diagram of the structure of a host provided in an embodiment of this application;
[0044] Figure 15 This is a schematic diagram of the composition structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0046] To facilitate understanding of the technical solutions provided in the embodiments of this application, some key terms used in the embodiments of this application will be explained below:
[0047] PCIe system: also known as PCIe domain. PCIe is a high-speed serial computer expansion bus standard, while PCI is a standard for defining local buses. The PCIe bus, as a local bus of the processor system, functions similarly to the PCI bus, primarily for connecting external devices within the processor system. The PCIe bus uses an end-to-end connection method; only one device can be connected to each end of a PCIe link, and these two devices act as the data sender and receiver for each other.
[0048] For a typical PCIe system, see Figure 1 As shown, it includes a root complex (RC), a PCIe switch, and PCIe endpoint devices. The RC can reside on the processor and occupies one bus number. The UP (upper port) of the PCIe switch is used to connect to the root port (RP) of the RC, and each DP can mount one PF (power field device).
[0049] Smart network interface card (NIC): also known as a data processing unit (DPU), is a new type of programmable processor with high processing power and high-performance network interface.
[0050] PF device: also known as a flexible device. In this application embodiment, a PF device refers to a device with hot-swappable characteristics, such as a hard drive, optical drive, printer, network card, graphics card, sound card, expansion card, etc. These devices can be added or removed during system operation without affecting the normal operation of the system.
[0051] A bridge (BR) is a bridge device in a PCIe or PCI topology. This application primarily involves PCIe-PCI bridges (PCIe bridges) and PCI-PCI bridges (PCI bridges). A PCIe bridge enables communication between PCIe bus protocol devices and PCI bus protocol devices, distributing control messages to lower-level PCI bridges. A PCI bridge enables communication between PCI bus protocol devices, distributing control messages to lower-level PCI bridges or PF devices. Thus, through the combination of PCIe and PCI bridges, multiple PF devices under the same DP can be made visible to the user, and multiple PF devices can be managed.
[0052] Cyclic Redundancy Check (CRC) algorithm: This is a fast algorithm that generates a short, fixed-length check code based on data such as network packets or computer files. It is a commonly used check code in the field of computer network communication. CRC utilizes the principle of division and remainders. CRC codes include a series of data encoding rules such as shifting and division. Its algorithm principle, algorithm program design and analysis can all be solved through corresponding software coding. Cyclic Redundancy Check is a software-based verification algorithm, therefore its verification speed is very fast, and its error rate is also low. The information transmission speed of the entire computer network communication is very high, and it has the advantages of clear principle and simple implementation.
[0053] The Standard Hot Plug Controller (SHPC) is used to manage and control hot-pluggable devices. The Peripheral Component Interconnect Special Interest Group (PCI SIG) developed the PCI hot-plug specification, which defined the platform, boards, and software elements necessary to support hot-plugging. Subsequently, the PCI SIG released the SHPC specification, which clarified the standard usage pattern of hot-plugging and strict register set requirements, and allowed operating system vendors to provide hot-plug support in addition to platform-specific software, thus gradually completing the standardization of hot-plugging.
[0054] System Identifier: This refers to the unique identifier of each device or module within a PCIe system. Taking a PF device as an example, the system identifier can be the unique identifier of the PF device within the PCIe system. Typically, the system identifier can be at least one of the following: the memory address assigned to the PF device, the base address register (BAR) address, or the bus number, device number, and function number combination (Bus-Device-Function (BDF) information.
[0055] Bare metal cloud servers are non-virtualized servers that run directly on the hardware and therefore have no intermediate virtualization layer. Compared with traditional virtualized servers, bare metal servers can provide higher performance and lower latency, thus providing enterprises with dedicated physical servers in the cloud. They offer superior computing performance and data security for core databases, critical application systems, high-performance computing, and big data services, allowing cloud service tenants to apply flexibly and use them on demand.
[0056] When bare metal cloud services are needed, tenants can choose elastic devices according to their own needs. However, in general, a computer system only supports one PCIe domain and provides 256 bus numbers for each PCIe domain. That is, the entire computer system only provides 256 bus numbers. Since the PCIe protocol is a point-to-point connection, the scale of PCIe devices that a PCIe domain can support is limited by the number of 256 bus numbers. It is impossible to support more PCIe devices in a PCIe domain. In practical applications, other PCIe devices besides PF devices will also occupy bus numbers. Therefore, the number of PF devices is always less than 256, making the number of PF devices available in the entire topology extremely limited.
[0057] To address the issue of insufficient PF quantity, related technologies have proposed using Single Root I / O Virtualization (SR-IOV) technology to virtualize virtual function (VF) devices to host elastic devices. However, VF devices generated using SR-IOV technology require setting feature bits and loading drivers after enabling SR-IOV, and then reversing the process during unloading. This is user-unfriendly, as errors in the order of operations can easily lead to failures in adding or deleting devices. Furthermore, modifying feature bits can intrude on the host image, making it difficult to implement in public cloud bare metal cloud server products.
[0058] Therefore, how to achieve hot-swapping of large-scale flexible devices remains an urgent problem to be solved.
[0059] Based on this, embodiments of this application provide a method for extending a PCIe system. In this method, a first number of first-type DPs and a second number of second-type DPs in the target PCIe system are allocated according to the target number of required PF devices to generate a target system topology corresponding to the target PCIe system. Then, the host configures the target PCIe system according to the topology. In the target system topology, the second type of DP connects multiple PF devices through its corresponding bridge group. Each bridge group includes a PCIe bridge bridging the second type DP and multiple PCI bridges that connect the PF devices. Each PCI bridge can also connect multiple PF devices. Thus, a single second type DP can connect multiple PF devices. It is evident that by using multi-level bridging, the number of PF devices that a single DP can connect is expanded, freeing the PCIe system from the limitations of the number of DPs and buses. This increases the number of PF devices available in the PCIe system, enabling hot-swappable configuration and management of large-scale flexible devices. It expands the user-visible hot-swappable devices and eliminates the need to generate VF devices through SR-IOV technology, significantly reducing system management complexity. Furthermore, since a second type DP can connect multiple PF devices, these connected PF devices can share the bus resources of that second type DP. In scenarios with large-scale PF devices, this greatly reduces bus resource overhead and improves bus resource utilization.
[0060] The above methods greatly expand the number of PF devices, but also bring about efficiency issues in PF device resource management. Therefore, in order to solve the problem of fast indexing of large-scale PF devices, this application provides a multi-level table entry query method. In this method, each table entry index consists of one or more index entries. If all entries corresponding to the table entry index are unavailable, the system switches to the next level table entry index for reading and writing, thereby meeting the indexing needs of a large number of PF devices.
[0061] Furthermore, this application embodiment also designs a highly reliable hot-swappable state machine to ensure that the final state of the PF device does not match the user's expected state, thus ensuring high availability. Specifically, the highly reliable hot-swappable state machine considers various scenarios such as host not being powered on, host restart, hot-swappable interrupt loss, and repeated hot-swapping by the user. It performs state machine rotation based on the user's expected state and the current state of the PF device to ensure that the device state presented to the user is the state expected by the user, thereby improving the user experience.
[0062] The following is a brief introduction to the application scenarios to which the technical solutions of the embodiments of this application are applicable. It should be noted that the application scenarios described below are only for illustrating the embodiments of this application and are not intended to limit the scope. In specific implementation, the technical solutions provided by the embodiments of this application can be flexibly applied according to actual needs.
[0063] The solution provided in this application can be applied to bare metal server scenarios. For example... Figure 2 The diagram shown is an application scenario provided by an embodiment of this application. In this scenario, a terminal device 10 and a cloud server 20 may be included.
[0064] Terminal device 10 can be any terminal device such as a mobile phone, tablet computer (PAD), laptop computer, desktop computer, smart TV, smart in-vehicle device, and smart wearable device. Terminal device 10 can have a target application installed. The target application should have the function of configuring and presenting bare metal cloud host. The application involved in the embodiments of this application can be a software client, or a web page, applet, or other client. The specific type of client is not limited.
[0065] Server 20 is used to deploy and manage bare-metal cloud servers and provide users with bare-metal cloud server-related services. Server 20 can be, for example, a standalone computer device or physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, i.e., Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited to these.
[0066] See Figure 2 As shown, server 20 provides bare metal cloud host services and may include a smart network card 201 and a host 202 that are interconnected. For example, smart network card 201 can be inserted into the corresponding card slot in host 202. Smart network card 201 has an independent computing unit. The PCIe system mentioned above is actually deployed and implemented by host 202. Host 202 contains a series of selectable PF devices, so the PCIe system can be deployed according to the user's selection.
[0067] It should be noted that the method for expanding the PCIe system in this embodiment can be implemented by combining the smart network interface card 201 and the host 202. Specifically, the smart network interface card 201 can determine the deployment topology based on the user's requirements for the PCIe system, such as indicating the number of PF devices required. When the number of PF devices exceeds the upper limit of DP, the traditional topology scheme cannot be implemented. Therefore, some DPs can be selected to mount multiple PF devices, generating a system topology containing this type of DP for the user and notifying the host 202 for corresponding deployment. In actual scenarios, the server 20 can include one smart network interface card 201 and multiple hosts 202, and this embodiment does not limit this.
[0068] In this embodiment, the terminal device 101 and the server 102 can communicate directly or indirectly through one or more networks. The network can be a wired network or a wireless network, such as a mobile cellular network or a Wireless-Fidelity (WIFI) network, or any other possible network. This embodiment does not limit the types of networks used.
[0069] like Figure 3 The diagram shown is an architecture diagram of a bare metal cloud host service system provided in an embodiment of this application. In this architecture, it can include a client part and a server part. The client part is used to provide users with a page for configuring and using cloud host services. The server part consists of a smart network interface card (NIC) and a host. The smart NIC is used to manage the cloud host service, and the host is the device that provides the actual cloud host service to the user. For example, if the user needs to configure multiple hot-swappable NICs, the host can be a device containing multiple NICs. The system can configure the required number of NICs for the user based on the user's selection to achieve the corresponding network service.
[0070] See Figure 3 As shown, the smart network interface card (NIC) includes a topology management module, a table query module, and a hot-swap status management module. The topology management module generates a corresponding PCIe system topology for the user based on the number of PF devices configured by the user, notifies the host to configure according to this topology, and presents the topology to the user, thus showcasing a large number of PF devices. The table query module provides a multi-level table query method based on a large number of PF devices. This method is used to quickly retrieve the relevant information of the required PF devices when device information, such as routing information, is needed. The hot-swap status management module manages the hot-swap status of PF devices to ensure that the status of each PF device is highly consistent with the user's expected status.
[0071] The host includes a CPU, SHPC, and PF devices. The PF devices are flexibly selected and configured based on the user's choices, and the SHPC is used to set the status of the PF devices.
[0072] In practical applications, users can configure the target PCIe system's system information, such as the target number of PF devices required by the system, on the configuration page provided by the client. The smart network interface card (NIC) can then generate the target system topology based on this target number and the extended PCIe system method provided in this application embodiment. The host is then notified to perform the corresponding configuration according to this target system topology to obtain the user's desired target PCIe system and provide related services. For example, when the number of PF devices required by the user is large, the extended PCIe system method provided in this application embodiment can be used. This involves a multi-level bridging approach, selecting some DPs to bridge multiple PF devices. PF devices connected to the same DP share the bus number corresponding to that DP, and a bridge group is connected to each DP to facilitate the management of these PF devices.
[0073] Considering the low efficiency of indexing large-scale PF devices, the table query module in this embodiment encodes the system identifier of the PF device as the table index for that PF device. When the table index is full, it jumps to the next level table index to continue storing the table entries. Similarly, when performing a table query, the system identifier of the PF device to be queried is encoded in the same way, and the table is searched level by level according to the obtained table index until the relevant information of the PF device is found. By using the index query method, the scope of the table search can be narrowed as much as possible, thereby improving the indexing efficiency of PF devices.
[0074] The following describes a method for extending a PCIe system provided by an exemplary embodiment of this application, in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.
[0075] See Figure 4 The diagram shown is a flowchart of a method for extending a PCIe system provided in an embodiment of this application. This method can be executed by a smart network interface card (NIC), and the specific implementation process is as follows:
[0076] Step 401: Obtain the target number of Physical Function PF devices required by the target PCIe system to be configured.
[0077] In this embodiment, the target quantity can be entered by the target object in the PCIe system's configuration page. For example, it can be entered when configuring a new PCIe system, or it can be entered when modifying an existing PCIe system. See [link to relevant documentation] Figure 5 The diagram shown is a schematic of the configuration page provided in an embodiment of this application. It includes configuration items for PF type selection and PF quantity. The target object can select the type of PF device and the required number of PF devices according to its own needs, and then initiate a configuration request for the target PCIe system. This allows the smart network card to obtain the target number of physical function PF devices required by the target PCIe system to be configured based on the configuration request.
[0078] It should be noted that the smart network interface card (NIC) can directly receive configuration requests sent by the target object, or it can receive the configuration request through other devices, such as receiving configuration requests forwarded by the access device of the cloud server. Furthermore, the target object refers to a user in the cloud host service platform, which can be a natural person or an account representing a user's identity, such as through an account identifier. This application embodiment does not impose many restrictions on the target object.
[0079] Step 402: Based on the target number and the pre-configured upper limit of the number of DPs, obtain the first number of first-type DPs with a single PF device mounted in the target PCIe system, and the second number of second-type DPs with multiple PF devices mounted.
[0080] The upper limit of the number of pre-configured downstream ports (DPs) can be, for example, the maximum number of DPs that a PCIe system can contain. For instance, a PCIe system can contain 32 DPs. According to the solutions in related technologies, only one power field (PF) device can be connected to each DP. Therefore, limited by the number of DPs and buses, the number of PF devices that a PCIe system can provide is limited, which obviously greatly restricts the number of PFs that can be presented in the system.
[0081] Therefore, this application provides a scalable hot-swappable technology based on the SHPC mechanism. By adopting a multi-level bridging approach, multiple PF devices are mounted on each DP to achieve high-capacity, high-density, and highly reliable flexible devices, directly presenting the PF devices to the user.
[0082] See Figure 6aThe diagram illustrates a scalable hot-swappable topology based on the SHPC mechanism provided in this application embodiment. To support SHPC hot-swapping, this application embodiment requires bridging a PCIe bridge under the DP and bridging a PCI bridge under the PCIe bridge. That is, the second type DP connects multiple PF devices through its corresponding bridge group. The bridge group includes a PCIe bridge bridging the second type DP and multiple PCI bridges that connect PF devices. Thus, for each PCI bridge, multiple hot-swappable PF devices can be connected under one PCI bridge, typically up to 32 PF devices. The target object can select multiple PCI bridges as needed, or the target object can input the target number of PF devices, and the smart network card will match an appropriate number of PCI bridges for the target object to achieve flexible expansion of elastic devices.
[0083] See Figure 6a As shown, DP includes two types: a first type DP that mounts a single PF device, and a second type DP that mounts multiple PF devices through multi-level bridging technology. In practical scenarios, the first number of the first type DP and the second number of the second type DP can be flexibly configured according to the target number to meet the target number of PF devices required by the target object.
[0084] In one possible implementation, in order to ensure the reliability and stability of the PCIe system, the number of first-type DPs can be increased as much as possible. That is, by maximizing the number of first-type DPs, the first number of first-type DPs and the second number of second-type DPs can be determined by combining the target number and the above-mentioned upper limit value.
[0085] For example, if the maximum number of DPs is 0, then if allowed, one of the DPs can be used as the second type DP, and the rest of the DPs can be the first type DPs. That is, the number of first type DPs is (0-1) and the number of second type DPs is 1.
[0086] In one possible implementation, in order to improve the reliability and stability of the PCIe system, the number of PF devices that can be attached to the second type of DP can also be constrained, that is, to avoid the number of PF devices attached to the same DP being too large. When allocating the first type of DP and the second type of DP, the first number of the first type of DP and the second number of the second type of DP can be obtained by combining the target number, the upper limit of the number of DPs, and the upper limit of the number of PF devices attached to each DP.
[0087] For example, when the total number of PF devices to be mounted is Q, the maximum number of DPs is O, and each DP can mount at most P PF devices, a traversal approach can be used to determine the first and second quantities. That is, firstly, when using 1 second-type DP, it can be traversed to see if the remaining (O-1) DPs can meet the mounting requirements of (QP) PF devices. If they can, the number of first-type DPs is (O-1) and the number of second-type DPs is 1. If they cannot, the number of second-type DPs is incremented by one, and the remaining (O-2) DPs are similarly determined to meet the mounting requirements of (Q-2P) PF devices. If they can, the number of first-type DPs is (O-2) and the number of second-type DPs is 2. If they cannot, the number of second-type DPs is incremented by one, and the above process is repeated until the mounting requirements are met.
[0088] In this embodiment, the second type DP connects multiple PF devices via multi-level bridging technology. This means that each second type DP connects to a bridge group, and the multiple PCI bridges in each bridge group adopt a tree topology. See [link to relevant documentation]. Figure 6b The diagram shown is a schematic of a tree topology provided in an embodiment of this application. The tree topology includes N levels of PCI bridges, and each level of PCI bridge contains at least one PCI bridge. As shown in Figure b, the first-level PCI bridge is used to bridge the PCIe bridge with its corresponding child node in the second-level PCI bridge, that is, to connect the PCIe bridge and the PCI bridge under the PCI bridge. The i-th level PCI bridge is used to bridge its corresponding parent node in the (i-1)-th level PCI bridge with its corresponding child node in the (i+1)-th level PCI bridge, where i is an integer greater than 1 and less than N. The N-th level PCI bridge bridges its corresponding parent node in the (N-1)-th level PCI bridge with at least one PF device attached to it. Thus, through multi-level bridging, a second-type DP can be attached to multiple PF devices.
[0089] In this embodiment, the number of levels N in the tree topology can be determined based on the total number of PF devices required to be mounted on the corresponding second type DP. This allows for flexible configuration of the number and level of PCI bridges according to the total number of PF devices required to be mounted, thereby meeting the target object's needs for PF devices, enabling the deployment and presentation of large-scale PF devices. Furthermore, when presenting the PCIe system to the target object through the aforementioned tree topology, the target object can clearly perceive the architecture of the entire PCIe system, thus facilitating the hot-plugging of PF devices.
[0090] In one possible implementation, when configuring the first and second quantities, the level N of the tree topology can also be considered simultaneously. The level N represents the depth of the tree topology to a certain extent. Increasing the depth will bring certain difficulties to subsequent management and indexing. Once you can increase the upper limit for the level N, then when determining the first and second quantities, you need to consider the upper limit of the large level N, which is similar to the process of considering the upper limit of DP mentioned above. Therefore, it will not be elaborated here.
[0091] Of course, the target object can also choose the ratio or preference of the first type DP and the second type DP, and the first and second quantities can be determined according to the selection of the target object.
[0092] See Figure 6a As shown, it illustrates a hot-swappable system topology with a level N of 1, in which only one level of PCI bridges is included. These PCI bridges are all connected to PCIe bridges and to PF devices.
[0093] See Figure 6c The diagram illustrates another scalable hot-swappable topology based on the SHPC mechanism. This is a hot-swappable system topology with a number of stages N of 2. In this system topology, for a second-type DP, its corresponding bridge group contains two levels of PCI bridges. The first-level PCI bridges are all connected to PCIe bridges and are connected to multiple PCI bridges in the second-level PCI bridges. The second-level PCI bridges are connected to multiple PF devices. By expanding the number of PCI bridge stages, the number of PF devices that can be mounted can be greatly increased, thereby increasing the number of PF devices displayed to the target object.
[0094] Step 403: Based on the first quantity and the second quantity, generate the target system topology corresponding to the target PCIe system; wherein, in the target system topology, the second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes PCIe bridges bridging the second type DP and multiple PCI bridges connected to PF devices.
[0095] In this embodiment of the application, after determining the first quantity and the second quantity, a target system topology corresponding to the target PCIe system can be generated based on the first quantity and the second quantity. This target system topology includes the first quantity of first-type DPs, each of which is connected to a PF device, and also includes the second quantity of second-type DPs, each of which is connected to a bridge group. Each bridge group is as follows: Figure 6b As shown, it includes a PCIe bridge that bridges a second type of DP and a PCI bridge that mounts multiple PF devices.
[0096] Step 404: Notify the host of the target system topology to instruct the host to configure the target PCIe system according to the target system topology.
[0097] Specifically, based on the target system topology generated above, the smart network interface card (NIC) generates corresponding system configuration information and sends it to the host side. This allows the host side to configure a target PCIe system as desired by the target object according to the target system topology. It should be noted that the system configuration information may include other configuration information besides the target system topology; this embodiment does not impose any limitations on this.
[0098] Correspondingly, on the host side, after receiving the target system topology, corresponding configurations can be performed. Therefore, the following section describes the host-side methodology. (See also...) Figure 7 The diagram shown is another flowchart illustrating a method for extending a PCIe system according to an embodiment of this application. This method can be executed by a host, and the specific implementation process is as follows:
[0099] Step 701: Receive the target system topology sent by the smart network card; wherein, in the target system topology, each second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect the PF devices.
[0100] See Figures 6a-6c As shown, the system topology provided in this application embodiment includes a first type DP that mounts a single PF device and a second type DP that mounts multiple PF devices. It can be flexibly configured in actual scenarios. For each second type DP, it can mount multiple PF devices through its corresponding bridge group. The bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that mount PF devices. For details, please refer to the relevant sections above, which will not be repeated here.
[0101] Step 702: Select PF devices that conform to the target system topology from the PF devices included in the system, and call the standard hot-swap controller (SHPC) of each PF device to initialize the state of each PF device.
[0102] In this embodiment, the host side includes a sufficient number of PF devices for the user to choose from. Once the user selects a certain number of PF devices, the PCIe system will be configured according to the selected number to meet the user's expectations.
[0103] Specifically, based on the received target system topology, the host side can determine which of its remaining PF devices can still be configured as the target system topology, and then allocate these PF devices to the target object and call the corresponding SHPC to initialize the state of each PF device, such as initializing the state of these PF devices to the inserted state.
[0104] Step 703: Assign a system identifier to each PF device and return the assigned system identifier to the smart network card; wherein, the system identifier is generated based on the target system topology and is a unique identifier for the target PF device in the target PCIe system.
[0105] The system identifier is used to distinguish different PF devices, and the PF device can be searched based on the system identifier in the future.
[0106] Generally, in a PCIe system, a 32-bit PCIe ID can be used to identify PCIe devices in the system, also known as PF devices. The 32 bits include the following:
[0107] (1) Domain information, i.e. PCIe domain identifier, occupies 1 bit.
[0108] (2) Bus ID, which occupies 8 bits.
[0109] (3) Device ID, which occupies 5 bits.
[0110] (4) Function ID, which occupies 3 bits.
[0111] Among them, bus id, device id, and function id are collectively referred to as BDF information.
[0112] In this embodiment, the process of assigning system identifiers to each PF device can be performed during initialization, or it can be assigned or updated during subsequent use, such as during a host restart or reset. Of course, the system identifier may also contain other information, and this embodiment does not impose any limitations on this.
[0113] In this embodiment of the application, to facilitate rapid device indexing during data plane interactions between PF devices or between PF devices and other devices, this embodiment also provides a method for rapid indexing of PF devices. See [link to relevant documentation]. Figure 8 The diagram shown is a flowchart illustrating the process of storing device information of a PF device according to an embodiment of this application.
[0114] Step 801: In response to the table entry update command, obtain the system identifier of the target PF device to be stored; wherein, the system identifier is generated based on the target system topology and is a unique identifier of the target PF device in the target PCIe system.
[0115] In this embodiment, to facilitate rapid retrieval of device information for each PF device during subsequent data processing, the smart network interface card (NIC) stores and maintains this device information. When this device information is needed, it is quickly indexed using the system identifier of the PF device. This device information may include, for example, the routing information or transmission queue number of the PF device.
[0116] Specifically, when the target PF device is initialized or its information changes, the corresponding table entry for the target PF device can be updated. The smart network card can then receive the table entry update command and obtain the system identifier of the target PF device.
[0117] In one possible implementation, when the host side allocates new device resource information to the PF device, it can notify the smart network card to update the table entries. Thus, the host side can send a table entry update command to the smart network card, carrying the system identifier of the target PF device. For example, the host side can output the system identifier to the smart network card through the Basic Input Output System (BIOS), such as the memory address, BAR address, and BDF information of the target PF device.
[0118] In one possible implementation, the table update command can also be triggered by the target object through the client. In this case, the smart network card can also obtain the system identifier of the target PF device from the BIOS on the host side based on the table update command.
[0119] Step 802: Encode the system identifier using a preset data encoding method to obtain the initial table entry index corresponding to the target PF device.
[0120] In this embodiment, the data encoding method can employ any possible data compression encoding method, such as CRC encoding or hash encoding. Taking CRC encoding as an example, the smart network interface card (NIC) can calculate the CRC index based on the system identifier of the received PF device, which is the entry index corresponding to the target PF device.
[0121] Step 803: Determine if the number of associated PF devices for the initial table entry index has reached the upper limit, i.e., determine if the initial table entry index is available.
[0122] Generally speaking, the number of table entries that a table entry index can associate and store is limited. This upper limit is called the depth of a table entry index. Therefore, after generating the initial table entry index for a PF device, it is necessary to determine whether the initial table entry index has reached the corresponding depth in order to determine the availability of the initial table entry index.
[0123] See Figure 9 The diagram shown illustrates the principle of a multi-level table entry query technique based on CRC encoding, as provided in an embodiment of this application. Each table entry index consists of one or more entries. When all entries corresponding to a given index are unavailable, the process switches to the next lower-level table entry index for reading and writing. Figure 9 As shown, each index corresponds to n entries, where n is the depth of each index. If each entry can be used to store the device information of a PF device, then each index can correspond to n associated PF devices.
[0124] Therefore, when it is necessary to store information about a PF device, the system identifier of the PF device is encoded using the CRC encoding method to obtain the initial table entry index, i.e. Figure 9 The index1 shown can be used to determine whether the n entries corresponding to index1 are available, that is, whether they have been occupied by other PF devices. If they are available, they can be used to store the device information of the current PF device. Otherwise, index1 cannot be used for storage, so a shift process is performed to obtain the next level index2, and the availability of index2 is checked again until an available index is found.
[0125] Step 804: If the result of step 803 is negative, then associate the target PF device with the initial table entry index, and update the table entry corresponding to the target table entry index based on the device information of the target PF device.
[0126] Specifically, when the upper limit is not reached, it indicates that the initial entry index is available, and all information corresponding to that index can be updated. For example, if the device information to be stored is the routing information of the target PF device, such as the bar address, pfindex, and valid flag, then if the initial entry index is available, this information is stored in the entry corresponding to that initial entry index. Of course, the device information can also include any other possible information, and this application embodiment does not limit this.
[0127] Step 805: If the result of step 803 is yes, then obtain the target table entry index from the table entry index cascaded from the initial table entry index, where the number of associated PF devices has not reached the upper limit.
[0128] Specifically, the process involves traversing the cascaded table entry indexes of the initial table entry index until a target table entry index is obtained where the number of associated PF devices has not reached the upper limit. Since each traversal process is similar, this traversal will be used as an example for explanation.
[0129] During each traversal, the current table entry index is shifted to obtain the next-level table entry index corresponding to the current table entry index. It is then determined whether the number of associated PF devices of the next-level table entry index has reached the upper limit. If it has not reached the upper limit, then the next-level table entry index is the target table entry index. If the number of associated PF devices of the next-level table entry index has reached the upper limit, then the next traversal process is entered, and the above process is repeated until a target table entry index that has not reached the upper limit is found.
[0130] In this case, when it is the first traversal, the current table entry index is the initial table entry index; when it is not the first traversal, the current table entry index is the table entry index obtained by shifting during the previous traversal.
[0131] Step 806: Associate the target PF device with the target table entry index, and update the table entry corresponding to the target table entry index based on the device information of the target PF device.
[0132] It should be noted that the table entry query module in the smart network card includes a software component, and the process of storing table entries described above can be implemented by the software component.
[0133] In real-world scenarios, if the system identifier or device information of a PF device changes, such as during a hot-plugging process, the previously stored content needs to be deleted.
[0134] In this embodiment, when the device information of the target device is stored, it can be retrieved by querying when needed later. Therefore, this embodiment provides a method for multi-level table entry query, see [link to relevant documentation]. Figure 10 The diagram shown is a flowchart of a device information query provided in an embodiment of this application. This method can also be implemented by a smart network card in an embodiment of this application.
[0135] Step 1001: In response to the device query command, obtain the system identifier of the target PF device to be queried.
[0136] In this embodiment, when device information stored in the smart network interface card (NIC) is needed, the smart NIC can be requested to query the device information. For example, when the PF device is a NIC, if NIC 1 needs to send a data packet to NIC 2, it needs to query the routing information of NIC 2 via the smart NIC to distribute the data packet to the corresponding packet queue. Similarly, when other devices need to send packets to the PF device, they also need to query the routing information of the PF device via the smart NIC to distribute the packets to the corresponding packet queue.
[0137] Step 1002: Use data encoding methods to encode the system identifier and obtain the query table entry index corresponding to the target PF device.
[0138] Similar to stored procedures, when performing table entry queries, the system identifier of the target PF device is also used as the basis for input. The same data encoding method as in stored procedures is used to encode the system identifier to obtain the query table entry index used in the query process.
[0139] Step 1003: Match the target PF device with the associated PF device in the query table entry index to determine if the match is successful.
[0140] Step 1004: If the match is successful, the device information of the target PF device is obtained based on the query table entry index. That is, the device information of the target PF device is stored in the table entry corresponding to the query table entry index, so the corresponding device information can be obtained from the table entry corresponding to the query table entry index.
[0141] Step 1005: If no match is found, perform a shift operation to obtain the next level table entry index.
[0142] Step 1006: Match the target PF device with the associated PF device in the next level table entry index to determine if the match is successful.
[0143] Step 1007: If a match is successful, obtain the device information of the target PF device based on the index of the successfully matched target table entry. If no match is found, proceed to step 1005 to continue execution.
[0144] Specifically, the table query process still adopts a multi-level table query method, that is, the target PF device is matched with the associated PF devices of the query table entry index or its cascaded table entry index in turn until a matching target table entry index is obtained, and the device information of the target PF device is obtained based on the target table entry index.
[0145] Taking the CRC encoding method as an example, the CRC is first calculated based on the system identifier to obtain the corresponding CRC code, which is then used as an index. The system determines whether the PF device can be successfully matched in the table entry corresponding to the index. If the match is unsuccessful, the CRC code is shifted to obtain a new index, and the matching is performed again. This process is repeated until the device information of the PF device is obtained.
[0146] In this embodiment, the flexible device hot-swapping process implemented by the scalable hot-swappable device based on the SHPC mechanism involves multiple components, is lengthy, and involves complex and diverse scenarios. To simplify the user application program interface (API) and improve the user experience, and to ensure that the hot-swapping state of each PF device meets the user's expectations, this embodiment also provides a highly reliable hot-swapping state machine. This state machine considers various scenarios such as the host not being powered on, the host restarting, hot-swapping interrupt loss, and repeated hot-swapping by the user. It performs state machine rotation based on the user's expected state and the current state of the PF device, ensuring that the hot-swapping state presented to the user is the user's expected state, thereby improving the user experience.
[0147] See Figure 11 The diagram shown is a flowchart of the hot-swap management process provided in an embodiment of this application. This method can be implemented by combining a smart network card and the host side.
[0148] Step 1101: The smart network card sends the target system topology to the target object corresponding to the target PCIe system so that the target system topology can be presented in the terminal device corresponding to the target object.
[0149] That is, after the target system topology is generated, or when the target object queries the topology information, the target system topology can be presented to the target object. Using the system topology provided in the embodiments of this application, a PCIe system containing a large number of PF devices can be directly presented to the target object, and the topology of the PCIe system is clearer and easier to manage.
[0150] Step 1102: The smart network card receives a hot-plug control request sent by the target object. The hot-plug control request is used to request that the target PF device in the target system topology be set to the target hot-plug state.
[0151] In practical applications, the target object can operate on the PF device to switch the hot-plug state of the PF device. For example, it can perform a hot-plug operation on the PF device to set it to the plugged-in state, or perform a hot-plug operation on the PF device to set it to the unplugged state.
[0152] Step 1103: The smart network card sends a hot-plug control message carrying the system identifier of the target PF device to the host. The hot-plug control message carries the system identifier of the target PF device and indicates that the target PF device should be set to the target hot-plug state.
[0153] Step 1104: Based on the target system topology and system identifier, the host side determines the target SHPC corresponding to the target PF device and calls the target SHPC to configure the target PF device to the target hot-swappable state.
[0154] In this embodiment, the host side can search step by step in the target system topology based on the system identifier until the corresponding target PF device is found, and then call its corresponding target SHPC to configure the target PF device to the target hot-swappable state. For example, the host side finds the corresponding PCIe bridge based on the system identifier, and then the PCIe bridge finds the PCI bridge corresponding to the current PF device from the PCI bridges it is connected to, and finally finds the current PF device.
[0155] Specifically, the highly reliable hot-swappable state machine in the smart network interface card stores the hot-swappable state of each PF device. When the smart network interface card receives a hot-swappable control request, it can combine the hot-swappable state of the target PF device stored in its own memory with the state on the host side to determine the hot-swappable control method for the target PF device.
[0156] See Figure 12 The diagram illustrates the state management principle of the high-reliability hot-swap state machine provided in this application embodiment. The hot-swap states of the PF device mainly include an uninserted state (Slot disabled), an inserted state (Slot enabled), and two intermediate states: hot adding and hot removing. Different processing methods are required when initiating a hot-swap operation in different states of the PF device, which will be explained below.
[0157] (1) The target object expects the target hot-plug state to be Slot enabled, that is, the operation initiated by the target object for the target PF device is a hot-plug operation:
[0158] For the first scenario, see [link / reference] Figure 12As shown, when the smart network interface card (NIC) determines that the hot-plug status of the target PF device is "Slot disabled" based on its own stored state information, it needs to detect the host's operating status. If the host is on, the smart NIC sends a hot-plug control message to the host, such as an MSI (message signal interrupt) signal, to indicate that the target PF device should be hot-plugged. At the same time, the smart NIC updates the hot-plug status of the target PF device in its stored state information to "hot adding". The host starts loading the PF device driver until it detects that the driver has been successfully loaded. Then, in response to the driver loading success indication, the smart NIC updates the hot-plug status of the target PF device to "Slot enabled", which is visible to the host.
[0159] For the second scenario, see [link / reference] Figure 12 As shown, based on its own stored state information, it determines that the hot-plug state of the target PF device is not inserted and the host is in a host-off state. For example, if the host is not powered on or restarted, the hot-plug state of the target PF device is directly updated to the target hot-plug state, and it waits for the host to be powered on. When the host is detected to be powered on, a state recovery message is sent to the host to instruct the hot-plug state of the target PF device to be restored to the target hot-plug state.
[0160] The third scenario, see [link / reference] Figure 12 As shown, if the target PF device is determined to be in the Slot enabled or hot adding state based on its own stored state information, it will maintain the current state and will not respond to the hot-plug operation performed by the target object.
[0161] (2) The target object expects the target hot-plug state to be the Slot disable state, that is, the operation initiated by the target object for the target PF device is a hot-plug (unplug) operation:
[0162] For the first scenario, see [link / reference] Figure 12 As shown, if the target PF device is determined to be in the Slot disable state or the hot removing state based on its own stored state information, it will maintain the current state and will not respond to the hot-plug operation performed by the target object.
[0163] For the second scenario, see [link / reference] Figure 12As shown, when the target PF device's hot-plug status is determined to be Slot enabled based on its own stored status information, the host's operating status needs to be checked. If the host is powered on, a hot-plug control message is sent to the host, such as an MSI interrupt signal, to instruct the target PF device to be hot-plugged. At the same time, the smart network card updates the target PF device's hot-plug status to hot removing. The host begins to uninstall the PF device's driver until it detects that the driver uninstallation of the target PF device is successful. Then, in response to the driver uninstallation success indication, the target PF device's hot-plug status is updated to Slot disabled, meaning the target PF device is removed and is no longer visible to the host.
[0164] The third scenario, see [link / reference] Figure 12 As shown, based on its own stored state information, if it determines that the hot-plug state of the target PF device is Slot enabled and the host side is not powered on, it will directly update the hot-plug state of the target PF device to the target hot-plug state and wait for the host side to be powered on. When it detects that the host side is powered on, it will send a state recovery message to the host side to instruct that the hot-plug state of the target PF device be restored to the target hot-plug state.
[0165] In addition, the smart network card will periodically detect the actual status of the PF device on the host side. When an anomaly occurs, such as the loss of the MSI interrupt signal, which causes the actual status on the host side to be inconsistent with the status stored on its own, the status on the host side will be adjusted. At the same time, if multiple hot-plug operations are received from the user in a short period of time, the card will respond selectively, such as responding to the first or last operation.
[0166] In summary, this application provides a scalable hot-swappable technology based on the SHPC mechanism, enabling high-capacity, high-density, and highly reliable elastic devices. It directly presents users with large-scale, elastic, hot-swappable devices. Compared to traditional hot-swappable mechanisms, it significantly reduces bus resource overhead, reducing bus resources to 1 / 30th of the original. A large number of PF devices can directly meet users' elastic device needs without relying on the VF provided by SRIOV technology, greatly reducing user experience and system management complexity. Furthermore, the multi-level table query technology provided in this application improves query efficiency in large-scale PF device scenarios. The addition of a state machine simplifies user operations, internally handles complex and diverse cloud host scenarios, and ensures that the final state of the elastic device meets expectations, improving user experience.
[0167] Please see Figure 13 Based on the same inventive concept, this application also provides a smart network card 130, including:
[0168] The receiving unit 1301 is used to obtain the target number of physical function PF devices required by the target PCIe system to be configured;
[0169] DP configuration unit 1302 is used to obtain, based on the target number and the pre-configured upper limit of the number of downstream port DPs, the first number of first type DPs with a single PF device mounted in the target PCIe system and the second number of second type DPs with multiple PF devices mounted.
[0170] Topology generation unit 1303 is used to generate a target system topology corresponding to the target PCIe system based on the first quantity and the second quantity; wherein, in the target system topology, the second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect the PF devices.
[0171] The sending unit 1304 is used to notify the host of the target system topology, so as to instruct the host to configure the target PCIe system according to the target system topology.
[0172] In one possible implementation, the multiple PCI bridges in each bridge group adopt a tree topology, which includes N levels of PCI bridges, and each level of PCI bridges contains at least one PCI bridge.
[0173] Among them, the first-level PCI bridge bridges the corresponding child nodes of the PCIe bridge and the second-level PCI bridge.
[0174] The i-th level PCI bridge connects its corresponding parent node in the (i-1)-th level PCI bridge to its corresponding child node in the (i+1)-th level PCI bridge, where i is an integer greater than 1 and less than N.
[0175] The Nth-level PCI bridge bridges the corresponding parent node of the (N-1)th-level PCI bridge with at least one PF device attached to it.
[0176] In one possible implementation, the number of levels N in the tree topology is determined based on the total number of PF devices required to be mounted on the corresponding second type of DP.
[0177] In one possible implementation, a multi-level lookup table unit 1305 is also included, for:
[0178] In response to the table entry update command, obtain the system identifier of the target PF device to be stored; wherein, the system identifier is generated based on the target system topology and is a unique identifier of the target PF device in the target PCIe system;
[0179] The system identifier is encoded using a preset data encoding method to obtain the initial table entry index corresponding to the target PF device;
[0180] If the number of associated PF devices in the initial table entry index has reached the limit, then obtain the target table entry index from the table entry index cascaded with the initial table entry index, where the number of associated PF devices has not reached the limit.
[0181] Associate the target PF device with the target table entry index, and update the table entry corresponding to the target table entry index based on the device information of the target PF device.
[0182] In one possible implementation, the multi-level lookup table unit 1305 is specifically used for:
[0183] Traverse the cascading indexes of the initial table entry index until the target table entry index is obtained before the upper limit of the number of associated PF devices is reached; one traversal includes the following steps:
[0184] The current table entry index is shifted to obtain the next level table entry index corresponding to the current table entry index; when it is the first traversal, the current table entry index is the initial table entry index, and when it is not the first traversal, the current table entry index is the table entry index obtained by the previous shifting process;
[0185] If the number of associated PF devices in the next level table entry index has reached the upper limit, then proceed to the next traversal process.
[0186] In one possible implementation, the multi-level lookup table unit 1305 is further configured to:
[0187] In response to a device query command, obtain the system identifier of the target PF device to be queried;
[0188] The system identifier is encoded using a data encoding method to obtain the query table entry index corresponding to the target PF device;
[0189] Match the target PF device with the associated PF device in the query table entry index. If the match is successful, obtain the device information of the target PF device based on the query table entry index.
[0190] If no match is found, the target PF device will be matched sequentially with the associated PF devices of the table entry indexes cascaded with the query table entry indexes until a matching target table entry index is obtained, and the device information of the target PF device will be obtained based on the target table entry index.
[0191] In one possible implementation, a state management unit 1306 is also included, for:
[0192] Send the target system topology to the target object corresponding to the target PCIe system so that the target system topology can be presented in the terminal device corresponding to the target object;
[0193] Receive hot-plug control requests sent by the target object. The hot-plug control requests are used to request that the target PF device in the target system topology be set to the target hot-plug state.
[0194] Send a hot-plug control message carrying the system identifier of the target PF device to the host, so that the host can switch the target PF device to the target hot-plug state based on the system identifier and the target system topology.
[0195] In one possible implementation, if the target hot-plug state is the insertion state, then the state management unit 1306 is specifically used for:
[0196] Based on its own stored state information, when the target PF device's hot-plug status is determined to be "not inserted", the host's operating status is detected.
[0197] If the host is powered on, a hot-insertion control message is sent to the host to instruct a hot-insertion operation on the target PF device and update the hot-insertion status of the target PF device to the hot-insertion in progress status.
[0198] After sending a hot-plug control message carrying the system identifier of the target PF device to the host, the method further includes:
[0199] In response to a successful driver loading indication for the target PF device, update the hot-plug status of the target PF device to the plugged-in state.
[0200] In one possible implementation, if the target hot-plug state is an uninserted state, then the state management unit 1306 is specifically used for:
[0201] Based on its own stored state information, when the hot-plug status of the target PF device is determined to be inserted, the host's operating status is detected.
[0202] If the host is powered on, a hot-swap control message is sent to the host to instruct the target PF device to be hot-swaped and to update the hot-swap status of the target PF device to the hot-swap in progress status.
[0203] After sending a hot-plug control message carrying the system identifier of the target PF device to the host, the method further includes:
[0204] In response to a successful driver uninstallation indication for the target PF device, the hot-plug status of the target PF device is updated to unplugged.
[0205] In one possible implementation, the state management unit 1306 is specifically used for:
[0206] If the host is detected to be in a powered-off state, then update the hot-plug status of the target PF device to the target hot-plug status;
[0207] When the host is detected to be powered on, a status recovery message is sent to the host to instruct that the hot-plug state of the target PF device be restored to the target hot-plug state.
[0208] This device can be used to execute the methods performed by the smart network card in the various embodiments of this application. Therefore, the functions that each functional module of this device can achieve can be referred to the description of the foregoing embodiments, and will not be repeated here.
[0209] Please see Figure 14 Based on the same inventive concept, this application also provides a host 140, including:
[0210] The receiving unit 1401 is used to receive the target system topology sent by the smart network card; wherein, in the target system topology, each second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect the PF devices.
[0211] PCIe system configuration unit 1402 is used to select PF devices that conform to the target system topology from the PF devices included in itself, and call the standard hot-swap controller (SHPC) of each PF device to initialize the state of each PF device.
[0212] Resource configuration unit 1403 is used to assign system identifiers to each of the PF devices and return the assigned system identifiers to the smart network interface card; wherein, the system identifier is generated based on the target system topology and is a unique identifier of the target PF device in the target PCIe system.
[0213] In one possible implementation, the PCIe system configuration unit 1402 is further configured to:
[0214] The system receives a hot-plug control message sent by the smart network card. The hot-plug control message carries the system identifier of the target PF device and is used to indicate that the target PF device is set to the target hot-plug state.
[0215] Based on the target system topology and the system identifier, the target SHPC corresponding to the target PF device is determined, and the target SHPC is invoked to configure the target PF device to the target hot-swappable state.
[0216] This device can be used to execute the methods executed by the host in the various embodiments of this application. Therefore, the functions that each functional module of this device can achieve can be referred to the description of the foregoing embodiments, and will not be repeated here.
[0217] Please see Figure 15 Based on the same technical concept, embodiments of this application also provide a computer device. In one embodiment, the computer device can be the aforementioned smart network card or a host, such as... Figure 15 As shown, it includes a memory 1501, a communication module 1503, and one or more processors 1502.
[0218] The memory 1501 is used to store computer programs executed by the processor 1502. The memory 1501 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0219] Memory 1501 may be volatile memory, such as random-access memory (RAM); memory 1501 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1501 may be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1501 may be a combination of the above-described memories.
[0220] Processor 1502 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 1502 is used to implement the above-described method of expanding the PCIe system when calling a computer program stored in memory 1501.
[0221] The communication module 1503 is used to communicate with terminal devices and other servers.
[0222] This application embodiment does not limit the specific connection medium between the memory 1501, communication module 1503, and processor 1502. This application embodiment... Figure 15 The memory 1501 and the processor 1502 are connected via a bus 1504, and the bus 1504 is in Figure 15The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 1504 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 15 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.
[0223] The memory 1501 stores a computer storage medium, which stores computer-executable instructions. The computer-executable instructions are used to implement the method of the extended PCIe system in the embodiments of this application. The processor 1502 is used to execute the method of the extended PCIe system in the above embodiments.
[0224] Based on the same inventive concept, embodiments of this application also provide a storage medium storing a computer program that, when run on a computer, causes the computer to perform the steps of the methods for extending the PCIe system according to various exemplary embodiments of this application described above.
[0225] In some possible implementations, various aspects of the methods for extending the PCIe system provided in this application can also be implemented in the form of a computer program product, which includes a computer program that, when run on a computer device, causes the computer device to perform the steps in the methods for extending the PCIe system according to various exemplary embodiments of this application described above. For example, the computer device can perform the steps of the various embodiments.
[0226] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0227] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on a computer device. However, the program product of this application is not limited thereto. In this application, the readable storage medium may be any tangible medium that contains or stores a program, and the computer program included therein may be used by or in conjunction with a command execution system, apparatus, or device.
[0228] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.
[0229] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0230] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages.
[0231] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0232] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0233] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0234] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0235] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method of scaling a PCIe system, the method comprising: Applied to smart network interface cards, the method includes: Obtain the target number of Physical Function PF devices required for the target PCIe system to be configured; Based on the target number and the pre-configured upper limit of the number of downstream port DPs, the first number of first type DPs with a single PF device mounted in the target PCIe system and the second number of second type DPs with multiple PF devices mounted are obtained. Based on the first quantity and the second quantity, a target system topology corresponding to the target PCIe system is generated; wherein, in the target system topology, the second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect PF devices. The target system topology is notified to the host to instruct the host to configure the target PCIe system according to the target system topology; In each bridge group, multiple PCI bridges adopt a tree topology structure, which includes N levels of PCI bridges. Each level of PCI bridge contains at least one PCI bridge. The first-level PCI bridge bridges the PCIe bridge with its corresponding child node in the second-level PCI bridge. The i-th level PCI bridge bridges the (i-1)-th level PCI bridge with its corresponding parent node in the (i-1)-th level PCI bridge with its corresponding child node in the (i+1)-th level PCI bridge, where i is an integer greater than 1 and less than N. The N-th level PCI bridge bridges the (N-1)-th level PCI bridge with its corresponding parent node and at least one PF device attached to it.
2. The method of claim 1, wherein, The number of levels N in the tree topology is determined based on the total number of PF devices required to be mounted on the corresponding second type of DP.
3. The method as described in claim 1 or 2, characterized in that, After notifying the host of the target system topology, the method further includes: In response to an entry update command, the system identifier of the target PF device to be stored is obtained; wherein, the system identifier is generated based on the target system topology and is a unique identifier of the target PF device in the target PCIe system; The system identifier is encoded using a preset data encoding method to obtain the initial table entry index corresponding to the target PF device; If the number of associated PF devices in the initial table entry index has reached the upper limit, then obtain the target table entry index from the table entry indexes cascaded with the initial table entry index where the number of associated PF devices has not reached the upper limit. Associate the target PF device with the target entry index, and update the entry corresponding to the target entry index based on the device information of the target PF device.
4. The method as described in claim 3, characterized in that, From the concatenated table entry indexes of the initial table entry index, obtain the target table entry index where the number of associated PF devices does not reach the upper limit, including: Traverse the cascaded table entry indexes of the initial table entry index until a target table entry index is obtained where the number of associated PF devices has not reached the upper limit; wherein, one traversal includes the following steps: The current table entry index is shifted to obtain the next-level table entry index corresponding to the current table entry index; wherein, when it is the first traversal, the current table entry index is the initial table entry index, and when it is not the first traversal, the current table entry index is the table entry index obtained by the previous shifting process; If the number of associated PF devices in the next level table entry index has reached the upper limit, then proceed to the next traversal process.
5. The method as described in claim 3, characterized in that, After updating the entry corresponding to the target entry index based on the device information of the target PF device, the method further includes: In response to a device query command, obtain the system identifier of the target PF device to be queried; The system identifier is encoded using the data encoding method described above to obtain the query table entry index corresponding to the target PF device; The target PF device is matched with the associated PF device in the query table index. If the match is successful, the device information of the target PF device is obtained based on the query table index. If no match is found, the target PF device is matched sequentially with the associated PF devices of the table entry index cascaded with the query table entry index until a matching target table entry index is obtained, and the device information of the target PF device is obtained based on the target table entry index.
6. The method as described in claim 1 or 2, characterized in that, After notifying the host of the target system topology, the method further includes: The target system topology is sent to the target object corresponding to the target PCIe system so that the target system topology can be presented in the terminal device corresponding to the target object; Receive a hot-plug control request sent by the target object, the hot-plug control request being used to request that the target PF device in the target system topology be set to the target hot-plug state; Send a hot-plug control message carrying the system identifier of the target PF device to the host, so that the host switches the target PF device to the target hot-plug state based on the system identifier and the target system topology.
7. The method as described in claim 6, characterized in that, If the target hot-plug state is the insertion state, a hot-plug control message carrying the system identifier of the target PF device is sent to the host, including: Based on its own stored state information, when it is determined that the hot-plug state of the target PF device is not inserted, the operating state of the host is detected. If the host is powered on, a hot-insertion control message is sent to the host to instruct a hot-insertion operation on the target PF device and update the hot-insertion status of the target PF device to the hot-insertion in progress status. After sending a hot-plug control message carrying the system identifier of the target PF device to the host, the method further includes: In response to a successful driver loading indication for the target PF device, the hot-plug status of the target PF device is updated to the inserted state.
8. The method as described in claim 6, characterized in that, If the target hot-plug state is not inserted, a hot-plug control message carrying the system identifier of the target PF device is sent to the host, including: Based on its own stored state information, when the hot-plug status of the target PF device is determined to be inserted, the operating status of the host is detected. If the host is powered on, a hot-swap control message is sent to the host to instruct a hot-swap operation to be performed on the target PF device, and the hot-swap status of the target PF device is updated to the hot-swap in progress status. After sending a hot-plug control message carrying the system identifier of the target PF device to the host, the method further includes: In response to a successful driver uninstallation indication for the target PF device, the hot-plug status of the target PF device is updated to an unplugged state.
9. The method as described in claim 6, characterized in that, Send a hot-plug control message carrying the system identifier of the target PF device to the host, including: If the host is detected to be in a powered-off state, then the hot-plug status of the target PF device is updated to the target hot-plug status; When the host is detected to be powered on, a status recovery message is sent to the host to instruct that the hot-plug state of the target PF device be restored to the target hot-plug state.
10. A method for extending a PCIe system, characterized in that, Applied to a host, the method includes: The system receives the target system topology corresponding to the target PCIe system sent by the smart network card; wherein, in the target system topology, each second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that are connected to PF devices. From the PF devices included in itself, select PF devices that conform to the target system topology, and call the standard hot-swap controller (SHPC) of each PF device to initialize the state of each PF device. Each PF device is assigned a system identifier, and the assigned system identifier is returned to the smart network interface card; wherein, the system identifier is generated based on the target system topology and is a unique identifier for each PF device in the target PCIe system; In each bridge group, multiple PCI bridges adopt a tree topology structure, which includes N levels of PCI bridges. Each level of PCI bridge contains at least one PCI bridge. The first-level PCI bridge bridges the PCIe bridge with its corresponding child node in the second-level PCI bridge. The i-th level PCI bridge bridges the (i-1)-th level PCI bridge with its corresponding parent node in the (i-1)-th level PCI bridge with its corresponding child node in the (i+1)-th level PCI bridge, where i is an integer greater than 1 and less than N. The N-th level PCI bridge bridges the (N-1)-th level PCI bridge with its corresponding parent node and at least one PF device attached to it.
11. The method as described in claim 10, characterized in that, After assigning system identifiers to each of the PF devices and returning the assigned system identifiers to the smart network interface card, the method further includes: Receive a hot-plug control message sent by the smart network card. The hot-plug control message carries the system identifier of the target PF device and is used to indicate that the target PF device is set to the target hot-plug state. Based on the target system topology and the system identifier, the target SHPC corresponding to the target PF device is determined, and the target SHPC is invoked to configure the target PF device to the target hot-swappable state.
12. A smart network interface card (NIC), characterized in that, include: The receiving unit is used to obtain the target number of Physical Function PF devices required by the target PCIe system to be configured; The DP configuration unit is used to obtain, based on the target number and the pre-configured upper limit of the number of downstream port DPs, the first number of first type DPs with a single PF device mounted in the target PCIe system and the second number of second type DPs with multiple PF devices mounted. A topology generation unit is used to generate a target system topology corresponding to the target PCIe system based on the first quantity and the second quantity; wherein, in the target system topology, the second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that connect the PF devices. A sending unit is configured to notify the host of the target system topology, so as to instruct the host to configure the target PCIe system according to the target system topology; In each bridge group, multiple PCI bridges adopt a tree topology structure, which includes N levels of PCI bridges. Each level of PCI bridge contains at least one PCI bridge. The first-level PCI bridge bridges the PCIe bridge with its corresponding child node in the second-level PCI bridge. The i-th level PCI bridge bridges the (i-1)-th level PCI bridge with its corresponding parent node in the (i-1)-th level PCI bridge with its corresponding child node in the (i+1)-th level PCI bridge, where i is an integer greater than 1 and less than N. The N-th level PCI bridge bridges the (N-1)-th level PCI bridge with its corresponding parent node and at least one PF device attached to it.
13. A host computer, characterized in that, include: A receiving unit is used to receive the target system topology corresponding to the target PCIe system sent by the smart network card; wherein, in the target system topology, each second type DP is connected to multiple PF devices through its corresponding bridge group, and the bridge group includes a PCIe bridge that bridges the second type DP and multiple PCI bridges that are connected to PF devices. The PCIe system configuration unit is used to select PF devices that conform to the target system topology from the PF devices it includes, and call the standard hot-swap controller (SHPC) of each PF device to initialize the state of each PF device. The resource configuration unit is used to assign system identifiers to each of the PF devices and return the assigned system identifiers to the smart network interface card; wherein, the system identifier is generated based on the target system topology and is a unique identifier for each PF device in the target PCIe system; In each bridge group, multiple PCI bridges adopt a tree topology structure, which includes N levels of PCI bridges. Each level of PCI bridge contains at least one PCI bridge. The first-level PCI bridge bridges the PCIe bridge with its corresponding child node in the second-level PCI bridge. The i-th level PCI bridge bridges the (i-1)-th level PCI bridge with its corresponding parent node in the (i-1)-th level PCI bridge with its corresponding child node in the (i+1)-th level PCI bridge, where i is an integer greater than 1 and less than N. The N-th level PCI bridge bridges the (N-1)-th level PCI bridge with its corresponding parent node and at least one PF device attached to it.
14. A cloud server, characterized in that, Includes the smart network interface card as described in claim 12 and the host as described in claim 13.
15. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9 or 10-11.
16. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 9 or 10 to 11.
Citation Information
Patent Citations
PCIe Switch system extension management method
CN113515478A