Fault transfer method and device of switching equipment, equipment and storage medium
By switching port group configurations and re-enumerating hosts on the PCIe Switch, the service interruption problem caused by single point of failure was resolved, the path transfer of PCIe devices and service continuity were realized, and the system reliability and hardware resource utilization were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-09
- Publication Date
- 2026-04-07
AI Technical Summary
In business areas such as servers, storage systems, and communication platforms, single points of failure can lead to service interruptions. How to perform failover of PCIe switches to ensure service continuity is an urgent problem to be solved.
By implementing multiple independent port groups on a PCIe Switch, with each port group connecting a host and multiple endpoint devices, when a host link anomaly is detected, the configuration of the target downlink port is switched from the first configuration to the backup configuration, and a failover completion event is reported to the target host, causing it to re-enumerate and initialize in order to restore data access and service processing.
It enables path transfer of PCIe devices in the event of a failure, ensuring service continuity, improving system reliability and hardware resource utilization, and reducing system costs.
Smart Images

Figure CN121814715A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device and storage medium for fault transfer of a switching device. Background Technology
[0002] In business areas such as servers, storage systems, and communication platforms, any single point of failure can lead to service interruption and significant losses. Therefore, system design typically considers redundancy and fault tolerance. In multi-processor servers, blade systems, or high-performance computing clusters, different computing nodes (hosts) need to flexibly and dynamically allocate and share peripheral I / O resources. By using PCIe switches that support multiple hosts and failover, PCIe devices such as GPUs, FPGAs, NVMe, and SSDs can be pooled, allowing multiple server nodes to share access on demand. This improves hardware resource utilization and provides redundant paths for devices. Therefore, how to implement PCIe switch failover has become a crucial technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0003] The purpose of this application is to provide a method, apparatus, device, and storage medium for failover of a switching device, which can complete the failover of a PCIe Switch and ensure service continuity.
[0004] To address the aforementioned technical problems, this application provides a method for fault transfer of a switching device, comprising: When the failover unit detects an abnormal host link status, it switches the configuration of the target downlink port in the first target port group from the first configuration to the second configuration; the first target port group is the port group with an abnormal host link status; both the first configuration and the second configuration include a port group identifier and a device identifier; The routing information of the target downlink port is deleted from the first target port group according to the device identifier in the first configuration, and the routing information of the target downlink port is added to the second target port group according to the device identifier in the second configuration; the second target port group is the port group corresponding to the port group identifier in the second configuration; A failover completion event is reported to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port; the target host is the host connected to the uplink port in the second target port group.
[0005] In some embodiments, the method further includes, before switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration: Determine whether the first target port group has the failover function enabled based on the port group control register; If the failover function is enabled in the first target port group, the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration.
[0006] In some embodiments, the method further includes, before switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration: Determine whether the target downlink port is enabled for failover based on the control register of the downlink port; If the target downlink port has the failover function enabled, the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration.
[0007] In some embodiments, the failover unit is associated with a port group and a downlink port, and the failover unit supports multiple fault triggering methods.
[0008] In some embodiments, the failover unit supports watchdog triggering, software configuration triggering, and general input / output interface input triggering.
[0009] In some embodiments, it also includes: When a failure of a primary link is detected, the corresponding primary link status indicator is set. When a failure is detected in a secondary link, the corresponding secondary link status indicator is set.
[0010] In some embodiments, the method further includes the following steps before reporting the failover completion event to the target host: The target downlink port is reset.
[0011] To address the aforementioned technical problems, this application also provides a failover device for switching equipment, comprising: The configuration switching module is used to switch the configuration of the target downlink port in the first target port group from the first configuration to the second configuration when the failover unit detects an abnormal host link status; the first target port group is the port group with an abnormal host link status; both the first configuration and the second configuration include a port group identifier and a device identifier; The routing information addition / deletion module is used to delete the routing information of the target downlink port in the first target port group according to the device identifier in the first configuration, and to add the routing information of the target downlink port in the second target port group according to the device identifier in the second configuration; the second target port group is the port group corresponding to the port group identifier in the second configuration; The event interruption reporting module is used to report a failover completion event to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port; the target host is the host connected to the uplink port in the second target port group.
[0012] To address the aforementioned technical problems, this application also provides a switching device for implementing the steps of the failover method for the switching device as described above.
[0013] To address the aforementioned technical problems, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the failover method for the switching device as described above.
[0014] The failover method for switching equipment provided in this application includes: when the failover unit detects an abnormal host link status, switching the configuration of the target downlink port in the first target port group from a first configuration to a second configuration; the first target port group is the port group with the abnormal host link status; both the first configuration and the second configuration include a port group identifier and a device identifier; deleting the routing information of the target downlink port in the first target port group according to the device identifier in the first configuration, and adding the routing information of the target downlink port in the second target port group according to the device identifier in the second configuration; the second target port group is the port group corresponding to the port group identifier in the second configuration; reporting a failover completion event to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port; the target host is the host connected to the uplink port in the second target port group.
[0015] As can be seen, the failover method for switching equipment provided in this application, based on the port group characteristics of a PCIe Switch, implements multiple independent port groups on a single PCIe Switch, with each port group connecting a host and multiple endpoint devices. When the failover unit detects a failure in the host link of a certain port group, it triggers all or part of the downlink ports within that port group to be transferred to a pre-configured backup port group. After the transfer is completed, a failover completion event is reported to the host of the backup port group, which then re-enumerates and initializes the port group and continues to provide service. This completes the path transfer of the PCIe devices within the failed port group, ensuring service continuity.
[0016] The failover device, switching equipment, and computer-readable storage medium provided in this application all have the aforementioned technical effects. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the prior art and embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a failover method for a switching device provided in an embodiment of this application; Figure 2 A schematic diagram of a port group provided in an embodiment of this application; Figure 3 A schematic diagram of a functional module provided in an embodiment of this application; Figure 4 This is a schematic diagram of a failover process provided in an embodiment of this application; Figure 5 A logic diagram provided for an embodiment of this application; Figure 6 This is a schematic diagram illustrating the effect of a failover completion provided in an embodiment of this application. Detailed Implementation
[0019] The core of this application is to provide a method, apparatus, device, and storage medium for failover of switching equipment, which can complete the failover of PCIe Switch and ensure service continuity.
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a failover method for a switching device provided in an embodiment of this application. (Refer to...) Figure 1 As shown, the method includes: S101: When the failover unit detects an abnormal host link status, it switches the configuration of the target downlink port in the first target port group from the first configuration to the second configuration; the first target port group is the port group with an abnormal host link status; both the first configuration and the second configuration include a port group identifier and a device identifier; S102: Delete the routing information of the target downlink port in the first target port group according to the device identifier in the first configuration, and add the routing information of the target downlink port in the second target port group according to the device identifier in the second configuration; the second target port group is the port group corresponding to the port group identifier in the second configuration; S103: Report a failover completion event to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port; the target host is the host connected to the uplink port in the second target port group.
[0022] The switching device is a PCIe Switch (a switching device based on the PCIe protocol). This embodiment utilizes the PortGroup feature of a PCIe Switch to implement multiple independent PortGroups on a single PCIe Switch. Each PortGroup connects a host and multiple EPs (End Points). The host can only enumerate EPs within the same PortGroup. When the failover unit detects a failure in a host within a PortGroup, it triggers the transfer of all or some DPs (Downstream Ports) within that PortGroup to a pre-configured backup PortGroup. After the transfer is complete, a failover completion event is reported to the host of the backup PortGroup. The host of the backup PortGroup then re-enumerates and initializes the DPs and continues to provide service. This completes the path transfer of PCIe devices within the failed PortGroup, ensuring service continuity. PortGroup refers to the virtual switch technology implemented by the PCIe Switch, which can divide a physical switch into multiple isolated PCIe domains.
[0023] A PCIe switch consists of multiple uplink ports and multiple downlink ports. Uplink ports connect to hosts, while downlink ports connect to endpoint devices. A PCIe switch is divided into multiple PortGroups, each independent and isolated from the others. Each uplink port belongs to one PortGroup, meaning there is a one-to-one correspondence between uplink ports and PortGroups. PortGroups allow for the addition and deletion of downlink ports, and routing between uplink and downlink ports is implemented within each PortGroup. A host connected to an uplink port can enumerate all downlink ports within that PortGroup and all endpoint devices connected to those downlink ports.
[0024] For example, refer to Figure 2 As shown, the PCIe Switch includes PortGroup0 and PortGroup1. The uplink port in PortGroup0 is connected to HOST0. Downlink port P1 in PortGroup0 is connected to endpoint device EP1, and downlink port P2 in PortGroup0 is connected to endpoint device EP2. The uplink port in PortGroup1 is connected to HOST1. Downlink port P4 in PortGroup1 is connected to endpoint device EP4, and downlink port P5 in PortGroup1 is connected to endpoint device EP5.
[0025] The PortGroup selection configuration (PortX_Ctrl.portgrp_id) is implemented in the control register of the downlink port to establish the association between the downlink port and the PortGroup. PortX is the downlink port number, and portgrp_id is the PortGroup identifier. For example, Port1_Ctrl.portgrp_1 indicates that the downlink port numbered Port1 is associated with PortGroup1, and the downlink port numbered Port1 belongs to PortGroup1.
[0026] The downlink port has a failover control register (PortX_Fo_Ctrl), which implements two sets of configurations: Primary and Secondary. The Primary (first) configuration and the Secondary (second) configuration respectively store the portgrp_id (port group identifier) and device_id (device identifier). The device identifier is the identifier of the downlink port. The port group identifier in the Secondary configuration is the identifier of the backup port group.
[0027] When the failover unit detects a failure in any host, it triggers a switch from the first configuration to the second configuration of the target downlink port in the first target port group; the first target port group is the port group to which the uplink port connected to the failed host belongs. The target downlink port is the downlink port in the first target port group.
[0028] For example, refer to Figure 2 As shown, when a fault is detected in HOST0, the configuration of downlink port P1 is switched from Primary to Secondary.
[0029] In some embodiments, the failover unit is associated with a port group and a downlink port, and the failover unit supports multiple fault triggering methods.
[0030] In some embodiments, the failover unit supports watchdog triggering, software configuration triggering, and general input / output interface input triggering.
[0031] In some embodiments, the method further includes, before switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration: Determine whether the first target port group has the failover function enabled based on the port group control register; If the failover function is enabled in the first target port group, the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration.
[0032] The PortGroup control register allows configuration of whether or not failover is enabled for port groups. If the PortGroup control register indicates that failover is enabled for the first target port group, then upon detecting a host failure on an uplink port in the first target port group, failover will occur, and the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration. If the PortGroup control register indicates that failover is not enabled for the first target port group, then upon detecting a host failure on an uplink port in the first target port group, failover will not occur, and therefore the configuration of the target downlink port in the first target port group will not be switched from the first configuration to the second configuration.
[0033] For example, if a failover enable configuration (GroupX_Ctrl.fo_en) is implemented in the PortGroup control register, then when the host connected to the uplink port in PortGroupX fails, the PortGroup control register determines that the failover function of PortGroupX is enabled, and then the configuration of the target downlink port in PortGroupX is switched from the first configuration to the second configuration.
[0034] In addition, the PortGroup control register can also be configured with a failover unit selection (GroupX_Ctrl.fo_id). When the failover function is enabled in PortGroup, the configured failover unit controls the state switching of PortGroup.
[0035] In some embodiments, the method further includes, before switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration: Determine whether the target downlink port is enabled for failover based on the control register of the downlink port; If the target downlink port has the failover function enabled, the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration.
[0036] The DP control register, specifically the downlink port control register, allows configuration to enable or disable failover for the downlink port. If the downlink port control register indicates that the target downlink port in the first target port group has failover enabled, the configuration of the target downlink port in the first target port group is switched from the first configuration to the second configuration. If the downlink port control register indicates that the target downlink port in the first target port group has not enabled failover, the configuration of the target downlink port in the first target port group is not switched from the first configuration to the second configuration.
[0037] For example, if the failover enable configuration (PortX_Ctrl.fo_en) is implemented in the DP control register, then when the host connected to the uplink port in the port group to which the downlink port PortX belongs fails, the DP control register determines that the failover function of PortX is enabled, and then switches the configuration of the downlink port PortX from the first configuration to the second configuration.
[0038] In addition, the DP control register can also be configured to select a failover unit (PortX_Ctrl.fo_id). When the failover function is enabled on the downstream port, the selected failover unit controls the configuration switching of the downstream port.
[0039] In some embodiments, the method further includes, before switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration: Determine whether the target downlink port is enabled for failover based on the control register of the downlink port; Determine whether the first target port group has the failover function enabled based on the port group control register; If failover is enabled on both the target downlink port and the first target port group, then the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration.
[0040] The PortGroup control register allows configuration to enable or disable failover for port groups, and the DP control register allows configuration to enable or disable failover for downlink ports. If both the target downlink port and the first target port group have failover enabled, failover occurs, and the configuration of the target downlink port in the first target port group is switched from the first configuration to the second configuration. Otherwise, failover does not occur.
[0041] After switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration, the device identifier deletes the routing information of the target downlink port in the first target port group and adds the routing information of the target downlink port in the second target port group according to the device identifier; the second target port group is the port group corresponding to the port group identifier in the second configuration.
[0042] For example, refer to Figure 2 As shown, when HOST0 fails, the configuration of downlink ports P1 and P2 in PortGroup0 is switched from the first configuration to the second configuration, where the port group identifier is PortGroup1. Based on the device identifier of downlink port P1, the routing information for downlink port P1 is deleted from PortGroup0; based on the device identifier of downlink port P2, the routing information for downlink port P2 is deleted from PortGroup0; based on the device identifier of downlink port P1, the routing information for downlink port P1 is added to PortGroup1; and based on the device identifier of downlink port P2, the routing information for downlink port P2 is added to PortGroup1.
[0043] A failover completion event is reported to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port; the target host is the host connected to the uplink port in the second target port group.
[0044] For example, refer to Figure 2 As shown, when HOST0 fails, the configuration of downlink ports P1 and P2 in PortGroup0 is switched from the first configuration to the second configuration. The port group in the second configuration is identified as PortGroup1, and the target host is HOST1. A failover completion event is reported to HOST1. After receiving the failover completion event, HOST1 re-enumerates and initializes the PCIe domain to discover the newly added downlink ports P1 and P2 and the connected EP devices (EP1 and EP2), thereby restoring data access and service processing for EP1 / EP2.
[0045] In some embodiments, it also includes: When a failure of a primary link is detected, the corresponding primary link status indicator is set. When a failure is detected in a secondary link, the corresponding secondary link status indicator is set.
[0046] The Primary and Secondary link status indicator flags (PortX_Sts.PFo and PortX_Sts.SFo) can be implemented in the DP's status register to identify whether a Primary link (the main link) or a Secondary link (the secondary link) has failed. When a failure of the Primary link is detected, the corresponding Primary link status indicator flag is set. When a failure of the Secondary link is detected, the corresponding Secondary link status indicator flag is set. A Primary link refers to a link configured under the Primary configuration, and a Secondary link refers to a link configured under the Secondary configuration.
[0047] In some embodiments, the method further includes the following steps before reporting the failover completion event to the target host: The target downlink port is reset.
[0048] As a specific implementation method, the failover method for switching equipment provided in this application embodiment is based on hardware logic.
[0049] As a specific implementation method, the failover method for switching equipment provided in this application embodiment can be implemented based on the DP control register, failover control register, DP status register, DP control module, PortGroup control module, failover unit, PortGroup control register, and failover module in a PCIe Switch.
[0050] refer to Figure 3 As shown, the PortGroup selection configuration (PortX_Ctrl.portgrp_id) is implemented in the DP's control register to establish the association between DP and PortGroup.
[0051] The failover enable configuration (PortX_Ctrl.fo_en) and failover unit selection configuration (PortX_Ctrl.fo_id) are implemented in the DP control register. When the failover function is enabled, the selected failover unit controls the DP configuration switching.
[0052] DP implements the failover control register (PortX_Fo_Ctrl), which contains two sets of configurations for Primary and Secondary, storing the portgrp_id and device_id of the Primary and Secondary configurations respectively.
[0053] The DP status register contains Primary and Secondary link status indicator flags (PortX_Sts.PFo and PortX_Sts.SFo) to indicate whether the Primary and Secondary links have failed.
[0054] The DP control module is used to control DP state transitions, PortGroup migrations, DP state returns, and DP automatic reset logic. DP states can include normal states, transition states, etc.
[0055] The PortGroup control module is used to control the state transitions and return to the previous state of a PortGroup, and to control the DP state switching. The state of a PortGroup can include normal state, fault state, change completed state, etc.
[0056] The failover unit is used to detect whether the host link status is normal. The failover unit supports multiple fault triggering methods such as watchdog triggering, software configuration triggering, and GPIO input triggering.
[0057] The failover enable configuration (GroupX_Ctrl.fo_en) and failover unit selection configuration (GroupX_Ctrl.fo_id) are implemented in the PortGroup control register. When the failover function is enabled, the configured failover unit controls the state switching of the PortGroup.
[0058] The failover module is used to centrally manage the failover units and output fault information to the DP control module and the PortGroup control module. In terms of timing, it ensures that after a fault occurs, the DP configuration switch is performed first, followed by the PortGroup state switch.
[0059] When a failover unit detects a failure in a host link, it will trigger a switch of all DP ports associated with that failover unit from Primary configuration to Secondary configuration. It will also trigger the PortGroup associated with that failover unit to add or remove DP ports and re-establish routes.
[0060] After the failover unit completes the switchover of the associated DP to its PortGroup, it notifies the UP of the backup PortGroup to report a failover completion event to the HOST host. Upon receiving the interruption, the HOST host re-enumerates the PCIe domain, allocates and initializes the newly added PCIe devices, and continues to provide data and service, thereby completing the backup link transfer operation of the PCIe devices.
[0061] refer to Figure 4As shown, the following describes a specific failover embodiment. Figure 4 The arrows with numbers in the middle indicate the steps to be performed during failover.
[0062] (1) Failover unit 0 detected a host failure.
[0063] (2) Failover unit 0 notifies the UP of PortGroup0 to generate a link abnormal event, marks the fault trigger status, and records the trigger type.
[0064] (3) Failover unit 0 notifies the DP control module to perform PortGroup switching.
[0065] (4) The DP control module switches the PortGroup to which DP0 belongs and the DeviceID of DP0 from the Primary configuration to the Secondary configuration.
[0066] (5) The DP control module returns the DP0 PortGroup configuration completion status.
[0067] (6) The failover module notifies the PortGroup control module to delete the routing information of DP0 in PortGroup0 using the DeviceID of DP0, and to add the routing information of DP0 in PortGroup1 using the DeviceID of DP0; PortGroup1 forwards packets from UP and DP0 based on the new routing information, thereby completing the PortGroup switchover of DP0.
[0068] (7) The PortGroup control module notifies the DP control module to reset the DP.
[0069] (8) PortGroup switching is complete, and the failover module is notified.
[0070] (9) The fault transfer module triggers a fault transfer event interrupt to the UP of PortGroup1 and reports it to the Host. The fault transfer is complete.
[0071] The logical relationship after the switch is completed is as follows: Figure 5 As shown.
[0072] refer to Figure 2 As shown, the following describes an embodiment of a 1+1 failover.
[0073] like Figure 2As shown, the system has two host scenarios (HOST0 and HOST1). The PCIeSwitch divides the two hosts into two domains, PortGroup0 and PortGroup1. The UP port of PortGroup0 is connected to HOST0, and two DP ports (P1 and P2) are connected to it. P1 and P2 are connected to EP devices EP1 and EP2, respectively. The UP port of PortGroup1 is connected to HOST1, and two DP ports (P4 and P5) are connected to it. P4 and P5 are connected to EP devices EP4 and EP5, respectively.
[0074] After HOST0 completes enumeration, it can enumerate devices such as UP, P1 / EP1, and P2 / EP2 in PortGroup0. After HOST1 completes enumeration, it can enumerate devices such as UP, P4 / EP4, and P5 / EP5 in PortGroup1.
[0075] Both PortGrp0 and P1 / P2 have failover enabled and are bound to failover unit 0. In the primary configuration of P1 and P2, portgrp is configured as Portgroup0, and in the secondary configuration, portgroup is configured as Portgroup1.
[0076] When failover unit 0 detects a fault in HOST0 (watchdog timeout, software trigger, IO trigger, etc.), it switches the portgroup ID configuration of the associated P1 and P2 from Primary to Secondary matching. After the switch is complete, the UP of Portgroup1 reports a failover completion event interrupt to HOST1. Upon receiving the interrupt, HOST1 re-enumerates and initializes the PCIe domain to discover the newly added P1 / P2 and its connected EP devices (EP1 / EP2), thereby restoring data access and service processing for EP1 / EP2. The completed result is as follows: Figure 6 As shown.
[0077] In summary, the failover method for switching equipment provided in this application, based on the port group characteristics of a PCIe switch, implements multiple independent port groups on a single PCIe switch, with each port group connecting one host and multiple endpoint devices. When the failover unit detects a host link failure in a port group, it triggers the transfer of all or part of the downlink ports within that port group to a pre-configured backup port group. After the transfer is complete, a failover completion event is reported to the host of the backup port group, which then re-enumerates and initializes the system and continues to provide service. This completes the path transfer of PCIe devices within the failed port group, ensuring service continuity. This method provides a solution for backup link switching of EP devices within a failed PCIe domain when a single host failure occurs in a multi-host PCIe domain scenario, improving system reliability, increasing EP device utilization, and reducing system costs. This method enables the deployment and fault backup of multi-CPU PCIe systems using a single PCIe switch. When a CPU fails, the connected PCIe EP devices can be transferred to the normal backup CPU PCIe system to continue providing data and business services, providing a low-cost and simple system backup solution for multi-CPU PCIe interconnection scenarios.
[0078] This application also provides a failover device for a switching device, the device comprising: The configuration switching module is used to switch the configuration of the target downlink port in the first target port group from the first configuration to the second configuration when the failover unit detects an abnormal host link status; the first target port group is the port group with an abnormal host link status; both the first configuration and the second configuration include a port group identifier and a device identifier; The routing information addition / deletion module is used to delete the routing information of the target downlink port in the first target port group according to the device identifier in the first configuration, and to add the routing information of the target downlink port in the second target port group according to the device identifier in the second configuration; the second target port group is the port group corresponding to the port group identifier in the second configuration; The event interruption reporting module is used to report a failover completion event to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port; the target host is the host connected to the uplink port in the second target port group.
[0079] In some embodiments, it also includes: The first judgment module is used to determine whether the first target port group has the failover function enabled based on the port group control register. If the failover function is enabled in the first target port group, the configuration switching module will switch the configuration of the target downlink port in the first target port group from the first configuration to the second configuration.
[0080] In some embodiments, it also includes: The second judgment module is used to determine whether the target downlink port is enabled for failover based on the control register of the downlink port. If the target downlink port has the failover function enabled, the configuration switching module will switch the configuration of the target downlink port in the first target port group from the first configuration to the second configuration.
[0081] In some embodiments, a failover unit is used to detect whether the host link status is abnormal; the failover unit is associated with a port group and a downlink port, and the failover unit supports multiple fault triggering methods.
[0082] In some embodiments, the failover unit supports watchdog triggering, software configuration triggering, and general input / output interface input triggering.
[0083] In some embodiments, it also includes: The first setting module is used to set the corresponding main link status indicator when a failure of the main link is detected. The second setting module is used to set the corresponding secondary link status indicator when a secondary link failure is detected.
[0084] In some embodiments, it also includes: The reset module is used to reset the target downlink port before reporting the failover completion event to the target host.
[0085] This application also provides a switching device for implementing the following steps: When the failover unit detects an abnormal host link status, it switches the configuration of the target downlink port in the first target port group from the first configuration to the second configuration. The first target port group is the port group with the abnormal host link status. Both the first configuration and the second configuration include a port group identifier and a device identifier. The unit deletes the routing information of the target downlink port in the first target port group according to the device identifier in the first configuration, and adds the routing information of the target downlink port in the second target port group according to the device identifier in the second configuration. The second target port group is the port group corresponding to the port group identifier in the second configuration. A failover completion event is reported to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port. The target host is the host connected to the uplink port in the second target port group.
[0086] For a description of the equipment provided in this application, please refer to the above method embodiments; further details will not be provided here.
[0087] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can perform the following steps: When the failover unit detects an abnormal host link status, it switches the configuration of the target downlink port in the first target port group from the first configuration to the second configuration. The first target port group is the port group with the abnormal host link status. Both the first configuration and the second configuration include a port group identifier and a device identifier. The unit deletes the routing information of the target downlink port in the first target port group according to the device identifier in the first configuration, and adds the routing information of the target downlink port in the second target port group according to the device identifier in the second configuration. The second target port group is the port group corresponding to the port group identifier in the second configuration. A failover completion event is reported to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port. The target host is the host connected to the uplink port in the second target port group.
[0088] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0089] For a description of the computer-readable storage medium provided in this application, please refer to the above method embodiments; further details will not be repeated here.
[0090] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatuses, devices, and computer-readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant details can be found in the method section.
[0091] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0092] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0093] The foregoing has provided a detailed description of the failover method, apparatus, device, and storage medium for the switching equipment provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A method for failover of a switching device, characterized in that, include: When the failover unit detects an abnormal host link status, it switches the configuration of the target downlink port in the first target port group from the first configuration to the second configuration. The first target port group is a port group with abnormal host link status; both the first configuration and the second configuration include a port group identifier and a device identifier; The routing information of the target downlink port is deleted from the first target port group according to the device identifier in the first configuration, and the routing information of the target downlink port is added to the second target port group according to the device identifier in the second configuration. The second target port group is the port group corresponding to the port group identifier in the second configuration; Report a failover completion event to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port; The target host is the host connected to the uplink port in the second target port group.
2. The fault transfer method for switching equipment according to claim 1, characterized in that, Before switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration, the following steps are also included: Determine whether the first target port group has the failover function enabled based on the port group control register; If the failover function is enabled in the first target port group, the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration.
3. The failover method for switching equipment according to claim 1, characterized in that, Before switching the configuration of the target downlink port in the first target port group from the first configuration to the second configuration, the following steps are also included: Determine whether the target downlink port is enabled for failover based on the control register of the downlink port; If the target downlink port has the failover function enabled, the configuration of the target downlink port in the first target port group will be switched from the first configuration to the second configuration.
4. The failover method for switching equipment according to claim 1, characterized in that, The failover unit is associated with the port group and the downlink port, and the failover unit supports multiple fault triggering methods.
5. The fault transfer method for switching equipment according to claim 4, characterized in that, The failover unit supports watchdog triggering, software configuration triggering, and general input / output interface input triggering.
6. The failover method for switching equipment according to claim 1, characterized in that, Also includes: When a failure of a primary link is detected, the corresponding primary link status indicator is set. When a failure is detected in a secondary link, the corresponding secondary link status indicator is set.
7. The failover method for switching equipment according to claim 1, characterized in that, Before reporting the failover completion event to the target host, the following also applies: The target downlink port is reset.
8. A failover device for a switching equipment, characterized in that, include: The configuration switching module is used to switch the configuration of the target downlink port in the first target port group from the first configuration to the second configuration when the failover unit detects an abnormal host link status; the first target port group is the port group with an abnormal host link status; both the first configuration and the second configuration include a port group identifier and a device identifier; The routing information addition / deletion module is used to delete the routing information of the target downlink port in the first target port group according to the device identifier in the first configuration, and to add the routing information of the target downlink port in the second target port group according to the device identifier in the second configuration. The second target port group is the port group corresponding to the port group identifier in the second configuration; The event interruption reporting module is used to report the failover completion event to the target host so that the target host can re-enumerate and initialize to restore data access and service processing to the target downlink port and the endpoint devices connected to the target downlink port. The target host is the host connected to the uplink port in the second target port group.
9. A switching device, characterized in that, The switching device is used to implement the steps of the failover method for the switching device as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the failover method for the switching device as described in any one of claims 1 to 7.
Citation Information
Patent Citations
PCIe topology switching method, device and equipment and readable storage medium
CN120950441A
Configurable PCI express switch
CN1694079A
Virtualized PCI switch
US20060242353A1
System and method for port migration in a pcie switch
US20170046295A1