Method for quickly recovering system after PCIE (Peripheral Component Interface Express) network host fails

By using heartbeat packets in PCIE network to monitor the status of the host and realize automatic slave switching, the problem of long system recovery time after the host failure is solved, and the system is quickly recovered and high reliability is achieved.

CN119922070APending Publication Date: 2025-05-02XIAN AVIATION COMPUTING TECH RES INST OF AVIATION IND CORP OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411966676.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

In PCIE network, the system needs a long time to recover after a host failure, resulting in service interruption and loss of service data.

Method used

The heartbeat packet monitors the online status of the host in real time. The slave automatically switches to the host when the host fails, takes over the equipment resources, and restores the failed machine to a backup machine to ensure the system is recovered quickly.

Benefits of technology

It realizes rapid recovery of the system after PCIE network host failure, avoids service interruption and data loss, and improves system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119922070A_ABST
    Figure CN119922070A_ABST
Patent Text Reader

Abstract

The invention discloses a method for quickly recovering a system after a host of a PCIE (Peripheral Component Interface Express) network fails, and aims at solving the problem of ensuring the timeliness and reliability of system recovery through a dual-host main-standby switching method when the host in the PCIE network fails. The method is established on a master-slave dual-computer communication structure taking PCIE (Peripheral Component Interconnect Express) exchange with a non-transparent bridge as a center, dual computers work normally at the same time and control own application domains, a slave computer monitors the online state of a master computer in real time through a heartbeat packet, and when the master computer fails, the slave computer automatically turns into the master computer state and takes over equipment resources managed by the master computer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of embedded computers, and in particular relates to a method for quickly recovering a system after a PCIE network host fails. Background Art

[0002] The PCIE bus uses a high-speed serial point-to-point channel to transmit data with a high transmission rate. Currently, many architectures are connected through the PCIE bus to form a network or system, so ensuring the reliability of the PCIE network is a very important task. Once the host device failure, operating system failure or software system failure in the system cannot be automatically restored in a short time, the entire system will be paralyzed and unable to operate normally. If the host is restarted after repair, it often takes a long time, resulting in the inability to process or lose business data in a timely manner. Summary of the invention

[0003] In order to solve the above problems, the present invention aims at the problem that the system needs a long time to recover after the above host failure, causing service interruption and business data cannot be processed or lost in time, and proposes a method for solving the problem of rapid recovery of the system after the PCIE network host failure. The online and offline status of the host is monitored in real time through the heartbeat packet. When the slave machine inquires that the host is offline, the fault is isolated in time, the device resources are taken over, and the faulty machine that has resumed normal operation is added to the system again as a standby machine. Through this method, the fault of the host is queried and isolated in time, so that the system runs uninterruptedly, and the reliability of the system is improved.

[0004] In view of this, according to one aspect of an embodiment of the present invention, a method for quickly recovering the system after a PCIE network host failure is provided, wherein the host node and the slave node are connected through a non-transparent port, and the host node and the slave node are configured with: heartbeat monitoring service, slave takeover service and faulty machine recovery service.

[0005] Optionally, the heartbeat monitoring service is configured as follows: the heartbeat packet is transmitted via the PCIE bus, the host node periodically sends the heartbeat packet to the shared area, and the slave node periodically obtains the heartbeat data of the shared area to monitor the online status of the host node;

[0006] Optionally, the host node and the slave node transmit heartbeat messages in a shared memory manner, and each node occupies a piece of shared memory, and the occupied shared memory is used to store heartbeat packets.

[0007] Optionally, after any of the host node and the slave node is powered on, the heartbeat message of the shared memory is periodically queried. If the heartbeat message does not change for multiple consecutive cycles, the opposite node is determined to be offline and the node itself is set as the host, otherwise the node itself is set as the slave.

[0008] Optionally, when the host node fails and cannot resume normal operation for a short time, the slave machine queries that the host is not online. The slave machine resets the configuration bridge chip based on the non-transparent port, switches itself to the host, re-enumerates and takes over the devices managed by the faulty machine, and the host machine switches to the faulty machine.

[0009] Optionally, the specific method for the slave to reset and configure the bridge chip based on the non-transparent port is: the slave sets the port connected to the bridge chip as an upstream port, sets the port connected to the host as a non-transparent port, and resets the bridge chip.

[0010] Optionally, after the faulty machine goes offline, the slave machine switches itself to the master and periodically sends heartbeat messages to the shared memory of the faulty machine.

[0011] Optionally, after the failed machine recovers and comes back online, it sets itself as a slave machine and monitors the online status of the master machine.

[0012] Optionally, after the failed machine comes back online, it periodically queries the heartbeat packets sent by the host, considers that the host is online, sets itself as a slave, and monitors the online status of the host in real time.

[0013] The advantages and effects of the present invention are:

[0014] (1) It effectively solves the problem that the system cannot operate normally when the host fails in the PCIE network, thereby improving the reliability of the system.

[0015] (2) Transmitting heartbeat packets in the PCIE bus can quickly detect host failures, isolate the fault recovery system in time, and effectively avoid the problem of service failure or data loss caused by host failures in the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Attached Figure 1 It is a communication diagram of the heartbeat monitoring process in the present invention.

[0017] Attached Figure 2 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] According to one aspect of an embodiment of the present invention, a method for quickly recovering a system after a PCIE network host failure is provided, in which a host node and a slave node are connected through a non-transparent port, and the host node and the slave node are configured with: a heartbeat monitoring service, a slave takeover service, and a faulty machine recovery service.

[0020] Further, the heartbeat monitoring service is configured as follows: the heartbeat packet is transmitted through the PCIE bus, the host node periodically sends the heartbeat packet to the shared area, and the slave node periodically obtains the heartbeat data of the shared area to monitor the online status of the host node;

[0021] Furthermore, the host node and the slave node transmit the heartbeat message in a shared memory manner, and each node occupies a piece of shared memory, and the occupied shared memory is used to store the heartbeat packet.

[0022] Furthermore, after any of the host node and the slave node is powered on, the heartbeat message of the shared memory is periodically queried. If the heartbeat message does not change after multiple consecutive cycles, the opposite node is determined to be offline and the node itself is set as the host, otherwise the node itself is set as the slave.

[0023] Furthermore, when the host node fails and cannot resume normal operation for a short time, the slave machine queries that the host is not online. The slave machine resets the configuration bridge chip based on the non-transparent port, switches itself to the host, re-enumerates and takes over the devices managed by the faulty machine, and the host switches to the faulty machine.

[0024] Furthermore, the specific method for the slave to reset and configure the bridge chip based on the non-transparent port is: the slave sets the port connected to the bridge chip as the upstream port, sets the port connected to the host as the non-transparent port, and resets the bridge chip.

[0025] Furthermore, after the faulty machine goes offline, the slave machine switches itself to the master and periodically sends heartbeat messages to the shared memory of the faulty machine.

[0026] Furthermore, after the faulty machine recovers and comes back online, it sets itself as a slave machine and monitors the online status of the master machine.

[0027] Furthermore, after the failed machine comes back online, it periodically queries the heartbeat packets sent by the host, considers that the host is online, sets itself as a slave, and monitors the online status of the host in real time.

[0028] The invention discloses a method for quickly recovering a system after a PCIE network host fails, aiming to solve the problem of ensuring the timeliness and reliability of system recovery by switching between dual machines when a host fails in a PCIE network.

[0029] The present invention is based on a master-slave dual-machine communication structure centered on a PCIE exchange with a non-transparent bridge. The dual machines work normally at the same time and control their own application domains. The slave machine monitors the online status of the host in real time through a heartbeat packet. When the host fails, the slave machine automatically switches to the host state and takes over the device resources managed by the host. This solution is implemented through the following technical solutions:

[0030] (1) Heartbeat monitoring service: A dual-machine heartbeat message sharing area is established in the PCIE network, which can be accessed by both the host and the slave when they are working normally. The host periodically sends a message to the heartbeat message sharing area to report the current online status, and the slave periodically checks the heartbeat message to monitor whether the host is offline in real time.

[0031] (2) Slave takeover service: Once the host encounters a short-term unrecoverable fault, the slave will detect that the host's heartbeat has stopped within a few cycles and will consider the host to be offline. The slave will then start the reset operation, switch the configuration of the non-transparent port of the bridge chip, automatically switch to the host, and re-enumerate the devices managed by the failed machine. Communication between devices in the PCIE network will proceed normally.

[0032] (3) Faulty machine recovery service: After the faulty machine returns to normal, the devices under the faulty machine node are enumerated, and the heartbeat packet is periodically queried to monitor the online status of the host.

[0033] The present invention can quickly discover host failures through the heartbeat monitoring service, providing a guarantee for timely recovery after system failures; fault isolation is achieved through the slave takeover service, and automatic switching from the slave to the host is completed without manual operation, ensuring uninterrupted operation of the host application domain equipment and improving the reliability of the system.

[0034] Example 1

[0035] A method for quickly recovering a system after a PCIE network host failure mainly includes the following steps:

[0036] Step 1: After the processor node is powered on, the PCIE bus is initialized and the devices under the node are enumerated. The two nodes apply for a shared memory space in their own private memory to support the communication of heartbeat messages between the two nodes.

[0037] Step 2: When any node of the dual machine is powered on, it periodically queries the heartbeat message of the shared memory. If the heartbeat message does not change for three consecutive cycles, it considers that the other node is offline and sets itself as the host. Otherwise, it sets itself as the slave.

[0038] Step 3: The two nodes determine the master-slave relationship. After normal operation, the master periodically writes heartbeat messages to the shared memory of the slave, and the slave periodically checks the heartbeat messages from the shared memory;

[0039] Step 4: Once the slave detects that the heartbeat message does not change for three consecutive cycles, it considers the host (faulty machine) to be faulty, sets itself as the host, reconfigures the upstream port and non-transparent port of the bridge chip, downgrades the port in the host from an upstream port to a non-transparent port, and upgrades the non-transparent port connected to the slave to an upstream port, resets the bridge chip to prevent the processor that has sent the fault from interacting with other parts of the system, and then re-enumerates the device.

[0040] Step 5: After the slave switches to the master, it manages the device resources under the node, maintains the normal operation of the system, and continues to send heartbeat messages to the faulty machine.

[0041] Step 6: After the faulty machine returns to normal, it queries the heartbeat message changes in the shared memory for three consecutive cycles, determines that the other end is the master, sets itself as a slave, and continues to monitor the online status of the master.

[0042] Example 2

[0043] like Figure 2 As shown, the implementation party takes two processors connected through a bridge chip as an example, namely processor A and processor B, processor A is connected to the upstream port of the bridge chip, and processor B is connected to the non-transparent port of the bridge chip. The method described in the present invention includes the following steps:

[0044] Step 1: Processor A is powered on first, while processor B is not powered on. Processor A enumerates the devices under it, allocates resources to the devices, applies for a memory space in the application program executed on processor A, and performs address translation on the non-transparent port of the bridge chip. This shared memory is used as a heartbeat message buffer, which is accessible to processor B.

[0045] Step 2: If processor A finds that the heartbeat message has not changed in three consecutive cycles, processor A is set as the master. After processor B is powered on, it performs the same work as step 1, and if it finds that the heartbeat message has changed, processor B is set as the slave.

[0046] Step 3: Power off the host or stop the software running on the host. The host becomes a faulty machine. The slave queries the heartbeat message for three consecutive cycles without change. Reconfigure the non-transparent port, configure the upstream port connected to the host as a non-transparent port, configure the non-transparent port connected to the slave as an upstream port, reset the bridge chip, switch itself to the host, and re-enumerate the device.

[0047] Step 4: The faulty machine is powered on again and reinitialized to perform the same work as step 1. The faulty machine finds that the other end is online, sets itself as a slave, and periodically queries the heartbeat packet to check whether the host is offline.

[0048] The above is only a specific embodiment of the present invention, and the present invention is described in detail. The unexplained part is a conventional technology. However, the protection scope of the present invention is not limited to this. Any changes or substitutions that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. The protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A method for quickly recovering a system after a PCIE network host failure, characterized in that: Connect the host node and the slave node through a non-transparent port, and configure the following services on both the host node and the slave node: heartbeat monitoring service, slave takeover service, and fault recovery service.

2. A method for quickly recovering a system after a PCIE network host failure according to claim 1, characterized in that: The heartbeat monitoring service is configured as follows: the heartbeat packet is transmitted through the PCIE bus, the host node periodically sends the heartbeat packet to the shared area, and the slave node periodically obtains the heartbeat data of the shared area to monitor the online status of the host node.

3. A method for quickly recovering a system after a PCIE network host failure according to claim 2, characterized in that: The host node and the slave node transmit the heartbeat message in a shared memory manner, and each node occupies a piece of shared memory, and the occupied shared memory is used to store the heartbeat packet.

4. A method for quickly recovering a system after a PCIE network host failure according to claim 2, characterized in that: After any of the host node and the slave node is powered on, the heartbeat message of the shared memory is periodically queried. If the heartbeat message does not change for multiple consecutive cycles, the opposite node is determined to be offline and the node itself is set as the host, otherwise the node itself is set as the slave.

5. A method for quickly recovering a system after a PCIE network host failure according to claim 1, characterized in that: When the host node fails and cannot resume normal operation for a short time, the slave machine queries that the host is not online. The slave machine resets the configuration bridge chip based on the non-transparent port, switches itself to the host, re-enumerates and takes over the devices managed by the faulty machine, and the host machine switches to the faulty machine.

6. A method for quickly recovering a system after a PCIE network host failure according to claim 5, characterized in that: The specific method for the slave to reset and configure the bridge chip based on the non-transparent port is as follows: the slave sets the port connected to the bridge chip as the upstream port, sets the port connected to the host as the non-transparent port, and resets the bridge chip.

7. A method for quickly recovering a system after a PCIE network host failure according to claim 5, characterized in that: After the faulty machine goes offline, the slave switches itself to the master and periodically sends heartbeat messages to the shared memory of the faulty machine.

8. A method for quickly recovering a system after a PCIE network host failure according to claim 1, characterized in that: After the faulty machine recovers and comes back online, it sets itself as a slave and monitors the online status of the master.

9. A method for quickly recovering a system after a PCIE network host failure according to claim 8, characterized in that: After the faulty machine comes back online, it periodically queries the heartbeat packets sent by the host, considers the host to be online, sets itself as a slave, and monitors the online status of the host in real time.

Citation Information

Patent Citations

  • Multi-device master-slave competition method and device, storage medium and electronic device

    CN117675531A

  • Dual-computer hot standby redundancy method and system of ultrahigh-speed maglev traffic central operation and control system

    CN118270078A

  • High available redundant terminal of encrypting based on nontransparent bridge of PCIE

    CN206807466U

  • Method, device, system and storage medium for implementing packet transmission in PCIE switching network

    US20140122768A1