Scale out network backup path fast switching notification system and method

By introducing an inspection FPGA module and a low-power switching chip into the AI ​​server, combined with a device power monitoring module, rapid switching of the scale-out network backup path was achieved, solving the problems of switching delay and power loss, and meeting the high real-time requirements of AI training tasks.

CN120935101AActive Publication Date: 2025-11-11SHANGHAI ORIENTAL COMPUTER TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511461105.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-11-11
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

In existing technologies for AI server scale-out networks, the switching latency of backup paths cannot meet the microsecond-level requirements, and the switching instructions are lost in power failure scenarios, resulting in wasted GPU computing power and task failure. The lag in state synchronization leads to erroneous or missed switching.

Method used

The inspection FPGA module is used to detect faults in real time and generate pre-installed switching instruction messages. Combined with low-power switching chips and equipment power monitoring modules, fault detection and message transmission are achieved through hardware acceleration, ensuring that the switching time does not exceed 1ms and maintaining power supply in power failure scenarios.

Benefits of technology

The backup path switching time has been reduced from 100ms to ≤1ms, solving the problem of AI training task interruption, improving the switching success rate and system stability, and avoiding message loss due to power outage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935101A_ABST
    Figure CN120935101A_ABST
Patent Text Reader

Abstract

The invention relates to a scale out network backup path fast switching notification system and method, and aims to solve the problem that the existing VRRP / BFD protocol switching time cannot meet the high real-time requirement of an AI server scale out network. The scheme is realized through a collaborative architecture of a CPU (Central Processing Unit), an inspection FPGA (Field Programmable Gate Array) and power supply monitoring: the CPU runs a VRRP / BFD (Bidirectional Forwarding Detection) protocol stack and maintains a state table, and the FPGA is responsible for detecting a power supply fault, a port fault and GPU alarm in real time, and generating and sending a pre-installed message in one main frequency period; and the power supply monitoring module ensures that power supply is maintained for more than or equal to 5ms after sudden power failure through the energy storage capacitor, so that complete transmission of fault messages is ensured. The system has the advantages that the system supports multi-scale out interface redundancy, the end-to-end delay of fault detection and backup path switching is smaller than or equal to 1ms by adopting a low-power-consumption switching chip and independent minimum power domain design, the fault response speed and reliability of an AI server scale out network are remarkably improved, and the system is suitable for guaranteeing the network continuity of a high-density AI training cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer systems and relates to the technology of AI server network communication, especially a fast switching notification system and method for scale-out network backup paths. Background Technology

[0002] With the rapid development of artificial intelligence technology, AI server clusters (especially scale-out horizontal architectures) are widely used in high-performance computing scenarios such as distributed training and deep learning inference. These clusters typically build high-speed interconnect networks using RoCE V2 or IB protocols, requiring microsecond-level communication latency between nodes to ensure task continuity. However, network link or node failures can lead to interruptions in training data transmission and even task failures. Therefore, the ability to quickly switch backup paths has become a key technical bottleneck.

[0003] In the current Ethernet field, VRRP (Virtual Router Redundancy Protocol) and BFD (Bidirectional Forwarding Detection Protocol) are the mainstream solutions for implementing path redundancy and fault detection. Among them:

[0004] The VRRP protocol uses virtual IP addresses and priority mechanisms to achieve primary and backup node switching, solving the single point of failure problem. However, it relies on CPU software to process protocol messages, and the switching delay is usually 50-100ms.

[0005] The BFD protocol monitors link status by periodically sending detection messages. The detection time can be configured to milliseconds, but message generation and decision-making still require CPU participation. In extreme scenarios (such as node power failure), message loss may occur due to power outage.

[0006] In AI server scale-out networks, the above techniques have significant limitations:

[0007] 1. Switching latency cannot meet the requirements: AI training tasks typically have a tolerance of less than 1ms for network interruptions, while the 100ms-level switching latency of VRRP / BFD can lead to problems such as wasted GPU computing power and gradient synchronization failure.

[0008] 2. Command loss during power failure: Traditional solutions rely on the CPU and switching chip for power supply. When the device suddenly loses power, the switching command message (such as VRRP Pri=0 announcement) cannot be sent in time, causing the upstream node to be unable to trigger the backup path switch.

[0009] 3. State synchronization lag: The state information (such as port link status and overall machine health) between the CPU and hardware modules is synchronized through software scheduling, which is delayed and can easily lead to state mismatch, resulting in incorrect or missed switching.

[0010] Furthermore, existing technologies are not optimized for the hardware characteristics of AI servers (such as DPU integration and multi-scale-out port design), making them difficult to adapt to high-density node interconnection scenarios. Therefore, there is an urgent need for a fast switching solution that can overcome CPU processing bottlenecks and ensure reliable instruction transmission under extreme scenarios. Summary of the Invention

[0011] The purpose of this invention is to solve the above-mentioned problems existing in the prior art and to provide a scale-out network backup path fast switching notification system.

[0012] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0013] A scale-out network backup path fast switching notification system includes:

[0014] The CPU module is configured to run the VRRP / BFD protocol stack, generate and maintain a state maintenance table containing port ID, link status and overall system status, and generate pre-installed handover instruction messages based on the state maintenance table. The inspection FPGA module is connected to the CPU module and is configured to detect power failures, scale-out port failures and GPU failures in real time. It synchronously stores the status maintenance table and determines whether to send the pre-installed switching instruction message according to a preset strategy. The equipment power monitoring module is configured to monitor the power supply status of the equipment, test the discharge maintenance time of the whole machine or some modules after power failure, and ensure the complete transmission of the pre-installed switching instruction message in the power failure scenario. The switching module, connected to the CPU module and the inspection FPGA module, is configured to support link switching with multiple scale-out ports; When the inspection FPGA module detects a fault, it sends the pre-installed switching instruction message to the upstream device through the switching module to trigger the backup path switching, and the switching response time does not exceed 1ms.

[0015] Preferably, the inspection FPGA module is further configured to: receive VRRP priority and BFD initial state information sent by the CPU module, and have a built-in state maintenance table that is updated synchronously with the CPU module, the state maintenance table including port ID, link status and overall system status.

[0016] Preferably, the device power monitoring module includes: The power input filter monitoring unit is configured to issue an alarm for abnormal power input at the output. The holding capacitor unit is configured to provide energy storage in the event of a power input abnormality, thereby maintaining the power supply to the inspection FPGA module and the switching module; The power domain management unit is configured to extend the discharge time after power failure through a minimum power domain design in high-power scenarios of the device.

[0017] Preferably, the switching module is an external low-power switching chip, which is configured to maintain the ability to send the pre-installed switching instruction message when the device loses power.

[0018] Another objective of this invention is to provide a method for quickly switching notifications for scale-out network backup paths.

[0019] To achieve the second objective mentioned above, the technical solution adopted by the present invention is as follows:

[0020] A method for quickly switching network backup paths in a scale-out manner includes the following steps: S1. Status Detection: Real-time detection of power status, scale-out port link status, and GPU status via FPGA inspection module; S2. Status Synchronization: The CPU module generates a status maintenance table containing port ID, link status, and overall system status, and synchronizes it to the inspection FPGA module. S3. Fault Judgment: The inspection FPGA module determines whether to send a pre-installed switching instruction message based on the preset strategy and the detected fault status, combined with the status maintenance table. S4. Fast switching: If the decision is made to send the pre-installed switching instruction message, the pre-installed switching instruction message is sent to the upstream device through the switching module to trigger the backup path switching, and the response time of the switching does not exceed 1ms; S5. Power Failure Protection: The power supply to the whole machine or some modules is maintained through the equipment power monitoring module after a power failure, ensuring the successful transmission of the pre-installed switching instruction message mentioned in step S4.

[0021] Preferably, in step S3, the preset strategy includes: when the inspection FPGA module detects a power failure, a scale-out port link failure, or a GPU failure, it immediately triggers the sending process of a pre-installed switching instruction message.

[0022] Preferably, in step S5, the device power monitoring module determines the energy storage capacity of the holding capacitor by testing the discharge time, ensuring that the power supply time of the inspection FPGA module and the switching module after power failure is not less than the sending cycle of the pre-installed switching instruction message.

[0023] Preferably, the pre-installed switchover instruction message is a customized message based on the VRRP / BFD protocol, which includes a fault node identifier, a backup path identifier, and a switchover execution instruction.

[0024] Due to the adoption of the above technical solution, the beneficial effects obtained by the present invention include:

[0025] 1. This invention accelerates fault detection and message transmission through FPGA hardware and combines it with hardware forwarding of low-power switching chips to reduce the backup path switching time from the existing 100ms level to ≤1ms, meeting the microsecond-level tolerance requirements of AI training tasks for network interruptions.

[0026] 2. This invention adopts a minimum power domain (powered only to the FPGA and switching module) and a holding capacitor design to ensure that the switching command message can still be sent completely when the device suddenly loses power, solving the problem of message loss caused by power interruption in traditional solutions and effectively improving the switching success rate.

[0027] 3. This invention avoids state mismatch caused by software scheduling by using real-time synchronization of the state tables of the CPU and FPGA and a hardware logic decision mechanism, thus effectively reducing the false switching rate. Attached Figure Description

[0028] Figure 1 This is a flowchart illustrating the structure of an embodiment of the scale-out network backup path fast switching notification system in this invention.

[0029] Figure 2 This is a functional configuration structure diagram of an embodiment of the CPU module and the inspection FPGA module in this invention.

[0030] Figure 3 This is a structural flowchart of an embodiment of the device power monitoring module in this invention.

[0031] Figure 4 This is a flowchart illustrating the structure of an embodiment of software decision-making in this invention.

[0032] Figure 5 This is a flowchart illustrating the structure of an embodiment of the scale-out network backup path fast switching notification method in this invention.

[0033] The attached figures are labeled as follows:

[0034] 1. CPU module; 2. Inspection FPGA module; 3. Equipment power monitoring module;

[0035] 4. Switching module. Detailed Implementation

[0036] Please see the appendix Figure 1-5As shown, this invention mainly relates to a fast switching notification scheme for scale-out network backup paths, primarily applied to AI servers. While existing VRRP and BFD protocols can solve backup and switching issues, their switching times are on the order of 100ms, which is insufficient for the needs of AI servers, especially in the event of sudden power outages or path failures. Therefore, this invention provides a fast switching notification system for scale-out network backup paths, comprising:

[0037] CPU module 1 is configured to run the VRRP / BFD protocol stack, generate and maintain a state maintenance table containing port ID, link status, and overall system status, and generate pre-installed handover instruction messages based on the state maintenance table. Specifically, a multi-core processor can be used to run the VRRP / BFD protocol stack, and the switching module integrated in the DPU (Data Processing Unit) can realize fast forwarding of protocol messages. The state maintenance table can adopt a hash table structure, update the port ID, link status, and overall system status every 10μs, and pre-generate handover instruction messages containing VRRP Pri=0 and BFD Down states, which are stored in the FPGA's high-speed cache.

[0038] The inspection FPGA module 2, which is connected to the CPU module 1, is configured to detect power failures, scale-out port failures, and GPU failures in real time, synchronize the storage status maintenance table, and determine whether to send a pre-installed switching instruction message according to a preset strategy.

[0039] During this period, the communication protocol between the FPGA and the CPU generally adopts a high-speed communication protocol, and PCIe, as the most commonly used high-speed communication interface of the CPU, is widely used for fast access and state synchronization between the FPGA and the CPU; therefore, the communication protocol can follow the high-speed PCIe IO standard, and its interrupt can also adopt the in-band interrupt method.

[0040] In this embodiment, the CPU needs to run both the VRRP and BFD protocol stacks simultaneously. The protocol stacks themselves interact with the FPGA primarily through the PORT / LINK maintenance table, corresponding to the VRRP / BFD parameter table (PRI / STA field definitions). The main content is the status of the PORT and the overall system status; the PORT status needs to correspond to the device's scale-out port.

[0041] The status maintenance table is used by the FPGA to determine which port or link partner to send a message to. The message can be an urgent, highest-priority message or a low-priority status report message. However, the message itself must conform to the VRRP / BFD frame format. Of course, the PORT ID and LINK status ultimately need to be converted into the VRRP / BFD frame structure definition.

[0042] The following is a brief explanation of the two protocols mentioned above, but this invention does not focus on these protocols; it is merely provided as background information. VRRP is a protocol used to provide gateway redundancy backup, primarily addressing the single point of failure issue of the default gateway. BFD, on the other hand, is a mechanism for quickly detecting network link failures, capable of completing detection within milliseconds. BFD itself does not perform switching actions; it only notifies associated upper-layer protocols, such as VRRP / OSPF, to trigger fault handling. A common approach is to use a VRRP+BFD linkage scheme to improve fault switching performance. BFD can compensate for the second-level detection limitation of VRRP, compressing the fault switching time to within 200ms. A typical scenario is that when an uplink failure occurs, BFD notifies VRRP to directly trigger a primary / backup switch, avoiding VRRP's default 3-second wait.

[0043] The detection of port status is related to the deployment of VRRP / BFD. The port status represents the different exits of the scaleout network. For example, if this device has four scaleout network ports, the status of these ports is usually monitored by the FPGA, such as port loss or intermittent disconnection. In this way, the CPU protocol stack can be notified through the port status, and the VRRP protocol stack can quickly switch to the backup port.

[0044] In addition, when the FPGA module detects an interruption due to a power failure alarm, it will determine that the device is about to lose power. Therefore, it will use the interrupt signal to directly trigger message transmission. After one main frequency cycle, the pre-installed message transmission module will be triggered, and the module will start the message transmission process to send the pre-installed message to the FPGA's message transmission module.

[0045] Furthermore, regarding message transmission in emergency situations, there are generally very stringent time requirements. Therefore, FPGAs are used to pre-install messages and then send them with a single click, directly to scale-out ports or switching modules. The most important field in the pre-installed message is the Priority field in the VRRP frame structure definition. The protocol stipulates that when this field is 0, it means that the device has stopped parsing VPPR messages, indicating that the node has failed. Previously, a field value of 0 was used to indicate that the device had experienced an emergency power condition and was about to go offline; after the peer device's VRRP protocol stack received this message, it could switch to the backup device. Device power monitoring module 3 is configured to monitor the device's power supply status and test the discharge sustaining time of the entire machine or some modules after a power outage, ensuring the complete transmission of the pre-installed switching command message in a power outage scenario.

[0046] In this embodiment, the device power monitoring module includes: The power input filtering monitoring unit uses an LC filter circuit to suppress input ripple and monitors the voltage through a comparator. When the input voltage is less than the threshold, it outputs a low-level alarm signal to the FPGA.

[0047] The holding capacitor unit is configured to provide energy storage in the event of a power input abnormality, thereby maintaining the power supply to the inspection FPGA module and the switching module; The power domain management unit is configured to extend the discharge time after power failure through a minimum power domain design in high power consumption scenarios of the device.

[0048] Furthermore, such as Figure 3 As shown, the power supply monitoring scheme in this embodiment is as follows:

[0049] 1. The Input stage is the first-level input of the power supply, mainly responsible for power input filtering and monitoring, and can output power input abnormality alarms.

[0050] 2. Hold-up capacitor is a power holding capacitor, mainly used to provide power storage to ensure the stability of power input. It can also be used to maintain normal power output for a period of time when the power input is abnormal, so as to complete some log recording or last words tasks.

[0051] 3. The On-Board Power Rail Active is the device's Level 2 power supply, primarily responsible for powering various hardware components. When the device requires significant power consumption, a dedicated minimum power domain needs to be designed to ensure that the minimum system can utilize hold-up time to operate for extended periods to complete its intended tasks.

[0052] 4. The Power monitor monitors the power status of the device and can also report power failures to the FPGA to trigger a rapid packet assembly and delivery process.

[0053] 5. FPGA / Packet / ports is the minimum packet sending path of the device, and it is necessary to consider completing the composition and transmission of the message within 1ms.

[0054] In this embodiment, the hardware subsystem of a typical AI server typically uses a 48V bus power supply design, and the subsequent secondary power supplies all operate based on a 48V input voltage. Therefore, the power detection circuit is based on 48V voltage detection; when the 48V bus voltage experiences an abnormal drop, such as falling to 36V, the secondary power supply will malfunction and be unable to output a normal voltage. Therefore, the power detection circuit will detect if the voltage is below 42V and will issue an alarm, indicating that the system will cease operation after time T. Based on actual power consumption measurements of a typical system described above, this time T is generally 2-10ms. If a longer hold-up time is required, a hold-up capacitor circuit will be used to extend the voltage hold-up time.

[0055] The size of the hold-up capacitor can be calculated using the formula: C = 2PΔt / (U1² - U2²), where P is the load power, Δt is the required hold time, U1 is the initial voltage, and U2 is the minimum operating voltage. This formula is based on the principle of energy conservation; the energy stored in the capacitor is E = ½C(U1² - U2²), while the energy consumed by the load is PΔt. For example, when the load power is 375W (considering 85% conversion efficiency) and the hold time is 9ms, the minimum calculated total series capacitance is 820μF.

[0056] Furthermore, a standard signal interrupt protocol is typically used between the FPGA and the power monitoring module. The power monitoring module is generally an ADC voltage acquisition module, which can pre-configure abnormal voltage thresholds via configuration resistors, internal ROM, or the I2C bus. For example, the abnormal voltage can be set to below 42V. Lowering the abnormal voltage setting reduces the time allotted for system response to anomalies, while higher settings increase the risk of false detections of voltage fluctuations, leading to system malfunctions. Therefore, upon detecting an abnormal bus voltage, the power monitoring module reports a hardware interrupt signal to the FPGA. This signal is typically high; a low signal indicates an abnormal alarm. Upon recognizing the power interrupt, the FPGA directly sends a pre-installed VRRP Priority=0 protocol message, generally for all ports; this represents the highest priority.

[0057] Furthermore, through its multi-level power management design, it can ensure that the device can still send complete switching command messages after a sudden power failure, solving the problem of message loss caused by power failure in traditional solutions and improving system stability in extreme scenarios.

[0058] In this embodiment, the switching module 4, connected to the CPU module and the inspection FPGA module, is configured to support link switching with multiple scale-out ports. Specifically, a low-power switching chip can be selected, supporting multiple scale-out ports, and a hardware forwarding engine is used to achieve line-rate forwarding of pre-installed switching instruction messages. Furthermore, the external low-power switching chip is configured to maintain the ability to send pre-installed switching instruction messages when the device loses power. The combination of the low-power switching chip and the power-loss retention design avoids interruption of message forwarding when the device loses power, ensuring reliable transmission of backup path switching instructions and improving the switching success rate.

[0059] When the inspection FPGA module 2 detects a fault, it sends a pre-installed switching instruction message to the upstream device through the switching module to trigger the backup path switching, and the switching response time is no more than 1ms. Furthermore, by accelerating fault detection and message transmission through FPGA hardware, combined with low-power components and minimum power domain design, the backup path switching time is shortened from the 100ms level of the existing technology to ≤1ms, which solves the problem of training task interruption caused by fault switching delay in the AI ​​server scale-out network and improves system reliability.

[0060] like Figure 2 As shown, in this embodiment, the inspection FPGA module 2 is also configured to: receive VRRP priority and BFD initial state information sent by the CPU module, and have a built-in state maintenance table that is updated synchronously with the CPU module. The state maintenance table includes port ID, link status, and overall system status. Specifically, the FPGA receives the VRRP priority (Pri=0 indicates forced switching) and BFD initial state sent by the CPU through the UART interface and stores these parameters in the on-chip BRAM. The state maintenance table adopts a dual-port RAM design. The CPU writes the PORT ID, link status, and overall system status through the AXI bus. The FPGA reads the updated data every 500ns to ensure seamless synchronization with the CPU status information, avoid state mismatch caused by software delay, effectively shorten the fault detection response time, and further ensure the immediacy of the switching command.

[0061] In this embodiment, the FPGA used in the test system can be the XLINK A7-100T series, with a main frequency of up to 250MHz, a SerDes rate of up to 6.25Gbps, and a power consumption of typically 5-8W.

[0062] Power monitoring circuits typically use TI / ADI's ADC detection circuits, equipped with hardware configuration resistors or I2C control interfaces to preset detection thresholds and parameters, as well as hardware interrupt signals to quickly report anomalies.

[0063] Actual test results:

[0064] VRRP / BFD without the present invention does not provide fast processing for power supply anomalies, and the time interval for detecting anomalies is typically 50ms-100ms.

[0065] Using the fast detection and processing mechanism of this invention, the VRRP / BFD switching time is 5-20ms (considering the VRRP response time of the peer, this time period is within expectations).

[0066] In this embodiment, the method for quickly switching network backup paths in a scale-out manner includes the following steps: S1. Status Detection: The power status, scale-out port link status and GPU status are detected in real time by the FPGA inspection module, and the detection results are written to the status maintenance table in real time. S2. Status Synchronization: The CPU module generates a status maintenance table containing port ID, link status and overall system status according to the preset VRRP / BFD strategy (such as port priority, number of fault retries), and synchronizes it to the inspection FPGA module. S3. Fault Judgment: The inspection FPGA module determines whether to send a pre-installed switching instruction message based on the preset strategy and the detected fault status, combined with the status maintenance table. Specifically, the FPGA has built-in decision logic (based on combinational logic circuits). When a power failure, port link down, or GPU alarm is detected, the pre-installed switching instruction message is immediately invoked, and a VRRP Pri=0 message or a BFD fault message is sent through the port module. Its standardized fault notification message format can ensure that upstream devices can quickly identify the switching intention and avoid protocol parsing delays.

[0067] S4. Fast switching: If the decision is to send a pre-installed switching instruction message, the pre-installed switching instruction message is sent to the upstream device through the switching module to trigger the backup path switching, and the switching response time does not exceed 1ms. S5. Power Failure Protection: The power supply module maintains power supply to the whole machine or some modules after a power failure, ensuring the successful transmission of the pre-installed switching instruction message in step S4. Specifically, when the power monitoring module detects that the input voltage drops below the preset value, it immediately notifies the FPGA through an interrupt. The FPGA starts sending the pre-installed switching instruction message within a certain time and automatically cuts off the power after completing the transmission by using the holding capacitor.

[0068] Furthermore, by accelerating the detection, decision-making, and transmission processes through hardware, the end-to-end switching time is reduced from the traditional 100ms level to the 1ms level, which is an effective improvement compared to traditional CPU software processing, meeting the low-latency switching requirements of AI server scale-out networks.

[0069] In this embodiment, the preset strategy in step S3 includes: when the FPGA inspection module detects a power failure, a scaleout port link failure, or a GPU failure, it immediately triggers the sending process of the pre-installed switching instruction message; its refined fault triggering conditions avoid misjudgment, and the hardware logic-implemented decision mechanism compresses the response time to the microsecond level, solving the problem of large decision delay in traditional software strategies.

[0070] In step S5, the device power monitoring module determines the energy storage capacity of the holding capacitor by testing the discharge time, ensuring that the power supply time of the inspection FPGA module and the switching module after power failure is not less than the transmission cycle of the pre-installed switching command message; through scientific discharge time testing and capacitor configuration, it ensures a 100% success rate of the switching command message transmission in the power failure scenario, eliminating the risk of switching failure due to insufficient power supply.

[0071] In this embodiment, the pre-installed switching instruction message is a customized message based on the VRRP / BFD protocol, which includes the fault node identifier, backup path identifier, and switching execution instruction. Its structured message format ensures that the upstream device can quickly parse the switching instruction. Combined with hardware-level path switching, it can achieve end-to-end fault recovery within 1ms and ensure the continuity of AI training tasks.

[0072] It should be noted that, through hardware acceleration and system co-design, three core technological effects were achieved in the AI ​​server scale-out network backup path switching:

[0073] 1. Switching time reduced from 100ms to 1ms: By replacing traditional CPU software processing with FPGA hardware logic, the end-to-end delay of fault detection, message generation and transmission is compressed to ≤1ms, meeting the high real-time requirements of AI training / inference tasks for network link continuity.

[0074] 2. Reliable notification capability under sudden failure: The power monitoring module, combined with the energy storage capacitor design, ensures that the equipment can still maintain power supply for ≥5ms after a sudden power failure, ensuring the complete transmission of VRRP / BFD switching messages and avoiding the loss of switching notifications due to power failure.

[0075] 3. Adaptive handling of multiple fault scenarios: Supports detection of multiple types of faults such as scale-out port faults, GPU alarms, and whole-machine power failures. Through the state maintenance table of FPGA and CPU collaboration and pre-installed message strategies, it realizes differentiated responses such as single-port switching and whole-machine shutdown, improving stability in complex network environments.

[0076] The foregoing descriptions and embodiments are provided to enable those skilled in the art to understand and apply the present invention. It will be apparent to those skilled in the art that various modifications can be easily made to these contents, and the general principles described herein can be applied to other embodiments without creative effort. Therefore, the present invention is not limited to the foregoing descriptions and embodiments. Improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from its scope should be within the protection scope of the present invention.

Claims

1. A scale-out network backup path fast switching notification system, characterized in that, include: The CPU module is configured to run the VRRP / BFD protocol stack, generate and maintain a state maintenance table containing port ID, link status and overall system status, and generate pre-installed handover instruction messages based on the state maintenance table. The inspection FPGA module is connected to the CPU module and is configured to detect power failures, scale-out port failures and GPU failures in real time. It synchronously stores the status maintenance table and determines whether to send the pre-installed switching instruction message according to a preset strategy. The equipment power monitoring module is configured to monitor the power supply status of the equipment, test the discharge maintenance time of the whole machine or some modules after power failure, and ensure the complete transmission of the pre-installed switching instruction message in the power failure scenario. The switching module, connected to the CPU module and the inspection FPGA module, is configured to support link switching with multiple scale-out ports; When the inspection FPGA module detects a fault, it sends the pre-installed switching instruction message to the upstream device through the switching module to trigger the backup path switching, and the switching response time does not exceed 1ms.

2. The system according to claim 1, characterized in that, The inspection FPGA module is also configured to receive VRRP priority and BFD initial state information sent by the CPU module, and to have a built-in state maintenance table that is updated synchronously with the CPU module. The state maintenance table includes port ID, link status and overall system status.

3. The system according to claim 1, characterized in that, The device power monitoring module includes: The power input filter monitoring unit is configured to issue an alarm for abnormal power input at the output. The holding capacitor unit is configured to provide energy storage in the event of a power input abnormality, thereby maintaining the power supply to the inspection FPGA module and the switching module; The power domain management unit is configured to extend the discharge time after power failure through a minimum power domain design in high-power scenarios of the device.

4. The system according to claim 1, characterized in that, The switching module is an external low-power switching chip, which is configured to maintain the ability to send the pre-installed switching instruction message when the device loses power.

5. A method for fast switching notification of scale-out network backup paths based on the system described in any one of claims 1-4, characterized in that, Includes the following steps: S1. Status Detection: Real-time detection of power status, scale-out port link status, and GPU status via FPGA inspection module; S2. Status Synchronization: The CPU module generates a status maintenance table containing port ID, link status, and overall system status, and synchronizes it to the inspection FPGA module. S3. Fault Judgment: The inspection FPGA module determines whether to send a pre-installed switching instruction message based on the preset strategy and the detected fault status, combined with the status maintenance table. S4. Fast switching: If the decision is made to send the pre-installed switching instruction message, the pre-installed switching instruction message is sent to the upstream device through the switching module to trigger the backup path switching, and the response time of the switching does not exceed 1ms. S5. Power Failure Protection: The power supply to the whole machine or some modules is maintained through the equipment power monitoring module after a power failure, ensuring the successful transmission of the pre-installed switching instruction message mentioned in step S4.

6. The method according to claim 5, characterized in that, In step S3, the preset strategy includes: when the inspection FPGA module detects a power failure, a scale-out port link failure, or a GPU failure, it immediately triggers the sending process of a pre-installed switching instruction message.

7. The method according to claim 5, characterized in that, In step S5, the device power monitoring module determines the energy storage capacity of the holding capacitor by testing the discharge time, ensuring that the power supply time of the inspection FPGA module and the switching module after power failure is not less than the sending cycle of the pre-installed switching instruction message.

8. The method according to claim 5, characterized in that, The pre-installed switchover instruction message is a customized message based on the VRRP / BFD protocol, which includes the fault node identifier, backup path identifier, and switchover execution instruction.

Citation Information

Patent Citations

  • Ethernet automatic protection link failure quick switching method

    CN101867495A

  • Master-slave upper port protective switching method and device based on UTN tunnel

    CN108156036A

  • Civil aviation VHF radio station main and standby switching method and system based on virtual IP

    CN118869458A

  • Equipment power failure prompting method based on FPGA capacitor discharge

    CN120455257A

  • Field programmable gate array (FPGA) and fast multi-protocol label switching (MPLS) operation administration and maintenance (OAM) protection switching method

    WO2012065301A1

Cited By

  • Fault processing method, device and equipment based on VRRP and BFD linkage and storage medium

    CN122120192A