Method for maintaining service continuity and network device

CN122247896APending Publication Date: 2026-06-19EVOC SMART IOT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EVOC SMART IOT TECH CO LTD
Filing Date
2026-04-17
Publication Date
2026-06-19

Smart Images

  • Figure CN122247896A_ABST
    Figure CN122247896A_ABST
Patent Text Reader

Abstract

This application relates to the field of network service technology and discloses a method and network device for maintaining service continuity. The method is applied to a network device including a main computing module, a data exchange module, and an independent management module. The method includes: the independent management module, in response to detecting a fault in the main computing module, sending an early warning signal to the data exchange module and, after waiting for a preset time, controlling the main computing module to perform a hard reboot; the data exchange module, in response to the early warning signal, clearing the network policies issued in real time by the main computing module and processing the network traffic accessed by the network device based on preset basic forwarding rules to maintain service continuity. Through the above method, this application achieves the maintenance of service continuity of the network device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network service technology, specifically to a method and network device for maintaining service continuity. Background Technology

[0002] With the convergence of high-performance computing and network communication, key network devices such as high-end routers, unified threat management devices, load balancers, and servers with integrated smart network interface cards constitute the core hub of modern information systems. However, the central processing unit (CPU) of such network devices is designed to perform the dual functions of computing and network forwarding simultaneously.

[0003] In scenarios with stringent reliability requirements, such as modern data centers, industrial automation, and telecommunications networks, if the CPU of such network equipment fails due to hardware or software problems, not only will the processing that relies on the CPU's computing power immediately stop, but the network services that rely on the CPU's real-time forwarding will also be interrupted simultaneously, causing a complete paralysis of business traffic, which in turn leads to service losses and economic risks.

[0004] To address CPU failures in such network devices, existing technologies widely employ out-of-band management solutions based on Baseboard Management Controllers (BMCs). This solution monitors the status of hardware sensors (such as temperature and voltage) and detects periodic heartbeat signals sent by the host operating system to remotely monitor device health. When a heartbeat times out, the system is deemed to be frozen, triggering a hardware-level restart of the main CPU.

[0005] However, the above solutions have limitations when dealing with the aforementioned high-reliability scenarios. Heartbeat monitoring cycles are typically on the order of minutes, failing to achieve second-level fault detection, resulting in a long time for upstream and downstream systems to synchronously perceive service interruptions. Furthermore, the recovery mechanisms offered by these solutions are singular and highly destructive; a hard reboot, as the sole recovery action, unconditionally interrupts all device functions, including basic network forwarding, during the reboot (often lasting several minutes). This fails to maintain any form of service continuity during CPU failures, thus amplifying component-level failures into complete site service interruptions.

[0006] Therefore, how to maintain the service continuity of such network devices has become a pressing technical problem. Summary of the Invention

[0007] In view of the above problems, embodiments of this application provide a method and network device for maintaining business continuity, which are used to solve the problem of poor business continuity in the prior art.

[0008] According to one aspect of the embodiments of this application, a method for maintaining service continuity is provided. The method is applied to a network device, which includes a main computing module, a data exchange module, and an independent management module. The method includes: in response to detecting a failure in the main computing module, the independent management module sends an early warning signal to the data exchange module and, after waiting for a preset time, controls the main computing module to perform a hard restart; in response to the early warning signal, the data exchange module clears the network policies issued in real time by the main computing module and processes the network traffic accessed by the network device based on preset basic forwarding rules to maintain service continuity.

[0009] In some embodiments, the method further includes: the main computing module periodically sending a heartbeat signal to the independent management module during operation; the independent management module monitors the heartbeat signal, and if it does not receive a heartbeat signal within a preset monitoring period, it identifies that the main computing module has malfunctioned.

[0010] In some embodiments, after the independent management module controls the main computing module to perform a hard reboot, the method further includes: if the main computing module recovers normally after the hard reboot, then reissues the network policy to the data exchange module; the data exchange module configures the reissued network policy to resume the execution of advanced service functions.

[0011] In some embodiments, after the independent management module controls the main computing module to perform a hard reboot, the method further includes: if the independent management module still does not detect a heartbeat signal within a preset recovery time, then performing a hard reboot on the main computing module no more than a first preset number of times, wherein if a heartbeat signal is detected within the preset recovery time after any hard reboot, then the subsequent hard reboot operation is terminated; if, after performing the first preset number of hard reboots, the independent management module does not detect a heartbeat signal within each preset recovery time, then the independent management module issues a fault alarm.

[0012] In some embodiments, if the independent management module fails to detect a heartbeat signal within a preset recovery time after performing a first preset number of hard reboots, the independent management module issues a fault alarm, including: if the independent management module fails to detect a heartbeat signal within a preset recovery time after performing a first preset number of hard reboots, the independent management module reads the power status signal of the main computing module; if the power status signal is abnormal, a first alarm signal containing power fault information is generated; if the power status signal is normal, a second alarm signal containing operating status fault information is generated.

[0013] In some embodiments, the method further includes: a data exchange module responding to the power-on of a network device by processing network traffic accessed by the network device based on basic forwarding rules.

[0014] In some embodiments, the method further includes: the independent management module responding to the power-on and power-on trigger commands of the network device to control the main computing module to power on and start; and if the independent management module detects a fault in the main computing module after the main computing module is powered on and before the system loading is completed, the main computing module is hard-rebooted.

[0015] In some embodiments, the method further includes: after the independent management module controls the main computing module to power on, if it does not receive a power-on self-test completion signal sent by the main computing module within a first time window, it determines that the main computing module has failed during the power-on self-test phase and performs no more than a second preset number of hard reboots. If a power-on self-test completion signal is received within the first time window after any hard reboot, the subsequent hard reboot operation is terminated. If, after performing the second preset number of hard reboots, no power-on self-test completion signal is received within the first time window after each hard reboot, the fault is identified as a power-on self-test phase fault.

[0016] In some embodiments, the method further includes: after receiving the power-on self-test completion signal, if the independent management module does not receive the operating system loading completion signal sent by the main computing module within a second time window, it determines that the main computing module has failed during the operating system loading phase and performs no more than a third preset number of hard reboots. If the operating system loading completion signal is received within the second time window after any hard reboot, the subsequent hard reboot operation is terminated. If, after performing the third preset number of hard reboots, the operating system loading completion signal is not received within the second time window after each hard reboot, the fault is identified as an operating system loading phase fault.

[0017] According to another aspect of the embodiments of this application, a network device is also provided, including a main computing module, a data exchange module and an independent management module, the network device being used to perform the service continuity maintenance method as described in any of the preceding claims.

[0018] This application embodiment separates computing and switching capabilities by simultaneously setting up a main computing module and a data switching module in the network device. This allows the switching functions that the main computing module should possess to be transferred to the data switching module. During a failure or restart of the main computing module, the data switching module no longer stagnates due to its reliance on the main computing module's policies. Instead, it proactively clears complex network policies that depend on the real-time issuance of the main computing module and maintains basic network traffic forwarding based on preset basic forwarding rules. This ensures that the network device can still maintain basic connectivity as a physical link node during a main computing module failure, limiting the impact of the failure to computing functions and maintaining the availability of network forwarding functions, thereby maintaining service continuity under main computing module failure. Furthermore, after identifying the fault, the independent management module does not immediately perform a hard restart. Instead, it first sends a warning signal to the data switching module and waits for a preset time before controlling the main computing module to restart. This provides necessary buffer time for the main computing module's restart, preventing sudden power outages and restarts from causing state instability in the data switching module, further ensuring network service continuity.

[0019] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0020] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This illustration shows an application scenario diagram of the business continuity maintenance method provided in the embodiments of this application; Figure 2 A schematic diagram of the network device provided in an embodiment of this application is shown; Figure 3 A schematic diagram of the power supply structure of the network device provided in an embodiment of this application is shown; Figure 4 A flowchart illustrating the business continuity maintenance method provided in an embodiment of this application is shown.

[0021] The reference numerals in the detailed embodiments are as follows: 1. Network equipment; 11. Main computing module; 12. Data exchange module; 13. Independent management module; 14. Power supply module. Detailed Implementation

[0022] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein.

[0023] Figure 1 This illustration shows an application scenario diagram of the business continuity maintenance method provided in the embodiments of this application, such as... Figure 1 As shown, network device 1 is located between the upstream and downstream devices and is used to forward data between the upstream and downstream devices.

[0024] Network device 1 is a key physical node in the network architecture that undertakes tasks such as data forwarding, computing, security protection, or traffic management.

[0025] Specifically, network device 1 can take the form of various types of devices to adapt to different application scenarios. First, network device 1 can be a network infrastructure device, such as a high-end router, switch, unified threat management device, or next-generation firewall. These devices are typically deployed at network backbone nodes or data center egress points, responsible for high-speed forwarding and security policy execution in massive data streams. They need to handle complex routing protocols and security algorithms, as well as handle line-speed forwarding. Second, network device 1 can be a traffic management and optimization device, such as a load balancer or WAN optimization controller. These devices deeply analyze business traffic and distribute or optimize it according to policies, requiring extremely high computing performance. If the CPU fails, a traditional system reboot will cause complete failure of traffic scheduling. Third, network device 1 can be a high-reliability computing system, such as a server in a data center with integrated smart network cards, or a general-purpose server carrying complex network functions. Furthermore, network device 1 can also be a dedicated device in industrial control or telecommunications networks, such as an industrial control computer in the industrial control field, a user plane function network element in a telecommunications network, or an edge computing node. In these scenarios, network devices 1 are generally designed to perform both computation and network forwarding functions simultaneously through their internal CPUs. However, these scenarios have extremely stringent reliability requirements and need to prevent communication service interruptions caused by single points of failure of the CPU.

[0026] Upstream and downstream devices refer to the sending and receiving ends of data flow relative to network device 1 in the network topology. An upstream device is a hardware entity or logical node that sends data traffic or request commands to network device 1. In different application scenarios, upstream devices can be core switches, backbone routers, user terminal devices (such as personal computers and mobile terminals), sensor nodes, or cloud servers. A downstream device is a hardware entity or logical node that receives data traffic or response results processed and forwarded by network device 1. Accordingly, downstream devices can be access switches, server clusters, storage devices, actuator nodes, or data center gateways.

[0027] For example, network device 1 is a next-generation firewall deployed in a data center, with a core switch as the upstream device and an application server cluster as the downstream device. When a user accesses an application service via the Internet, the data traffic first reaches the core switch (upstream device) and is then sent to network device 1. Network device 1 performs deep packet inspection on the data packets according to security policies. After determining that the packets are legitimate, it forwards them to the corresponding application server (downstream device).

[0028] Figure 2 A schematic diagram of the network device provided in an embodiment of this application is shown, such as... Figure 2 As shown, the network device 1 includes a main computing module 11, a data exchange module 12, and an independent management module 13. The network device 1 is used to perform a service continuity maintenance method.

[0029] Network device 1 may also include a motherboard, with the main computing module 11, data exchange module 12 and independent management module 13 all located on the motherboard of network device 1.

[0030] The main computing module 11 is the core hardware unit in network device 1 responsible for running the operating system, executing complex network protocol processing, and performing business logic operations. Its hardware components include a CPU, memory (such as double-rate synchronous dynamic random access memory), and high-speed memory (such as a solid-state drive). The main computing module 11 undertakes complex computing tasks in network device 1, such as routing protocol calculations, deep packet inspection, and encryption / decryption operations. Under normal operating conditions, the main computing module 11 is also responsible for issuing complex business configuration rules to the data exchange module 12, such as access control lists, quality of service policies, and dynamic routing table entries.

[0031] The data exchange module 12 is a hardware unit in network device 1 responsible for high-speed data packet forwarding. Its hardware components include a switching chip, a network interface (such as multiple gigabit or 10-gigabit Ethernet ports), and independent storage (used to store forwarding table entries).

[0032] The data exchange module 12 is capable of independent operation, and its internal functions include basic forwarding functions and advanced service functions. Basic forwarding functions include Layer 2 Media Access Control Address (MAC) learning and forwarding, basic virtual LAN processing, etc. This part of the function can run independently after the switching chip is powered on and initialized, without relying on the software configuration of the main computing module 11. Advanced service functions include access control lists, quality of service policies, and complex routing policies issued by the main computing module 11. During the failure or restart of the main computing module 11, the basic forwarding function of the data exchange module 12 can still function normally, thereby maintaining basic network connectivity.

[0033] The independent management module 13 is an independent hardware unit in network device 1 responsible for status monitoring and fault recovery control. Its hardware components include a microcontroller unit (MCU), a monitoring interface, and a control interface.

[0034] The independent management module 13 does not participate in the processing of business data, but is only responsible for status monitoring and fault response control. Unlike the traditional out-of-band management scheme based on BMC, which uses multiple dedicated chips such as power timing controller, watchdog timer, front panel button controller and complex programmable logic device to implement management functions, the embodiment of this application uses a general-purpose microcontroller to replace all the above-mentioned dedicated devices, achieving a high degree of integration design.

[0035] The independent management module 13 features a multi-functional integrated design. It includes an external power identification circuit for recognizing the connection of an external power source or the physical power-on of network device 1. This circuit is an interface circuit composed of optocoupler isolation devices. Its input side is connected to the rectified output of the external AC power supply, and its output side is connected to the microcontroller's wake-up pin (e.g., PA0) via a signal line. When an external power source is connected, the optocoupler conducts, the signal line level changes, and the microcontroller is woken up. Specifically, the microcontroller of the independent management module 13 detects the connection status of the external AC power source through the optocoupler isolation circuit. The input side (primary side) of the optocoupler is connected to the rectified output of the AC power supply, and the output side (secondary side) is connected to the microcontroller's wake-up pin. When an external power source is connected, the optocoupler conducts, the microcontroller wakes up from low-power standby mode and begins initialization to ensure electrical safety isolation between the high-voltage side and the microcontroller's low-voltage side.

[0036] The independent management module 13 is also used to identify the power-on command of network device 1. The independent management module 13 completes the identification of the power-on trigger command by scanning multiple signal lines through the internal logic of the microcontroller. For example, the signal line connected to the front panel button is connected to the microcontroller pin after passing through a resistor-capacitor filtering circuit. The microcontroller firmware scans the level of the signal line, debouncing it to confirm it as a power-on command. The automatic startup configuration is determined by reading the status of the signal line corresponding to the configuration jumper or memory.

[0037] Furthermore, the microcontroller in the independent management module 13 also supports three power-on command recognition methods for network device 1: front panel physical button triggering, remote network command triggering, and incoming call auto-start configuration triggering. These are managed uniformly through priority arbitration logic within the firmware. For example, the front panel power button of network device 1 is connected to the microcontroller's input pin after hardware debouncing via a resistor-capacitor low-pass filter circuit. The microcontroller firmware periodically scans this pin, and after detecting a continuous low level and debouncing it, it is confirmed as a valid power-on trigger event.

[0038] Regarding the connection between the independent management module 13 and the main computing module 11, a serial communication bus connects them. This serial communication bus physically includes two signal lines: a serial data line (SDA) and a serial clock line (SCL). These two signal lines are connected via board-level traces to the I2C controller pins of the independent management module 13 and the main computing module 11. Based on this serial communication bus, the independent management module 13 can receive startup status signals sent by the main computing module 11. For example, after the main computing module 11 completes its power-on self-test or the operating system loads, it sends a predefined bytecode (such as 0x55 or 0xAA) via SDA. The independent management module 13 monitors the startup phase by parsing this bytecode.

[0039] A general-purpose input / output line is also connected between the independent management module 13 and the main computing module 11. This connection line physically includes multiple independent signal lines, including a runtime heartbeat signal line and a power enable signal line. The runtime heartbeat signal line is a physical wire that connects a general-purpose output pin of the main computing module 11 to a general-purpose input pin of the independent management module 13. The main computing module 11 outputs periodic square wave pulses as a heartbeat signal through this signal line, and the independent management module 13 monitors the operating status by detecting level transitions on this signal line. The power enable signal line is a physical wire that connects a general-purpose output pin of the independent management module 13 to the control terminal (such as the gate of a P-MOSFET) of the power supply circuit of the main computing module 11. The independent management module 13 physically connects or disconnects the power supply path of the main computing module 11 by changing the level state (high level or low level) on this signal line, thereby achieving hard reboot control.

[0040] Regarding the connection between the independent management module 13 and the data exchange module 12, a hardware warning line connects them. This connection line physically includes multiple independent signal lines, including a collaborative warning signal line and a service recovery signal line. The collaborative warning signal line is a physical wire that connects a general-purpose output pin of the independent management module 13 to a general-purpose input pin of the data exchange module 12. When the independent management module 13 determines that the main computing module 11 is faulty, it sends a specific pulse signal (e.g., a 50-millisecond high-level pulse) through this signal line. After the data exchange module 12 captures the pulse on this signal line, it immediately triggers its internal emergency mode switching logic without software protocol interaction, thus achieving a microsecond-level response speed. The service recovery signal line is also a physical wire that connects a general-purpose output pin of the independent management module 13 to the reset or mode configuration pin of the data exchange module 12. When the main computing module 11 recovers, the independent management module 13 notifies the data exchange module 12 to exit the emergency mode through this signal line.

[0041] Regarding the connection between the main computing module 11 and the data exchange module 12, they are respectively configured with high-speed data bus interfaces (such as PCIe interfaces) and management interfaces (such as Ethernet management interfaces). These interfaces are connected to multiple parallel differential signal lines. The main computing module 11 issues complex service configuration rules to the data exchange module 12 through these interfaces. In terms of hardware design, the switching function of the main computing module 11 is stripped away, and the data exchange module 12 has an independent forwarding hardware engine. It only relies on the main computing module 11 when there are advanced policy requirements. In the basic forwarding mode, the service data channels and control channels of the two are logically decoupled to ensure that the failure of the main computing module 11 does not affect the basic forwarding.

[0042] Figure 3 A schematic diagram of the power supply structure of the network device provided in an embodiment of this application is shown, such as... Figure 3 As shown, network device 1 includes a power module 14, which supplies power to the main computing module 11, the data switching module 12, and the independent management module 13. Network device 1 also internally includes three independent power rails, each powered through an independent physical power interface. The first is a standby power rail, which supplies power to the independent management module 13 through an independent power interface; its physical circuitry is energized upon connection to an external AC power source. The second is the main power rail, which supplies power to the main computing module 11 through a controlled power interface; the on / off state of this interface is controlled by the aforementioned power enable signal line. The third is the switching chip power rail, which supplies power to the data switching module 12 through an independent power interface. It is directly connected to the output of power module 14 and is not affected by the power control signal lines of the main computing module 11 or the independent management module 13, thus ensuring physical isolation of the power supply to each module.

[0043] The power supply timing of the main computing module 11, data exchange module 12, and independent management module 13 is described below. When an external AC power supply is connected to the power module 14, the standby power rail immediately powers on the independent management module 13 through the power interface. Subsequently, the switching chip power rail powers on the data exchange module 12 through its independent power interface. At this time, the power interface of the main power rail is disconnected. After the independent management module 13 completes initialization, it identifies the power-on command issuance method by recognizing the power-on command issuance methods such as front panel physical button triggering, remote network command triggering, and power-on automatic start. When the power-on command is confirmed, the independent management module 13 pulls the power enable signal line high to connect the power interface of the main power rail and supply power to the main computing module 11. During the startup of the main computing module 11, the independent management module 13 monitors the startup status of the main computing module 11 through the serial communication bus. If the main computing module 11 fails, the independent management module 13 disconnects the main power rail by controlling the power enable signal line to achieve a restart. During this process, the data exchange module 12 continues to operate because its power interface is independent.

[0044] Figure 4 A flowchart illustrating a service continuity maintenance method provided in an embodiment of this application is shown, the method being executed by network device 1. Figure 4 As shown, the method includes the following steps S210~S220: S210, in response to the detection of a fault in the main computing module 11, the independent management module 13 sends a warning signal to the data exchange module 12, and after waiting for a preset time, controls the main computing module 11 to perform a hard restart.

[0045] In S210, a failure in the main computing module 11 indicates a failure during its operation. For runtime failures, the independent management module 13 identifies the failure via a runtime heartbeat signal line. If the independent management module 13 does not detect a level transition signal via this signal line within a preset heartbeat timeout period (e.g., 5 seconds), it determines that the main computing module 11 has experienced a runtime failure. Specifically, during the operation phase of the main computing module 11, the independent management module 13 also identifies whether the main computing module 11 has failed; that is, the method further includes steps S208-S209: S208, the main computing module 11 periodically sends heartbeat signals to the independent management module 13 during operation.

[0046] S209, the independent management module 13 monitors the heartbeat signal. If no heartbeat signal is received within the preset monitoring time, it identifies that the main computing module 11 has malfunctioned.

[0047] The microcontroller of the independent management module 13 maintains a countdown timer. During normal operation, the operating system kernel or daemon of the main computing module 11 sends a pulse signal (e.g., a rising edge) through the runtime heartbeat signal line at regular intervals (e.g., every 5 seconds). Upon detecting a valid edge on this signal line, the independent management module 13 automatically resets its internal timer. When the main computing module 11 stops sending heartbeat pulses due to software crash or hardware failure, the internal timer of the independent management module 13, unable to receive the reset signal, times out. Based on this, the microcontroller's hardware logic determines that the main computing module 11 has malfunctioned.

[0048] The warning signal is an electrical signal sent by the independent management module 13 to the data exchange module 12 to switch states before performing a destructive recovery operation such as a hard reboot. The warning signal is transmitted through the coordinated warning signal line. The signal is usually a pulse level of a specific width (such as a 50-millisecond high-level pulse) or a level flip, which is used to trigger the emergency mode switching logic inside the data exchange module 12.

[0049] After determining that the main computing module 11 is faulty, the microcontroller of the independent management module 13 does not immediately pull down the power enable signal line of the main computing module 11. Instead, it first controls the pin connected to the collaborative warning signal line on the independent management module 13 to output a predefined pulse waveform to control the level state of the general-purpose output pin of the data exchange module 12, thus completing the transmission of the warning signal. Since the collaborative warning signal line is a physical wire directly connected to the general-purpose input pin of the data exchange module 12, the signal transmission delay is extremely low, and the data exchange module 12 captures the warning signal almost in real time.

[0050] After sending the warning signal, the independent management module 13 starts an internal delay timer. At this time, the interface connected to the power enable signal line on the main computing module 11 remains enabled (high level), and no operation is performed; it is in a waiting state. When the delay timer count reaches the preset duration (e.g., 3 seconds), the microcontroller executes a restart control action: by controlling the pin connected to the power enable signal line on the independent management module 13 to output a low level, the P-MOSFET in the power supply circuit of the main computing module 11 is turned off, physically cutting off the main power rail of the main computing module 11. After maintaining the power-off state for a period of time (e.g., 1 second), the microcontroller pulls the level on the signal line high again to restore the power supply to the main computing module 11, thereby completing the hard restart process. This timing ensures that the data exchange module 12 is ready to maintain basic traffic processing before the main computing module 11 is powered off.

[0051] The preset duration is usually set to 3 seconds, which is sufficient for the data exchange module 12 to complete the clearing of complex rules and the switching of internal states, without excessively prolonging the response time for fault recovery.

[0052] S220, in response to the warning signal, the data exchange module 12 clears the network policy issued in real time by the main computing module 11, and processes the network traffic accessed by the network device 1 based on the preset basic forwarding rules to maintain business continuity.

[0053] The network policy refers to the complex forwarding rules issued by the main computing module 11 and dependent on the computing capabilities of the control plane. These rules are stored in the independent storage or on-chip cache of the data exchange module 12 and include access control lists, quality of service policies, dynamic routing table entries, etc. These policies are characterized by requiring the main computing module 11 to participate in the calculation, updating, and maintenance in real time, and may become unreliable or fail during the failure of the main computing module 11.

[0054] Among them, the basic forwarding rules refer to the forwarding logic built into the data exchange module 12 that can be executed without the intervention of the main computing module 11. It mainly refers to MAC address learning and forwarding, basic virtual LAN pass-through, etc. These rules are solidified in the hardware logic or microcode of the switching chip, or rely solely on the locally maintained static MAC address table.

[0055] The data exchange module 12 clears the network policies issued in real time by the main computing module 11. Specifically, this involves the internal register operations and entry management of the data exchange module 12 after receiving the warning signal. When the data exchange module 12 captures the warning signal through the cooperative warning signal line, its internal logic immediately triggers the emergency mode, marking dynamic entries (such as routing table entries) and complex policy rules (such as access control list matching rules) that depend on the main computing module 11 in the forwarding table as invalid or removing them directly from the hardware forwarding table. This is to prevent the data exchange module 12 from causing abnormal traffic forwarding or packet loss due to the execution of outdated or incorrect policies during the restart of the main computing module 11.

[0056] After clearing complex policies, the data exchange module 12 only enables basic forwarding functions. At this time, the data exchange module 12 no longer performs Layer 3 route lookup or Access Control List (ACL) filtering on data packets received from the network interface. Instead, it performs simple forwarding or broadcasting based on the MAC address table. This allows network device 1 to degenerate into a Layer 2 switch with basic connectivity during the failure of the main computing module 11, maintaining physical link connectivity between upstream and downstream devices.

[0057] Through the emergency mode of data exchange module 12, although some computationally-dependent advanced services (such as deep inspection and complex routing) are temporarily interrupted, basic traffic forwarding does not stop. For upstream and downstream devices, the network link remains uninterrupted, and TCP connections are not immediately broken due to link interruption, thus maintaining business continuity.

[0058] The service reconstruction process after a hard reboot of the main computing module 11 is described. In some embodiments, after the independent management module 13 controls the main computing module 11 to perform a hard reboot, the method further includes steps S231-S232: S231, if the main computing module 11 recovers to normal after a hard reboot, the network policy is reissued to the data exchange module 12.

[0059] After the independent management module 13 controls the main computing module 11 to complete a hard reboot, the main computing module 11 re-executes the power-on self-test and operating system loading process. The independent management module 13 monitors the startup phase signals via the serial communication bus. If the main computing module 11 successfully sends the bytecode signal indicating that the operating system loading is complete and resumes pulse output on the runtime heartbeat signal line, the independent management module 13 determines that the main computing module 11 has returned to normal. At this time, the main computing module 11 actively initiates a connection to the data exchange module 12 via the high-speed data bus and reissues the previously cleared access control lists, quality of service policies, and dynamic routing table entries, etc., as network policies.

[0060] S232, the data exchange module 12 configures the reissued network policy to resume the execution of advanced service functions.

[0061] Data exchange module 12 receives and parses configuration messages, then writes the received policy rules into its internal forwarding memory, overwriting the default configuration in emergency mode. After the configuration takes effect, data exchange module 12 exits emergency mode and re-enables advanced service processing logic (such as Layer 3 route lookup and ACL filtering), restoring its ability to process complex network services. At this point, network device 1 is fully restored to its normal operating state before the failure, achieving lossless service reconstruction.

[0062] For example, during the operation of network device 1, the main computing module 11 stops working due to an operating system kernel crash. First, the independent management module 13 monitors the heartbeat signal through the runtime heartbeat signal line. After no heartbeat pulse is detected for more than 5 seconds, the independent management module 13 determines that the main computing module 11 has failed. At this time, the independent management module 13 does not immediately cut off the power to the main computing module 11, but instead sends a high-level pulse lasting 50 milliseconds as a warning signal to the data exchange module 12 through the cooperative warning signal line. The data exchange module 12 captures this signal through an interrupt and immediately executes its internal logic: clearing the access control list and quality of service policy entries, retaining only the basic Layer 2 forwarding function, and entering emergency mode.

[0063] At this time, the user data packet arrives at network device 1. Data exchange module 12 forwards it according to the basic Layer 2 rules, and network traffic is not interrupted. After the independent management module 13 sends a warning signal, it initiates a 3-second delay. After the delay ends, the independent management module 13 pulls the power enable signal line low, cutting off the power to the main computing module 11. After waiting for 1 second, it pulls the signal line high again to restore power. The main computing module 11 restarts, and after the power-on self-test and operating system loading phases, it reports its startup status to the independent management module 13 via the serial communication bus. After startup, the main computing module 11 re-establishes heartbeat communication and sends the complete network policy to the data exchange module 12. The data exchange module 12 resumes normal forwarding mode, completing the recovery.

[0064] Service reconstruction after a hard reboot of the main computing module 11 may fail. Therefore, preferably, after the independent management module 13 controls the main computing module 11 to perform a hard reboot, the method further includes steps S233-S234: S233, if the independent management module 13 still does not detect a heartbeat signal within the preset recovery time, then the main computing module 11 will be hard-rebooted no more than the first preset number of times. If a heartbeat signal is detected within the preset recovery time after any hard-reboot, then the subsequent hard-reboot operation will be terminated.

[0065] After the independent management module 13 completes a hard reboot of the main computing module 11, it starts a timer with a preset recovery duration. If the independent management module 13 receives a valid heartbeat signal from the main computing module 11 via the runtime heartbeat signal line or the serial communication bus within this time period, the reboot recovery is considered successful, and the process jumps to steps S231-S232. If the independent management module 13 does not receive a valid heartbeat signal from the main computing module 11 via the runtime heartbeat signal line or the serial communication bus within this time period, the reboot recovery is considered unsuccessful. The register inside the independent management module 13 records the reboot count and increments it. If the current reboot count has not reached the first preset count (e.g., 3 times), the independent management module 13 executes the hard reboot loop again via the control power enable signal line. If the current reboot count has reached the first preset count, step S234 is executed.

[0066] By performing multiple hard reboots to power off the system and restart it, system hangs caused by transient soft failures can be prevented.

[0067] The preset recovery duration can be a period of time when the main computing module 11 starts timing after completing a hard reboot, that is, a period of time when the main computing module 11 starts timing after loading the main system, such as 5 seconds. Alternatively, it can be a period of time when the independent management module 13 pulls the level on the power enable signal line low and cuts off the power to the main computing module 11, such as 30 seconds. The specific setting can be set according to the actual usage.

[0068] S234, if after performing the first preset number of hard reboots, the independent management module 13 does not detect a heartbeat signal within each preset recovery time, the independent management module 13 will issue a fault alarm.

[0069] When the number of restarts recorded by the independent management module 13 reaches the upper limit of the first preset number (e.g., 3 times), and no heartbeat signal is received within the preset recovery time after the last restart, it indicates that the main computing module 11 has a serious hardware failure or a persistent software failure and cannot be recovered by restarting. At this time, the independent management module 13 locks the fault status, stops the automatic restart cycle, and sends a fault alarm signal through the indicator light, buzzer, or remote management interface on the network device 1 to notify the operation and maintenance personnel to perform manual intervention.

[0070] Furthermore, the fault diagnosis process for a serious hardware failure or persistent software failure in the main computing module 11 is described. Preferably, S234 includes sub-steps S234a to S234c: S234a, if after performing the first preset number of hard reboots, the independent management module 13 does not detect a heartbeat signal within each preset recovery time, then the independent management module 13 reads the power status signal of the main computing module 11.

[0071] A power detection signal line is also connected between the independent management module 13 and the main computing module 11. This power line is connected to the power management circuit of the main computing module 11. The independent management module 13 is used to read the level status on this power line and identify whether the main power rail of the main computing module 11 is outputting normally, so as to complete the reading of the power status signal of the main computing module 11.

[0072] S234b If the power status signal is abnormal, a first alarm signal containing power failure information is generated.

[0073] If the independent management module 13 reads a low level (or abnormal level) on the power detection signal line, the voltage on the main power rail of the main computing module 11 will drop or the power management chip will report an error. At this time, the independent management module 13 generates a first alarm signal. The data frame of this signal contains a specific fault code, indicating that the fault type is a power failure, which helps maintenance personnel to quickly locate the problem of the power module 14 or the motherboard power supply circuit.

[0074] S234c If the power status signal is normal, a second alarm signal containing operating status fault information is generated.

[0075] If the independent management module 13 reads a high level on the power detection signal line, it indicates that the power supply is normal, but the main computing module 11 still cannot output a heartbeat signal. This indicates that the fault is not caused by the power supply, but by damage to the processor core, memory, or boot firmware of the main computing module 11. At this time, the independent management module 13 generates a second alarm signal. The data frame of this signal contains a fault code indicating a hardware failure of the computing core or a boot failure, prompting maintenance personnel to replace the core board of the main computing module 11 or check the storage media.

[0076] For example, the independent management module 13 determines that the main computing module 11 has malfunctioned if no pulse is detected on the running heartbeat signal line within 5 seconds. First, the independent management module 13 notifies the data exchange module 12 to enter emergency mode via the collaborative early warning signal line, and then restarts the main computing module 11 by power-off via the power enable signal line. After restarting, wait for 5 seconds. If no heartbeat is received, restart again. After 3 consecutive restarts, if the main computing module 11 still does not respond, the independent management module 13 reads the power detection signal line. If the signal line is found to be low, a power module 14 fault alarm is reported via the remote interface; if the signal line is found to be high, a motherboard CPU fault alarm is reported, and automatic restart is stopped. Network device 1 remains in the basic forwarding mode of the data exchange module 12, waiting for manual repair.

[0077] By periodically sending and monitoring heartbeat signals, runtime fault detection can be achieved at the second level, and the multi-level restart and retry mechanism can effectively deal with transient faults and avoid complete service interruption due to a single startup failure. By reading the power status signal line, the fault type can be located. Even if the main computing module 11 is completely damaged and cannot be recovered, the data exchange module 12 can still maintain basic network connectivity in emergency mode, achieving degraded service continuity assurance and improving the reliability and maintainability of network device 1.

[0078] Preferably, the method further includes: S201, the data exchange module 12 responds to the power-on of the network device 1 by processing the network traffic accessed by the network device 1 based on the basic forwarding rules.

[0079] An external AC power supply is connected to the power module 14 of network device 1 to power on network device 1.

[0080] When an external AC power source is connected, the power module 14 immediately outputs power to the power rail of the switching chip. Since this power rail is independent of the main power rail and is not controlled by the independent management module 13, the switching chip of the data switching module 12 receives power and automatically performs hardware initialization. The hardware logic or microcode embedded in the switching chip of the data switching module 12 is then activated, and the basic forwarding function is activated.

[0081] The data switching module 12 already has packet processing capabilities the moment network device 1 is plugged in. At this time, if network traffic enters from the network interface, the switching chip directly performs table lookup and forwarding based on preset basic forwarding rules (such as default VLAN configuration and MAC address self-learning).

[0082] Furthermore, the data exchange module 12 can be configured to respond to the power-on and power-on trigger commands of the network device 1 according to actual usage needs, and process the network traffic accessed by the network device 1 based on the basic forwarding rules. In this case, the data exchange module 12 responds to the power-on event and needs to wait for the power-on trigger command. The power-on trigger command is a logical trigger signal generated when the user presses the power button on the front panel or when the system detects that the automatic start-up configuration is valid.

[0083] The timing control logic of the independent management module 13 is described below. Preferably, the method further includes steps S202-S203: S202, the independent management module 13 responds to the power-on and power-on trigger commands of the network device 1 and controls the main computing module 11 to power on and start.

[0084] After the external AC power supply is connected, the standby power rail becomes active immediately, and the microcontroller of the independent management module 13 receives power and completes initialization. At this time, the microcontroller sets the power enable signal line to a low level, ensuring that the main power rail of the main computing module 11 is in the off state, and the main computing module 11 remains powered off. Subsequently, the microcontroller enters a cyclic scanning state, waiting for instructions through the power-on trigger instruction recognition circuit (such as scanning the signal line level connected to the front panel button or reading the configuration jumper status). When the user presses the front panel power button (resulting in a low-level transition on the signal line) or detects that the power-on self-start configuration is valid, the microcontroller confirms it as a valid power-on trigger event. In response to this instruction, the microcontroller outputs a high level through the power enable signal line, controlling the power switch of the main power rail to close, supplying power to the main computing module 11, thereby triggering the startup process of the main computing module 11.

[0085] S203, after the main computing module 11 is powered on and before the system loading is completed, if the independent management module 13 detects that the main computing module 11 has failed, then the main computing module 11 will be hard-rebooted.

[0086] The independent management module 13 identifies faults through a serial communication bus and a preset timer mechanism. If the independent management module 13 does not receive the expected stage completion signal within a preset time window after the main computing module 11 is powered on, it determines that the main computing module 11 has failed during the startup phase. The recovery strategy at this time is to perform a hard reboot. For example, the independent management module 13 controls the power enable signal line to pull low to cut off the main power supply, maintains it for a certain period of time (e.g., 1 second), and then pulls it high to restore power supply and retry the startup.

[0087] Specifically, if a fault occurs during the power-on self-test phase, preferably, the method further includes steps S204~S205: S204, after the independent management module 13 controls the main computing module 11 to power on, if it does not receive the power-on self-test completion signal sent by the main computing module 11 within the first time window, it determines that the main computing module 11 has failed during the power-on self-test phase and performs a hard reboot no more than the second preset number of times. If the power-on self-test completion signal is received within the first time window after any hard reboot, the subsequent hard reboot operation is terminated.

[0088] The first time window is the time length reserved for the main computing module 11 to perform the power-on self-test (POST), for example, 15 seconds; the power-on self-test completion signal refers to a specific byte code (e.g., 0x55) sent by the BIOS or Bootloader of the main computing module 11 through the serial communication bus after completing basic hardware initialization such as memory detection and bus scanning.

[0089] In response to the power-on trigger command, the independent management module 13 immediately starts its internal Timer_POST timer as the timing reference for the first time window the moment the main power is turned on. Simultaneously, the independent management module 13 listens for data on the bus via the serial data line and serial clock line. If the main computing module 11 is running normally, it will send a predefined bytecode via the I2C bus after completing hardware initialization. The independent management module 13 receives the predefined bytecode sent by the main computing module 11 via the I2C bus, indicating that it has received the power-on self-test completion signal from the main computing module 11. If the timer times out (i.e., the end of the first time window is reached) and the bytecode is not received, the independent management module 13 determines that the main computing module 11 is stuck in the power-on self-test phase (e.g., memory failure or BIOS corruption), and then triggers a hard reboot retry mechanism.

[0090] After the independent management module 13 completes a hard reboot of the main computing module 11, it starts the Timer_POST timer. If the independent management module 13 receives predefined bytecode from the main computing module 11 via the I2C bus within this timer period, it determines that the reboot recovery was successful, and the main computing module 11 enters the operating system loading stage. The independent management module 13 further monitors whether a fault occurred in the main computing module 11 during the operating system loading stage, that is, it jumps to execute S206~S207. If the independent management module 13 does not receive predefined bytecode from the main computing module 11 via the I2C bus within this timer period, it determines that the reboot recovery failed, and the registers inside the independent management module 13 record the reboot count and increment it. If the current reboot count has not reached the second preset number (e.g., 3 times), the independent management module 13 executes the hard reboot loop again via the control power enable signal line. If the current reboot count has reached the second preset number, it executes S205.

[0091] S205. If, after performing the second preset number of hard reboots, no power-on self-test completion signal is received within the first time window after each hard reboot, the fault is locked as a power-on self-test stage fault.

[0092] The independent management module 13 maintains a restart counter internally. Each time a restart occurs due to a timeout in the first time window, the counter increments by one. If the counter reaches a preset upper limit (e.g., 3 times), and no power-on self-test completion signal is received after each restart, the independent management module 13 determines that the fault is a permanent hardware fault and cannot be repaired by restarting. At this point, the independent management module 13 stops the restart cycle, locks the fault status, and reports the power-on self-test phase fault via indicator lights or a remote interface. Simultaneously, the data exchange module 12 remains in basic forwarding mode to ensure uninterrupted basic network operation.

[0093] The independent management module 13 receives a power-on self-test completion signal from the main computing module 11 within the first time window after power-on, indicating that the main computing module 11 has successfully completed the power-on self-test. Alternatively, if the independent management module 13 does not receive a power-on self-test completion signal from the main computing module 11 within the first time window after power-on, but receives a valid power-on self-test completion signal during restart if the current restart count has not reached the second preset number (e.g., 3 times), in both cases, the main computing module 11 enters the operating system loading stage, and the independent management module 13 further monitors whether a fault occurs in the main computing module 11 during the operating system loading stage. Therefore, preferably, the method further includes S206~S207: S206, after receiving the power-on self-test completion signal, if the independent management module 13 does not receive the operating system loading completion signal sent by the main computing module 11 within the second time window, it determines that the main computing module 11 has failed during the operating system loading stage and performs a hard reboot no more than the third preset number of times. If the operating system loading completion signal is received within the second time window after any hard reboot, the subsequent hard reboot operation is terminated.

[0094] The second time window is the time reserved for the main computing module 11 to load the operating system kernel and driver, for example, 60 seconds; the operating system loading completion signal is another specific byte code (e.g., 0xAA) sent through the serial communication bus after the operating system has finished starting and the network protocol stack has been initialized.

[0095] Once the independent management module 13 successfully receives the power-on self-test completion signal, it immediately stops the timer corresponding to the first time window and starts the TimerOSLoad timer as the timing reference for the second time window. During this period, the operating system kernel of the main computing module 11 performs operations such as decompression, driver loading, and file system mounting.

[0096] Once all startup tasks are complete, an operating system daemon sends an operating system loading completion signal via the I2C bus. If the second time window timer expires without receiving this signal, the operating system loading is deemed to have failed (e.g., due to file system corruption or critical driver crash), and the independent management module 13 executes a hard reboot retry mechanism.

[0097] After the independent management module 13 completes a hard reboot of the main computing module 11, it starts the TimerOSLoad timer. If the independent management module 13 receives an operating system loading completion signal from the main computing module 11 via the I2C bus within the timer period, it determines that the reboot recovery was successful, and the main computing module 11 enters the running phase. The independent management module 13 further monitors whether the main computing module 11 has failed during the running phase, that is, it executes S208~S209 to monitor whether the main computing module 11 has failed, and if the main computing module 11 has failed, it executes S210~S220. If the independent management module 13 does not receive an operating system loading completion signal from the main computing module 11 via the I2C bus within the timer period, it determines that the reboot recovery has failed, and the register inside the independent management module 13 records the reboot count and increments it. If the current reboot count has not reached the third preset number (e.g., 3 times), the independent management module 13 executes the hard reboot cycle again via the control power enable signal line. If the current reboot count has reached the third preset number, it executes S207.

[0098] S207. If, after performing the third preset number of hard reboots, no operating system loading completion signal is received within the second time window after each hard reboot, the fault is locked as an operating system loading phase fault.

[0099] If the restart counter reaches the preset limit (e.g., 3 times) and the fault still exists during the operating system loading phase, the independent management module 13 determines it to be a persistent software or firmware fault, locks the fault, stops restarting, and reports the fault information. At this time, network device 1 is in a degraded available state, data exchange module 12 maintains basic forwarding, and maintenance personnel can quickly repair system software or storage problems based on the fault location information.

[0100] After receiving the power-on self-test completion signal, the independent management module 13 also receives the operating system loading completion signal sent by the main computing module 11 within the second time window, meaning that the main computing module 11 has successfully completed the operating system loading; or, after receiving the power-on self-test completion signal, the independent management module 13 does not receive the operating system loading completion signal sent by the main computing module 11 within the second time window, but during restart, if the current restart count has not reached the third preset number (e.g., 3 times), the independent management module 13 receives a valid operating system loading completion signal. In both cases, the independent management module 13 further monitors whether the main computing module 11 has failed during the operation phase, that is, it executes S208~S209 to monitor whether the main computing module 11 has failed, and executes S210~S220 when the main computing module 11 fails.

[0101] It should be noted that the first, second, and third preset counts can be the same or different, and their specific settings can be configured according to actual usage. This application does not impose any specific restrictions on this.

[0102] By refining monitoring to two specific stages—power-on self-test and operating system loading—the granularity of monitoring can be improved, and the accuracy of fault location can be enhanced. By setting independent time windows and specific handshake signals, fine-grained control of the startup process is achieved. Multiple hard reboots are performed for faults at different stages, effectively addressing transient faults (such as poor memory contact or occasional data loading errors). If rebooting is ineffective, fault locking is implemented, preventing equipment overheating or power supply oscillations caused by endless reboots. Combined with the basic forwarding maintenance function of the data exchange module 12, it ensures that network device 1 maintains link-layer connectivity even if the main computing module 11 fails to start, thus gaining valuable time for maintenance personnel to troubleshoot and quickly restore services, further maintaining business continuity.

[0103] The logical flow in this application embodiment can be divided into three core stages: power-on standby and basic service readiness, power-on triggering and startup monitoring, and runtime monitoring and collaborative recovery.

[0104] Phase 1: Power-on standby and basic service readiness phase, corresponding to S201 above. This phase begins with the connection of external AC power and ends before the power-on trigger command arrives.

[0105] When an external AC power source is connected to network device 1, power module 14 outputs power to the standby power rail and the switching chip power rail. Independent management module 13 is powered by the standby power rail and completes initialization. After initialization, independent management module 13 sets the power enable signal line to low level, ensuring that the main computing module 11 remains powered off. Simultaneously, data switching module 12 is directly powered by the switching chip power rail, and its switching chip automatically completes hardware initialization. At this time, data switching module 12 can activate basic forwarding functions. Although the main computing module 11 has not yet started, data switching module 12 can already process network traffic based on preset basic forwarding rules, and network device 1, as a physical link node, has basic connectivity. Independent management module 13 enters a cyclic scanning state, waiting for the power-on trigger command.

[0106] The second stage: the power-on trigger and startup monitoring stage, corresponding to S202~S207 above. This stage begins with the arrival of the power-on trigger command and ends when the main computing module 11 completes the loading of the operating system.

[0107] When the user presses the power button on the front panel or the system detects that the automatic start configuration is valid upon receiving a call, the independent management module 13 recognizes a valid power-on trigger command. In response to this command, the independent management module 13 connects the main power rail and supplies power to the main computing module 11 by pulling the power enable signal line high.

[0108] Subsequently, a multi-level monitoring process is initiated: First, the Power-On Self-Test (POST) monitoring is performed, and the independent management module 13 starts a first time window timer. If a POST completion signal is received from the main computing module 11 via the serial communication bus within the first time window, the POST is considered successful, and the process proceeds to the next stage. If no signal is received within the timeout period, the main computing module 11 is considered to have failed during the POST phase. In this case, the independent management module 13 performs a hard reboot no more than a second preset number (e.g., 3 times). If the hard reboot still fails after the second preset number of times, the fault is identified as a POST phase fault, and the reboot process is stopped.

[0109] Next, the operating system loading monitoring is executed. After a successful POST, the independent management module 13 starts a second time window timer. If an operating system loading completion signal is received within the second time window, the system is considered to have started successfully. If no signal is received within the time window, the main computing module 11 is considered to have failed during the operating system loading phase. In this case, the independent management module 13 performs a hard reboot no more than the third preset number of times. If the reboot still fails, the fault is identified as an operating system loading phase fault, and the reboot process is stopped.

[0110] During this phase, if a fault lockout occurs, the data exchange module 12 will remain operational to maintain basic network services.

[0111] Phase 3: Runtime Monitoring and Collaborative Recovery Phase. This phase begins when the main computing module 11 enters normal operating condition.

[0112] After the main computing module 11 operates normally, it periodically sends a heartbeat signal to the independent management module 13. The independent management module 13 continuously monitors this signal. If it does not receive a heartbeat signal within a preset monitoring period, it determines that the main computing module 11 has experienced a runtime failure. The independent management module 13 then triggers a collaborative recovery mechanism: first, it sends a warning signal to the data exchange module 12, notifying it to enter emergency mode (clearing complex policies and retaining basic forwarding), and after waiting for a preset period, it controls the main computing module 11 to perform a hard reboot. After a successful reboot, the main computing module 11 reissues the network policy and restores full functionality.

[0113] For example, the network administrator plugs the power cord of network device 1 into a power outlet. The power module 14 immediately starts working, the independent management module 13 powers on and initializes, and sets the signal line controlling the main power supply low. The main computing module 11 remains powered off. At the same time, the switching chip of the data switching module 12 powers on and automatically activates the basic Layer 2 forwarding function. At this time, the network interface indicator light on the panel of network device 1 lights up, the physical link between upstream and downstream devices is connected, and basic network traffic can pass normally. However, the power button is not pressed at this time, and the main computing module 11 is not working.

[0114] The network administrator walks to the device and presses the power button on the front panel. The independent management module 13 scans the signal and confirms it as a valid power-on trigger command. In response to this command, the independent management module 13 controls the power enable signal line to output a high level, the main power rail is connected, and the main computing module 11 receives power and begins startup.

[0115] After power-on, the main computing module 11 begins executing the BIOS hardware self-test. Simultaneously, the independent management module 13 starts the first time window timer. Assuming that during this boot process, the memory module experiences a poor connection, causing the self-test to stall, and the independent management module 13 does not receive a POST completion signal after 15 seconds, the independent management module 13 determines a POST stage failure and immediately performs a hard reboot, powering off and restarting the main computing module 11. If the self-test passes successfully on the second reboot, the BIOS sends a power-on self-test completion signal via the I2C bus. The independent management module 13 receives the signal, determines POST success, and proceeds to the next stage. If three consecutive reboots fail, the fault is locked and an alarm is triggered; the device remains in basic forwarding mode.

[0116] After a successful POST, the independent management module 13 starts the second time window timer. The main computing module 11 begins loading the operating system kernel from the hard drive. Assuming the operating system files are complete, the system will boot up in 40 seconds. The operating system daemon will send an operating system loading completion signal via I2C. Upon receiving the signal, the independent management module 13 determines that the operating system has been successfully loaded. If no signal is received during this timeout period, the system will restart repeatedly. If three consecutive restarts fail, a fault will be locked and an alarm will be triggered, and the device will remain in basic forwarding mode.

[0117] The main computing module 11 enters normal operation, begins sending heartbeat signals, and distributes complex access control lists and routing policies to the data exchange module 12. Network device 1 operates at full capacity. After a period of time, assuming the main computing module 11 crashes due to a software malfunction, the heartbeat signal stops. The independent management module 13 detects a heartbeat timeout, determines an operational fault, and immediately sends a warning pulse to the data exchange module 12 via a dedicated signal line. Upon receiving this signal, the data exchange module 12 instantly clears the complex policies and reverts to basic Layer 2 forwarding mode to ensure uninterrupted network operation. Subsequently, the independent management module 13 restarts the main computing module 11. After a successful restart, the main computing module 11 resumes heartbeats and reissues policies, and the device fully recovers to normal operation.

[0118] Throughout the process, for users, apart from brief interruptions to advanced services, basic network connectivity was never interrupted due to the failure or restart of the main computing module 11.

[0119] This application embodiment separates computing and switching capabilities by simultaneously setting up a main computing module and a data switching module in the network device. This allows the switching functions that the main computing module should possess to be transferred to the data switching module. During a failure or restart of the main computing module, the data switching module no longer stagnates due to its reliance on the main computing module's policies. Instead, it proactively clears complex network policies that depend on the real-time issuance of the main computing module and maintains basic network traffic forwarding based on preset basic forwarding rules. This ensures that the network device can still maintain basic connectivity as a physical link node during a main computing module failure, limiting the impact of the failure to computing functions and maintaining the availability of network forwarding functions, thereby maintaining service continuity under main computing module failure. Furthermore, after identifying the fault, the independent management module does not immediately perform a hard restart. Instead, it first sends a warning signal to the data switching module and waits for a preset time before controlling the main computing module to restart. This provides necessary buffer time for the main computing module's restart, preventing sudden power outages and restarts from causing state instability in the data switching module, further ensuring network service continuity.

[0120] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for maintaining business continuity.

[0121] This application provides a computer program that can be executed by a processor to implement the above-described method for maintaining business continuity.

[0122] This application provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method for maintaining business continuity.

[0123] In the several embodiments provided in this application, any function, if implemented as a software functional module / unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of this application can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or other electronic device) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing computer program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0124] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of this application are not directed to any particular programming language. It should be understood that the content of this application described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of this application.

[0125] It should be noted that the above embodiments are illustrative of this application and not restrictive, and those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In claims enumerating several means, several units or modules of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

[0126] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method of maintaining service continuity, characterized by, Applied to a network device, the network device including a main computing module, a data exchange module, and an independent management module, the method includes: In response to detecting a fault in the main computing module, the independent management module sends an early warning signal to the data exchange module and, after waiting for a preset time, controls the main computing module to perform a hard reboot. In response to the warning signal, the data exchange module clears the network policy issued in real time by the main computing module and processes the network traffic accessed by the network device based on the preset basic forwarding rules to maintain business continuity.

2. The method of claim 1, wherein, The method further includes: The main computing module periodically sends heartbeat signals to the independent management module during runtime; The independent management module monitors the heartbeat signal. If the heartbeat signal is not received within a preset monitoring period, the main computing module is identified as malfunctioning.

3. The method of claim 2, wherein, After the independent management module controls the main computing module to perform a hard reboot, the method further includes: If the main computing module recovers to normal after a hard reboot, the network policy is reissued to the data exchange module; The data exchange module configures the reissued network policy to resume the execution of advanced service functions.

4. The method of claim 2, wherein, After the independent management module controls the main computing module to perform a hard reboot, the method further includes: If the independent management module fails to detect the heartbeat signal within the preset recovery time, it performs a hard reboot on the main computing module for no more than a first preset number of times. If the heartbeat signal is detected within the preset recovery time after any hard reboot, the subsequent hard reboot operation is terminated. If, after performing the first preset number of hard reboots, the independent management module fails to detect the heartbeat signal within the preset recovery time for each reboot, the independent management module will issue a fault alarm.

5. The method according to claim 4, characterized in that, If, after performing the first preset number of hard reboots, the independent management module fails to detect the heartbeat signal within the preset recovery time for each reboot, the independent management module issues a fault alarm, including: If, after performing the first preset number of hard reboots, the independent management module does not detect the heartbeat signal within the preset recovery time for each reboot, the independent management module reads the power status signal of the main computing module. If the power status signal is abnormal, a first alarm signal containing power fault information is generated. If the power status signal is normal, a second alarm signal containing operational status fault information is generated.

6. The method according to claim 1, characterized in that, The method further includes: The data exchange module responds to the power-on of the network device and processes the network traffic accessed by the network device based on the basic forwarding rules.

7. The method according to claim 6, characterized in that, The method further includes: The independent management module responds to the power-on and power-on trigger commands of the network device, and controls the main computing module to power on and start. If the independent management module detects a fault in the main computing module after the main computing module is powered on and before the system loading is completed, it will perform a hard reboot on the main computing module.

8. The method according to claim 7, characterized in that, The method further includes: If the independent management module does not receive a power-on self-test completion signal from the main computing module within the first time window after controlling the main computing module to power on, it determines that the main computing module has failed during the power-on self-test phase and performs a hard reboot no more than the second preset number of times. If the power-on self-test completion signal is received within the first time window after any hard reboot, the subsequent hard reboot operation is terminated. If, after performing the second preset number of hard reboots, no power-on self-test completion signal is received within the first time window after each hard reboot, the fault is identified as a power-on self-test stage fault.

9. The method according to claim 8, characterized in that, The method further includes: If the independent management module does not receive the operating system loading completion signal from the main computing module within the second time window after receiving the power-on self-test completion signal, it determines that the main computing module has failed during the operating system loading stage and performs a hard reboot no more than a third preset number of times. If the operating system loading completion signal is received within the second time window after any hard reboot, the subsequent hard reboot operation is terminated. If, after performing the third preset number of hard reboots, no operating system loading completion signal is received within the second time window after each hard reboot, the fault is identified as an operating system loading phase fault.

10. A network device, characterized in that, The network device includes a main computing module, a data exchange module, and an independent management module, and is used to perform the service continuity maintenance method as described in any one of claims 1 to 9.