Router and method based on collaborative warm backup, storage medium and program product

By adopting a collaborative warm standby router architecture, and combining multi-dimensional redundancy design of the control plane, data plane and management plane, the problems of imbalance between router fault recovery speed and power consumption and the inability to integrate security and reliability are solved, achieving a balance between high reliability and high security and reducing equipment costs.

CN121750554APending Publication Date: 2026-03-27PURPLE MOUNTAIN LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing routers struggle to balance fault recovery speed and power consumption, and cannot achieve a unified balance between security and reliability, posing significant challenges, especially in space environments.

Method used

A router architecture based on collaborative warm standby is adopted. Through multi-dimensional redundancy design of the control plane, data plane and management plane, dynamic heterogeneous redundancy is achieved. Combined with anomaly detection and cleaning mechanisms, the execution and management units are dynamically scheduled to achieve high reliability and high security.

Benefits of technology

At the system level, a balance is achieved between fault recovery speed and operating power consumption, and the router's resilience and self-healing capabilities are improved, reducing the requirements for hardware reliability and significantly reducing equipment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750554A_ABST
    Figure CN121750554A_ABST
Patent Text Reader

Abstract

The invention discloses a router and method based on collaborative warm standby, a storage medium and a program product, the router comprises a control plane, a data plane and a management plane, the control plane comprises at least two execution bodies, one execution body is in a main state, at least one execution body is in a standby state, and the data plane is in the standby state. The executor in the main state sends routing information to the data plane according to a routing protocol; the data plane comprises at least two data forwarding groups, one data forwarding group is in a main state, at least one data forwarding group is in a standby state, the data forwarding group in the main state executes data forwarding according to routing information, the management plane comprises at least two management units, one management unit is in the main state, and the other management unit is in the standby state. At least one management unit is in the standby state, and the management unit in the main state carries out hardware state abnormity monitoring and abnormity cleaning on the data surface and the control surface. According to the invention, the balance between the fault recovery speed and the operation power consumption is realized, and the safety and the reliability are improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to communication technology, and more particularly to a router, method, storage medium, and program product based on collaborative temperature backup. Background Technology

[0002] As the core hub of network data transmission, the reliability of routers directly determines the reliability and stability of the entire network. Due to the massive amount of code involved in the design and implementation of routers, they are prone to numerous potential vulnerabilities, making it difficult to completely eliminate the possibility of backdoors. Furthermore, the harsh environment of space, with its strong radiation and large temperature differences, poses a significant challenge to the reliability of router devices. To improve the reliability of router equipment, existing technologies mainly employ homogeneous backup redundancy schemes, with common backup schemes including cold backup, hot backup, and warm backup.

[0003] Although existing technologies have improved the reliability of router devices to some extent, the following core technical problems still exist:

[0004] 1. There is a contradiction between fault recovery speed and power consumption. Traditional protocols such as VRRP can achieve cold standby and failover for routing nodes, but the switchover time for cold standby is too long, usually taking tens of seconds to several minutes, which will lead to service interruption. Hot standby for routing nodes can achieve rapid failover, but hot standby requires the primary / standby devices to maintain consistent status in real time and operate under continuous high load, which places high demands on hardware reliability and will lead to a significant increase in power consumption, posing a great challenge to power supply and heat dissipation in the space environment. Most routing node warm standby solutions are simple redundancy with one primary and one standby at the whole machine level, lacking a coordination mechanism, unable to achieve dynamic scheduling of devices, and not optimized for self-healing recovery after failure, which may still lead to service interruption, resulting in an imbalance between fault recovery speed and power consumption.

[0005] 2. There is a contradiction between achieving security and reliability in a unified way. Existing technologies achieve high reliability through homogeneous redundancy, which can tolerate random failures to a certain extent, but cannot defend against unknown security threats, especially network attacks based on unknown vulnerabilities and backdoors. Existing technologies only focus on fault switching and lack systematic and integrated resilience and self-healing capabilities, resulting in a failure to achieve security and reliability in a unified way. Summary of the Invention

[0006] To address the problems existing in the prior art, the purpose of this invention is to provide a router, method, storage medium, and program product based on collaborative thermal backup, which solves the problems of unbalanced fault recovery speed and power consumption, and the inability to integrate security and reliability in existing routers.

[0007] To achieve the above-mentioned objectives, the present invention provides the following technical solution:

[0008] In a first aspect, the present invention provides a router based on cooperative warm standby, comprising:

[0009] The control plane includes at least two execution entities, one of which is in a primary state and at least one of which is in a standby state. The execution entity in the primary state sends routing information to the data plane according to the routing protocol. The execution entity in the standby state switches to the primary state if the execution entity in the primary state malfunctions.

[0010] The data plane includes at least two data forwarding groups, one of which is in a primary state and at least one of which is in a backup state. The data forwarding group in the primary state performs data forwarding according to the routing information received from the control plane. The data forwarding group in the backup state synchronizes data with the data forwarding group in the primary state and switches to the primary state in case of an anomaly in the data forwarding group in the primary state.

[0011] The management plane includes at least two management units, one of which is in a primary state and at least one of which is in a standby state. The management unit in the primary state performs hardware status anomaly monitoring and anomaly cleaning on the data plane and the control plane. The management unit in the standby state switches to the primary state if the primary management unit in the primary state malfunctions.

[0012] Furthermore, the data forwarding group includes:

[0013] A forwarding control unit is used to control the forwarding unit and send the routing information received from the control plane to the forwarding unit;

[0014] A forwarding unit is used to forward data based on the received routing information;

[0015] In this configuration, the forwarding control unit and the forwarding unit in the primary state operate simultaneously, the forwarding control unit in the standby state synchronizes the data with the forwarding control unit in the primary state, and the forwarding unit in the standby state stops operating.

[0016] Furthermore, the forwarding control unit in the primary state also monitors the software status of the execution unit and performs primary / backup switching of the execution unit according to preset scheduling rules. The preset scheduling rules include:

[0017] If an abnormality is detected in the execution unit, the forwarding control unit in the primary state sends a cleaning control command to the management unit in the primary state, so that the management unit cleans the abnormal execution unit according to the cleaning control command and sets the execution unit's state to offline state; if the cleaned execution unit is in the primary state, the forwarding control unit in the primary state selects an execution unit from the execution units in the standby state to switch to the primary state;

[0018] When the preset scheduling period is reached, the forwarding control unit in the primary state switches the executor in the primary state to the standby state, and selects one executor from the standby executors to switch to the primary state.

[0019] Furthermore, in the initial state, the data forwarding group sets the primary / standby status by reading a predefined configuration. If no predefined configuration exists, it negotiates with other data forwarding groups to determine the primary / standby status.

[0020] Furthermore, in the initial state, the executor sends an online request to the data forwarding group in the primary state, and sets the primary / backup state according to the returned online response. The data forwarding group in the primary state determines whether the executor in the online request is the first one. If the executor in the online request is the first one, the returned online response is in the primary state; if the executor in the online request is not the first one, the returned online response is in the backup state.

[0021] Furthermore, in the initial state, the management unit decides its primary / standby state by negotiating with other management units. If it does not receive negotiation information from other management units within a preset time, it automatically enters the primary state.

[0022] Furthermore, the management unit in the main state is specifically used to perform the following steps:

[0023] Hardware status monitoring is performed on the data plane and the control plane;

[0024] If an abnormality is detected in the hardware status, it is determined whether the number of cleaning cycles for the abnormal unit has reached the maximum number within a preset time threshold.

[0025] If the maximum number of cleaning cycles is reached, the unit that experienced the abnormality will be deactivated.

[0026] If the maximum number of cleaning cycles has not been reached, the abnormal unit is cleaned, and the cleaning cycle is updated.

[0027] Furthermore, the router is provided with a data channel and a management channel. The data channel is a data transmission channel between the execution entities and between the execution entities and the data forwarding group. The management channel is a data transmission channel between the management unit and the execution entities, and between the management unit and the data forwarding group.

[0028] Furthermore, the standby state of the execution entity is any one of hot standby, warm standby, and cold standby, the standby state of the data forwarding group is warm standby, and the standby state of the management unit is hot standby.

[0029] Secondly, the present invention provides a routing method based on cooperative warm backup, the method comprising:

[0030] In the control plane, the executor in the primary state sends routing information to the data plane according to the routing protocol, and the executor in the backup state switches to the primary state in the event of an anomaly in the executor in the primary state; wherein, the control plane includes at least two executors, one of which is in the primary state and at least one of which is in the backup state;

[0031] In the data plane, the data forwarding group in the primary state performs data forwarding according to the routing information received from the control plane. The data forwarding group in the backup state synchronizes data with the data forwarding group in the primary state and switches to the primary state in case of an anomaly in the data forwarding group in the primary state. The data plane includes at least two data forwarding groups, one of which is in the primary state and at least one of which is in the backup state.

[0032] The management unit in the primary state of the management plane performs hardware status anomaly monitoring and anomaly cleaning on the data plane and the control plane. The management unit in the standby state switches to the primary state when the primary management unit in the primary state is abnormal. The management plane includes at least two management units, one of which is in the primary state and at least one of which is in the standby state.

[0033] Thirdly, the present invention provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method described in the second aspect.

[0034] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the method described in the second aspect.

[0035] Fifthly, the present invention provides a computer program product comprising a computer program that, when executed by a processor, implements the method described in the second aspect.

[0036] Compared with the prior art, the beneficial effects of this invention are as follows: This invention implements routing protocol execution through the execution body of the control plane, implements data forwarding control through the data forwarding group of the data plane, and implements anomaly detection and anomaly cleaning of the data plane and control plane through the management unit of the management plane. Combined with the multi-dimensional redundancy of the primary and backup states of the control plane, data plane, and management plane, it realizes the collaborative warm backup and resilient self-healing of the router's overall internal architecture. It achieves high reliability and high security in an integrated manner at the whole machine level, and balances fault recovery speed and operating power consumption. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the structure of a router based on collaborative temperature backup provided in an embodiment of the present invention;

[0038] Figure 2 This is a schematic diagram of the control surface structure provided in an embodiment of the present invention, which is composed of a physical execution entity and a virtual execution entity.

[0039] Figure 3 This is a flowchart illustrating the routing method based on collaborative warm backup provided in an embodiment of the present invention;

[0040] Figure 4 This is a flowchart of the management unit master / slave status initialization process provided by an embodiment of the present invention;

[0041] Figure 5 This is a flowchart of the cleaning operation process performed by the management unit of the management surface according to an embodiment of the present invention;

[0042] Figure 6 This is a flowchart of the data plane forwarding group primary / standby state initialization process provided according to an embodiment of the present invention;

[0043] Figure 7 This is a flowchart of the control plane execution body master / slave state initialization according to an embodiment of the present invention;

[0044] Figure 8 This is a flowchart of the dynamic scheduling process of the control plane execution body according to an embodiment of the present invention;

[0045] Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0046] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0047] Example 1

[0048] This invention provides a router based on collaborative temperature backup, such as... Figure 1 As shown, it includes a control plane 101, a data plane 102, and a management plane 103.

[0049] The router control plane 101 is complex and contains many hidden risks and vulnerabilities, making it a key area requiring protection. This invention, based on a Dynamic Heterogeneous Redundancy (DHR) architecture, constructs a redundant heterogeneous control plane 101, enabling it to withstand unknown security threats. In this embodiment, the control plane 101 consists of n (n≥2) multi-dimensional heterogeneous functionally equivalent execution entities. These multi-dimensional heterogeneity includes, but is not limited to, CPU heterogeneity, operating system heterogeneity, and application heterogeneity. Specifically, it can also be composed of multiple homogeneous functionally equivalent execution entities. In this case, the router device's ability to withstand common-mode defects decreases, but its resilience and self-healing capabilities remain unaffected.

[0050] In this embodiment, all executors have identical functions, forming a functionally equivalent heterogeneous executor pool. Each executor runs various routing protocols and distributes routing information, such as routing table entries, to the forwarding unit through the forwarding control unit, enabling the forwarding unit to correctly forward packets and other data based on the routing information. During operation, only one executor is in the primary state (primary executor), while the others are in standby state (standby executors) or offline. Each executor can operate in any of the following modes: hot standby, warm standby, or cold standby. Hot standby ensures complete consistency between the primary and standby states, and the switch between primary and standby executors is imperceptible to the outside world. Cold standby has the slowest switchover, and a restart of the routing protocol will cause a interruption of forwarding services. Warm standby has a switchover time between hot and cold standby, and will also cause a short-term interruption of forwarding services. During device operation, the primary and standby states of the executors can be dynamically scheduled and switched. In the event of an anomaly in the primary executor, the standby executor switches to the primary state, achieving control plane 101 resilience self-healing. The control plane 101 is based on a dynamic heterogeneous redundancy (DHR) architecture, which enables it to achieve high security at the architectural level and has the ability to detect, defend against, recover from and adapt to random failures and threats.

[0051] In the initial state (i.e., immediately after power-on), the executor sends an online request to the data forwarding group in the primary state and sets the primary / standby state according to the returned online response. The data forwarding group in the primary state determines whether the executor requesting the online request is the first one. If the executor requesting the online request is the first one, the returned online response is in the primary state. If the executor requesting the online request is not the first one, the returned online response is in the standby state.

[0052] Each execution unit in control plane 101 can be an independent physical execution unit, or it can be composed of a combination of several physical execution units and several virtual execution units. In this embodiment, a physical execution unit refers to an execution unit whose functions run directly on the operating system of an independent physical unit. In this embodiment, a virtual execution unit refers to multiple execution unit functions running on the operating system of a physical unit in a virtualized manner. The virtualization method includes, but is not limited to, virtual machines, container technology, Linux namespaces, etc. An example of physical and virtual execution units jointly forming a heterogeneous redundant execution unit pool is shown below. Figure 2 As shown, this example has a total of n executors, including m-1 physical executors and n-m+1 virtual executors. Clearly, Figure 2 This is merely an example and does not limit the composition of the functionally equivalent implementer described in this invention.

[0053] In this embodiment, the data plane 102 includes at least two data forwarding groups. For example... Figure 1 As shown, taking the construction of a redundant data plane 102 with two redundancies as an example, the data plane 102 includes two functionally equivalent data forwarding groups, each of which includes a heterogeneous forwarding control unit and a heterogeneous forwarding unit. In particular, it can also be composed of two functionally equivalent homogeneous forwarding control units and forwarding units. In this case, the router device's ability to resist common-mode defects is reduced, but its resilience and self-healing ability are not affected.

[0054] The forwarding control unit of data plane 102 in this embodiment is responsible for caching the configuration and routing information issued by the execution body of control plane 101 and issuing it to the forwarding units, controlling the power-on initialization of the forwarding units corresponding to this forwarding group, and scheduling the execution body of control plane 101. The forwarding units of data plane 102 in this embodiment are responsible for completing the router's data forwarding according to the routing information.

[0055] The two data forwarding groups on data plane 102 operate in a warm standby mode. During operation, one data forwarding group is in the primary state, with its forwarding control unit and forwarding units operating at full functionality. The other data forwarding group is in the standby state, where its forwarding control unit only operates core modules, such as the data and status synchronization module. This module communicates with the primary forwarding control unit to ensure that the data and status of the standby forwarding control unit are consistent with the primary forwarding control unit, facilitating rapid primary / standby failover. Its forwarding units cease operation and are in a powered-off state. In the event of an anomaly in the primary data forwarding group, the standby data forwarding group switches to the primary state, achieving resilient self-healing.

[0056] In practice, in the initial state, the data forwarding group can set the primary / standby status by reading the predefined configuration. If there is no predefined configuration, it will negotiate with other data forwarding groups to decide on the primary / standby status.

[0057] The forwarding control unit in the primary state can also monitor the software status of the executor and perform primary / backup switching of the executor according to preset scheduling rules. Preset scheduling rules include, but are not limited to, periodic scheduling of executors and scheduling when an executor experiences a decision anomaly. The periodic scheduling refers to switching the executor to primary / backup status according to a set time period, which aims to achieve the dynamism of the control plane 101, thereby improving its security. The decision anomaly refers to a system functional abnormality of the executor detected collaboratively by the executor of the control plane 101 and the forwarding control unit of the data plane 102, based on the Dynamic Heterogeneous Redundancy (DHR) architecture. The abnormal executor may be either the primary or backup executor. This embodiment achieves high reliability, high security, and strong resilience of the control plane 101 through dynamic scheduling of executors.

[0058] Specifically, the preset scheduling rules can be as follows: when an abnormality is detected in the execution unit, the forwarding control unit in the primary state sends a cleaning control command to the management unit in the primary state, so that the management unit cleans the abnormal execution unit according to the cleaning control command and sets the execution unit's state to offline; when the cleaned execution unit is in the primary state, the forwarding control unit in the primary state selects an execution unit from the standby execution units to switch to the primary state; when the preset scheduling cycle is reached, the forwarding control unit in the primary state switches the execution units in the primary state to the standby state and selects an execution unit from the standby execution units to switch to the primary state. This achieves the dynamism of the control plane 101 and improves the security of the control plane 101.

[0059] The management surface 103 in this embodiment includes at least two management units, such as... Figure 1 As shown, taking the construction of a dual-redundant management plane 103 as an example, the management plane 103 consists of two functionally equivalent heterogeneous management units. In particular, it can also be composed of two functionally equivalent homogeneous management units. In this case, the router device's ability to resist common-mode defects decreases, but its resilience and self-healing ability are not affected.

[0060] Each management unit operates in hot standby mode. During operation, only one management unit is in the primary state (primary management unit), while other primary management units are in standby state (standby management units) or offline. The primary management unit is responsible for sequentially activating the data plane 102 forwarding control unit and the control plane 101 execution unit. The primary management unit is responsible for real-time monitoring of the hardware status of the control plane 101 execution unit, the data plane 102 forwarding control unit, and the forwarding units. Hardware status monitoring includes, but is not limited to, current sensor signals, voltage sensor signals, and temperature sensor signals. When a hardware anomaly is detected, the corresponding unit is subjected to a soft reset or power-on / off cleaning operation ("cleaning" refers to a soft reset / power-on / off operation, meaning that this operation attempts to eliminate the existing anomaly). If the number of cleaning operations for a unit exceeds the set maximum number of cleaning operations within a set time range, the unit is powered off. The primary management unit receives control signals from the data plane 102 forwarding control unit and provides soft reset / power-on / off control functions for the control plane 101 execution unit. When the primary management unit malfunctions, the standby management unit switches to the primary state, fully taking over the functions of the original primary management unit, thus achieving resilience and self-healing of the management plane 103.

[0061] When the router powers on, the management unit of management plane 103 is started first, followed by the main management unit sequentially starting the forwarding control unit of data plane 102 and the execution unit of control plane 101. In the initial state (i.e., when the management unit powers on), the management unit decides its primary or backup state by negotiating with other management units. If it does not receive negotiation information from other management units within a preset time, it automatically enters the primary state.

[0062] In other embodiments, the router may also include a data channel and a management channel. The data channel is a data transmission channel between the execution units of the control plane 101 and between the execution units of the control plane 101 and the forwarding control unit of the data plane 102. The management channel is a transmission channel for the management unit of the management plane 103 to detect hardware status information of the execution units of the control plane 101, the forwarding control unit of the data plane 102, and the forwarding unit, and to transmit hardware control commands such as soft reset / power-on / off.

[0063] It should be noted that, depending on the hardware implementation, the data channel and management channel described in this embodiment can be either independent physical communication channels or they can share the same physical communication channel. This invention only logically divides them into data channels and management channels.

[0064] This invention implements routing protocol execution through the control plane execution unit, data forwarding control through the data plane data forwarding group, and anomaly detection and mitigation for both the data and control planes through the management plane management unit. Combined with multi-dimensional redundancy, real-time monitoring, and dynamic scheduling of the primary and backup states of the control, data, and management planes, it achieves collaborative warm standby and resilient self-healing of the router's overall internal architecture. This integrated approach at the system level achieves high reliability and high security, balancing fault recovery speed and power consumption. Based on an integrated security architecture and dynamic heterogeneous redundancy design, this invention achieves high reliability at the architecture level, significantly reducing the reliability requirements of individual internal components. This allows the use of common COTS (Chip-on-Ship) devices in the fabrication of spaceborne routers, significantly reducing equipment costs and achieving an optimal balance between cost and reliability.

[0065] Example 2

[0066] This embodiment provides a routing method based on collaborative temperature backup, such as... Figure 3 As shown, it includes the following steps:

[0067] S201. The execution entity in the control plane in the primary state sends routing information to the data plane according to the routing protocol. The execution entity in the backup state switches to the primary state in the event of an error in the execution entity in the primary state.

[0068] The control plane includes at least two actuators, one of which is in a master state and at least one of which is in a standby state.

[0069] S202. The data forwarding group in the primary state in the data plane performs data forwarding according to the routing information received from the control plane. The data forwarding group in the backup state synchronizes data with the data forwarding group in the primary state, and switches to the primary state in the event of an anomaly in the data forwarding group in the primary state.

[0070] The data plane includes at least two data forwarding groups, one of which is in a primary state and at least one of which is in a standby state.

[0071] S203. The management unit in the primary state in the management plane performs hardware status anomaly monitoring and anomaly cleaning on the data plane and the control plane. The management unit in the standby state switches to the primary state in the event of an anomaly in the primary management unit in the primary state.

[0072] The management interface includes at least two management units, one of which is in a master state and at least one of which is in a standby state.

[0073] In this embodiment, the management unit status of the management plane is defined as: [primary status, backup status, offline status].

[0074] When the management unit in the management plane starts up, it performs state initialization, and the process is as follows: Figure 4 As shown, the specific steps are described below:

[0075] The first step is to start the management unit and initialize its status to offline.

[0076] The second step involves reading the predefined configuration if one exists, setting the status of the management unit according to the predefined configuration, and ending the initialization process. If no predefined configuration exists, the management unit negotiates the primary / standby status with another management unit and proceeds to the next step of the process.

[0077] The third step is to determine the primary / backup status of this management unit based on the negotiation situation. If no negotiation information is received from the other management unit after the set time threshold is exceeded, it will automatically enter the primary status.

[0078] The fourth step is to start the data plane forwarding control unit and the management plane execution units in sequence after entering the main state.

[0079] The standby management unit does not monitor the hardware status of the data plane and control plane units, nor does it respond to control commands sent by the forwarding control unit. It only monitors the status of the primary management unit in real time. When an anomaly is detected in the primary management unit, the standby management unit automatically switches to the primary state and takes over the functions of the primary management unit, achieving resilience and self-healing of the management plane.

[0080] The management unit in master mode monitors the hardware status of each unit on the data plane and control plane in real time through the management channel. When an abnormal hardware signal is detected in a unit, a soft reset or power-on / off operation is performed on the unit according to preset rules to achieve anomaly cleaning and status restoration. The processing flow is as follows: Figure 4 As shown, the specific steps are described below:

[0081] The first step is that the main management unit detects an abnormal hardware status signal for a certain unit.

[0082] The second step is to determine whether the number of cleaning cycles for the unit has reached the maximum set number within the set time threshold.

[0083] The third step is to power off the unit if the maximum number of cleaning cycles has been reached, thus ending the processing procedure.

[0084] Fourth step: If the maximum number of cleaning cycles has not been reached, perform a soft reset or power-on / off cleaning operation on the unit, update the cleaning operation count, and the process ends.

[0085] If a unit remains in an abnormal state after a set number of soft resets or power-on / off cleaning operations within a set time threshold, the unit will be powered off. The purpose of this operation is to prevent the unit from continuing to malfunction and causing damage and interference to the overall hardware and software functionality of the router device.

[0086] The management unit in the master state receives control signals sent by the data plane forwarding control unit through the management channel, and performs soft reset or power-on / off operations on the relevant data plane forwarding units and control plane actuators according to the control signals.

[0087] The data plane data forwarding group status is defined as: [primary status, backup status, offline status].

[0088] The data plane forwarding group initializes its state upon startup, as follows: Figure 6 As shown, the specific steps are described below:

[0089] The first step is to start the forwarding control unit inside the data forwarding group and initialize its status to offline.

[0090] The second step is to read the predefined configuration if it exists, set the status of this forwarding group according to the predefined configuration, and then proceed to the fourth step; if no predefined configuration exists, negotiate the primary / backup status with another data forwarding group.

[0091] The third step is to determine the primary / backup status of this forwarding group based on the negotiation situation. If no negotiation information from another forwarding group is received after the set time threshold is exceeded, the group will automatically enter the primary status.

[0092] Fourth step: If this data forwarding group is in standby mode, perform data synchronization operation with the main forwarding control unit, but do not perform power-on initialization operation on the corresponding forwarding unit of this forwarding group. This forwarding group is running in warm standby mode, and the initialization is completed.

[0093] Fifth step: If this data forwarding group is in the master state, then send the power-on command of the forwarding unit corresponding to this forwarding group to the master management unit through the management channel to perform the power-on and initialization operations of the forwarding unit corresponding to this forwarding group. This forwarding group runs in the full-function state and the initialization is completed.

[0094] The control plane execution state is defined as: [primary state, backup state, offline state].

[0095] The control plane execution body initializes its state upon startup, as follows: Figure 7 As shown, the specific steps are described below:

[0096] The first step is that when the executor starts, the state is initialized to offline and a request to come online is sent to the main forwarding control unit of the data plane.

[0097] The second step is for the data plane main forwarding control unit to receive the execution unit online request. If it is the first execution unit to come online, the execution unit's status is set to the master status and the current execution unit status is immediately announced. If it is not the first execution unit to come online, the execution unit's status is set to the standby status and the current execution unit status is immediately announced.

[0098] Third, the executor enters either the primary or backup state based on the state set by the forwarding control unit. If the executor enters the backup state, it performs data synchronization with the primary executor, and initialization ends; if the executor enters the primary state, initialization ends.

[0099] The data plane primary forwarding control unit dynamically schedules control plane executors according to scheduling rules, including but not limited to periodic scheduling of executors and scheduling when an executor experiences a decision anomaly. Periodic scheduling refers to switching the primary executor at set time intervals, which aims to achieve dynamic control plane performance and thus improve its security. Decision anomalies refer to system malfunctions detected collaboratively by the control plane executors and the data plane forwarding control unit based on a Dynamic Heterogeneous Redundancy (DHR) architecture. The malfunctioning executor could be either the primary or backup executor. This embodiment achieves high reliability, high security, and strong resilience of the control plane through dynamic scheduling of executors.

[0100] The processing flow after triggering the scheduling rule is as follows: Figure 8 As shown, the specific steps are described below:

[0101] The first step is that the execution body scheduling rules are triggered, and the main forwarding control unit switches the control plane execution body scheduling according to the scheduling rules.

[0102] The second step involves the main forwarding control unit sending a control command to the main management unit via the management channel to schedule the currently scheduled execution unit offline for cleaning and set its status to offline.

[0103] The third step is to select an executor from the current backup executors as the new master executor and set the status of that executor to master.

[0104] The fourth step is for the main forwarding control unit to notify the execution entity of its status and complete the scheduling.

[0105] This invention implements routing protocol execution through the control plane execution unit, data forwarding control through the data plane data forwarding group, and anomaly detection and mitigation for both the data and control planes through the management plane management unit. Combined with multi-dimensional redundancy, real-time monitoring, and dynamic scheduling of the primary and backup states of the control, data, and management planes, it achieves collaborative warm standby and resilient self-healing of the router's overall internal architecture. This integrated approach at the system level achieves high reliability and high security, balancing fault recovery speed and power consumption. Based on an integrated security architecture and dynamic heterogeneous redundancy design, this invention achieves high reliability at the architecture level, significantly reducing the reliability requirements of individual internal components. This allows the use of common COTS (Chip-on-Ship) devices in the fabrication of spaceborne routers, significantly reducing equipment costs and achieving an optimal balance between cost and reliability.

[0106] Example 3

[0107] This invention provides a computer device to provide services for implementing the method of Embodiment 2 above. For example... Figure 9 As shown, the device may include: a memory 301 storing a computer-executable program; a processor 302 coupled to the memory 301; the processor 302 calls the computer-executable program stored in the memory 301 to perform the steps in the method described in Embodiment 2.

[0108] Memory 301 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The device may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, memory 301 may be used to read and write non-removable, non-volatile magnetic media (commonly referred to as a "hard disk drive"). A program / utility having a set (at least one) of program modules may be stored, for example, in memory 301. Such program modules include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The computer-executable program of the program modules typically performs the functions and / or methods described in the embodiments of the present invention.

[0109] The processor 302 executes various functional applications and data processing by running programs stored in the memory 301, such as implementing the method provided in Embodiment 2 of the present invention.

[0110] The code of a computer executable program can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages.

[0111] Example 4

[0112] This invention provides a storage medium containing a computer-executable program, which, when executed by a computer processor, is used to perform the method of Embodiment 2.

[0113] The storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0114] Of course, the computer-executable program in the storage medium provided in the embodiments of the present invention is not limited to the above-described method operations, but can also perform related operations in the methods provided in any embodiment of the present invention.

[0115] Example 5

[0116] This invention provides a computer program product, such as an app on a mobile phone or tablet, or an installer on a computer. The product includes a computer program / instructions that, when executed by a processor, implement the method described in Embodiment 2. The code for the computer-executable program used to perform the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0117] It should be understood that the embodiments and descriptions above are only the principles, main features and advantages of the present invention. Various changes and modifications can be made to the present invention without departing from the spirit and scope of the invention, and all such changes and modifications fall within the protection scope of the present invention.

Claims

1. A router based on collaborative warm standby, characterized in that, include: The control plane includes at least two execution entities, one of which is in a primary state and at least one of which is in a standby state. The execution entity in the primary state sends routing information to the data plane according to the routing protocol. The execution entity in the standby state switches to the primary state if the execution entity in the primary state malfunctions. The data plane includes at least two data forwarding groups, one of which is in a primary state and at least one of which is in a backup state. The data forwarding group in the primary state performs data forwarding according to the routing information received from the control plane. The data forwarding group in the backup state synchronizes data with the data forwarding group in the primary state and switches to the primary state in case of an anomaly in the data forwarding group in the primary state. The management plane includes at least two management units, one of which is in a primary state and at least one of which is in a standby state. The management unit in the primary state performs hardware status anomaly monitoring and anomaly cleaning on the data plane and the control plane. The management unit in the standby state switches to the primary state if the primary management unit in the primary state malfunctions.

2. The router based on collaborative temperature backup according to claim 1, characterized in that, The data forwarding group includes: A forwarding control unit is used to control the forwarding unit and send the routing information received from the control plane to the forwarding unit; A forwarding unit is used to forward data based on the received routing information; In this configuration, the forwarding control unit and the forwarding unit in the primary state operate simultaneously, the forwarding control unit in the standby state synchronizes its data with the forwarding control unit in the primary state, and the forwarding unit in the standby state stops operating.

3. The router based on collaborative warm backup according to claim 2, characterized in that, The forwarding control unit, which is in the primary state, is also used to monitor the software state anomalies of the execution unit and to perform primary / backup switching of the execution unit according to preset scheduling rules.

4. The router based on collaborative warm backup according to claim 3, characterized in that, The preset scheduling rules include: If an abnormality is detected in the execution unit, the forwarding control unit in the primary state sends a cleaning control command to the management unit in the primary state, so that the management unit cleans the abnormal execution unit according to the cleaning control command and sets the execution unit's state to offline state; if the cleaned execution unit is in the primary state, the forwarding control unit in the primary state selects an execution unit from the execution units in the standby state to switch to the primary state; When the preset scheduling period is reached, the forwarding control unit in the primary state switches the executor in the primary state to the standby state, and selects one executor from the standby executors to switch to the primary state.

5. The router based on collaborative temperature backup according to claim 1, characterized in that, In the initial state, the executor sends an online request to the data forwarding group in the primary state, and sets the primary / backup state according to the returned online response. The data forwarding group in the primary state determines whether the executor in the online request is the first one. If the executor in the online request is the first one, the returned online response is in the primary state. If the executor in the online request is not the first one, the returned online response is in the backup state.

6. The router based on collaborative temperature backup according to claim 1, characterized in that, The router is provided with a data channel and a management channel. The data channel is a data transmission channel between the execution entities and between the execution entities and the data forwarding group. The management channel is a data transmission channel between the management unit and the execution entities, and between the management unit and the data forwarding group.

7. The router based on collaborative warm backup according to claim 1, characterized in that, The standby state of the execution unit is any one of hot standby, warm standby, and cold standby. The standby state of the data forwarding group is warm standby. The standby state of the management unit is hot standby.

8. A routing method based on cooperative warm standby, characterized in that, The method includes: In the control plane, the executor in the primary state sends routing information to the data plane according to the routing protocol, and the executor in the backup state switches to the primary state in the event of an anomaly in the executor in the primary state; wherein, the control plane includes at least two executors, one of which is in the primary state and at least one of which is in the backup state; In the data plane, the data forwarding group in the primary state performs data forwarding according to the routing information received from the control plane. The data forwarding group in the backup state synchronizes data with the data forwarding group in the primary state and switches to the primary state in case of an anomaly in the data forwarding group in the primary state. The data plane includes at least two data forwarding groups, one of which is in the primary state and at least one of which is in the backup state. The management unit in the primary state of the management plane performs hardware status anomaly monitoring and anomaly cleaning on the data plane and the control plane. The management unit in the standby state switches to the primary state when the primary management unit in the primary state is abnormal. The management plane includes at least two management units, one of which is in the primary state and at least one of which is in the standby state.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the method of claim 8.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method of claim 8.