Programmable network fault rapid detection and recovery method and device based on multi-dimensional heartbeat detection information
By adopting multi-dimensional heartbeat detection information and voting mechanism methods in the data plane, the problems of high delay and low efficiency of network fault detection and recovery in the prior art are solved, and fast and accurate fault detection and recovery are achieved, improving the reliability and stability of the network.
Patent Information
- Application Number
- CN202410278232.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-03-12
AI Technical Summary
The prior art has problems of high latency and low efficiency in network failure detection and recovery, especially in complex and unstable network environments.
A programmable network fault fast detection and recovery method based on multi-dimensional heartbeat detection information is adopted. By selecting as a fault detection switch in the data plane, heartbeat detection between adjacent data transmission switches and between fault detection switches and data transmission switches is realized. The fault type is judged in combination with the voting mechanism and fault recovery is carried out.
It realizes rapid detection and autonomous classification identification of switch and link failures without passing through the control plane, reducing false alarms and missed alarms, improving the accuracy and efficiency of fault detection, shortening fault response time, and ensuring network reliability and stability.
Smart Images

Figure CN118175014B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software-defined network fault detection, and in particular to a method and device for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information. Background Art
[0002] With the emergence of many new network applications such as VR (Virtual Reality), AR (Augmented Reality), Industrial Internet and the highly anticipated Metaverse, the network scale has been significantly expanded and the network structure has become more complex. These applications have unprecedented requirements for network stability and reliability, making network fault detection and recovery critical.
[0003] In production environments, link or switch failures are common, mainly due to device crashes, reboots, connection problems, hardware or firmware defects, and power problems. These failures not only lead to degraded network performance, but may also cause service interruptions, inconvenience and economic losses to users. When traditional network architectures fail, they rely on routers to exchange distributed routing protocol messages to communicate changes in the network topology. However, the convergence process of distributed routing protocols may cause inconsistencies in the forwarding table, which can cause packet loss, especially in large networks, where this loss may last for a long time.
[0004] Furthermore, software-defined networks (SDNs) centrally manage programmable switch networks through a central controller, which can quickly detect and respond to faults when they occur. However, SDN still faces some technical challenges in practical applications. First, the process of the controller obtaining fault information from the network switch may introduce additional delays. Second, after the controller receives the fault notification and calculates the global rule changes, it is also a technical challenge to apply these changes to the running network. In addition, due to the communication efficiency issues between the data plane and the control plane, the programming and rule installation process in SDN may also generate delays of tens of milliseconds to several minutes.
[0005] Meanwhile, in fault detection schemes, the three-heartbeat mechanism is a common strategy. When a device or connection fails, the system sends multiple (usually three) heartbeat signals to verify whether the fault actually exists. If no response is received for multiple consecutive times (for example, three times), the system assumes that the device or connection has failed. However, a major disadvantage of this mechanism is that it may cause higher latency because it needs to wait for multiple responses to heartbeat signals to confirm the fault. This latency may be more significant in complex or unstable network environments.
[0006] In summary, the detection and recovery of network faults are crucial to ensuring network stability and reliability. Although existing technologies can achieve this goal to a certain extent, they still have problems such as high latency and low efficiency. Therefore, how to provide a more reliable and rapid fault detection mechanism based on existing programmable networks is a research issue that has both economic value and technical challenges. Summary of the invention
[0007] The purpose of the present invention is to provide a method and device for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information in view of the deficiencies of the prior art. The present invention can realize autonomous classification, identification and recovery of switch and link faults only in the data plane without passing through the control plane.
[0008] The objective of the present invention is achieved through the following technical solutions: In a first aspect, an embodiment of the present invention provides a method for rapid fault detection and recovery of a programmable network based on multi-dimensional heartbeat detection information, wherein the programmable network includes a plurality of data transmission switches, and a fault detection switch is selected from the plurality of data transmission switches, wherein the fault detection switch satisfies the following conditions: it is directly connected to the data transmission switch under its jurisdiction and the remaining bandwidth is reserved; the method for rapid fault detection and recovery includes the following steps:
[0009] (1) Implementing heartbeat detection between adjacent data transmission switches based on the heartbeat detection protocol: Each data transmission switch maintains a network topology adjacency table, generates a detection heartbeat data packet based on the heartbeat detection protocol to interact with neighboring nodes, and detects each other's heartbeat status in real time to obtain heartbeat detection results. The data transmission switch simultaneously sends its own operating status information and the heartbeat detection results between it and the fault detection switch to the other party. Finally, the heartbeat status of the neighboring node is identified based on whether the returned heartbeat data packet is received within the set time.
[0010] (2) Implementing heartbeat detection between the fault detection switch and the data transmission switch based on the heartbeat detection protocol: The fault detection switch and the data transmission switch maintain the interaction of heartbeat data packets, wherein the fault detection switch simultaneously sends the operating status information of itself and its neighboring nodes, and the data transmission switch sends the heartbeat detection results between itself and its neighboring nodes to the fault detection switch through the heartbeat data packets;
[0011] (3) Determine the fault type based on the voting mechanism: The fault detection switch and the data transmission switch determine the fault detection result of the node according to the node operation status they read; the fault detection switch and the data transmission switch determine whether a fault occurs based on the multi-dimensional fault detection results detected by each of them and their respective voting mechanisms, and further determine whether the fault type is a link fault or a switch fault;
[0012] (4) Fault recovery based on fault recovery mechanism: After detecting fault information and its type, the data transmission switch performs fault recovery according to the fault recovery strategy configured in the switch. If there is no corresponding fault recovery strategy, the affected data flow is redirected to the fault detection switch for transmission. After detecting fault information and its type, the fault detection switch promptly sends the fault information to the relevant switches including the upstream multi-hop data transmission switch of the affected data flow path and the data transmission switches on the backup path, so that corresponding measures can be taken to perform fault recovery.
[0013] Furthermore, the heartbeat detection protocol uses a standard Ethernet header, and the heartbeat detection protocol includes a heartbeat protocol type heartbeat_type, a signature of the sender of the heartbeat data packet signature, an identifier of the detected party Monitored_switch, a next protocol protoc_type, and n groups of switch state structure information, wherein n is the number of neighbor nodes corresponding to the detected party of the heartbeat data packet, and the switch state structure includes a switch label switch_id_i, a corresponding switch state switch_state_i, and whether it is the last bit is_last.
[0014] Furthermore, each of the data transmission switches also maintains a series of data transmission switch detection status registers data_switch_state_reg and corresponding data transmission switch detection time registers data_switch_state_time_reg, which are respectively used to record the node operation status detected by the data transmission switch and the last detection timestamp; at the same time, it also maintains a fault detection switch detection status register faultDetect_state_reg and a corresponding fault detection switch detection time register faultDetect_state_time_reg, which are used to record the node operation status detected by the fault detection switch and the last detection timestamp.
[0015] Furthermore, the specific implementation process of the heartbeat detection between adjacent data transmission switches is as follows:
[0016] The data transmission switch periodically sends a detection heartbeat data packet to its neighbor node, which carries its own operation status information and the heartbeat detection result between it and the fault detection switch. When the detection heartbeat data packet is sent, the operation status of the neighbor node in the data transmission switch detection status register data_switch_state_reg is updated to inactive, and the timestamp of the corresponding neighbor node in the data transmission switch detection time register data_switch_state_time_reg is updated;
[0017] When the neighbor node of the data transmission switch receives the detection heartbeat data packet, it updates its own data transmission switch detection state register data_switch_state_reg according to the heartbeat data packet sender signature signature field and the switch state structure in the detection heartbeat data packet, and clears the value in the switch state structure field in the detection heartbeat data packet, then adds the heartbeat detection result between itself and the fault detection switch, and modifies the heartbeat protocol type heartbeat_type field in the detection heartbeat data packet to REPLY, and returns it to the sending node;
[0018] After receiving the returned detection heartbeat data packet, the sending node updates the state of the corresponding neighbor node in the data transmission switch detection state register data_switch_state_reg to alive, and updates the timestamp of the corresponding neighbor node in the data transmission switch detection time register data_switch_state_time_reg to realize heartbeat detection between adjacent data transmission switches.
[0019] Furthermore, the specific implementation process of the heartbeat detection between the fault detection switch and the data transmission switch is as follows:
[0020] The fault detection switch sends detection heartbeat data packets to the data transmission switches within its jurisdiction one by one, reads the operating status of all neighboring nodes of the current target data transmission switch from the fault detection switch detection state register faultDetect_state_reg one by one, forms a switch state structure corresponding to all neighboring nodes, adds it to the detection heartbeat data packet, and sends the detection heartbeat data packet to the current target data transmission switch; at the same time, the fault detection switch updates the operating status of the data transmission switch in the fault detection switch detection state register faultDetect_state_reg to inactive, and records the corresponding sending timestamp in the fault detection switch detection time register faultDetect_state_time_reg;
[0021] When the target data transmission switch receives the detection heartbeat data packet, it reads the switch state structure field in the detection heartbeat data packet, and updates the fault detection switch detection state register faultDetect_state_reg in the target data transmission switch one by one according to the field information, records the heartbeat detection result between it and the fault detection switch and the heartbeat detection result between the transmitted fault detection switch and the neighboring node; at the same time, it reads the running state of the neighboring node stored in the data transmission switch detection state register data_switch_state_reg, updates it to the detection heartbeat data packet header, modifies the heartbeat protocol type heartbeat_type field to REPLY, and returns the modified detection heartbeat data packet to the fault detection switch;
[0022] When the fault detection switch receives the reply detection heartbeat data packet, it reads and stores the operating status information of the neighbor node of the data transmission switch carried in the detection heartbeat data packet, modifies the operating status of the corresponding data transmission switch in the fault detection switch detection state register faultDetect_state_reg, and updates the corresponding timestamp in the fault detection switch detection time register faultDetect_state_time_reg.
[0023] Furthermore, the specific implementation process of judging the fault type based on the voting mechanism is as follows:
[0024] In each heartbeat cycle, the fault detection result of the data detection switch on the neighbor node is read according to the data transmission switch detection state register data_switch_state_reg and its corresponding data transmission switch detection time register data_switch_state_time_reg: if the running state of the node is inactive, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the data transmission switch detection time register data_switch_state_time_reg and the current timestamp is less than or equal to the preset first time threshold, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the data transmission switch detection time register data_switch_state_time_reg and the current timestamp is greater than the preset first time threshold, the trust state is faulty;
[0025] The fault detection result of the fault detection switch on the neighbor node is read according to the fault detection switch detection state register faultDetect_state_reg and its corresponding fault detection switch detection time register faultDetect_state_time_reg: if the running state of the node is inactive, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the fault detection switch detection time register faultDetect_state_time_reg and the current timestamp is less than or equal to the preset second time threshold, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the fault detection switch detection time register faultDetect_state_time_reg and the current timestamp is greater than the preset second time threshold, the trust state is faulty;
[0026] In the voting mechanism for fault detection switches, for any switch, the fault detection results of the fault detection switch and the switch and its n neighboring nodes are obtained to obtain n+1-dimensional fault detection results; different weights are set for fault detection results from different sources according to their credibility. The detection result value is calculated based on the weight accumulation. If the detection result value is less than or equal to the preset ratio threshold α, it is considered that the switch is faulty; if the detection result value is greater than the preset ratio threshold α, it is considered that the link between the neighboring switch that has detected the fault and the switch is faulty;
[0027] In the voting mechanism for detection by the data transmission switch, for any neighbor node, the fault detection results of the data transmission switch itself and the fault detection switch on the neighbor node are obtained. If both fault detection results are normal, the neighbor node and the connected link status are considered normal; if both fault detection results are faulty, the neighbor node is considered to have a fault; if the data transmission switch itself detects a fault on the neighbor node and the fault detection switch detects a normal fault on the neighbor node, the link between the neighbor node and the data transmission switch is considered to have a fault.
[0028] Furthermore, the specific implementation process of determining the fault type based on the voting mechanism also includes: when the data transmission switch detects that the fault detection switch may have a fault, the data transmission switch will automatically switch to the three-heartbeat mechanism for detection, and obtain the operating status information by continuously sending and confirming three heartbeat signals.
[0029] Furthermore, the specific implementation process of performing fault recovery based on the fault recovery mechanism is as follows:
[0030] After detecting the fault information and its type, the data transmission switch recovers according to the fault recovery strategy configured in the switch. If there is no corresponding fault recovery strategy, the affected data flow is redirected to the fault detection switch for transmission. After receiving the affected data flow, the fault detection switch selectively transmits the data flow to the multi-hop switch or the destination switch under the original path according to the different fault types of the link and the switch;
[0031] After detecting the fault information and its type, the fault detection switch promptly sends the fault information to relevant switches including the upstream multi-hop data transmission switch of the affected data flow path and the data transmission switch on the backup path, so that corresponding measures can be taken to recover from the fault.
[0032] A second aspect of an embodiment of the present invention provides a programmable network fault rapid detection and recovery device based on multi-dimensional heartbeat detection information, comprising one or more processors and a memory, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information.
[0033] A third aspect of an embodiment of the present invention provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it is used to implement the above-mentioned programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information.
[0034] The beneficial effects of the present invention are as follows: the programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information of the present invention runs directly on the data plane, and can detect and process faults in real time on the transmission path of network traffic, avoiding communication delays between the control plane and the data plane; the present invention comprehensively judges faults through a multi-dimensional fault detection mechanism and a voting mechanism located in the data plane, and can more comprehensively and accurately identify faults in the network, including switch faults, link faults, etc., which helps to reduce false alarms and missed alarms, improve the accuracy of fault detection, make fault detection more efficient, and effectively shorten the fault response time; the present invention can quickly restore connectivity when a network fault occurs through a multi-dimensional fault detection method and a fault recovery mechanism, thereby ensuring service continuity and network reliability; the present invention has shown significant beneficial effects in improving the accuracy and efficiency of fault detection, and enhancing the reliability and stability of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flow chart of a method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information provided by an embodiment of the present invention;
[0036] Figure 2is a schematic diagram of a heartbeat detection protocol provided by an embodiment of the present invention;
[0037] Figure 3 It is a structural schematic diagram of a programmable network fault rapid detection and recovery device based on multi-dimensional heartbeat detection information provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0038] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0039] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0040] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0041] The present invention is described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the features of the following embodiments and implementations can be combined with each other.
[0042] The programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information of the present invention divides the fault detection switch and the data transmission switch, and combines the fault detection mechanism of the multi-dimensional heartbeat detection information, through the heartbeat detection between adjacent data transmission switches, the heartbeat detection between the fault detection switch and the data transmission switch, the voting mechanism and the classification and identification of the fault type and the fault recovery mechanism, thereby realizing the autonomous classification and identification and recovery of switch and link faults only in the data plane without passing through the control plane.
[0043] Combination Figure 1The present invention describes a method for rapid fault detection and recovery of a programmable network based on multi-dimensional heartbeat detection information, wherein the programmable network includes multiple data transmission switches, and a fault detection switch is selected from the multiple data transmission switches. The fault detection switch meets the following conditions: it is directly connected to the data transmission switch under its jurisdiction and retains a certain amount of residual bandwidth for heartbeat detection and fault recovery. Figure 1 As shown, the fault rapid detection and recovery method specifically includes the following steps:
[0044] (1) Heartbeat detection between adjacent data transmission switches is implemented based on the heartbeat detection protocol: Each data transmission switch maintains a network topology adjacency table, generates a detection heartbeat data packet based on the heartbeat detection protocol to interact with neighboring nodes, and detects each other's heartbeat status in real time to obtain the heartbeat detection result. The data transmission switch simultaneously sends its own operating status information and the heartbeat detection result between it and the fault detection switch to the other party. Finally, the heartbeat status of the neighboring node is identified based on whether the returned heartbeat data packet is received within the set time.
[0045] It should be understood that when implementing the heartbeat detection between adjacent data transmission switches based on the heartbeat detection protocol, the heartbeat detection protocol is embedded in the data packet to form a heartbeat data packet, whose structure is as follows: Figure 2 After switch a sends a heartbeat packet to switch b, as long as it successfully receives the return heartbeat packet from switch b REPLY, it means that switch a detects that switch b is alive, which is the heartbeat state. In this way, the heartbeat states of each other can be obtained.
[0046] Furthermore, the heartbeat detection protocol uses a standard Ethernet header. Figure 2 As shown, the heartbeat detection protocol includes the heartbeat protocol type heartbeat_type, the signature of the sender of the heartbeat data packet signature, the detected identifier Monitored_switch, the next protocol protoc_type and n groups of switch state structures, where n is the number of neighbor nodes corresponding to the detected heartbeat data packet. The switch state structure includes the switch number switch_id_i, the corresponding switch state switch_state_i and whether it is the last bit is_last. The switch state structure is used to transmit the activity information of the switch. The heartbeat detection protocol can be expressed as:
[0047] Γ:[ω,σ,μ,Ω,(∈0,ξ0,ζ n ),…,(∈ n ,ξ n ,ζ n )]
[0048] Among them, ω represents the heartbeat protocol type heartbeat_type, σ represents the signature of the heartbeat data packet sender signature, μ represents the detected identifier Monitored_switch, Ω represents the next protocol protoc_type, (∈ i ,ξ i ,ζ i ) represents the switch state structure of the i-th neighbor node of the heartbeat data packet being detected, ∈ i Indicates the switch number switch_id_i of the i-th neighbor node, ξ i Indicates the switch state corresponding to the i-th neighbor node switch_state_i,ζ i Indicates whether the i-th neighbor node is the last one is_last, i is a non-negative integer, 0≤i≤n.
[0049] It should be understood that the switch state structure may be a state structure of a data transmission switch or a state structure of a fault detection switch.
[0050] Furthermore, each data transmission switch also maintains a series of data transmission switch detection status registers data_switch_state_reg and corresponding data transmission switch detection time registers data_switch_state_time_reg, which are respectively used to record the node operation status detected by the data transmission switch and the last detection timestamp; at the same time, it also maintains a fault detection switch detection status register faultDetect_state_reg and a corresponding fault detection switch detection time register faultDetect_state_time_reg, which are used to record the node operation status detected by the fault detection switch and the last detection timestamp.
[0051] Furthermore, the specific implementation process of the heartbeat detection between adjacent data transmission switches is as follows:
[0052] (1.1) The data transmission switch periodically sends a detection heartbeat data packet to its neighbor node. The detection heartbeat data packet carries its own operating status information and the heartbeat detection result between it and the fault detection switch. When the detection heartbeat data packet is sent, the operating status of the neighbor node in the data transmission switch detection status register data_switch_state_reg is updated to inactive (NOALIVE), and the timestamp of the corresponding neighbor node in the data transmission switch detection time register data_switch_state_time_reg is updated at the same time.
[0053] It should be understood that when the detection heartbeat data packet is sent, the operating status of the neighbor node in the data transmission switch detection status register data_switch_state_reg is updated to inactive, which is a state of preparation for failure. Because the register can only be operated through data packets, the operating status value and timestamp are combined to jointly determine whether there is a fault, so as to record the status before the returned heartbeat data packet is received; if there is no fault, the returned heartbeat data packet will change this state back to ALIVE.
[0054] Specifically, in Figure 1 In the programmable network shown in the figure, s1~s6 are data transmission switches, h7 is a fault detection switch, and according to the network topology adjacency table maintained by s6, the neighbor nodes of s6 are Neighbor(s6)={s1,s2,s3,s4}, and s6 sends a detection heartbeat data packet to s2: Γ:[ASK,s6,s 2, 0x0800, (h7, ALIVE, 1)], the detection heartbeat data packet contains the operating status information of s6 itself and the heartbeat detection result between it and h7. When the detection heartbeat data packet is sent, the operating status of the neighbor node in the data transmission switch detection status register data_switch_state_reg is updated to be inactive, that is, data_switch_state_reg (s2) = NOALIVE, and the timestamp of the corresponding neighbor node in the data transmission switch detection time register data_switch_state_time_reg data_switch_state_time_reg (s2) = timestamp (sending time).
[0055] (1.2) When the neighbor node of the data transmission switch receives the detection heartbeat data packet, it updates its own data transmission switch detection state register data_switch_state_reg according to the heartbeat data packet sender signature signature field and the switch state structure in the detection heartbeat data packet, and clears the value in the switch state structure field in the detection heartbeat data packet. It then adds the heartbeat detection result between itself and the fault detection switch, and modifies the heartbeat protocol type heartbeat_type field in the detection heartbeat data packet to return (REPLY), and returns it to the issuing node.
[0056] Specifically, Figure 1As shown, when s6's neighbor node s2 receives the detection heartbeat data packet, it updates its own data transmission switch detection state register data_switch_state_reg according to the heartbeat data packet sender signature signature field and the switch state structure in the detection heartbeat data packet, that is, data_switch_state_reg(s6)=ALIVE, data_switch_state_reg(h7)=ALIVE. Clear the value in the switch state structure field in the detection heartbeat data packet, then add the heartbeat detection result between itself and the fault detection switch, and modify the heartbeat protocol type heartbeat_type field in the detection heartbeat data packet to return (REPLY), and finally return the modified detection heartbeat data packet, i.e., Γ:[REPLY,s6,s2,0x0800,(h7,ALIVE,1)] to the sending node s6.
[0057] It should be understood that when the neighbor node s2 of s6 receives the detection heartbeat data packet, it indicates that the running state of s6 is alive, and the data transmission switch detection state register data_switch_state_reg is updated accordingly.
[0058] (1.3) After receiving the returned detection heartbeat data packet, the sending node updates the state of the corresponding neighbor node in the data transmission switch detection state register data_switch_state_reg to alive (ALIVE), and at the same time updates the timestamp of the corresponding neighbor node in the data transmission switch detection time register data_switch_state_time_reg, thereby realizing heartbeat detection between adjacent data transmission switches.
[0059] (2) Heartbeat detection between the fault detection switch and the data transmission switch is implemented based on the heartbeat detection protocol: the fault detection switch and the data transmission switch maintain the interaction of heartbeat data packets, wherein the fault detection switch simultaneously sends the operating status information of itself and its neighboring nodes, and the data transmission switch sends the heartbeat detection results between itself and its neighboring nodes to the fault detection switch via heartbeat data packets.
[0060] Furthermore, the specific implementation process of the heartbeat detection between the fault detection switch and the data transmission switch is as follows:
[0061] (2.1) The fault detection switch sends detection heartbeat data packets to the data transmission switches within its jurisdiction one by one, reads the operating status of all neighbor nodes of the current target data transmission switch from the fault detection switch detection state register faultDetect_state_reg one by one, constructs a switch state structure corresponding to all neighbor nodes, adds it to the detection heartbeat data packet, and sends the detection heartbeat data packet to the current target data transmission switch; at the same time, the fault detection switch updates the operating status of the data transmission switch in the fault detection switch detection state register faultDetect_state_reg to inactive (NOALIVE), and records the corresponding sending timestamp in the fault detection switch detection time register faultDetect_state_time_reg.
[0062] For example, Figure 1 As shown, h7's jurisdiction is Partition(h7) = {s1, s2, s3, s4, s5, s6}, and h7 will generate 6 detection heartbeat packets. When it is sent to s6, the network topology adjacency table of s6 is read as Neighbor(s6) = {s1, s2, s3, s4}. h7 reads the operating status of the corresponding neighbor nodes s1, s2, s3, s4 one by one from the fault detection switch detection state register faultDetect_state_reg, puts it into the switch state structure field, and generates the corresponding detection heartbeat packet as shown in the following formula:
[0063]
[0064] The detection heartbeat data packet shown in the above formula is sent to the current target data transmission switch s6. At the same time, the operation state of s6 in the fault detection switch detection state register faultDetect_state_time_reg is updated to NOALIVE, and the corresponding sending timestamp of s6 is recorded in the fault detection switch detection time register faultDetect_state_time_reg.
[0065] (2.2) When the target data transmission switch receives the detection heartbeat data packet, it reads the switch state structure field in the detection heartbeat data packet, and updates the fault detection switch detection state register faultDetect_state_reg in the target data transmission switch one by one according to the field information, and records the heartbeat detection result between the target data transmission switch and the fault detection switch and the heartbeat detection result between it and the neighboring node. At the same time, it reads the running state of the neighboring node stored in the data transmission switch detection state register data_switch_state_reg, updates it to the detection heartbeat data packet header, modifies the heartbeat protocol type heartbeat_type field to REPLY, as shown in the following formula, and returns the modified detection heartbeat data packet to the fault detection switch.
[0066]
[0067] It should be understood that the switch status structure field in the detection heartbeat data packet shown in the above formula records the heartbeat detection result of the neighbor node, which is the information detected by the fault detection switch. In this way, s6 has both the neighbor node information detected by itself and the information sent by the fault detection switch. Combining these two sets of data facilitates the subsequent judgment of whether the neighbor node is normal and whether the link is normal.
[0068] (2.3) When the fault detection switch receives the reply detection heartbeat data packet, it reads and stores the operating status information of the neighbor node of the data transmission switch carried in the detection heartbeat data packet, and modifies the operating status of the corresponding data transmission switch in the fault detection switch detection state register faultDetect_state_reg, and updates the corresponding timestamp in the fault detection switch detection time register faultDetect_state_time_reg to record the time of this successful detection.
[0069] It should be understood that after the fault detection switch receives the reply detection heartbeat data packet, it indicates that the fault detection switch has successfully interacted with the data transmission switch, so its operating state is determined to be ALIVE, and the operating state of the corresponding data transmission switch in the fault detection switch detection state register faultDetect_state_reg is modified accordingly.
[0070] (3) Determine the fault type based on the voting mechanism: The fault detection switch and the data transmission switch determine the fault detection result of the node according to the node operation status they read; the fault detection switch and the data transmission switch determine whether a fault occurs based on the multi-dimensional fault detection results detected by each of them combined with their respective voting mechanisms, and further determine whether the fault type is a link fault or a switch fault.
[0071] Furthermore, the specific implementation process of judging the fault type based on the voting mechanism is as follows:
[0072] (3.1) Within each heartbeat cycle, the fault detection result of the data detection switch on the neighbor node is read according to the data transmission switch detection state register data_switch_state_reg and its corresponding data transmission switch detection time register data_switch_state_time_reg: if the running state of the node is inactive (NOALIVE), the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the data transmission switch detection time register data_switch_state_time_reg and the current timestamp is less than or equal to the preset first time threshold, the trust state is normal; if the running state of the neighbor node of the node is inactive and the difference between the timestamp in the data transmission switch detection time register data_switch_state_time_reg and the current timestamp is greater than the preset first time threshold, the trust state is faulty.
[0073] Similarly, the fault detection result of the fault detection switch on the neighbor node is read according to the fault detection switch detection state register faultDetect_state_reg and its corresponding fault detection switch detection time register faultDetect_state_time_reg: if the running state of the node is inactive, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the fault detection switch detection time register faultDetect_state_time_reg and the current timestamp is less than or equal to the preset second time threshold, the trust state is normal; if the running state of the neighbor node of the node is inactive and the difference between the timestamp in the fault detection switch detection time register faultDetect_state_time_reg and the current timestamp is greater than the preset second time threshold, the trust state is faulty.
[0074] (3.2) In the voting mechanism for fault detection switches, for any switch v, the fault detection results of the fault detection switch and switch v and its n neighbor nodes are obtained based on step (3.1), and n+1-dimensional fault detection results are obtained; different weights are set for fault detection results from different sources according to their credibility. The detection result value is calculated based on the weight accumulation. If the detection result value is less than or equal to the preset ratio threshold α, it is considered that the switch v has a fault; if the detection result value is greater than the preset ratio threshold α, it is considered that the link between the neighboring switch that has detected the fault and the switch v has a fault.
[0075] Furthermore, factors to consider for credibility include load conditions, network status, etc.
[0076] It should be noted that, for any switch v, the method in step (3.1) is used to read the data transmission switch detection state register data_switch_state_reg and its corresponding data transmission switch detection time register data_switch_state_time_reg, as well as the fault detection switch detection state register faultDetect_state_reg and its corresponding fault detection switch detection time register faultDetect_state_time_reg in the switch v, so as to obtain a fault detection result of the fault detection switch and the switch v, as well as the fault detection results of the fault detection switch and the n neighboring nodes of the switch v, that is, the n+1-dimensional fault detection result.
[0077] For example, in the case of n+1-dimensional fault detection results, if the fault detection switch is more reliable, its detection result can be set to ×3, a neighboring switch is set to ×2, and other switches are set to ×1, then the accumulated detection result value is 6. If the preset ratio threshold α is 3, then only the result of the fault detection switch can be used to determine that there is no fault. If the fault detection switch shows a fault, the normal results of the other three switches are required to determine that the switch is normal.
[0078] (3.3) In the voting mechanism for the data transmission switch to perform detection, for any neighbor node w, based on step (3.1), the fault detection results of the data transmission switch itself and the fault detection switch on the neighbor node w are obtained. If both fault detection results are normal, then the neighbor node w and the connected link status are considered normal; if both fault detection results are faulty, then the neighbor node w is considered to have a fault; if the data transmission switch itself detects that the neighbor node w is a fault and the fault detection result of the fault detection switch on the neighbor node w is normal, then the link between the neighbor node w and the data transmission switch is considered to have a fault.
[0079] It should be noted that step (3.2) and step (3.3) are independent of each other and are performed on the fault detection switch and the data transmission switch respectively. The fault type is determined through their respective voting mechanisms. After step (3.2) and step (3.3), the fault information and its type can be obtained.
[0080] Furthermore, judging the fault type based on the voting mechanism also includes: when the data transmission switch detects that the fault detection switch may have a fault, the data transmission switch will automatically switch to the three-heartbeat mechanism for detection, and obtain the operating status information by continuously sending and confirming three heartbeat signals.
[0081] Specifically, when the data transmission switch interacts with the fault detection switch, it can obtain the detection results of the fault detection switch. If the results of multiple detections are all faults, it is determined that the link between the fault detection switch and the data transmission switch is faulty, and automatically switches to the traditional three-heartbeat mechanism for detection. The data transmission switches exchange their respective detection results of the fault detection switch with each other, set the ratio threshold θ and different weights The data transmission switch calculates the detection result according to the weight. When the result is not higher than θ, the result that the fault detection switch fails is adopted, and all data transmission switches automatically switch to the traditional three-heartbeat mechanism for detection.
[0082] (4) Fault recovery based on fault recovery mechanism: After detecting fault information and its type, the data transmission switch performs fault recovery according to the fault recovery strategy configured in the switch. If there is no corresponding fault recovery strategy, the affected data flow is redirected to the fault detection switch for transmission. After detecting fault information and its type, the fault detection switch promptly sends the fault information to the relevant switches including the upstream multi-hop data transmission switch of the affected data flow path and the data transmission switches on the backup path, so that they can take corresponding measures to recover from the fault.
[0083] Furthermore, the specific implementation process of fault recovery based on the fault recovery mechanism is as follows:
[0084] (4.1) After detecting the fault information and its type, the data transmission switch recovers according to the fault recovery strategy configured in the switch. If there is no corresponding fault recovery strategy, the affected data flow is redirected to the fault detection switch for transmission. After receiving the affected data flow, the fault detection switch selectively transmits the data flow to the multi-hop switch or the destination switch under the original path according to the different fault types of the link and the switch.
[0085] Specifically, after detecting the fault information and its type, the data transmission switch performs recovery according to the fault recovery strategy configured in the switch, such as Figure 1 As shown in Figure 1, when s2 fails, the data flow is forwarded to s5 and quickly sent to s6. If there is no corresponding fault recovery strategy, the affected data flow is redirected to the fault detection switch for transmission, such as Figure 1As shown, when s2 fails, the data flow can be quickly sent to s6 through the fault detection switch h7. After receiving the affected data flow, the fault detection switch can selectively transmit the data flow to the next few switches in the original path or the destination switch according to the different fault types of the link and the switch.
[0086] (4.2) After detecting the fault information and its type, the fault detection switch promptly sends the fault information to relevant switches including the upstream multi-hop data transmission switch of the affected data flow path and the data transmission switch on the backup path, so that they can take corresponding measures to recover from the fault.
[0087] Corresponding to the aforementioned embodiment of the method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information, the present invention also provides an embodiment of an apparatus for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information.
[0088] See also Figure 3 An embodiment of the present invention provides a programmable network fault rapid detection and recovery device based on multi-dimensional heartbeat detection information, including one or more processors and a memory, and the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information in the above embodiment.
[0089] The embodiment of the programmable network fault rapid detection and recovery device based on multi-dimensional heartbeat detection information of the present invention can be applied to any device with data processing capability, and the device with data processing capability can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capability in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From the hardware level, if Figure 3 As shown in the figure, it is a hardware structure diagram of any device with data processing capability where the programmable network fault rapid detection and recovery device based on multi-dimensional heartbeat detection information of the present invention is located. Figure 3 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0090] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0091] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0092] An embodiment of the present invention further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the program implements the programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information in the above embodiment.
[0093] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capability and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0094] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information, characterized in that: The programmable network includes a plurality of data transmission switches, and a fault detection switch is selected from the plurality of data transmission switches. The fault detection switch meets the following conditions: it is directly connected to the data transmission switch under its jurisdiction and the remaining bandwidth is reserved. The fault rapid detection and recovery method includes the following steps: (1) Implementing heartbeat detection between adjacent data transmission switches based on the heartbeat detection protocol: Each data transmission switch maintains a network topology adjacency table, generates a detection heartbeat data packet based on the heartbeat detection protocol to interact with neighboring nodes, and detects each other's heartbeat status in real time to obtain heartbeat detection results. The data transmission switch simultaneously sends its own operating status information and the heartbeat detection results between it and the fault detection switch to the other party. Finally, the heartbeat status of the neighboring node is identified based on whether the returned heartbeat data packet is received within the set time. (2) Implementing heartbeat detection between the fault detection switch and the data transmission switch based on the heartbeat detection protocol: The fault detection switch and the data transmission switch maintain the interaction of heartbeat data packets, wherein the fault detection switch simultaneously sends the operating status information of itself and its neighboring nodes, and the data transmission switch sends the heartbeat detection results between itself and its neighboring nodes to the fault detection switch through the heartbeat data packets; (3) Determine the fault type based on the voting mechanism: The fault detection switch and the data transmission switch determine the fault detection result of the node according to the node operation status they read; the fault detection switch and the data transmission switch determine whether a fault occurs based on the multi-dimensional fault detection results detected by each of them and their respective voting mechanisms, and further determine whether the fault type is a link fault or a switch fault; The specific implementation process of judging the fault type based on the voting mechanism is as follows: In each heartbeat cycle, the fault detection result of the data detection switch on the neighbor node is read according to the data transmission switch detection state register data_switch_state_reg and its corresponding data transmission switch detection time register data_switch_state_time_reg: if the running state of the node is inactive, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the data transmission switch detection time register data_switch_state_time_reg and the current timestamp is less than or equal to the preset first time threshold, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the data transmission switch detection time register data_switch_state_time_reg and the current timestamp is greater than the preset first time threshold, the trust state is faulty; The fault detection result of the fault detection switch on the neighbor node is read according to the fault detection switch detection state register faultDetect_state_reg and its corresponding fault detection switch detection time register faultDetect_state_time_reg: if the running state of the node is inactive, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the fault detection switch detection time register faultDetect_state_time_reg and the current timestamp is less than or equal to the preset second time threshold, the trust state is normal; if the running state of the node is inactive and the difference between the timestamp in the fault detection switch detection time register faultDetect_state_time_reg and the current timestamp is greater than the preset second time threshold, the trust state is faulty; In the voting mechanism for fault detection switches, for any switch, the fault detection results of the fault detection switch and the switch and its n neighboring nodes are obtained to obtain n+1-dimensional fault detection results; different weights are set for fault detection results from different sources according to their credibility. The detection result value is calculated based on the weight accumulation. If the detection result value is less than or equal to the preset ratio threshold α, it is considered that the switch is faulty; if the detection result value is greater than the preset ratio threshold α, it is considered that the link between the neighboring switch that has detected the fault and the switch is faulty; In the voting mechanism for the data transmission switch to perform detection, for any neighbor node, the fault detection results of the data transmission switch itself and the fault detection switch on the neighbor node are obtained. If both fault detection results are normal, the neighbor node and the connected link status are considered normal; if both fault detection results are faulty, the neighbor node is considered to be faulty; if the data transmission switch itself detects the neighbor node as faulty and the fault detection result of the fault detection switch on the neighbor node is normal, the link between the neighbor node and the data transmission switch is considered to be faulty; (4) Fault recovery based on fault recovery mechanism: After detecting fault information and its type, the data transmission switch performs fault recovery according to the fault recovery strategy configured in the switch. If there is no corresponding fault recovery strategy, the affected data flow is redirected to the fault detection switch for transmission. After detecting fault information and its type, the fault detection switch promptly sends the fault information to the relevant switches including the upstream multi-hop data transmission switch of the affected data flow path and the data transmission switches on the backup path, so that corresponding measures can be taken to perform fault recovery.
2. The method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information according to claim 1, characterized in that: The heartbeat detection protocol uses a standard Ethernet header. The heartbeat detection protocol includes a heartbeat protocol type heartbeat_type, a signature of the sender of the heartbeat data packet, an identifier of the detected party Monitored_switch, a next protocol protoc_type, and n groups of switch state structure information, where n is the number of neighbor nodes corresponding to the detected party of the heartbeat data packet, and the switch state structure includes a switch label switch_id_i, a corresponding switch state switch_state_i, and whether it is the last bit is_last.
3. The method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information according to claim 1, characterized in that: Each of the data transmission switches also maintains a series of data transmission switch detection status registers data_switch_state_reg and corresponding data transmission switch detection time registers data_switch_state_time_reg, which are respectively used to record the node operation status detected by the data transmission switch and the last detection timestamp; at the same time, it also maintains a fault detection switch detection status register faultDetect_state_reg and a corresponding fault detection switch detection time register faultDetect_state_time_reg, which are used to record the node operation status detected by the fault detection switch and the last detection timestamp.
4. The method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information according to claim 1, characterized in that: The specific implementation process of the heartbeat detection between adjacent data transmission switches is as follows: The data transmission switch periodically sends a detection heartbeat data packet to its neighbor node, which carries its own operation status information and the heartbeat detection result between it and the fault detection switch. When the detection heartbeat data packet is sent, the operation status of the neighbor node in the data transmission switch detection status register data_switch_state_reg is updated to inactive, and the timestamp of the corresponding neighbor node in the data transmission switch detection time register data_switch_state_time_reg is updated; When the neighbor node of the data transmission switch receives the detection heartbeat data packet, it updates its own data transmission switch detection state register data_switch_state_reg according to the heartbeat data packet sender signature signature field and the switch state structure in the detection heartbeat data packet, and clears the value in the switch state structure field in the detection heartbeat data packet, then adds the heartbeat detection result between itself and the fault detection switch, and modifies the heartbeat protocol type heartbeat_type field in the detection heartbeat data packet to REPLY, and returns it to the sending node; After receiving the returned detection heartbeat data packet, the sending node updates the state of the corresponding neighbor node in the data transmission switch detection state register data_switch_state_reg to alive, and updates the timestamp of the corresponding neighbor node in the data transmission switch detection time register data_switch_state_time_reg to realize heartbeat detection between adjacent data transmission switches.
5. The method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information according to claim 1, characterized in that: The specific implementation process of the heartbeat detection between the fault detection switch and the data transmission switch is as follows: The fault detection switch sends detection heartbeat data packets to the data transmission switches within its jurisdiction one by one, reads the operating status of all neighboring nodes of the current target data transmission switch from the fault detection switch detection state register faultDetect_state_reg one by one, forms a switch state structure corresponding to all neighboring nodes, adds it to the detection heartbeat data packet, and sends the detection heartbeat data packet to the current target data transmission switch; at the same time, the fault detection switch updates the operating status of the data transmission switch in the fault detection switch detection state register faultDetect_state_reg to inactive, and records the corresponding sending timestamp in the fault detection switch detection time register faultDetect_state_time_reg; When the target data transmission switch receives the detection heartbeat data packet, it reads the switch state structure field in the detection heartbeat data packet, and updates the fault detection switch detection state register faultDetect_state_reg in the target data transmission switch one by one according to the field information, records the heartbeat detection result between it and the fault detection switch and the heartbeat detection result between the transmitted fault detection switch and the neighboring node; at the same time, it reads the running state of the neighboring node stored in the data transmission switch detection state register data_switch_state_reg, updates it to the detection heartbeat data packet header, modifies the heartbeat protocol type heartbeat_type field to REPLY, and returns the modified detection heartbeat data packet to the fault detection switch; When the fault detection switch receives the reply detection heartbeat data packet, it reads and stores the operating status information of the neighbor node of the data transmission switch carried in the detection heartbeat data packet, modifies the operating status of the corresponding data transmission switch in the fault detection switch detection state register faultDetect_state_reg, and updates the corresponding timestamp in the fault detection switch detection time register faultDetect_state_time_reg.
6. The method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information according to claim 1, characterized in that: The specific implementation process of judging the fault type based on the voting mechanism also includes: when the data transmission switch detects that the fault detection switch may have a fault, the data transmission switch will automatically switch to the three-heartbeat mechanism for detection, and obtain the operation status information by continuously sending and confirming three heartbeat signals.
7. The method for rapid detection and recovery of programmable network faults based on multi-dimensional heartbeat detection information according to claim 1, characterized in that: The specific implementation process of fault recovery based on the fault recovery mechanism is as follows: After detecting the fault information and its type, the data transmission switch recovers according to the fault recovery strategy configured in the switch. If there is no corresponding fault recovery strategy, the affected data flow is redirected to the fault detection switch for transmission. After receiving the affected data flow, the fault detection switch selectively transmits the data flow to the multi-hop switch or the destination switch under the original path according to the different fault types of the link and the switch; After detecting the fault information and its type, the fault detection switch promptly sends the fault information to relevant switches including the upstream multi-hop data transmission switch of the affected data flow path and the data transmission switch on the backup path, so that corresponding measures can be taken to recover from the fault.
8. A programmable network fault rapid detection and recovery device based on multi-dimensional heartbeat detection information, comprising one or more processors and a memory, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, it is used to implement the programmable network fault rapid detection and recovery method based on multi-dimensional heartbeat detection information as described in any one of claims 1-7.
Citation Information
Patent Citations
Method for detecting single-channel fault of ring-type network
CN101043383A
BFD detection device and method
CN103916275A