A network sub-health state detection method and device, electronic equipment and storage medium

By sending probe information to the probe nodes in the cluster and analyzing the response information, and by comprehensively considering the network status of the probe nodes, the problem of low accuracy in detecting sub-healthy network conditions in distributed systems is solved, and more accurate network status judgment is achieved.

CN118659993BActive Publication Date: 2025-11-21JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410864609.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-11-21
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of nodes in distributed systems in detecting sub-healthy network conditions is low, making it difficult to accurately determine whether the network in which their own nodes are located is in a sub-healthy state.

Method used

By sending probe information to probe nodes in the cluster, including the network status data of the node itself, and receiving response information from the probe nodes, which includes enhanced probe information and the network status data of the probe nodes, the network status of the node itself is determined by comprehensively analyzing the network situation of the probe nodes.

Benefits of technology

It improves the accuracy of detecting sub-healthy network conditions, enabling more accurate judgment of the network status of its own nodes and reducing false alarms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118659993B_ABST
    Figure CN118659993B_ABST
Patent Text Reader

Abstract

The application discloses a network sub-health state detection method and device, electronic equipment and a storage medium, and is applied to the technical field of computers, and aims at solving the problem that the network sub-health state detection of nodes is inaccurate in related technologies. The method is applied to any node in a cluster, and comprises the following steps: sending detection information to a detection node in the cluster; the detection information comprises network state data of the node itself, and the detection node is different from the node itself; receiving reply information returned by the detection node; the reply information comprises enhanced detection information and the detection information, and the enhanced detection information comprises network state data of the detection node; and determining whether the network state of the node itself is a network sub-health state according to the reply information of the detection node. In the detection of the network sub-health state, not only the network condition of the node itself is considered, but also the network condition of the detection node, so that the accuracy of the network sub-health state detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and computer-readable storage medium for detecting sub-health status of networks. Background Technology

[0002] The operating state of the network in a cluster, such as a distributed system, can be divided into three types: "healthy", "unhealthy" and "sub-healthy". When the network in which a node is located is in a sub-healthy state, it usually leads to the distributed system entering a state of low performance. Therefore, it is necessary to detect the sub-healthy state of the network.

[0003] In related technologies, when detecting the sub-healthy state of a network, the local node in a distributed system usually judges whether the network it is in is in a sub-healthy state based on the information of the probe packets it sends. However, it is difficult to accurately detect whether the network it is in is in a sub-healthy state based solely on the probe packets of its own node.

[0004] Therefore, how to improve the accuracy of detecting sub-health status on the internet is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a method, apparatus, electronic device, and computer-readable storage medium for detecting sub-health status of networks, which can improve the accuracy of detecting sub-health status of networks.

[0006] To address the aforementioned technical problems, the embodiments of the present invention provide the following technical solutions:

[0007] One embodiment of the present invention provides a method for detecting network sub-health status, applicable to any node in a cluster, comprising:

[0008] Send probe information to probe nodes in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself.

[0009] Receive the response information returned by the probe node; the response information includes enhanced probe information and the probe information, the enhanced probe information including the network status data of the probe node;

[0010] Based on the response information from the probe node, determine whether the network status of your own node is in a sub-healthy state.

[0011] In some embodiments, there are multiple detection nodes;

[0012] The step of determining whether the network status of its own node is in a sub-healthy state based on the response information from the probe node includes:

[0013] For each of the probe nodes, the corresponding network packet loss rate and network latency are calculated based on the response information;

[0014] If at least one of the network packet loss rates is greater than the packet loss rate threshold, or if at least one of the network delays is greater than the delay threshold, for each of the probe nodes, the corresponding network detection result is determined based on the response information returned by the probe node.

[0015] Based on the network detection results corresponding to each probe node, determine whether the network status of its own node is in a sub-healthy state.

[0016] In some embodiments, determining whether the network status of a node is in a sub-healthy state based on the network detection results corresponding to each probe node includes:

[0017] If at least one of the network detection results includes a case where the node itself is abnormal, then the network status of the node is determined to be a sub-healthy state.

[0018] In some embodiments, the network status data in the detection information includes at least one of a first transmission error statistics value, a first transmission packet loss statistics value, a first transmission timestamp, a first reception error statistics value, a first reception packet loss statistics value, and a first reception timestamp;

[0019] The network status data in the enhanced detection information includes at least one of the following: second transmission error statistics, second transmission packet loss statistics, second transmission timestamp, second reception error statistics, second reception packet loss statistics, and second reception timestamp.

[0020] In some embodiments, when at least one of the network packet loss rates is greater than a packet loss rate threshold among all network packet loss rates or at least one of the network delays is greater than a delay threshold among all network delays, for each probe node, determining the corresponding network detection result based on the response information returned by the probe node includes:

[0021] If at least one of the network packet loss rates is greater than the packet loss rate threshold among all network packet loss rates, for each of the probe nodes, the network detection result is determined based on the first transmission error statistics, the first transmission packet loss statistics, the first reception error statistics, the first reception packet loss statistics, the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics in the response information.

[0022] Alternatively, if at least one of the network delays is greater than the delay threshold among all network delays, for each of the probe nodes, the network detection result is determined based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information.

[0023] In some embodiments, determining the network detection result based on the first transmission error statistics, first transmission packet loss statistics, first reception error statistics, first reception packet loss statistics, second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the response information includes:

[0024] Based on the first transmission error statistics and the first transmission packet loss statistics, and combined with the first transmission error statistics and the first transmission packet loss statistics in the probe data sent to the probe node in the last time, it is determined whether the increment of the first transmission error statistics or the increment of the first transmission packet loss statistics is greater than zero.

[0025] If the increment of the first transmission error statistics value or the increment of the first transmission packet loss statistics value is greater than zero, then the network detection result is determined to be an anomaly of its own node.

[0026] If neither the increment of the first transmission error statistics nor the increment of the first transmission packet loss statistics is greater than zero, then based on the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics, combined with the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics in the previously received reply data returned by the probe node, it is determined whether the increment of the second transmission error statistics or the increment of the second transmission packet loss statistics or the increment of the second reception error statistics or the increment of the second reception packet loss statistics is greater than zero.

[0027] If the increment of the second transmission error statistics value, the increment of the second transmission packet loss statistics value, the increment of the second reception error statistics value, or the increment of the second reception packet loss statistics value is greater than zero, then the network detection result is determined to be that the probe node is abnormal.

[0028] If the increment of the second transmission error statistics value, the increment of the second transmission packet loss statistics value, the increment of the second reception error statistics value, and the increment of the second reception packet loss statistics value are all not greater than zero, then based on the first reception error statistics value and the first reception packet loss statistics value, combined with the first reception error statistics value and the first reception packet loss statistics value in the previous detection data sent to the detection node, it is determined whether the increment of the first reception error statistics value or the increment of the first reception packet loss statistics value is greater than zero.

[0029] If the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, then the network detection result is determined to be an anomaly of its own node.

[0030] If neither the increment of the first received error statistics nor the increment of the first received packet loss statistics is greater than zero, then the network detection result is determined to be that neither the node itself nor the probe node has been identified as abnormal.

[0031] In some embodiments, the detection information further includes a first system load value; the method further includes:

[0032] If the increment of the first transmission error statistics value or the increment of the transmission packet loss statistics value is greater than zero, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the network of its own node has a transmission failure.

[0033] Alternatively, if the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the network of its own node has a receiving fault.

[0034] In some embodiments, determining the network detection result based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information includes:

[0035] Determine whether the first difference between the second received timestamp and the first sent timestamp exceeds a first time threshold.

[0036] If the first difference exceeds the first time threshold, the network detection result is determined to be an anomaly of its own node;

[0037] If the first difference does not exceed the first time threshold, determine whether the second difference between the second sending timestamp and the second receiving timestamp is greater than the second time threshold.

[0038] If the second difference is greater than the second time threshold, then the network detection result is determined to be that the probe node is abnormal;

[0039] If the second difference is not greater than the second time threshold, then determine whether the third difference between the first receiving timestamp and the second sending timestamp is greater than the third time threshold.

[0040] If the third difference is greater than the third time threshold, then the network detection result is determined to be an anomaly of its own node;

[0041] If the third difference is not greater than the third time threshold, then the network detection result is determined to be that the network itself and the probe node have not been identified as abnormal.

[0042] In some embodiments, the detection information further includes a first system load value; the method further includes:

[0043] If the first difference exceeds the first time threshold, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the latency from its own node to the probe node is abnormal.

[0044] Alternatively, if the third difference is greater than the third time threshold, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the latency from the probe node to its own node is abnormal.

[0045] In some embodiments, before sending probe information to the probe nodes in the cluster, the method further includes:

[0046] Multiple probe nodes corresponding to the node itself are determined from the cluster using a skip selection method.

[0047] In some embodiments, determining multiple probe nodes corresponding to the node itself from the cluster includes:

[0048] Number all nodes in the cluster according to the order of their Internet Protocol (IP) addresses;

[0049] The node located immediately following its own node is designated as the first probe node;

[0050] The number of the second probe node is determined based on the number of the first probe node and the total number of nodes in the cluster, combined with the first calculation formula.

[0051] The number of the third probe node is determined based on the number of the first probe node and the total number of nodes in the cluster, combined with the second calculation formula; the number of the third probe node is greater than the number of the second node.

[0052] The second probe node is determined from the cluster based on its number, and the third probe node is determined from the cluster based on its number.

[0053] In some embodiments, the first calculation formula is: n2 = M + INT(N / 3);

[0054] The second calculation formula is: n3 = M + 2 * INT(N / 3);

[0055] Where n2 represents the number of the second probe node, n3 represents the number of the third probe node, M represents the number of the first probe node, N represents the total number of nodes in the cluster, and INT() represents the floor function.

[0056] Another embodiment of the present invention provides a network sub-health state detection device, applied to any node in a cluster, comprising:

[0057] A sending module is used to send probe information to probe nodes in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself.

[0058] A receiving module is configured to receive response information returned by the probe node; the response information includes enhanced probe information and the probe information, wherein the enhanced probe information includes network status data of the probe node;

[0059] The detection module is used to determine whether the network status of its own node is in a sub-healthy state based on the response information from the detection node.

[0060] Another aspect of the present invention provides an electronic device, comprising:

[0061] Memory, used to store computer programs;

[0062] A processor is used to execute the computer program to implement the steps of the network sub-health state detection method as described above.

[0063] Another embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the network sub-health state detection method described above.

[0064] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows:

[0065] This invention provides a method for detecting network sub-health status, applicable to any node in a cluster, comprising: sending probe information to a probe node in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself; receiving response information returned by the probe node; the response information includes enhanced probe information and probe information, the enhanced probe information including the network status data of the probe node; and determining whether the node itself is in a network sub-health state based on the response information of the probe node.

[0066] Therefore, in this invention, when a node detects whether its network is in a sub-healthy state, it sends probe information containing its own network status data to probe nodes in the cluster that are different from itself. After receiving the probe information, the probe nodes return the probe information and enhanced probe information carrying their own network status data. After receiving the response information returned by each probe node, the node determines whether its own network is in a sub-healthy state based on its own probe information and the enhanced probe information of each probe node. In this embodiment of the invention, the detection of sub-healthy network status not only considers the network status of the node itself, but also the network status of the probe nodes, thus more accurately determining the sub-healthy state of the node and improving the accuracy of network sub-healthy state detection. Attached Figure Description

[0067] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0068] Figure 1 A flowchart illustrating a method for detecting sub-health status of a network provided in an embodiment of the present invention. Figure 2

[0069] Figure 2 A flowchart illustrating another method for detecting sub-health status of networks provided in this embodiment of the invention. Figure 2

[0070] Figure 3 This invention provides a schematic diagram of information interaction between nodes and probe nodes in a cluster. Figure 2

[0071] Figure 4 A schematic diagram of network detection results provided in an embodiment of the present invention. Figure 2

[0072] Figure 5 This invention provides a schematic diagram of a network detection process when the network packet loss rate exceeds a packet loss rate threshold. Figure 2

[0073] Figure 6 This invention provides another network detection process for cases where the network packet loss rate exceeds a packet loss rate threshold. Figure 2

[0074] Figure 7This invention provides a schematic diagram of a network detection process when network latency exceeds a latency threshold. Figure 2

[0075] Figure 8 This invention provides another network detection process for cases where network latency exceeds a latency threshold, as illustrated in this embodiment. Figure 2

[0076] Figure 9 A schematic diagram of the structure of a network sub-health state detection device provided in an embodiment of the present invention. Figure 2

[0077] Figure 10 This is a schematic diagram of a network sub-health state detection device provided in an embodiment of the present invention. Detailed Implementation

[0078] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0079] The terms "comprising" and "having," and any variations thereof, in the specification and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may include steps or units not listed.

[0080] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0081] The operational state of a network in a cluster, such as a distributed system, can be categorized into three types: "healthy," "unhealthy," and "sub-healthy." A "healthy" state refers to a network that operates normally and recovers quickly after external shocks; a "sub-healthy" state refers to a state where the network is paralyzed and unable to operate normally; and a "sub-healthy" state refers to a state where the network normally operates but has extremely low resilience to risks. When a node's network is in a sub-healthy state, it typically leads to poor performance in the distributed system, thus requiring detection of this sub-healthy state. This invention addresses the problem of low accuracy in current network sub-health detection methods by proposing a network sub-health detection method that improves detection accuracy.

[0082] Next, we will describe in detail a method for detecting network sub-health status provided by an embodiment of the present invention. Figure 1 This is a flowchart illustrating a method for detecting network sub-health status according to an embodiment of the present invention. This method is applied to any node in a cluster. In practical applications, any node can execute the method provided in this embodiment. The node executing this method can be referred to as its own node. This embodiment describes the technical solution in detail from the perspective of its own node. The method includes:

[0083] S 110: Send probe information to the probe nodes in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself;

[0084] It should be noted that when any node in the cluster needs to detect the sub-health status of its network, it can send probe information to probe nodes in the cluster that are different from itself. This probe information can include the node's own network status data, which is data information that reflects the node's network status. Specifically, probe nodes targeting that particular node can be pre-selected from the cluster.

[0085] S 120: Receive the response information returned by the probe node; the response information includes enhanced probe information and probe information, the enhanced probe information including the network status data of the probe node;

[0086] Specifically, after receiving network status data, the probe node will return a response message. The response message includes the probe information received by the probe node and the probe node's enhanced probe information, which includes the probe node's network status data.

[0087] S130: Based on the response information from the probe node, determine whether the network status of its own node is in a sub-healthy state.

[0088] Specifically, after receiving the response information from the probe node, the node can further judge whether the network status of the network in which the node is located is in a sub-healthy state based on the probe information related to the node and the enhanced probe information related to the probe node in the response information.

[0089] Therefore, in this invention, when a node detects whether its network is in a sub-healthy state, it sends probe information containing its own network status data to probe nodes in the cluster that are different from itself. After receiving the probe information, the probe nodes return the probe information and enhanced probe information carrying their own network status data. After receiving the response information returned by each probe node, the node determines whether its own network is in a sub-healthy state based on its own probe information and the enhanced probe information of each probe node. In this embodiment of the invention, the detection of sub-healthy network status not only considers the network status of the node itself, but also the network status of the probe nodes, thus more accurately determining the sub-healthy state of the node and improving the accuracy of network sub-healthy state detection.

[0090] Based on the above embodiments, the present invention further explains and introduces the technical solutions as follows:

[0091] In this invention, there can be multiple detection nodes. Please refer to... Figure 2 , Figure 2 This is a flowchart illustrating another method for detecting sub-health status of a network provided in an embodiment of the present invention. The method includes:

[0092] S210: Send probe information to the probe nodes in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself.

[0093] It should be noted that in this embodiment of the invention, the node itself sends probe information to multiple probe nodes in the cluster. Specifically, it can send probe information to each probe node once at a preset time interval, wherein each of the multiple probe nodes is different from the node itself.

[0094] like Figure 3 As shown, the nodes in the cluster are A, B, C, D, and E. Taking node A as an example, node A's probe nodes are B, C, and D. Therefore, node A sends probe information to nodes B, C, and D respectively. Taking node B as an example, node B's probe nodes are C, D, and E. Therefore, node B sends probe information to nodes C, D, and E respectively.

[0095] S220: Receive the response information returned by the probe node; the response information includes enhanced probe information and probe information, the enhanced probe information including the network status data of the probe node;

[0096] Specifically, in this embodiment of the invention, each detection node returns corresponding response information after receiving detection information. The response information includes the detection information received by the detection node and the enhanced detection information of the detection node.

[0097] S230: For each probe node, calculate the corresponding network packet loss rate and network latency based on the response information;

[0098] It should be noted that each probe node returns a response after receiving the probe information. In this embodiment of the invention, to further improve detection efficiency, for each probe node, the node itself can calculate the corresponding network packet loss rate and network latency based on the response information returned by the probe node, thus obtaining a set of network packet loss rate and network latency for each probe node. Specifically, the network latency can be calculated based on the sending timestamp of the probe information in the returned response information and the receiving timestamp when the response information is received; if no response information is received within a preset time period, it is considered that the information is lost. Therefore, the network packet loss rate and network latency can be calculated by sending probe information to the probe node multiple times and receiving response information multiple times.

[0099] For example, if node A is itself, and the probe nodes are B, C, and D, A can calculate a set of network packet loss rate and network latency based on the response information returned by B, A can calculate a set of network packet loss rate and network latency based on the response information returned by C, and A can calculate a set of network packet loss rate and network latency based on the response information returned by D.

[0100] S240: If at least one network packet loss rate is greater than the packet loss rate threshold among all network packet loss rates or at least one network latency is greater than the latency threshold among all network latency, for each probe node, the corresponding network detection result shall be determined based on the response information returned by the probe node.

[0101] It should be noted that a node's network is in a sub-healthy state, which can be divided into three reasons based on the physical connection between nodes. One is that the node itself has a network sub-healthy fault, another is that the link has a network sub-healthy fault, and the third is that the related probe node has a network sub-healthy fault.

[0102] Therefore, in order to further rule out that the sub-healthy state of a node is caused by a network sub-healthy fault in the link or the network sub-healthy state of the related probe node, in the above-calculated network packet loss rate for each probe node, if at least one network packet loss rate is greater than the packet loss rate threshold or at least one network delay is greater than the delay threshold, for each probe node, the corresponding network detection result is determined based on the response information returned by the probe node. Specifically, the network status data included in the probe information and the network status data included in the enhanced probe information in the response information can be analyzed to obtain the network detection result corresponding to the node itself and the probe node.

[0103] Understandably, through the above calculations, a set of network packet loss rates and network latency will be obtained for each probe node. If there are m probe nodes, then m network packet loss rates and m network latencies can be obtained. If at least one of the m network packet loss rates is greater than a packet loss rate threshold, or at least one of the m network latencies is greater than a latency threshold, then the network in which the node is located can be considered to have a certain risk. At this point, further network detection analysis can be performed on the first probe node based on the response information returned by that probe node to determine the network detection results corresponding to the node and the first probe node; ...; for the i-th probe node, network detection analysis can be performed based on the response information returned by that probe node to determine the network detection results corresponding to the node and the i-th probe node; ...; for the m-th probe node, network detection analysis can be performed based on the response information returned by that probe node to determine the network detection results corresponding to the node and the m-th probe node. Therefore, m network detection results can be obtained.

[0104] S250: Based on the network detection results corresponding to each probe node, determine whether the network status of its own node is in a sub-healthy state.

[0105] It should be noted that, in this embodiment of the invention, after obtaining the network detection results corresponding to each probe node, all network detection results can be combined to determine whether the network status of the node itself is in a sub-healthy state. By comprehensively considering the network status of the probe nodes, it is possible to accurately detect whether the network status of the node itself is in a sub-healthy state.

[0106] In some embodiments, the process of determining whether the network state of a node is in a sub-healthy state based on the network detection results corresponding to each probe node in S250 above may include:

[0107] If at least one network detection result in the case of an abnormality of its own node, the network status of its own node is determined to be a sub-healthy state.

[0108] It should be noted that, as the above analysis shows, for each probe node, a corresponding network detection result is obtained. This network detection result is obtained by analyzing the network status data of the node itself and the network status data of the probe node. The network detection result may include the node itself being abnormal, or the probe node being abnormal, or the node itself and the corresponding probe node not being identified as being abnormal.

[0109] In this embodiment of the invention, if at least one of the obtained network detection results includes an abnormality of its own node, then the network status of its own node can be determined to be a sub-healthy network state.

[0110] It's understandable that the node itself and the probe nodes are grouped together. For example, if the node itself is A, and the probe nodes are B, C, and D, then AB is one group, AC is another, and AD is yet another. Figure 4 As shown, for each group of nodes, the network detection results can be divided into three categories: one is that the node itself is abnormal, one is that the probe node is abnormal, and one is that neither the node itself nor the probe node is abnormal. For example, for group A and B, if the node itself is abnormal, it is A abnormal; if the probe node is abnormal, it is B abnormal; and if neither the node itself nor the probe node is abnormal, it is neither A nor B.

[0111] If the network detection results for group AB are A abnormal, the network detection results for group AC are A abnormal, and the network detection results for group AD are A abnormal, then it can be determined that the network status of node A is a sub-healthy state.

[0112] In some embodiments, the method may further include:

[0113] If all network detection results indicate that neither the node itself nor its corresponding probe node is abnormal, then the network link corresponding to the node itself is determined to be abnormal.

[0114] That is, if the network detection results for each probe node are all that the node itself was not identified and that the corresponding probe node is abnormal, such as Figure 4 As shown, if the network detection results of group AB show that no abnormalities were detected in its own node and the corresponding probe node, the network detection results of group AC show that no abnormalities were detected in its own node and the corresponding probe node, and the network detection results of group AD show that no abnormalities were detected in its own node and the corresponding probe node, then it can be determined that the network link corresponding to its own node A is abnormal. That is, the network abnormality is caused by the network link abnormality, and it is not that the network of node A is in a sub-healthy state.

[0115] To avoid impacting business operations, after determining that the network link corresponding to a node is abnormal, a failover can be performed on the network port corresponding to the node. For example, if the node has only one network port, then that network port should be taken down; if the node has multiple network ports, a backup network port can be identified from the other network ports, and the node's network port can be switched to the backup network port so that the node's network can return to normal.

[0116] In some embodiments, in order to more accurately detect whether the network where the node is located is in a sub-healthy state, the network status data in the detection information sent by each detection node may include at least one of the following: first transmission error statistics, first transmission packet loss statistics, first transmission timestamp, first reception error statistics, first reception packet loss statistics, and first reception timestamp.

[0117] The network status data in the enhanced probe information returned by the probe node may include at least one of the following: second transmission error statistics, second transmission packet loss statistics, second transmission timestamp, second reception error statistics, second reception packet loss statistics, and second reception timestamp.

[0118] like Figure 3 As shown in the example, in this embodiment of the invention, the self node is taken as node A and the probe node is node B. The network status data in the probe information sent by node A may include at least one of the following: the sending error statistics value corresponding to node A's tx_error, the sending packet loss statistics value corresponding to node A's tx_drop, the sending timestamp T1 corresponding to node A's sending timestamp, the receiving error statistics value corresponding to node A's rx_error, the receiving packet loss statistics value corresponding to node A's rx_drop, and the receiving timestamp T4 corresponding to node A's receiving timestamp.

[0119] The network status data in the enhanced probe information of node B may include at least one of the following: the transmission error statistics corresponding to node B's tx_error, the transmission packet loss statistics corresponding to node B's tx_drop, the transmission timestamp T3 corresponding to node B's transmission timestamp, the reception error statistics corresponding to node B's rx_error, the reception packet loss statistics corresponding to node B's rx_drop, and the reception timestamp T2 corresponding to node B's reception timestamp.

[0120] It should be noted that, in order to further distinguish the network status data of the self node and the probe node in this embodiment of the invention, the parameters in the network status data of the self node are respectively referred to as the first transmission error statistics, the first transmission packet loss statistics, the first transmission timestamp, the first reception error statistics, the first reception packet loss statistics, and the first reception timestamp; and the parameters in the network status data of the probe node are respectively referred to as the second transmission error statistics, the second transmission packet loss statistics, the second transmission timestamp, the second reception error statistics, the second reception packet loss statistics, and the second reception timestamp.

[0121] In this configuration, the first sending timestamp (T1) corresponding to the node itself is the timestamp for sending probe information to the probe node; the second receiving timestamp (T2) corresponding to the probe node is the timestamp for the probe node to receive the probe information; the second sending timestamp (T3) corresponding to the probe node is the timestamp for the probe node to return reply information; and the first receiving timestamp (T4) corresponding to the node itself is the timestamp for the node itself to receive the reply information. In some embodiments, the network status data in the probe information sent by the probe node may include each of the following: a first sending error statistics value, a first sending packet loss statistics value, a first sending timestamp, a first receiving error statistics value, a first receiving packet loss statistics value, and a first receiving timestamp.

[0122] The network status data in the enhanced probe information returned by the probe node may include each of the following: second transmission error statistics, second transmission packet loss statistics, second transmission timestamp, second reception error statistics, second reception packet loss statistics, and second reception timestamp.

[0123] Accordingly, in S240 above, if at least one network packet loss rate is greater than a packet loss rate threshold among all network packet loss rates, or if at least one network delay is greater than a delay threshold among all network delays, the process of determining the corresponding network detection result for each probe node based on the response information returned by the probe node may include:

[0124] If at least one network packet loss rate exceeds the packet loss rate threshold among all network packet loss rates, for each probe node, the network detection result is determined based on the first transmission error statistics, first transmission packet loss statistics, first reception error statistics, first reception packet loss statistics, second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the response information.

[0125] Alternatively, if at least one network delay exceeds a delay threshold among all network delays, for each probe node, the network detection result is determined based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information.

[0126] It should be noted that when at least one network packet loss rate exceeds the packet loss rate threshold, the first transmission error statistics, first transmission packet loss statistics, first reception error statistics, and first reception packet loss statistics can reflect the network transmission status of the node itself, while the second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics of the probe node can reflect the network transmission status of the probe node. Therefore, in this embodiment of the invention, when at least one network packet loss rate exceeds the packet loss rate threshold, for each probe node, based on the first transmission error statistics, first transmission packet loss statistics, first reception error statistics, first reception packet loss statistics, second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the response information returned by the probe node, the network detection result of each group of nodes can be determined more accurately.

[0127] When at least one network delay exceeds a delay threshold among all network delays, the first sending timestamp and the first receiving timestamp reflect the network delay status of the node itself, while the second sending timestamp and the second receiving timestamp reflect the network delay status of the probe node. Therefore, in this embodiment of the invention, when at least one network delay exceeds a delay threshold among all network delays, for each probe node, the network detection result of each group of nodes can be determined more accurately based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information.

[0128] In some embodiments, please refer to Figure 5 If at least one network packet loss rate exceeds a certain threshold, further measures can be taken for each probe node, such as... Figure 5 The method shown enhances packet loss detection to accurately determine network detection results. In this embodiment, node A is used as the node itself, and node B is used as the probe node for detailed explanation.

[0129] The process of determining the network detection result based on the first transmission error statistics, first transmission packet loss statistics, first reception error statistics, first reception packet loss statistics, second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the response information may include:

[0130] S 501: Based on the first transmission error statistics and the first transmission packet loss statistics, and combined with the first transmission error statistics and the first transmission packet loss statistics in the probe data sent to the probe node last time, calculate the increment of the first transmission error statistics and the increment of the first transmission packet loss statistics.

[0131] It should be noted that, in this embodiment of the invention, the increment of the first transmission error statistics (i.e., Atx_error increment) and the increment of the first transmission packet loss statistics (i.e., Atx_drop increment) can be calculated based on the first transmission error statistics and the first transmission packet loss statistics in the previous probe data sent to the probe node, and the first transmission error statistics and the first transmission packet loss statistics in the probe information sent to the probe node this time.

[0132] S 502: Determine whether the increment of the first transmission error statistics value or the increment of the first transmission packet loss statistics value is greater than zero; if the increment of the first transmission error statistics value or the increment of the first transmission packet loss statistics value is greater than zero, proceed to S503; if neither the increment of the first transmission error statistics value nor the increment of the first transmission packet loss statistics value is greater than zero, proceed to S504.

[0133] It should be noted that, in the implementation of this invention, after obtaining the first transmission error statistical value increment (i.e., Atx_error increment) and the first transmission packet loss statistical value increment (i.e., Atx_drop increment), it can be further determined whether the Atx_error increment or the Atx_drop increment is greater than 0. If at least one of them is greater than 0, then proceed to S503; if neither increment is greater than 0, then proceed to S504.

[0134] S503: The network detection result indicates that the node itself is abnormal;

[0135] Understandably, in practical applications, if at least one of the increments of A_tx_error or A_tx_drop of node A is greater than 0, it indicates that there may be an anomaly in the network of node A. In this case, it can be determined that the network detection result of node A for probe node B is that node A is abnormal.

[0136] S504: Based on the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics, and combined with the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics in the reply data returned by the probe node in the previous reception, calculate the increment of the second transmission error statistics, the increment of the second transmission packet loss statistics, the increment of the second reception error statistics, and the increment of the second reception packet loss statistics.

[0137] It should be noted that if both the increments of A_tx_error and A_tx_drop are not greater than 0, it can be determined that the transmission of node A itself is not abnormal, but it cannot be determined whether node A itself has other abnormalities. Therefore, in order to further improve the detection accuracy, the increments of the second transmission error statistics (B_tx_error increment), the second transmission packet loss statistics (B_tx_drop increment), the second reception error statistics (Br_error increment), and the second reception packet loss statistics (Br_drop increment) corresponding to node B can be calculated based on the second transmission error statistics (B_tx_error increment), the second transmission packet loss statistics (B_drop increment), the second reception error statistics (Br_error increment), and the second reception packet loss statistics (Br_drop increment) in the reply data sent by node B in the current reception data.

[0138] S505: Determine whether the increment of the second transmission error statistic, the increment of the second transmission packet loss statistic, the increment of the second reception error statistic, or the increment of the second reception packet loss statistic is greater than zero; if the increment of the second transmission error statistic, the increment of the second transmission packet loss statistic, the increment of the second reception error statistic, or the increment of the second reception packet loss statistic is greater than zero, proceed to S506; if the increments of the second transmission error statistic, the increment of the second transmission packet loss statistic, the increment of the second reception error statistic, and the increment of the second reception packet loss statistic are all not greater than zero, proceed to S507;

[0139] It should be noted that, in order to more accurately determine the network status of probe node B in this embodiment of the invention, after calculating the second transmission error statistical value increment (B tx_error increment), the second transmission packet loss statistical value increment (B tx_drop increment), the second reception error statistical value increment (Br_error increment), and the second reception packet loss statistical value increment (Br_drop increment) corresponding to probe node B, it can be further determined whether Btx_error increment, Btx_drop increment, Br_error increment, or Br_drop increment is greater than 0. If at least one of Btx_error increment, Btx_drop increment, Br_error increment, and Br_drop increment is greater than 0, then proceed to S506; otherwise, proceed to S507.

[0140] S560: The network detection result indicates that the probe node is abnormal;

[0141] It is understandable that if at least one of the increments of B_tx_error, B_tx_drop, B_r_error, and B_drop is greater than 0, it indicates that the network state of probe node B is abnormal, and the network detection result can be obtained as probe node abnormality.

[0142] S507: Based on the first reception error statistics and the first reception packet loss statistics, and combined with the first reception error statistics and the first reception packet loss statistics in the probe data sent to the probe node last time, calculate the increment of the first reception error statistics and the increment of the first reception packet loss statistics.

[0143] Understandably, if the increments of B_tx_error, B_tx_drop, B_r_error, and B_r_drop are all not greater than 0, then it can be assumed that probe node B is not abnormal, meaning no anomalies have been detected in probe node B. In this case, it is necessary to further check whether there are any other anomalies in node A itself. The increments of the first received error statistics (A_error increment) and the first received packet loss statistics (A_drop increment) can be calculated based on the first received error statistics and the first received packet loss statistics from the previous probe data sent to the probe node, and based on the first received error statistics and the first received packet loss statistics calculated from the response information returned by probe node B.

[0144] S508: Determine whether the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero; if the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, proceed to S509; if neither the increment of the first received error statistics value nor the increment of the first received packet loss statistics value is greater than zero, proceed to S510.

[0145] Understandably, after obtaining the Ar_error increment and Ar_drop increment, it is possible to further determine whether the Ar_error increment or Ar_drop increment is greater than zero.

[0146] S509: The network detection result indicates that the node itself is abnormal;

[0147] It should be noted that if at least one of the increments of Ar_error and Ar_drop is greater than 0, it indicates that node A itself is still abnormal. In this case, it can be determined that the network detection result is that node A itself is abnormal.

[0148] S510: The network detection result indicates that neither its own node nor the probe node was detected as abnormal.

[0149] It should be noted that in this embodiment of the invention, if it is determined that neither the Ar_error increment nor the Ar_drop increment is greater than 0, then it can be considered that there is no abnormality in its own node A, that is, no abnormality has been identified in its own node.

[0150] It is understood that the embodiments of the present invention comprehensively consider the first transmission error statistics, the first transmission packet loss statistics, the first reception error statistics, and the first reception packet loss statistics of the node itself, as well as the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics of each probe node. Therefore, even when at least one network packet loss rate is greater than the packet loss rate threshold among all network packet loss rates, the specific network detection result can be accurately determined, which is beneficial to further improve the accuracy of network sub-health state detection.

[0151] In Figure 5 Based on the corresponding embodiments, please refer to Figure 6 In this embodiment of the invention, the detection information may further include a first system load value, and the enhanced detection information may further include a second system load value. System load is a measure of the workload of the CPU (Central Processing Unit) in a node, and the system load value is the average number of threads in the run queue over a specific time interval. In this embodiment of the invention, to distinguish between the system load value of the node itself and the system load value of the detection node, the system load value in the detection information of the node itself is referred to as the first system load value, and the system load value in the enhanced detection information of the detection node is referred to as the second system load value.

[0152] The method may also include:

[0153] If the increment of the first transmission error statistics value or the increment of the transmission packet loss statistics value is greater than zero, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the network of its own node has a transmission failure.

[0154] Alternatively, if the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, determine whether the first system load value is greater than the second load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the network of its own node has a receiving fault.

[0155] It should be noted that when the increment of the first transmission error statistics value or the increment of the first transmission packet loss statistics value is greater than zero, or when the increment of the first reception error statistics value or the increment of the first reception packet loss statistics value is greater than zero, it can be determined that the node itself is abnormal. In this embodiment of the invention, in order to further determine the cause of the node's abnormality in these two cases and take corresponding actions in a timely manner to better ensure the health of the network, when the increment of the first transmission error statistics value or the increment of the first transmission packet loss statistics value is greater than zero, it can be determined whether the first system load value is greater than the first load threshold. If the first system load value is greater than the first load threshold, it is determined that the central processing unit of the node itself is abnormal, and the network interface of the node itself can be switched over. If the first system load value is not greater than the first load threshold, it is determined that the network of the node itself has a transmission fault, and the network interface of the node itself can also be switched over to bring the network to a better state and ensure the normal operation of services.

[0156] Understandably, if the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, it can also be determined whether the first system load value is greater than the second load threshold. If the first system load value is greater than the second load threshold, it can be determined that the central processing unit of its own node is abnormal, and at this time, the network interface of its own node can be switched over. If the first system load value is not greater than the second load threshold, it can be determined that the network of its own node has a receiving fault, and at this time, the network interface of its own node can be switched over.

[0157] It should be noted that the first load threshold and the second load threshold in the embodiments of the present invention can both be empirical values, or they can be determined according to actual needs. In addition, for the process of fault switching of the network interface of the node in the embodiments of the present invention, it can be further determined whether the node has a backup network interface. If a backup network interface exists, the node's network interface (current network interface) can be switched to the backup network interface so that the network of the node can be restored to normal. If no backup network interface exists, the network interface can be shut down.

[0158] If the increment of the second transmission error statistic, the increment of the second transmission packet loss statistic, the increment of the second reception error statistic, or the increment of the second reception packet loss statistic is greater than zero, it can be determined that the probe node is abnormal. In order to further determine the cause of the abnormality of the probe node, it can be determined whether the second system load value is greater than the third load threshold. If the second system load value is greater than the third load threshold, it can be determined that the central processing unit of the probe node is abnormal. If the second system load value is not greater than the third load threshold, it can be determined that the probe node has transmission packet loss or reception packet loss. In this case, there is no need to switch the network interface of its own node.

[0159] In some embodiments, please refer to Figure 7 and Figure 8 If at least one network latency exceeds a latency threshold among all network latencies, then for each probe node, the following approach can be adopted: Figure 7 The method shown enhances latency detection to accurately determine network detection results. In this embodiment, node A is used as the node itself, and node B is used as the probe node for detailed explanation.

[0160] The process of determining the network detection result based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information may include:

[0161] S701: Determine whether the first difference between the second received timestamp and the first sent timestamp exceeds the first time threshold; if the first difference exceeds the first time threshold, proceed to S702; if the first difference does not exceed the first time threshold, proceed to S703.

[0162] It should be noted that, in this embodiment of the invention, a first difference can be calculated based on the second receiving timestamp T2 and the first sending timestamp T1, and the first difference T2-T1 can be compared with a first time threshold shold4.

[0163] S702: The network detection result indicates that the node itself is abnormal;

[0164] Understandably, when T2-T1 > shold4, the network detection result indicates an anomaly in the node itself. To further determine the cause of this node anomaly, the detection information can also include the first system load value; please refer to [reference needed]. Figure 8 It can be further determined whether the first system load value A sysload is greater than the first load threshold shold5. If it is, it is determined that the central processing unit of its own node is abnormal, and the network interface of its own node can be switched over. If not, it is determined that the latency from its own node to the probe node is abnormal, and the network interface of its own node can be switched over.

[0165] S703: Determine whether the second difference between the second sending timestamp and the second receiving timestamp is greater than the second time threshold; if the second difference is greater than the second time threshold, proceed to S704; if the second difference is not greater than the second time threshold, proceed to S705.

[0166] It should be noted that, after determining that the first difference T2-T1 does not exceed the first time threshold shold4, it is possible to further determine whether the second difference T3-T2 between the second sending timestamp T3 and the second receiving timestamp T2 is greater than the second time threshold shold6.

[0167] S704: The network detection result indicates that the probe node is abnormal;

[0168] Specifically, if the second difference T3-T2 is determined to be greater than the second time threshold shold6, it can be determined that the probe node is abnormal, and at this time there is no need to operate the network port of its own node.

[0169] For further details, please refer to Figure 8 The enhanced detection information in this embodiment of the invention also includes a second system load value. In order to further determine the cause of the abnormality of the detection node, if the second difference T3-T2 is greater than the second time threshold shold6, it can be determined whether the second system load value (such as B sysload) is greater than the third load threshold shold7. If so, it is determined that the central processing unit (CPU) of the detection node is abnormal; if not, it is determined that the detection node is abnormal.

[0170] S705: Determine whether the third difference between the first received timestamp and the second sent timestamp is greater than the third time threshold; if the third difference is greater than the third time threshold, proceed to S706; if the third difference is not greater than the third time threshold, proceed to S707.

[0171] It should be noted that in this embodiment of the invention, if the second difference is not greater than the second time threshold, it can be determined that no abnormality has been detected in the probe node. In order to further identify whether there are other abnormalities in its own node, it can be further determined whether the third difference T4-T3 between the first receiving timestamp T4 and the second sending timestamp T3 is greater than the third time threshold shold8.

[0172] S706: The network detection result indicates that the node itself is abnormal;

[0173] Specifically, if the third difference T4-T3 is determined to be greater than the third time threshold shold8, then the node itself can be identified as abnormal. To further determine the cause of this abnormality, such as... Figure 8 As shown, it can be further determined whether the first system load value (such as A sysload) is greater than the second load threshold shold9. If so, it is determined that the central processing unit of its own node is abnormal, and the network interface of its own node can be switched over. If not, it is determined that the latency from the probe node to its own node is abnormal, and the network interface of its own node can be switched over.

[0174] S707: The network detection result indicates that neither its own node nor the probe node was detected as abnormal.

[0175] It should be noted that in this embodiment of the invention, if the third difference T4-T3 is not greater than the third time threshold shold8, it can be determined that the abnormality of its own node has not been identified. That is, the network detection result can be concluded that the abnormality of its own node and the probe node has not been identified.

[0176] It should also be noted that the specific values ​​of each threshold in the embodiments of the present invention can be determined according to actual needs, and the embodiments of the present invention do not impose any special limitations on this.

[0177] In addition, regarding the process of switching the network port of its own node in the embodiments of the present invention, it can be further determined whether the own node has a backup network port. If a backup network port exists, the network port (current network port) of the own node can be switched to the backup network port so that the network of the own node can be restored to normal. If no backup network port exists, the network port can be shut down.

[0178] In some embodiments, before sending probe information to the probe nodes in the cluster in S110 above, the method may further include:

[0179] Multiple probe nodes corresponding to the node itself are determined from the cluster using a skip selection method.

[0180] It should be noted that if all nodes probe each other, excessive probe information will consume bandwidth resources. Therefore, in this embodiment of the invention, in order to reduce bandwidth consumption while ensuring detection accuracy, multiple probe nodes can be selected from the cluster. Considering that nodes with adjacent IP addresses are more likely to use the same hardware resources, in order to avoid the selected probe nodes having the same network problems, a skip-selection method can be used to select probe nodes with non-adjacent IP addresses from the cluster.

[0181] In some embodiments, the process of determining multiple probe nodes corresponding to a node in a cluster may include:

[0182] Number all nodes in the cluster according to the order of their Internet Protocol (IP) addresses;

[0183] The node located immediately following its own node is designated as the first probe node;

[0184] The number of the second probe node is determined based on the number of the first probe node and the total number of nodes in the cluster, combined with the first calculation formula.

[0185] The number of the third probe node is determined based on the number of the first probe node and the total number of nodes in the cluster, combined with the second calculation formula; the number of the third probe node is greater than the number of the second node.

[0186] The second probe node is determined from the cluster based on its number, and the third probe node is determined from the cluster based on its number.

[0187] It should be noted that in this embodiment of the invention, three probe nodes can be selected from the cluster. Specifically, all nodes in the cluster can be numbered according to the order of their Internet Protocol (IP) addresses. The node immediately following the first probe node is then used as the first probe node. Based on the number of the first probe node and the total number of nodes in the cluster, the number of the second probe node is determined using a first calculation formula. Then, based on the number of the first probe node and the total number of nodes in the cluster, the number of the third probe node is determined using a second calculation formula. Finally, based on the number of the second probe node, the second probe node is selected from the cluster, and based on the number of the third probe node, the third probe node is selected from the cluster, thus selecting three probe nodes. Wherein:

[0188] The first calculation formula is: n2 = M + INT(N / 3);

[0189] The second calculation formula is: n3 = M + 2 * INT(N / 3); where n2 represents the number of the second probe node, n3 represents the number of the third probe node, M represents the number of the first probe node, N represents the total number of nodes in the cluster, and INT() represents the floor function.

[0190] For example, in a cluster with N nodes, numbered from 1 to N based on their IP addresses, a node first selects the node M next to its own number as its first probe node. Next, it selects the node at position M+INT(N / 3) as another probe node, and finally, it selects the node at position M+2*INT(N / 3) as its second probe node. It's important to note that if the position calculated using M+2*INT(N / 3) exceeds N, the remainder of M+2*INT(N / 3) divided by N is used, and the node at the position corresponding to this remainder is selected as the third probe node.

[0191] by Figure 3 Taking the nodes shown as an example, ABCDE are numbered 12345. A is selected from the nodes at positions 2, 3 (which is 2+INT(5 / 3)=3), and 4 (2+2*INT(5 / 3)=4), which are the nodes BCD.

[0192] Regarding the timeliness of the method provided in the embodiments of the present invention, Figure 3 Taking five nodes as an example, as shown in Table 1:

[0193] Table 1 Timeliness Analysis Table

[0194] unit time A B C D E 1 Information A 2 Information A Information A Information A 3 Information A Information A Information A Information A Information A

[0195] Understandably, during information propagation, information from node A can spread to all nodes within 3 units of time. Assuming there are N nodes in the cluster, for any node, it takes t units of time to spread to all nodes, where t = t1 + 1, and t1 can be obtained by rounding up N / 3.

[0196] Therefore, it can be seen that the probe information sent by any node in this embodiment of the invention can be gradually diffused to all nodes based on the principle of chain diffusion and the principle of network sub-health interaction. In this embodiment of the invention, the node itself can more accurately determine whether the network sub-health fault exists in itself, the link, or other probe nodes based on the enhanced probe information returned by the probe nodes and the probe information sent by itself. This improves the detection of whether the network it is in is in a sub-healthy state while consuming as few resources as possible. The detection process consumes few resources, is highly efficient, and can obtain more accurate judgments of network sub-health faults, which helps to reduce the complexity of analyzing network anomalies and improve the efficiency of analyzing network sub-health problems.

[0197] This invention also provides a corresponding device for detecting network sub-health status, further enhancing the practicality of the method. The device can be described from both a functional module perspective and a hardware perspective. The following describes the network sub-health status detection device provided by this invention, which is used to implement the network sub-health status detection method provided by this invention. In this embodiment, the network sub-health status detection device may include or be divided into one or more program modules. These program modules are stored in a storage medium and executed by one or more processors to complete the network sub-health status detection method disclosed in the above embodiments. The program module referred to in this invention is a series of computer program instruction segments capable of performing specific functions, which is more suitable than the program itself for describing the execution process of the network sub-health status detection device in the storage medium. The following description will specifically introduce the functions of each program module in this embodiment. The network sub-health status detection device described below can be referred to in correspondence with the network sub-health status detection method described above.

[0198] From the perspective of functional modules, see Figure 9 , Figure 9 This is a structural diagram of the network sub-health state detection device provided by the present invention in a specific embodiment. The device is applied to any node in a cluster and may include:

[0199] The sending module 11 is used to send probe information to the probe nodes in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself;

[0200] The receiving module 12 is used to receive the response information returned by the probe node; the response information includes enhanced probe information and probe information, and the enhanced probe information includes the network status data of the probe node;

[0201] The detection module 13 is used to determine whether the network status of its own node is in a sub-healthy state based on the response information from the probe node.

[0202] In some embodiments, the detection nodes are multiple; the detection module 13 includes:

[0203] The first calculation unit is used to calculate the corresponding network packet loss rate and network latency for each of the probe nodes based on the response information;

[0204] The first determining unit is configured to determine the corresponding network detection result for each probe node based on the response information returned by the probe node when at least one of the network packet loss rates is greater than the packet loss rate threshold or at least one of the network delays is greater than the delay threshold among all the network delays.

[0205] The second determining unit is used to determine whether the network status of its own node is in a sub-healthy state based on the network detection results corresponding to each probe node.

[0206] In some embodiments, the second determining unit is configured to:

[0207] If at least one of the network detection results includes a case where the node itself is abnormal, then the network status of the node is determined to be a sub-healthy state.

[0208] In some embodiments, the network status data in the detection information includes at least one of a first transmission error statistics value, a first transmission packet loss statistics value, a first transmission timestamp, a first reception error statistics value, a first reception packet loss statistics value, and a first reception timestamp;

[0209] The network status data in the enhanced detection information includes at least one of the following: second transmission error statistics, second transmission packet loss statistics, second transmission timestamp, second reception error statistics, second reception packet loss statistics, and second reception timestamp.

[0210] In some embodiments, the first determining unit includes:

[0211] The first determining subunit is configured to, when at least one of the network packet loss rates is greater than a packet loss rate threshold among all network packet loss rates, determine the network detection result for each of the probe nodes based on the first sending error statistics, the first sending packet loss statistics, the first receiving error statistics, the first receiving packet loss statistics, the second sending error statistics, the second sending packet loss statistics, the second receiving error statistics, and the second receiving packet loss statistics in the response information.

[0212] Alternatively, the second determining subunit is configured to, for each of the probe nodes, determine the network detection result based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information, when at least one of the network delays is greater than a delay threshold among all network delays.

[0213] In some embodiments, the first determining subunit includes:

[0214] The first judgment subunit is used to determine whether the increment of the first transmission error statistics value or the increment of the transmission packet loss statistics value is greater than zero, based on the first transmission error statistics value and the first transmission packet loss statistics value, combined with the first transmission error statistics value and the first transmission packet loss statistics value in the previous probe data sent to the probe node; if the increment of the first transmission error statistics value or the increment of the transmission packet loss statistics value is greater than zero, then the third determination subunit is triggered; if neither the increment of the transmission error statistics value nor the increment of the transmission packet loss statistics value is greater than zero, then the second judgment subunit is triggered.

[0215] The third determining subunit is used to determine whether the network detection result indicates that its own node is abnormal;

[0216] The second judgment subunit, based on the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics, and in conjunction with the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics in the previously received response data returned by the probe node, determines whether the increment of the second transmission error statistics or the increment of the second transmission packet loss statistics or the increment of the second reception error statistics or the increment of the second reception packet loss statistics is greater than zero; if the increment of the second transmission error statistics or the increment of the second transmission packet loss statistics or the increment of the second reception error statistics or the increment of the second reception packet loss statistics is greater than zero, then the fourth determination subunit is triggered; if the increments of the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics are all not greater than zero, then the third judgment subunit is triggered.

[0217] The fourth determining subunit is used to determine that the network detection result indicates that the probe node is abnormal;

[0218] The third judgment subunit is used to determine whether the increment of the first receiving error statistics value or the increment of the first receiving packet loss statistics value is greater than zero, based on the first receiving error statistics value and the first receiving packet loss statistics value, combined with the first receiving error statistics value and the first receiving packet loss statistics value in the probe data sent to the probe node last time; if the increment of the first receiving error statistics value or the increment of the first receiving packet loss statistics value is greater than zero, the fifth determination subunit is triggered; if neither the increment of the first receiving error statistics value nor the increment of the first receiving packet loss statistics value is greater than zero, the sixth determination subunit is triggered.

[0219] The fifth determining subunit is used to determine whether the network detection result indicates that its own node is abnormal;

[0220] The sixth determining subunit is used to determine that the network detection result is that the node itself was not identified and the detected node is abnormal.

[0221] In some embodiments, the detection information further includes a first system load value; the device further includes:

[0222] The fourth determination subunit is used to determine whether the first system load value is greater than the first load threshold when the increment of the first transmission error statistics value or the increment of the transmission packet loss statistics value is greater than zero. If yes, the seventh determination subunit is triggered; if no, the eighth determination subunit is triggered.

[0223] The seventh determination subunit is used to determine the CPU anomaly of its own node;

[0224] The eighth determining subunit is used to determine whether there is a transmission failure in the network of its own node;

[0225] Alternatively, the fifth determination subunit is used to determine whether the first system load value is greater than the first load threshold when the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero. If yes, the ninth determination subunit is triggered; if no, the tenth determination subunit is triggered.

[0226] The ninth determination subunit is used to determine the CPU abnormality of its own node;

[0227] The tenth determining subunit is used to determine if there is a receiving failure in the network of its own node.

[0228] In some embodiments, the second determining subunit includes:

[0229] The sixth judgment subunit is used to determine whether the first difference between the first received timestamp and the first sent timestamp exceeds the first time threshold; if the first difference exceeds the first time threshold, the eleventh determination subunit is triggered; if the first difference does not exceed the first time threshold, the seventh judgment subunit is triggered.

[0230] The eleventh determination subunit is used to determine whether the network detection result indicates that its own node is abnormal;

[0231] The seventh judgment subunit is used to determine whether the second difference between the second sending timestamp and the first receiving timestamp is greater than the second time threshold; if the second difference is greater than the second time threshold, the twelfth determination subunit is triggered; if the second difference is not greater than the second time threshold, the eighth judgment subunit is triggered.

[0232] The twelfth determining subunit is used to determine that the network detection result indicates that the probe node is abnormal;

[0233] The eighth determination subunit is used to determine whether the third difference between the second received timestamp and the second sent timestamp is greater than the third time threshold; if the third difference is greater than the third time threshold, the thirteenth determination subunit is triggered; if the third difference is not greater than the third time threshold, the fourteenth determination subunit is triggered.

[0234] The thirteenth determining subunit is used to determine whether the network detection result indicates that its own node is abnormal;

[0235] The fourteenth determining subunit is used to determine that the network detection result is that the self-node was not identified and the detected node is abnormal.

[0236] In some embodiments, the detection information further includes a first system load value; the device further includes:

[0237] The ninth determination subunit is used to determine whether the first system load value is greater than the first load threshold when the first difference exceeds the first time threshold. If yes, the fifteenth determination subunit is triggered; if no, the sixteenth determination subunit is triggered.

[0238] The fifteenth determination subunit is used to determine the CPU abnormality of its own node;

[0239] The sixteenth determining subunit is used to determine the time delay anomaly from its own node to the probe node;

[0240] Alternatively, the tenth determination subunit is used to determine whether the first system load value is greater than the first load threshold when the third difference is greater than the third time threshold. If yes, the seventeenth determination subunit is triggered; if no, the eighteenth determination subunit is triggered.

[0241] The seventeenth determination subunit is used to determine the CPU abnormality of its own node;

[0242] The eighteenth determining subunit is used to determine the time delay anomaly from the probe node to its own node.

[0243] In some embodiments, the device further includes:

[0244] The selection module is used to identify multiple probe nodes corresponding to its own node from the cluster using a skip selection method.

[0245] In some embodiments, the selection module includes:

[0246] The sorting unit is used to number all nodes in the cluster according to the order of their Internet Protocol addresses;

[0247] The third determining unit is used to identify the node located next to itself as the first probe node;

[0248] The fourth determining unit is used to determine the number of the second probe node based on the number of the first probe node and the total number of nodes in the cluster, combined with the first calculation formula.

[0249] The fifth determining unit is used to determine the number of the third probe node based on the number of the first probe node and the total number of nodes in the cluster, combined with the second calculation formula; the number of the third probe node is greater than the number of the second node;

[0250] The sixth determining unit is used to determine the second probe node from the cluster based on the number of the second probe node, and to determine the third probe node from the cluster based on the number of the third probe node.

[0251] In some embodiments, the first calculation formula is: n2 = M + INT(N / 3);

[0252] The second calculation formula is: n3 = M + 2 * INT(N / 3);

[0253] Where n2 represents the number of the second probe node, n3 represents the number of the third probe node, M represents the number of the first probe node, N represents the total number of nodes in the cluster, and INT() represents the floor function.

[0254] It should be noted that the description of the features in the corresponding embodiments of the network sub-health status detection device provided in the present invention can be found in the relevant descriptions of the corresponding embodiments in the above-described method implementation, and will not be repeated here.

[0255] The network sub-health status detection device mentioned above is described from the perspective of functional modules. Furthermore, the present invention also provides an electronic device, which is described from the perspective of hardware. Figure 10 A structural diagram of an electronic device provided in an embodiment of the present invention, such as... Figure 10 As shown, the electronic device includes: a memory 60 for storing computer programs;

[0256] The processor 61 is used to execute computer programs to implement the steps of the network sub-health state detection method as described in the above embodiments.

[0257] The electronic devices provided in this embodiment may include, but are not limited to, smartphones, tablets, laptops, or desktop computers.

[0258] The processor 61 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 61 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 61 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 61 may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.

[0259] The memory 60 may include one or more computer-readable storage media, which may be non-transitory. The memory 60 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 60 is used to store at least the following computer program 601, which, after being loaded and executed by the processor 61, is capable of implementing the relevant steps of the network sub-health state detection method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. The operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include, but is not limited to, network sub-health state detection results.

[0260] In some embodiments, the electronic device may further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.

[0261] Those skilled in the art will understand that Figure 10 The structures shown do not constitute a limitation on electronic devices and may include more or fewer components than those shown.

[0262] It is understood that if the network sub-health status detection method in the above embodiments is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the current technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, magnetic disk, or optical disk, and other media capable of storing program code.

[0263] Based on this, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the network sub-health state detection method described above.

[0264] In addition, embodiments of the present invention also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the network sub-health state detection method described above.

[0265] The foregoing has provided a detailed description of a method, apparatus, electronic device, and computer-readable storage medium for detecting network sub-health status according to embodiments of the present invention. The various embodiments are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0266] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0267] The present invention has provided a detailed description of a method, apparatus, electronic device, and computer-readable storage medium for detecting network sub-health status. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of these embodiments are merely illustrative of the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to the invention without departing from its principles, and these improvements and modifications also fall within the scope of protection of the claims.

Claims

1. A method for detecting sub-health status on the internet, characterized in that, Applied to any node in the cluster, including: Send probe information to probe nodes in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself. Receive the response information returned by the probe node; the response information includes enhanced probe information and the probe information, the enhanced probe information including the network status data of the probe node; Based on the response information from the probe nodes, determine whether the network status of your own node is in a sub-healthy state; wherein, there are multiple probe nodes; The step of determining whether the network status of its own node is in a sub-healthy state based on the response information from the probe node includes: Based on the first transmission error statistics and the first transmission packet loss statistics in the reply information, and combined with the first transmission error statistics and the first transmission packet loss statistics in the previous probe information sent to the probe node, determine whether the increment of the first transmission error statistics or the increment of the first transmission packet loss statistics is greater than zero; If the increment of the first transmission error statistics value or the increment of the first transmission packet loss statistics value is greater than zero, then the network detection result is determined to be an anomaly of its own node. If neither the increment of the first transmission error statistics nor the increment of the first transmission packet loss statistics is greater than zero, then based on the second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the reply information, combined with the second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the reply information returned by the probe node in the previous reception, it is determined whether the increment of the second transmission error statistics or the increment of the second transmission packet loss statistics or the increment of the second reception error statistics or the increment of the second reception packet loss statistics is greater than zero. If the increment of the second transmission error statistics value, the increment of the second transmission packet loss statistics value, the increment of the second reception error statistics value, or the increment of the second reception packet loss statistics value is greater than zero, then the network detection result is determined to be that the probe node is abnormal. If the increment of the second transmission error statistics value, the increment of the second transmission packet loss statistics value, the increment of the second reception error statistics value, and the increment of the second reception packet loss statistics value are all not greater than zero, then based on the first reception error statistics value and the first reception packet loss statistics value in the reply information, combined with the first reception error statistics value and the first reception packet loss statistics value in the previous probe information sent to the probe node, it is determined whether the increment of the first reception error statistics value or the increment of the first reception packet loss statistics value is greater than zero. If the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, then the network detection result is determined to be an anomaly of its own node. If neither the increment of the first received error statistics nor the increment of the first received packet loss statistics is greater than zero, then the network detection result is determined to be that the network itself and the probe node have not been identified as abnormal. Determine whether the network status of your node is in a sub-healthy state based on the network detection results.

2. The method for detecting sub-health status of networks according to claim 1, characterized in that, The step of determining whether the network status of its own node is in a sub-healthy state based on the response information from the probe node includes: For each of the probe nodes, the corresponding network packet loss rate and network latency are calculated based on the response information; If at least one of the network packet loss rates is greater than the packet loss rate threshold, or if at least one of the network delays is greater than the delay threshold, for each of the probe nodes, the corresponding network detection result is determined based on the response information returned by the probe node. Based on the network detection results corresponding to each probe node, determine whether the network status of its own node is in a sub-healthy state.

3. The method for detecting sub-health status of networks according to claim 2, characterized in that, The step of determining whether the network status of its own node is in a sub-healthy state based on the network detection results corresponding to each detection node includes: If at least one of the network detection results includes a case where the node itself is abnormal, then the network status of the node is determined to be a sub-healthy state.

4. The method for detecting sub-health status of networks according to claim 2, characterized in that, The network status data in the detection information includes at least one of the following: first transmission error statistics, first transmission packet loss statistics, first transmission timestamp, first reception error statistics, first reception packet loss statistics, and first reception timestamp; The network status data in the enhanced detection information includes at least one of the following: second transmission error statistics, second transmission packet loss statistics, second transmission timestamp, second reception error statistics, second reception packet loss statistics, and second reception timestamp.

5. The method for detecting sub-health status of networks according to claim 4, characterized in that, In the case where at least one of the network packet loss rates is greater than a packet loss rate threshold, or at least one of the network delays is greater than a delay threshold, for each probe node, the corresponding network detection result is determined based on the response information returned by the probe node, including: If at least one of the network packet loss rates is greater than the packet loss rate threshold among all network packet loss rates, for each of the probe nodes, the network detection result is determined based on the first transmission error statistics, the first transmission packet loss statistics, the first reception error statistics, the first reception packet loss statistics, the second transmission error statistics, the second transmission packet loss statistics, the second reception error statistics, and the second reception packet loss statistics in the response information. Alternatively, if at least one of the network delays is greater than the delay threshold among all network delays, for each of the probe nodes, the network detection result is determined based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information.

6. The method for detecting sub-health status of networks according to claim 1, characterized in that, The detection information also includes a first system load value; the method further includes: If the increment of the first transmission error statistics value or the increment of the transmission packet loss statistics value is greater than zero, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the network of its own node has a transmission failure. Alternatively, if the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the network of its own node has a receiving fault.

7. The method for detecting sub-health status of networks according to claim 5, characterized in that, The step of determining the network detection result based on the first sending timestamp, the first receiving timestamp, the second sending timestamp, and the second receiving timestamp in the response information includes: Determine whether the first difference between the second received timestamp and the first sent timestamp exceeds a first time threshold. If the first difference exceeds the first time threshold, the network detection result is determined to be an anomaly of its own node; If the first difference does not exceed the first time threshold, determine whether the second difference between the second sending timestamp and the second receiving timestamp is greater than the second time threshold. If the second difference is greater than the second time threshold, then the network detection result is determined to be that the probe node is abnormal; If the second difference is not greater than the second time threshold, then determine whether the third difference between the first receiving timestamp and the second sending timestamp is greater than the third time threshold. If the third difference is greater than the third time threshold, then the network detection result is determined to be an anomaly of its own node; If the third difference is not greater than the third time threshold, then the network detection result is determined to be that the network itself and the probe node have not been identified as abnormal.

8. The method for detecting sub-health status of networks according to claim 7, characterized in that, The detection information also includes a first system load value; the method further includes: If the first difference exceeds the first time threshold, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the latency from its own node to the probe node is abnormal. Alternatively, if the third difference is greater than the third time threshold, determine whether the first system load value is greater than the first load threshold. If yes, determine that the central processing unit of its own node is abnormal; if no, determine that the latency from the probe node to its own node is abnormal.

9. The method for detecting sub-health status of the internet according to any one of claims 1 to 8, characterized in that, Before sending probe information to the probe nodes in the cluster, the method further includes: Multiple probe nodes corresponding to the node itself are determined from the cluster using a skip selection method.

10. The method for detecting sub-health status of networks according to claim 9, characterized in that, The process of identifying multiple probe nodes corresponding to its own node from the cluster includes: Number all nodes in the cluster according to the order of their Internet Protocol (IP) addresses; The node located immediately following its own node is designated as the first probe node; The number of the second probe node is determined based on the number of the first probe node and the total number of nodes in the cluster, combined with the first calculation formula. The number of the third probe node is determined based on the number of the first probe node and the total number of nodes in the cluster, combined with the second calculation formula; the number of the third probe node is greater than the number of the second node. The second probe node is determined from the cluster based on its number, and the third probe node is determined from the cluster based on its number.

11. The method for detecting sub-health status of networks according to claim 10, characterized in that, The first calculation formula is: n2 = M + INT(N / 3); The second calculation formula is: n3 = M + 2 * INT(N / 3); Where n2 represents the number of the second probe node, n3 represents the number of the third probe node, M represents the number of the first probe node, N represents the total number of nodes in the cluster, and INT() represents the floor function.

12. A device for detecting sub-health status of internet users, characterized in that, Applied to any node in the cluster, including: A sending module is used to send probe information to probe nodes in the cluster; the probe information includes the network status data of the node itself, and the probe node is different from the node itself. A receiving module is configured to receive response information returned by the probe node; the response information includes enhanced probe information and the probe information, wherein the enhanced probe information includes network status data of the probe node; The detection module is used to determine whether the network status of its own node is in a sub-healthy state based on the response information from the detection nodes; wherein, there are multiple detection nodes; The detection module is used for: Based on the first transmission error statistics and the first transmission packet loss statistics in the reply information, and combined with the first transmission error statistics and the first transmission packet loss statistics in the previous probe information sent to the probe node, determine whether the increment of the first transmission error statistics or the increment of the first transmission packet loss statistics is greater than zero; If the increment of the first transmission error statistics value or the increment of the first transmission packet loss statistics value is greater than zero, then the network detection result is determined to be an anomaly of its own node. If neither the increment of the first transmission error statistics nor the increment of the first transmission packet loss statistics is greater than zero, then based on the second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the reply information, combined with the second transmission error statistics, second transmission packet loss statistics, second reception error statistics, and second reception packet loss statistics in the reply information returned by the probe node in the previous reception, it is determined whether the increment of the second transmission error statistics or the increment of the second transmission packet loss statistics or the increment of the second reception error statistics or the increment of the second reception packet loss statistics is greater than zero. If the increment of the second transmission error statistics value, the increment of the second transmission packet loss statistics value, the increment of the second reception error statistics value, or the increment of the second reception packet loss statistics value is greater than zero, then the network detection result is determined to be that the probe node is abnormal. If the increment of the second transmission error statistics value, the increment of the second transmission packet loss statistics value, the increment of the second reception error statistics value, and the increment of the second reception packet loss statistics value are all not greater than zero, then based on the first reception error statistics value and the first reception packet loss statistics value in the reply information, combined with the first reception error statistics value and the first reception packet loss statistics value in the previous probe information sent to the probe node, it is determined whether the increment of the first reception error statistics value or the increment of the first reception packet loss statistics value is greater than zero. If the increment of the first received error statistics value or the increment of the first received packet loss statistics value is greater than zero, then the network detection result is determined to be an anomaly of its own node. If neither the increment of the first received error statistics nor the increment of the first received packet loss statistics is greater than zero, then the network detection result is determined to be that the network itself and the probe node have not been identified as abnormal. Determine whether the network status of your node is in a sub-healthy state based on the network detection results.

13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the network sub-health state detection method as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the network sub-health state detection method as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and system for detecting network sub-health state of client node

    CN113132160A

  • Network failure diagnosis method and apparatus, device and storage medium

    WO2022134742A1