Link sub-health detection method and device, and electronic equipment
By adding a response identifier to the service message, initial detection is performed using the service message and response message, and secondary detection is performed by combining the probe message. This solves the problems of low efficiency and insufficient accuracy in network sub-health detection in the existing technology, and achieves efficient and accurate link sub-health detection.
Patent Information
- Application Number
- CN202410853180.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-06-27
AI Technical Summary
In a dual-active data center environment, existing technologies that detect network sub-health via ping packets generate a large number of probe packets that consume network bandwidth, affecting the processing of normal business packets. Furthermore, the fixed threshold method is prone to false negatives and false negatives.
By adding response identifiers to business messages, network information can be determined based on business messages and response messages within a specified time period. After the initial detection of suspected sub-health conditions, probe messages are constructed for secondary detection, avoiding additional probe messages and improving detection efficiency and accuracy.
It reduces the impact of probe messages on normal services, improves the efficiency and accuracy of link sub-health detection, and adapts to network sub-health detection in complex networking environments.
Smart Images

Figure CN118694685B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network communication, and in particular to a link sub-health detection method and device and an electronic device. BACKGROUND
[0002] In a complex networking environment for distributed storage, such as a dual-active data center environment, the sub-health of a network communication link (which can be referred to as network sub-health) has a great influence on the stability and performance of the dual-active data center. The dual-active data center refers to a networking environment constructed by two data centers, one of which is a master data center (which can be used to bear user services, etc.), and the other is a backup data center (which can be used to backup data, configurations or services of the master data center, etc.). The two data centers are master and backup to each other and perform real-time data backup. Network sub-health refers to an intermediate state in which the network can still be applied but the performance is lower than expected, such as high packet loss rate, delay and other phenomena in the transmission process.
[0003] Currently, in actual application, the ping packet method is usually used, that is, the network sub-health conditions such as packet loss and delay anomaly of the network communication link are detected by actively sending probe packets to perform related network fault processing. However, the above method of detecting link sub-health will generate a large number of probe packets, which will occupy unnecessary network bandwidth and affect the processing of normal service packets. SUMMARY
[0004] Therefore, the present application provides a link sub-health detection method, device and electronic device to avoid the situation that a large number of probe packets generated in the link sub-health detection process occupy network bandwidth and affect the processing of normal service packets.
[0005] The present application provides a link sub-health detection method, which is applied to any node in any data center in a dual-data center. The method comprises:
[0006] For the communication link corresponding to the master network port of the node, the first network information of the communication link in a specified time period is determined according to each service packet sent by the node through the communication link and the response packet of each service packet in the specified time period. Any service packet carries a response identifier. The response identifier is used to indicate that the target end of the service packet replies the response packet of the service packet when receiving the service packet.
[0007] If the state of the communication link is detected as a suspected sub-health state according to the first network information, a probe packet for probing the communication link is constructed, and a probe node for probing the communication link is determined, so as to determine whether the state of the communication link is a sub-health state based on the probe packet and the probe node.
[0008] The embodiment of the application further provides a link sub-health detection device, the device is arranged in any node in any data center in the double data centers, the device comprises:
[0009] A determining module is configured to determine first network information of a communication link corresponding to a primary network port on the node in a specified time period according to each service packet sent by the node through the communication link and the response packet of each service packet in the specified time period; any service packet carries a response identifier; the response identifier is used to indicate that the target end of the service packet replies the response packet when receiving the service packet;
[0010] A detecting module is configured to construct a detection packet used for detecting the communication link and determine a detection node used for detecting the communication link if the state of the communication link is detected as a suspected sub-health state according to the first network information, so as to determine whether the state of the communication link is a sub-health state based on the detection packet and the detection node.
[0011] The embodiment of the application further provides an electronic device, which comprises:
[0012] A processor; and
[0013] A memory in which computer program instructions are stored, the computer program instructions, when executed by the processor, causing the processor to perform the steps of the above method.
[0014] The embodiment of the application further provides a computer readable storage medium, the computer readable storage medium stores computer program instructions, the computer program instructions, when executed by the processor, causing the processor to perform the steps in the above method.
[0015] As can be seen from the above technical solutions, in the embodiment, the state of the communication link corresponding to the primary network port on the node is first detected according to each service packet sent by the node through the communication link and the response packet of each service packet in a specified time period, so that the sub-health detection of the link does not need to use additional detection packets in the initial detection, avoiding the case that a large number of detection packets occupy network bandwidth and affect the processing of normal service packets, and improving the efficiency of the sub-health detection of the link. Further, the state of the communication link is detected again based on the constructed detection packet and the determined detection node if the initial detection result is a suspected sub-health state, which can effectively improve the accuracy of the sub-health detection of the link. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are incorporated into and form a part of the specification, illustrate one embodiment consistent with the present application and, together with the description, serve to explain the principles of the application.
[0017] Figure 1 Method flowchart provided for the embodiments of the present application.
[0018] Figure 2 Message schematic diagram provided for the embodiments of the present application.
[0019] Figure 3 Service message interaction flowchart provided for the embodiments of the present application.
[0020] Figure 4 Networking schematic diagram provided for the embodiments of the present application.
[0021] Figure 5 Data center internal network sub-health detection method flowchart provided for the embodiments of the present application.
[0022] Figure 6 Inter-data center network sub-health detection method flowchart provided for the embodiments of the present application.
[0023] Figure 7 Device structure diagram provided for the embodiments of the present application.
[0024] Figure 8 Electronic device structure schematic diagram provided for the embodiments of the present application. DETAILED DESCRIPTION
[0025] In order to facilitate the understanding of the present scheme, before describing the present scheme, the technical problems existing in the networking environment for distributed storage will be described first:
[0026] In the distributed storage system, in order to ensure that the data of the storage cluster is not lost and normal read-write services are provided in the case of local failure of the distributed storage system. The distributed storage system deals with the problem of partial failure of the system through computing or hardware redundancy mechanism. The main failure types of the distributed storage system generally include disk failure, node failure, cabinet failure, network failure, etc. For the explicit failure types, such as disk failure and node failure, data redundancy technology can generally be used to prevent storage data loss and unavailability; network failure can generally be solved by isolating the failure node or retry mechanism.
[0027] In contrast, the sub-health state of the hardware is difficult to find because of the difficulty in identification, and the impact on the storage cluster is greater. Hardware sub-health refers to an intermediate state in which the hardware can still operate normally but the performance is lower than expected. The causes of hardware sub-health also include device temperature, hardware environment, configuration errors, etc.
[0028] The distributed storage system achieves fault redundancy and load balancing through network card binding, network plane separation and other technologies in view of the high availability requirement of the network. The internal network of the storage or the internal data exchange network is connected together through a gigabit or a kilobit network to form a dedicated internal network of the storage pool node, which is responsible for data exchange and coordination. In the distributed storage system, multiple storage pools can be included, and each storage pool includes multiple storage nodes. Network sub-health can also affect the performance of the distributed storage system. Common network sub-health scenarios include high packet loss rate, high packet error rate, high delay, and flicker, and the distributed storage system needs to monitor, alarm, and perform fault switching and isolation for various network sub-health scenarios.
[0029] The current network sub-health detection method can collect network information of each node periodically through a management module to perform statistics, and calculate whether there is packet loss and packet error in read / write and bandwidth of the network card to perform network fault processing. Alternatively, the network link packet loss and delay anomaly can be identified through active probe packets without affecting the business, and then network fault processing is performed. However, the above method still has the following problems:
[0030] 1. When the cluster size becomes larger, the network can be detected by ping packets, which can generate a large number of probe packets, thereby affecting the processing of normal business packets and occupying unnecessary network bandwidth.
[0031] 2. In a complex architecture networking environment (such as a large-scale cluster, a dual-active data center, and a multi-site networking environment), the network connection of different data centers has different characteristics. Using a fixed threshold method (such as a fixed delay threshold and a packet loss rate threshold) to process network sub-health can easily lead to false positives and false negatives of network sub-health, thereby affecting normal business of users. Here, the multi-site can be understood as multiple data centers.
[0032] In order to enable those skilled in the art to better understand the technical solutions provided by the embodiments of the present application, and to make the above-mentioned purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.
[0033] Referring to Figure 1 , Figure 1 The method flowchart provided by the embodiments of the present application. The method is applied to any node in any data center in a dual data center. As an embodiment, the node here is an electronic device such as a server, and the embodiments are not specifically limited. The dual data center here refers to the dual-active data center described above.
[0034] As Figure 1 shown, the flow can include the following steps:
[0035] In step 101, the first network information of the communication link corresponding to the main network port of the node in the specified time period is determined according to the service messages and the reply messages of the service messages sent by the node through the communication link in the specified time period.
[0036] In this embodiment, any service message carries a reply identifier. The reply identifier is used to indicate that the reply message of the service message is returned when the target end receives the service message. Alternatively, the reply identifier can be flexibly set based on actual needs, such as being a specified field. This embodiment does not make specific limitations on the reply identifier.
[0037] As an example, the node adds the reply identifier in the data area of the service message when receiving the service message, and then sends the service message carrying the reply identifier to the target end. Then, the target end sends the reply message of the service message to the node immediately when receiving the service message and identifying that the service message carries the reply identifier. The specific content of the reply message is not limited here.
[0038] As an example, the first network information at least includes service delay information and service packet loss information. The service delay information is determined based on the delay parameters of the service messages sent by the node through the communication link in the specified time period. The service packet loss information is determined based on the packet loss parameters of the service messages sent by the node through the communication link in the specified time period. Here, the specified time period can be flexibly set based on actual needs, such as 5 seconds, 10 seconds, etc. Alternatively, the delay parameter of each service message can be the round-trip delay corresponding to each service message. The packet loss parameter of each service message can be the packet loss rate corresponding to each service message. This embodiment does not make specific limitations.
[0039] Based on the above description, the first network information of the communication link in the specified time period is determined according to the service messages and the reply messages of the service messages sent by the node through the communication link in the specified time period. In specific implementation, there are various ways, such as, in determining the service delay information, first, for each service message, if the reply message of the service message is found, the round-trip delay corresponding to the service message is determined based on the sending time of the service message and the receiving time of the reply message of the service message. Then, the service delay information is determined based on the round-trip delays corresponding to each service message. In determining the packet loss delay information, the service packet loss information corresponding to N new service messages is calculated according to the number of service messages and the number of reply messages of service messages. For example, the ratio between the number of reply messages of service messages and the number of service messages is taken as the service packet loss information (i.e. the packet loss rate).
[0040] In the embodiment, the service delay information is determined based on the round-trip delays corresponding to the service packets, for example, a maximum round-trip delay corresponding to the service packets is selected as the service delay information, or an average of the round-trip delays corresponding to the service packets is determined as the service delay information.
[0041] In step 102, if it is detected according to the first network information that the state of the communication link is a suspected sub-health state, a probe packet for probing the communication link is constructed, and a probe node for probing the communication link is determined, so as to determine whether the state of the communication link is a sub-health state based on the probe packet and the probe node.
[0042] In the embodiment, the state of the communication link is detected as a suspected sub-health state according to the first network information, for example, it is checked whether the service delay information in the first network information reaches a service delay requirement corresponding to the communication link, and / or, it is checked whether the service packet loss information in the first network information reaches a service packet loss requirement corresponding to the communication link. If the service delay information does not reach the service delay requirement corresponding to the communication link, and / or, the service packet loss information does not reach the service packet loss requirement corresponding to the communication link, it is determined that the state of the communication link is a suspected sub-health state.
[0043] Optionally, the service delay requirement can be, for example, that the service delay information is less than a set delay threshold, and the service packet loss requirement can be, for example, that the service packet loss information is less than a set packet loss rate threshold. Here, the set delay threshold and the set packet loss rate threshold can be set based on actual needs, and the embodiment is not limited in particular.
[0044] How to determine the probe node for probing the communication link, and how to determine whether the state of the communication link is a sub-health state based on the probe packet and the probe node will be described by way of example below, and will not be described here in detail.
[0045] At this point, the process shown in the flowchart is completed. Figure 1
[0046] By Figure 1 As can be seen from the flow, in the embodiment, the state of the communication link corresponding to the primary network port on the node is first detected according to the service packets sent by the node through the communication link in a specified time period and the response packets of the service packets, so that the sub-health detection of the link does not need to be performed through additional probe packets during the first detection, which avoids the case that a large number of probe packets occupy network bandwidth and affect the processing of normal service packets, and improves the efficiency of the sub-health detection of the link. Further, in the case that the first detection result is a suspected sub-health state, the state of the communication link is detected again based on the constructed probe packet and the determined probe node, which can effectively improve the accuracy of the sub-health detection of the link.
[0047] The following describes how to determine the probe node for detecting the communication link:
[0048] In a specific implementation, the communication link corresponding to the primary network port on the node can be generally divided into two categories, one is the communication link between the node and other storage nodes in the data center, and the other is the communication link between the node and nodes in the opposite data center.
[0049] Based on the above description, as an embodiment, when the communication link is the communication link between the node and other nodes in the data center, the determination of the probe node for detecting the communication link can include, in a specific implementation, selecting at least two nodes from the other nodes in the data center as the probe nodes.
[0050] Of course, as another embodiment, when the communication link is the communication link between the node and nodes in the opposite data center, the determination of the probe node for detecting the communication link can include, in a specific implementation, selecting at least two nodes from the nodes in the opposite data center as the probe nodes.
[0051] The following describes how to determine whether the state of the communication link is a sub-health state based on the probe packet and the probe node:
[0052] As an embodiment, there are many specific implementation manners for determining whether the state of the communication link is the sub-healthy state based on the probe packet and the probe node, for example, first, for each probe node, second network information corresponding to the reference communication link is determined according to the probe packet and the response packet of the probe packet received through the reference communication link between the node and the probe node; wherein the second network information at least includes: probe delay information and probe packet loss information; the probe delay information is determined based on the time delay parameter of each probe packet sent by the node through the reference communication link; the probe packet loss information is determined based on the packet loss parameter of each probe packet sent by the node through the reference communication link. Then, whether the state of the communication link is the sub-healthy state is detected according to the determined second network information.
[0053] In the embodiment, how to determine the second network information corresponding to the reference communication link according to the probe packet and the response packet of the probe packet received through the reference communication link between the node and the probe node is similar to how to determine the first network information of the communication link in the specified time period according to each service packet and the response packet of each service packet sent by the node through the communication link in the specified time period in the above step 101, which will not be repeated here.
[0054] It should be noted that the construction of the probe packet is not specifically limited in the embodiment, and the construction rule of the probe packet can be set based on actual needs to construct the corresponding probe packet based on the construction rule.
[0055] As to how to detect whether the state of the communication link is the sub-healthy state according to the determined second network information, it is similar to how to detect that the state of the communication link is the suspected sub-healthy state according to the first network information in the above step 102, which will not be repeated here.
[0056] As an embodiment, there are many specific implementation manners for detecting whether the state of the communication link is the sub-healthy state according to the determined second network information, for example, first, a target delay information is determined based on the probe delay information in each second network information, and a target packet loss information is determined based on the probe packet loss information in each second network information; then, whether the target delay information meets the service delay requirement corresponding to the communication link is checked; and / or, whether the target packet loss information meets the service packet loss requirement corresponding to the communication link is checked. If the target delay information does not meet the service delay requirement corresponding to the communication link, and / or, the target packet loss information does not meet the service packet loss requirement corresponding to the communication link, it is determined that the state of the communication link is the sub-healthy state.
[0057] In the embodiment, the target delay information is determined based on the probe delay information in each second network information, for example, a maximum probe delay information in the probe delay information in each second network information is selected as the target delay information, or an average of the probe delay information in each second network information is taken as the target delay information. As to how to determine the target packet loss information based on the probe packet loss information in each second network information, it is similar to how to determine the target delay information based on the probe delay information in each second network information, and thus is not described herein.
[0058] As an example, because the physical distance between the nodes corresponding to different types of communication links is different, the fixed delay and packet loss threshold mode is used for link sub-health detection, which is easy to cause the situation of missing or false reporting of link sub-health events. Therefore, in order to avoid this situation, the embodiment sets different requirements (i.e., service delay requirement and service packet loss requirement) for link sub-health detection corresponding to different types of communication links, that is, each type of communication link sets the service delay requirement and the service packet loss requirement (i.e., delay and packet loss threshold) based on its own network connection characteristics and actual needs.
[0059] The following further describes the processing procedure after determining that the state of the communication link is the sub-health state:
[0060] As an example, in the case that the communication link is between the node and other nodes in the data center, after determining that the state of the communication link is the sub-health state, the embodiment isolates the primary network port and enables the standby network port corresponding to the primary network port as the primary network port.
[0061] As another example, in the case that the communication link is between the node and the node in the peer data center, after determining that the state of the communication link is the sub-health state, the embodiment reports the link sub-health event indicating that the communication link is in the sub-health state to an arbitration node connected with the dual data centers. Wherein, when the arbitration node determines that the number of target nodes corresponding to each data center reaches a set number threshold according to all the link sub-health events stored locally, the arbitration node isolates the backup data center in the dual data centers; here, the target node corresponding to any data center refers to the node in the data center that has reported the link sub-health event.
[0062] In the embodiment, the arbitration node detects whether there is a target link sub-health event with a time interval between the event reporting time point and the current time point greater than or equal to a set threshold every set period, and if there is, the target link sub-health event is cleared. Here, the set period and the set threshold can be set based on actual needs, and the embodiment is not limited specifically.
[0063] As an embodiment, in the case that the communication link is the second communication link, the above detection of whether the communication link is in the network sub-health state according to the first network detection information corresponding to the communication link and the set threshold condition for network detection corresponding to the communication link can include: if it is found that the first network detection information corresponding to the second communication link satisfies the set threshold condition corresponding to the second communication link, at least two second probe storage nodes are selected from the storage nodes in the peer data center; K second probe packets are sent to each second probe storage node through the second communication link; K is greater than 1; the third network detection information corresponding to the second communication link is determined according to the K second probe packets and the response packets corresponding to the received K second probe packets; the third network detection information at least includes the time delay information and the packet loss information corresponding to the K second probe packets; whether the second communication link is in the network sub-health state is detected according to the third network detection information.
[0064] As an embodiment, if it is determined that the second communication link is in the network sub-health state, a network sub-health event indicating that the second communication link is in the network sub-health state is reported to an arbitration node connected with the dual-active data center; wherein, when the arbitration node determines that the number of target nodes corresponding to each data center reaches a set number threshold according to all the network sub-health events stored locally, the backup data center in the dual-active data center is isolated; the target node corresponding to any data center is a node in the data center that has reported a network sub-health event.
[0065] In order to facilitate the understanding of the specific implementation process of the above link sub-health detection method, the following will be described by specific embodiments.
[0066] The embodiment proposes a high-efficiency network sub-health detection method based on grouping adaptive adjustment isolation strategy, which adds a statistical field (i.e. response identifier) to the normal service packet, and counts the packet loss and time delay of the network on the service side (i.e. service packet sending end), and reports to the network management module (i.e. module for managing and processing network sub-health events) for packet sending detection after meeting certain conditions. For complex networking environment, by grouping the network, the network communication link inside the data center and the core network between the data centers (i.e. network communication link between the two data centers) are distinguished for different sub-health processing, so as to accurately identify and process the sub-health network in large-scale cluster and complex networking scenario.
[0067] First, the service packet is modified. Referring to Figure 2As shown, taking a common TCP / IP service message as an example, a byte field (i.e. Tag) is added in the data segment of the service message without modifying the specific message format, so that the new field (i.e. response identifier) is carried in the normal service message.
[0068] Then, referring to Figure 3 As shown, the service interaction process of the service modules (e.g. the modules configured on the nodes in the data center for data service processing, i.e. service module 1 and service module 2) is as follows: Figure 3 As shown, the service interaction process of the service modules (e.g. the modules configured on the nodes in the data center for data service processing, i.e. service module 1 and service module 2) is as follows:
[0069] S301, service module 1 sends a service message to service module 2;
[0070] S302, service module 2 immediately replies a response message to service module 1 when it identifies that the service message carries a Tag;
[0071] S303, service module 2 processes the service message;
[0072] S304, service module 2 returns a processing result message of the service message to service module 1.
[0073] In this way, each module with message interaction can count the link delay and packet loss information within a certain period.
[0074] After that, when there is a dual-active or multi-site scenario for grouping detection, referring to Figure 4 As shown, in this networking scenario, there are an arbitration site 401, a router 402, a site 1 403, and a site 2 404, wherein the arbitration site has no service and contains an arbitration node (i.e. node 5); the nodes in site 1 and site 2 communicate with the arbitration node through the router, and the nodes in site 1 and site 2 also communicate through the router. Taking this network as an example, the network can be divided into three groups, the network in site 1 (e.g. the communication link between node 1 and node 2), the network in site 2 (e.g. the communication link between node 3 and node 4), and the core interaction network between site 1 and site 2 (e.g. the communication link between node 1 and node 2). Different sub-health detection and processing strategies are adopted for different groups.
[0075] The sub-health detection method flow of the data center internal network (denoted as the first communication link) is described as follows:
[0076] Referring to Figure 5 As shown, the method is applied to any node in any data center in a dual-data center, and the method includes the following steps:
[0077] S501, statistics packet loss and latency information corresponding to service packets sent through the first communication link within a certain period of time through the configured service module.
[0078] Here, the service packets are normal service packets such as heartbeat, data interaction, etc.
[0079] S502, whether the packet loss and latency information exceeds the set threshold; if yes, execute S503; otherwise, return to execute S501.
[0080] S503, report the suspected network sub-health event to the configured BM module.
[0081] Here, the BM module refers to the network management module described above, which is configured in the service module.
[0082] S504, randomly select a number of nodes in the data center through the BM module to perform packet sending detection, and obtain the packet sending detection result.
[0083] Here, the packet sending detection result refers to the detection packet loss information and detection latency information as described above. For how to obtain the packet sending detection result, please refer to the above description, which will not be repeated here.
[0084] S505, based on the packet sending detection result, determine whether the set requirement corresponding to the first communication link is reached; if not, execute S506.
[0085] Here, the set requirement refers to the set latency requirement and set packet loss requirement corresponding to the communication link.
[0086] S506, isolate the current primary network port, and start the standby network port of the primary network port for communication.
[0087] If the set requirement is reached, it is considered that the suspected network sub-health event is a false positive sub-health detection event, and no isolation of the network port and reporting of the alarm is performed.
[0088] The sub-health detection method flow of the inter-data center network (referred to as the second communication link) is described as follows:
[0089] Referring to Figure 6 The method is applied to any node in any data center in a dual-data center, and the method comprises the following steps:
[0090] S601, the service module for cross-data center interaction will count the packet loss and latency information of the link within a certain period of time during normal message interaction including heartbeat, data interaction, etc. When a certain threshold is exceeded, it is reported to the BM module as a suspected network sub-health event.
[0091] S601, statistics packet loss and latency information corresponding to a service message sent through the second communication link in a certain period by a configured service module.
[0092] S602, whether the packet loss and latency information exceeds a set threshold; if yes, S603 is executed; otherwise, S601 is returned to be executed.
[0093] S603, report a suspected network sub-health event to a configured BM module.
[0094] S604, randomly select a plurality of nodes in a peer data center for packet sending detection by the BM module to obtain corresponding packet sending detection results.
[0095] S605, determine whether the set requirement corresponding to the second communication link is reached based on the packet sending detection results; if not, S606 is executed.
[0096] S606, report a network sub-health event indicating that the second communication link is in a network sub-health state to an arbitration site.
[0097] In the embodiment, the arbitration site collects network sub-health events. If a network sub-health event reported by site 1 to site 2 and a network sub-health event reported by site 2 to site 1 exist at the same time, it is considered that a sub-health problem occurs in the core network between the data centers, triggering the keep-alive processing of the arbitration site (that is, isolating the backup data center in the dual data centers). The arbitration site periodically cleans up the stored network sub-health events. How to clean up can be referred to the above description, which will not be described here.
[0098] As can be seen from the above, the embodiment adds a statistical field in the normal TCP message to statistically count the packet loss and latency of the network on the service side. When a certain condition is met, the network management module is reported for packet sending detection. The passive packet sending detection saves normal business resource messages. The network link state can be counted through normal business message interaction. There is no detection message outside the business message in the normal business interaction, which is not affected by the common detection message cycle. There is a detection message only in the suspected network sub-health scenario, which improves the efficiency of network sub-health detection and reduces resource consumption. Further, the embodiment further detects the reported suspected sub-health event by the BM module to prevent false reporting of network sub-health events, thereby improving the detection accuracy. The embodiment further detects the network sub-health in a grouping manner, which can detect core network faults in a dual-active or multi-site scenario. Different isolation thresholds are set based on different groups for different processing, thereby improving the accuracy and flexibility of network sub-health detection.
[0099] Thus, the method provided by the embodiment of the present application is described, and the device provided by the embodiment of the present application is described as follows: Thus, the method provided by the embodiment of the present application is described, and the device provided by the embodiment of the present application is described as follows:
[0100] As an embodiment, the embodiment also provides a link sub-health detection device. Referring to Figure 7 , Figure 7 The device structure diagram provided by the embodiment of the present application is shown in the figure. As shown in the figure, the device is configured in any node in any data center in the dual data center, and the device 700 comprises: Figure 7
[0101] The determining module 701 is configured to determine, for the communication link corresponding to the primary network port of the node, first network information of the communication link in a specified time period according to each service packet sent by the node through the communication link and the reply packet of each service packet in the specified time period; wherein any service packet carries a reply identifier; the reply identifier is used to indicate that the target end of the service packet replies to the service packet when the target end receives the service packet;
[0102] The detecting module 702 is configured to, if it is detected according to the first network information that the state of the communication link is a suspected sub-health state, construct a detection packet for detecting the communication link and determine a detection node for detecting the communication link, so as to determine whether the state of the communication link is a sub-health state based on the detection packet and the detection node.
[0103] As an embodiment, the first network information at least comprises service delay information and service packet loss information; the service delay information is determined based on the time delay parameter of each service packet sent by the node through the communication link in the specified time period; and the service packet loss information is determined based on the packet loss parameter of each service packet sent by the node through the communication link in the specified time period.
[0104] The detection that the state of the communication link is a suspected sub-health state according to the first network information comprises:
[0105] checking whether the service delay information in the first network information reaches the service delay requirement corresponding to the communication link; and / or,
[0106] checking whether the service packet loss information in the first network information reaches the service packet loss requirement corresponding to the communication link.
[0107] If not, it is determined that the state of the communication link is a suspected sub-health state.
[0108] As an embodiment, the communication link is a communication link between the node and other nodes in the data center.
[0109] The determination of the detection node for detecting the communication link comprises:
[0110] selecting at least two nodes from the other nodes in the data center as the detection node.
[0111] As an embodiment, after determining that the state of the communication link is the sub-healthy state, the apparatus further comprises:
[0112] The isolation module is configured to isolate the active network port and enable the standby network port corresponding to the active network port as the active network port.
[0113] As an embodiment, the communication link is a communication link between the node and a node in the peer data center;
[0114] The determination of the detection node for detecting the communication link comprises: selecting at least two nodes from the nodes in the peer data center as the detection node.
[0115] As an embodiment, after determining that the state of the communication link is the sub-healthy state, the apparatus further comprises:
[0116] The reporting module is configured to report a link sub-healthy event indicating that the communication link is in the sub-healthy state to an arbitration node connected with the dual data centers;
[0117] The arbitration node is configured to isolate the backup data center in the dual data centers when determining that the number of target nodes corresponding to each data center reaches a set number threshold according to all the link sub-healthy events stored locally; and the target node corresponding to any data center is a node in the data center that has reported a link sub-healthy event.
[0118] As an embodiment, the determination of whether the state of the communication link is the sub-healthy state based on the detection packet and the detection node comprises:
[0119] For each detection node, the second network information corresponding to the reference communication link is determined according to the detection packet and the response packet of the detection packet received through the reference communication link between the node and the detection node; wherein the second network information at least comprises: detection delay information and detection packet loss information; the detection delay information is determined based on the time delay parameter of each detection packet sent by the node through the reference communication link; the service packet loss information is determined based on the packet loss parameter of each detection packet sent by the node through the reference communication link; and the state of the communication link is detected according to the determined second network information.
[0120] Thus, the structure of the apparatus is described. Figure 7 The functions and effects of each unit in the above apparatus are achieved in the implementation process of the corresponding steps in the above method, which will not be described here.
[0121] The functions and effects of each unit in the above apparatus are achieved in the implementation process of the corresponding steps in the above method, which will not be described here.
[0122] For the apparatus embodiment, since it basically corresponds to the method embodiment, the relevant part can be seen from the part of the method embodiment. The apparatus embodiment described above is only illustrative, wherein the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, i.e., can be located in one place or distributed on multiple network modules. Part or all of the modules can be selected to achieve the purpose of the application according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0123] Please refer to Figure 8 A hardware structure schematic diagram of an electronic device is provided for an exemplary embodiment of the present application. The electronic device can include a processor 801, a communication interface 802, a memory 803 and a communication bus 804. The processor 801, the communication interface 802 and the memory 803 complete the communication with each other through the communication bus 804. Among them, the memory 803 stores computer program instructions; the processor 801 can execute the steps of the method described in the above embodiment by executing the computer program instructions stored in the memory 803. The electronic device can also include other hardware according to the actual function of the electronic device, which will not be described here.
[0124] Correspondingly, the present application also provides a computer readable storage medium, which stores a plurality of computer program instructions, and the computer program instructions can implement the method disclosed in the above exemplary embodiments of the present application when executed by a processor.
[0125] Exemplarily, the computer readable storage medium can be any electronic, magnetic, optical or other physical storage apparatus, which can contain or store information such as executable instructions, data, etc. For example, the computer readable storage medium can be RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard drive), solid state disk, any type of storage disk (such as optical disk, dvd, etc.), or similar storage medium, or combination thereof. The processor and the memory can be supplemented by or incorporated into special logic circuit.
[0126] The above is only a preferred embodiment of the present application, and is not used to limit the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method of link sub- health detection, the method comprising: The method is applied to any node in any data center of the dual data centers, and the method comprises: For a communication link corresponding to a primary network port on the node, first network information of the communication link in a specified time period is determined according to each service packet sent by the node through the communication link in the specified time period and a reply packet of each service packet, wherein any service packet carries a reply identifier; the reply identifier is used to indicate that a target end of the service packet replies to the reply packet when the target end receives the service packet; If a state of the communication link is detected as a suspected sub-healthy state according to the first network information, a detection packet for detecting the communication link is constructed, and a detection node for detecting the communication link is determined, so as to determine whether the state of the communication link is a sub-healthy state based on the detection packet and the detection node.
2. The method of claim 1, wherein The first network information at least comprises service delay information and service packet loss information; the service delay information is determined based on a delay parameter of each service packet sent by the node through the communication link in a specified time period; and the service packet loss information is determined based on a packet loss parameter of each service packet sent by the node through the communication link in a specified time period; The detection of the state of the communication link as a suspected sub-healthy state according to the first network information comprises: checking whether the service delay information in the first network information reaches a service delay requirement corresponding to the communication link; and / or, checking whether the service packet loss information in the first network information reaches a service packet loss requirement corresponding to the communication link; If the service delay information in the first network information does not reach the service delay requirement corresponding to the communication link, and / or if the service packet loss information in the first network information does not reach the service packet loss requirement corresponding to the communication link, it is determined that the state of the communication link is a suspected sub-healthy state.
3. The method of claim 1, wherein, The communication link is a communication link between the node and other nodes in the data center; The determination of the detection node for detecting the communication link comprises: selecting at least two nodes from the other nodes in the data center as the detection node.
4. The method of claim 3, wherein, After determining that the state of the communication link is a sub-healthy state, the method further comprises: isolating the primary network port, and enabling a standby network port corresponding to the primary network port as a primary network port.
5. The method of claim 1, wherein, The communication link is a communication link between the node and a node in a peer data center; The determination of the detection node for detecting the communication link comprises: selecting at least two nodes from the nodes in the peer data center as the detection node.
6. The method of claim 5, wherein, After determining that the state of the communication link is a sub-healthy state, the method further comprises: reporting a link sub-healthy event indicating that the communication link is in a sub-healthy state to an arbitration node connected with the dual data centers; The arbitration node isolates the backup data center in the dual data center when determining that the number of target nodes corresponding to each data center reaches a set number threshold according to all link sub-health events stored locally; the target node corresponding to any data center refers to a node in the data center that has reported a link sub-health event.
7. The method of claim 1, wherein, The method further includes: For each probe node, the second network information corresponding to the reference communication link is determined according to the probe packet and a response packet of the probe packet received through the reference communication link between the node and the probe node; the second network information at least includes probe delay information and probe packet loss information; the probe delay information is determined based on a delay parameter of each probe packet sent by the node through the reference communication link; and the probe packet loss information is determined based on a packet loss parameter of each probe packet sent by the node through the reference communication link. The state of the communication link is detected to be a sub-health state according to the determined second network information.
8. A link sub-health detection apparatus, characterized in that, The device is configured in any node in any data center in the dual data center, and the device includes: A determination module is configured to determine, for a communication link corresponding to a main network port of the node, first network information of the communication link in a specified time period according to each service packet sent by the node through the communication link and a response packet of each service packet in the specified time period; any service packet carries a response identifier; the response identifier is used to indicate that the target end of the service packet replies to the response packet of the service packet when receiving the service packet. A detection module is configured to, if it is detected that the state of the communication link is a suspected sub-health state according to the first network information, construct a probe packet for detecting the communication link and determine a probe node for detecting the communication link, so as to determine whether the state of the communication link is a sub-health state based on the probe packet and the probe node.
9. The apparatus of claim 8, wherein, The first network information at least includes service delay information and service packet loss information; the service delay information is determined based on a delay parameter of each service packet sent by the node through the communication link in the specified time period; and the service packet loss information is determined based on a packet loss parameter of each service packet sent by the node through the communication link in the specified time period. The detecting the state of the communication link as a suspected sub-health state according to the first network information comprises: checking whether service delay information in the first network information reaches a service delay requirement corresponding to the communication link; and / or checking whether service packet loss information in the first network information reaches a service packet loss requirement corresponding to the communication link; if the service delay information in the first network information does not reach the service delay requirement corresponding to the communication link, and / or if the service packet loss information in the first network information does not reach the service packet loss requirement corresponding to the communication link, determining that the state of the communication link is a suspected sub-health state; and / or The determining whether the state of the communication link is a sub-health state based on the probe packet and the probe node comprises: for each probe node, determining second network information corresponding to a reference communication link between the node and the probe node according to the probe packet and a response packet of the probe packet received through the reference communication link; wherein the second network information at least comprises probe delay information and probe packet loss information; the probe delay information is determined based on a delay parameter of each probe packet sent by the node through the reference communication link; the probe packet loss information is determined based on a packet loss parameter of each probe packet sent by the node through the reference communication link; and detecting whether the state of the communication link is a sub-health state according to the determined second network information.
10. The apparatus of claim 8, wherein, The communication link is a communication link between the node and other nodes in the data center; The determining the probe node for probing the communication link comprises: selecting at least two nodes from other nodes in the data center as the probe node; The device further comprises an isolation module configured to isolate the primary network port and enable a standby network port corresponding to the primary network port as a primary network port after determining that the state of the communication link is a sub-health state.
11. The apparatus of claim 8, wherein, The communication link is a communication link between the node and a node in a peer data center; The determining the probe node for probing the communication link comprises: selecting at least two nodes from nodes in the peer data center as the probe node; The device further comprises a reporting module configured to report a link sub-health event indicating that the communication link is in a sub-health state to an arbitration node connected with the dual data centers after determining that the state of the communication link is a sub-health state; wherein the arbitration node is configured to isolate a backup data center in the dual data centers when determining that a number of target nodes corresponding to each data center reaches a set number threshold according to all stored link sub-health events; and a target node corresponding to any data center is a node in the data center that has reported a link sub-health event.
12. An electronic device, comprising: The electronic device comprises: a processor; and a memory having computer program instructions stored therein, the computer program instructions, when executed by the processor, causing the processor to perform the steps in the method of any one of claims 1 to 7. The electronic device comprises: a processor; and a memory having computer program instructions stored therein, the computer program instructions, when executed by the processor, causing the processor to perform the steps in the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Method, device and system for health examination in load balancing
CN112738238A
Network sub-health detection method and device and medium
CN116016253A