Distributed cluster-oriented high-precision low-delay node fault detection method
By analyzing the node's business operation information and historical data in a large-scale distributed cluster, predicting the communication traffic value, and combining the network topology diagram to generate the optimal deployment location of the auxiliary monitoring node, the accuracy and real-time problems of the existing fault detection methods under the differences in network latency and communication traffic are solved, and high-precision and low-latency node fault detection are achieved.
Patent Information
- Application Number
- CN202510438166.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing fault detection methods based on heartbeat and timeout are not accurately judged in large-scale distributed clusters due to network delay differences and communication traffic differences, which is due to node failure or network delay, resulting in the accuracy and real-time accuracy of node failure detection.
By analyzing the service operation information reported by each node in the distributed cluster, referring to historical service operation data, predicting the predicted communication traffic value of each node at each moment in the target period, and integrating it into the network topology diagram to build a network communication traffic topology diagram. Based on the predicted communication traffic value and the number of heartbeat monitoring nodes of the auxiliary monitoring node, the optimal deployment location information of the auxiliary monitoring node is generated, thereby improving the accuracy and real-timeness of node failure detection.
By analyzing the predicted communication traffic value and the number of heartbeat monitoring nodes, the impact of network delay caused by communication traffic and communication distance of different nodes in a large-scale distributed cluster on node failure detection can be accurately summarized, and the optimal deployment location of auxiliary monitoring nodes can be given, which significantly improves the accuracy and real-timeness of node failure detection.
Smart Images

Figure CN119945893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed cluster node fault detection, and in particular to a high-precision and low-latency node fault detection method for distributed clusters. Background Art
[0002] A large-scale distributed cluster is a system composed of a large number of computing nodes (such as servers, virtual machines, etc.) interconnected through a network. These nodes work together to provide powerful computing power, storage capacity and the ability to process large amounts of data. It can provide business processing support for big data processing, cloud computing platforms, artificial intelligence and machine learning, as well as online services of large e-commerce and social media platforms.
[0003] In traditional high-availability large-scale distributed clusters, since several nodes in the cluster have data distribution dependence, task allocation collaboration, service dependence and cascading effects, the failure of a single node may affect the availability and performance of the entire cluster. Existing fault detection methods based on heartbeats and timeouts are affected by dynamic networks in practical applications. Since different nodes perform different business operations at different times, the communication traffic of communication network links at different times is different, which in turn causes the network delay of communication network links at different times to have obvious differences. At the same time, the network communication distance between the monitoring node and the monitored node will also affect the execution of heartbeat monitoring due to the delay. When using heartbeat monitoring to judge node faults, it is impossible to judge whether the cause of the heartbeat monitoring timeout is caused by node failure or network delay, resulting in the accuracy and real-time performance of node fault detection being less than ideal.
[0004] Therefore, how to consider the impact of network delays caused by business operation communication traffic and heartbeat monitoring network communication distance of different nodes in large-scale distributed clusters on node fault detection, and improve the accuracy and real-time performance of node fault detection in large-scale distributed clusters, is a technical problem that needs to be solved urgently. Summary of the invention
[0005] The present invention provides a high-precision and low-latency node fault detection method for a distributed cluster, aiming to solve at least one of the above technical problems.
[0006] To achieve the above object, the present invention provides a high-precision and low-latency node fault detection method for a distributed cluster, the method comprising the following steps: S1: Obtaining business operation information reported by each node in the distributed cluster; wherein the business operation information includes the node identification of each node and the business operation content in the target period; S2: using the node identifier and the service operation content, querying a historical service operation database to determine a predicted communication flow value of each node at each moment in a target period; S3: Call the network topology diagram of the distributed cluster, write the predicted communication flow value of each node at each moment in the target period into each topological node in the network topology diagram, and construct the network communication flow topology diagram of the distributed cluster at each moment in the target period; S4: Generate optimal deployment location information of auxiliary monitoring nodes of the distributed cluster at each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period and the limit of the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target period recorded in the network communication traffic topology structure diagram; S5: driving each node in the distributed cluster to perform a node heartbeat monitoring action according to the optimal deployment location information to obtain a node fault detection result of the distributed cluster.
[0007] Optionally, step S1 specifically includes: S11: When each node receives a business operation task, it extracts the business operation content and business operation time in the business operation task, determines whether the business operation time is within the target period, and if so, reports the business operation content of the business operation task to the central node of the distributed cluster; S12: The central node receives the business operation content reported by each node and extracts the node identifier of the corresponding node, and generates business operation information of each node.
[0008] Optionally, step S2 specifically includes: S21: extracting the node identifier in the service operation information, and using the node identifier to query the historical service operation data of each node in the historical service operation database; S22: extracting the business operation content in the business operation information, using the business operation content to match the historical communication traffic change sequence of each node in the historical business operation data, and determining the predicted communication traffic value of each node at each moment in the target time period based on the historical communication traffic change sequence.
[0009] Optionally, step S22 specifically includes: S221: extracting the service operation content in the service operation information, and matching several groups of historical communication traffic change sequences of the service operation content of the type to which each node belongs according to the historical communication traffic change sequences corresponding to the different types of service operation contents recorded in the historical service operation data; S222: averaging the historical communication traffic at each moment in the plurality of groups of historical communication traffic change sequences to determine the predicted communication traffic value of each node at each moment in the target time period.
[0010] Optionally, step S3 specifically includes: S31: calling a network topology diagram of a distributed cluster; wherein the network topology diagram includes a topology node composed of a plurality of cluster nodes and a network link connecting two topology nodes; S32: configure a network topology diagram for each moment in the target period, and write the predicted communication flow value of each node at each moment in the target period into the topology node in the corresponding network topology diagram, and construct a network communication flow topology diagram of the distributed cluster at each moment in the target period.
[0011] Optionally, step S4 specifically includes: S41: when obtaining each topological node as an auxiliary monitoring node, the limit on the number of heartbeat monitoring nodes at each moment in the target period; S42: Based on the predicted communication traffic value of each topological node at each moment in the target period, considering the predicted communication traffic value and the limit on the number of heartbeat monitoring nodes as constraints, and taking the number of switching times of the auxiliary monitoring nodes as the optimization target, generate the optimal deployment location information of the auxiliary monitoring nodes.
[0012] Optionally, step S41 specifically includes: S411: querying the usage status information of the device hardware resources when each topological node executes the corresponding business operation content at each moment in the target period; wherein the usage status information of the device hardware resources includes the device CPU usage ratio and memory usage ratio; S412: Based on the device hardware usage status information and the upper limit of the device hardware resource usage status, considering the standard usage ratio of hardware resources for each heartbeat monitoring obtained when executing the test, estimate the limit on the number of heartbeat monitoring nodes at each moment in the target time period when each topological node is used as an auxiliary monitoring node.
[0013] Optionally, step S42 specifically includes: S421: deploying a number of auxiliary monitoring nodes at corresponding topological node positions and associated topological nodes for each auxiliary monitoring node to perform heartbeat monitoring in a network topological structure diagram corresponding to each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period; Among them, the deployment of several auxiliary monitoring nodes and associated topological nodes corresponding to each moment in the target time period needs to meet the following requirements: the sum of the predicted communication flow values of the network communication topological routes between each auxiliary monitoring node and the associated topological nodes is less than the preset communication flow value, the number of associated topological nodes for which each auxiliary monitoring node performs heartbeat monitoring is less than the limit on the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target time period, and the sum of the number of topological nodes for which the auxiliary monitoring node has deployment changes between every two adjacent moments in several moments in the target time period is the smallest; S422: Generate optimal deployment location information of the auxiliary monitoring node based on the deployment location of the auxiliary monitoring node and the deployment location of the associated topology node at each moment in the target time period in the network topology structure diagram.
[0014] Optionally, step S5 specifically includes: S51: extracting the deployment positions of the auxiliary monitoring nodes and the deployment positions of the associated topological nodes at each moment in the target period from the optimal deployment position information, and sending them to each node in the distributed cluster; S52: At each moment in the target time period, each node in the distributed cluster switches to the node type corresponding to the moment, and drives the auxiliary monitoring node and the corresponding associated topology node to perform node heartbeat monitoring actions to obtain the node fault detection result of the distributed cluster.
[0015] Optionally, the method further includes S6: after obtaining the node fault detection result of the distributed cluster, performing business operation content transfer and data redistribution and recovery actions on the node determined to be faulty in the node fault detection result.
[0016] The beneficial effects of the present invention are as follows: a high-precision and low-latency node fault detection method for distributed clusters is proposed, which estimates the predicted communication flow value of each node at each moment in the target period by analyzing the business operation information reported by each node in the distributed cluster, referring to the historical business operation data, and integrating the predicted communication flow value into each topological node in the network topology structure diagram of the distributed cluster, constructing the network communication flow topology structure diagram, considering the predicted communication flow value and the number of heartbeat monitoring nodes of each auxiliary monitoring node, and generating the optimal deployment location information of the auxiliary monitoring node. Thus, by analyzing the business operation information of each node and predicting the communication flow value, considering the number of heartbeat monitoring nodes of the predicted communication flow value and the number of switching times of the auxiliary monitoring node, and generating the optimal deployment location information of the auxiliary monitoring node, it is possible to summarize the influence of the network delay caused by the communication flow and communication distance of different nodes in the large-scale distributed cluster on the node fault detection, and give the optimal deployment location of the auxiliary monitoring node, thereby improving the accuracy and real-time performance of the large-scale distributed cluster node fault detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 The present invention is a flowchart of an embodiment of a high-precision and low-latency node fault detection method for distributed clusters.
[0018] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0020] The embodiment of the present invention provides a high-precision and low-latency node fault detection method for distributed clusters. Figure 1 , Figure 1 The present invention is a flowchart of an embodiment of a high-precision and low-latency node fault detection method for distributed clusters.
[0021] In this embodiment, a high-precision and low-latency node fault detection method for a distributed cluster is provided, the method comprising the following steps: S1: Obtaining business operation information reported by each node in the distributed cluster; wherein the business operation information includes the node identification of each node and the business operation content in the target period; S2: using the node identifier and the service operation content, querying a historical service operation database to determine a predicted communication flow value of each node at each moment in a target period; S3: Call the network topology diagram of the distributed cluster, write the predicted communication flow value of each node at each moment in the target period into each topological node in the network topology diagram, and construct the network communication flow topology diagram of the distributed cluster at each moment in the target period; S4: Generate optimal deployment location information of auxiliary monitoring nodes of the distributed cluster at each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period and the limit of the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target period recorded in the network communication traffic topology structure diagram; S5: driving each node in the distributed cluster to perform a node heartbeat monitoring action according to the optimal deployment location information to obtain a node fault detection result of the distributed cluster.
[0022] It should be noted that in traditional high-availability large-scale distributed clusters, due to the data distribution dependence, task allocation collaboration, service dependence and cascading effect of several nodes in the cluster, the failure of a single node may affect the availability and performance of the entire cluster. The existing fault detection methods based on heartbeats and timeouts will be affected by dynamic networks in practical applications. Since different nodes execute different business operations at different times, the communication traffic of communication network links at different times is different, which in turn causes the network delay of communication network links at different times to have obvious differences. At the same time, the network communication distance between the monitoring node and the monitored node will also affect the execution of heartbeat monitoring due to the delay. When using heartbeat monitoring to judge node faults, it is impossible to judge whether the cause of the heartbeat monitoring timeout is caused by node failure or network delay, resulting in the accuracy and real-time performance of node fault detection being less than ideal.
[0023] In order to solve the above problems, this embodiment analyzes the business operation information reported by each node in the distributed cluster, estimates the predicted communication flow value of each node at each moment in the target period and constructs a network communication flow topology diagram, considers the predicted communication flow value and the number of heartbeat monitoring nodes of each auxiliary monitoring node, and generates the optimal deployment location information of the auxiliary monitoring node. Thus, by analyzing the predicted communication flow value of each node, the number of heartbeat monitoring nodes and the number of switching times of the auxiliary monitoring nodes, it is possible to summarize the impact of the network delay caused by the communication flow and communication distance of different nodes in a large-scale distributed cluster on node fault detection, give the best deployment location of the auxiliary monitoring node, and improve the accuracy and real-time performance of node fault detection in large-scale distributed clusters.
[0024] In a preferred embodiment, step S1 specifically includes: S11: When each node receives a business operation task, it extracts the business operation content and business operation time in the business operation task, determines whether the business operation time is within the target period, and if so, reports the business operation content of the business operation task to the central node of the distributed cluster; S12: The central node receives the business operation content reported by each node and extracts the node identifier of the corresponding node, and generates business operation information of each node.
[0025] In this embodiment, first, in the business operation task receiving period of the node, the business operation time of the received business operation content is considered, and the business operation content of the business operation task whose business operation time falls within the target period is reported to the central node of the distributed cluster. The central node then extracts the node identifier and saves it together with the business operation content as the business operation information of each node. In this way, the central node can summarize the business operation content of all business operation tasks within the target period, and provide data support for the prediction of subsequent communication traffic values.
[0026] In a preferred embodiment, step S2 specifically includes: S21: extracting the node identifier in the service operation information, and using the node identifier to query the historical service operation data of each node in the historical service operation database; S22: extracting the business operation content in the business operation information, using the business operation content to match the historical communication traffic change sequence of each node in the historical business operation data, and determining the predicted communication traffic value of each node at each moment in the target time period based on the historical communication traffic change sequence.
[0027] Furthermore, step S22 specifically includes: S221: extracting the service operation content in the service operation information, and matching several groups of historical communication traffic change sequences of the service operation content of the type to which each node belongs according to the historical communication traffic change sequences corresponding to the different types of service operation contents recorded in the historical service operation data; S222: averaging the historical communication traffic at each moment in the plurality of groups of historical communication traffic change sequences to determine the predicted communication traffic value of each node at each moment in the target time period.
[0028] In this embodiment, the central node uses the received node identifier to query the historical business operation data of each node in the historical business operation database, and then queries the historical communication traffic change sequence corresponding to the same type of business operation content, and uses the historical communication traffic to calculate the average value to predict the predicted communication traffic value of each node at each moment when executing the business operation content.
[0029] In a preferred embodiment, step S3 specifically includes: S31: calling a network topology diagram of a distributed cluster; wherein the network topology diagram includes a topology node composed of a plurality of cluster nodes and a network link connecting two topology nodes; S32: configure a network topology diagram for each moment in the target period, and write the predicted communication flow value of each node at each moment in the target period into the topology node in the corresponding network topology diagram, and construct a network communication flow topology diagram of the distributed cluster at each moment in the target period.
[0030] In this embodiment, after obtaining the predicted communication flow value of each node at each moment, it is written into each node of the network topology structure diagram of the distributed cluster to form a network communication flow topology structure diagram. The network communication flow topology structure diagram is used to reflect the network communication flow value of the shortest network link between any two topological nodes. Through these network communication flow values, the link congestion status of each network link due to data forwarding and processing volume can be quantified, and it can be used as an influencing factor for the generation of the optimal deployment location information of the subsequent auxiliary monitoring nodes.
[0031] In a preferred embodiment, step S4 specifically includes: S41: when obtaining each topological node as an auxiliary monitoring node, the limit on the number of heartbeat monitoring nodes at each moment in the target period; S42: Based on the predicted communication traffic value of each topological node at each moment in the target period, considering the predicted communication traffic value and the limit on the number of heartbeat monitoring nodes as constraints, and taking the number of switching times of the auxiliary monitoring nodes as the optimization target, generate the optimal deployment location information of the auxiliary monitoring nodes.
[0032] Furthermore, step S41 specifically includes: S411: querying the usage status information of the device hardware resources when each topological node executes the corresponding business operation content at each moment in the target period; wherein the usage status information of the device hardware resources includes the device CPU usage ratio and memory usage ratio; S412: Based on the device hardware usage status information and the upper limit of the device hardware resource usage status, considering the standard usage ratio of hardware resources for each heartbeat monitoring obtained when executing the test, estimate the limit on the number of heartbeat monitoring nodes at each moment in the target time period when each topological node is used as an auxiliary monitoring node.
[0033] Furthermore, step S42 specifically includes: S421: deploying a number of auxiliary monitoring nodes at corresponding topological node positions and associated topological nodes for each auxiliary monitoring node to perform heartbeat monitoring in a network topological structure diagram corresponding to each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period; Among them, the deployment of several auxiliary monitoring nodes and associated topological nodes corresponding to each moment in the target time period needs to meet the following requirements: the sum of the predicted communication flow values of the network communication topological routes between each auxiliary monitoring node and the associated topological nodes is less than the preset communication flow value, the number of associated topological nodes for which each auxiliary monitoring node performs heartbeat monitoring is less than the limit on the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target time period, and the sum of the number of topological nodes for which the auxiliary monitoring node has deployment changes between every two adjacent moments in several moments in the target time period is the smallest; S422: Generate optimal deployment location information of the auxiliary monitoring node based on the deployment location of the auxiliary monitoring node and the deployment location of the associated topology node at each moment in the target time period in the network topology structure diagram.
[0034] In this embodiment, by querying the usage status information of the equipment hardware resources when each topological node executes the corresponding business operation content at each moment in the target period, and considering the standard usage proportion of the hardware resources of each heartbeat monitoring obtained during the execution test, the number limit of the heartbeat monitoring nodes at each moment in the target period is estimated for each topological node as an auxiliary monitoring node. After that, considering the predicted communication flow value and the number limit of the heartbeat monitoring nodes of each auxiliary monitoring node, by analyzing the business operation information of each node and predicting the communication flow value, considering the predicted communication flow value, the number of heartbeat monitoring nodes and the number of switching of the auxiliary monitoring nodes, the optimal deployment location information of the auxiliary monitoring node is generated, which can summarize the influence of the network delay caused by the communication flow and communication distance of different nodes in the large-scale distributed cluster on the node fault detection. At the same time, in order to ensure the stability of the system and the continuity of business execution, it is necessary to limit the number of switching of the auxiliary monitoring node to a minimum, so as to avoid the influence of excessive switching of node roles on the normal operation of the cluster caused by network configuration adjustment and node re-coordination, and finally, the optimal deployment location of the auxiliary monitoring node is given, so as to improve the accuracy and real-time performance of the large-scale distributed cluster node fault detection.
[0035] In a preferred embodiment, step S5 specifically includes: S51: extracting the deployment positions of the auxiliary monitoring nodes and the deployment positions of the associated topological nodes at each moment in the target period from the optimal deployment position information, and sending them to each node in the distributed cluster; S52: At each moment in the target time period, each node in the distributed cluster switches to the node type corresponding to the moment, and drives the auxiliary monitoring node and the corresponding associated topology node to perform node heartbeat monitoring actions to obtain the node fault detection result of the distributed cluster.
[0036] It should be noted that the method further includes S6: after obtaining the node fault detection result of the distributed cluster, performing business operation content transfer and data redistribution and recovery actions on the node determined to be faulty in the node fault detection result.
[0037] In this embodiment, by extracting the deployment position of the auxiliary monitoring node and the deployment position of the associated topological node at each moment in the target time period and sending them to each node in the distributed cluster, the distributed cluster is driven to perform the heartbeat monitoring of the associated topological node by the auxiliary monitoring node according to the role of each node at each moment. Due to the location deployment of each auxiliary monitoring node and the corresponding associated topological node, the influence of the network delay caused by the communication traffic and communication distance of different nodes in the large-scale distributed cluster on the node fault detection is considered. Therefore, the network delay when each auxiliary monitoring node and the corresponding associated topological node perform heartbeat monitoring can be maintained in the optimal state of the system, thereby improving the accuracy and real-time performance of large-scale distributed cluster node fault detection. After performing node heartbeat monitoring, according to the node fault detection result, the business operation content transfer and data redistribution and recovery actions are performed on the node determined to be faulty, so as to realize the redeployment of the distributed cluster.
[0038] In addition, the present invention also proposes a high-precision, low-latency node fault detection device for distributed clusters, and the high-precision, low-latency node fault detection device for distributed clusters includes: a memory, a processor, and a high-precision, low-latency node fault detection program for distributed clusters stored in the memory and runnable on the processor. When the high-precision, low-latency node fault detection program for distributed clusters is executed by the processor, the steps of the high-precision, low-latency node fault detection method for distributed clusters as described above are implemented.
[0039] The specific implementation of the high-precision and low-latency node fault detection device for distributed clusters in the present application is basically the same as the above-mentioned embodiments of the high-precision and low-latency node fault detection method for distributed clusters, and will not be repeated here.
[0040] It is understood that, in the description of this specification, the description with reference to the terms "one embodiment", "another embodiment", "other embodiments", or "first to Nth embodiments" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0041] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0042] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A high-precision and low-latency node fault detection method for distributed clusters, characterized in that: The method comprises the following steps: S1: Obtaining business operation information reported by each node in the distributed cluster; wherein the business operation information includes the node identification of each node and the business operation content in the target period; S2: using the node identifier and the service operation content, querying a historical service operation database to determine a predicted communication flow value of each node at each moment in a target period; S3: Call the network topology diagram of the distributed cluster, write the predicted communication flow value of each node at each moment in the target period into each topological node in the network topology diagram, and construct the network communication flow topology diagram of the distributed cluster at each moment in the target period; S4: Generate optimal deployment location information of auxiliary monitoring nodes of the distributed cluster at each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period and the limit of the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target period recorded in the network communication traffic topology structure diagram; S5: driving each node in the distributed cluster to perform a node heartbeat monitoring action according to the optimal deployment location information to obtain a node fault detection result of the distributed cluster.
2. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S1 specifically includes: S11: When each node receives a business operation task, it extracts the business operation content and business operation time in the business operation task, determines whether the business operation time is within the target period, and if so, reports the business operation content of the business operation task to the central node of the distributed cluster; S12: The central node receives the business operation content reported by each node and extracts the node identifier of the corresponding node, and generates business operation information of each node.
3. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S2 specifically includes: S21: extracting the node identifier in the service operation information, and using the node identifier to query the historical service operation data of each node in the historical service operation database; S22: extracting the business operation content in the business operation information, using the business operation content to match the historical communication traffic change sequence of each node in the historical business operation data, and determining the predicted communication traffic value of each node at each moment in the target time period based on the historical communication traffic change sequence.
4. The high-precision and low-latency node fault detection method for distributed clusters according to claim 3, characterized in that: Step S22 specifically includes: S221: extracting the service operation content in the service operation information, and matching several groups of historical communication traffic change sequences of the service operation content of the type to which each node belongs according to the historical communication traffic change sequences corresponding to the different types of service operation contents recorded in the historical service operation data; S222: averaging the historical communication traffic at each moment in the plurality of groups of historical communication traffic change sequences to determine the predicted communication traffic value of each node at each moment in the target time period.
5. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S3 specifically includes: S31: calling a network topology diagram of a distributed cluster; wherein the network topology diagram includes a topology node composed of a plurality of cluster nodes and a network link connecting two topology nodes; S32: configure a network topology diagram for each moment in the target period, and write the predicted communication flow value of each node at each moment in the target period into the topology node in the corresponding network topology diagram, and construct a network communication flow topology diagram of the distributed cluster at each moment in the target period.
6. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S4 specifically includes: S41: when obtaining each topological node as an auxiliary monitoring node, the limit on the number of heartbeat monitoring nodes at each moment in the target period; S42: Based on the predicted communication traffic value of each topological node at each moment in the target period, considering the predicted communication traffic value and the limit on the number of heartbeat monitoring nodes as constraints, and taking the number of switching times of the auxiliary monitoring nodes as the optimization target, generate the optimal deployment location information of the auxiliary monitoring nodes.
7. The high-precision and low-latency node fault detection method for distributed clusters according to claim 6, characterized in that: Step S41 specifically includes: S411: querying the usage status information of the device hardware resources when each topological node executes the corresponding business operation content at each moment in the target period; wherein the usage status information of the device hardware resources includes the device CPU usage ratio and memory usage ratio; S412: Based on the device hardware usage status information and the upper limit of the device hardware resource usage status, considering the standard usage ratio of hardware resources for each heartbeat monitoring obtained when executing the test, estimate the limit on the number of heartbeat monitoring nodes at each moment in the target time period when each topological node is used as an auxiliary monitoring node.
8. The high-precision and low-latency node fault detection method for distributed clusters according to claim 6, characterized in that: Step S42 specifically includes: S421: deploying a number of auxiliary monitoring nodes at corresponding topological node positions and associated topological nodes for each auxiliary monitoring node to perform heartbeat monitoring in a network topological structure diagram corresponding to each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period; Among them, the deployment of several auxiliary monitoring nodes and associated topological nodes corresponding to each moment in the target time period needs to meet the following requirements: the sum of the predicted communication flow values of the network communication topological routes between each auxiliary monitoring node and the associated topological nodes is less than the preset communication flow value, the number of associated topological nodes for which each auxiliary monitoring node performs heartbeat monitoring is less than the limit on the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target time period, and the sum of the number of topological nodes for which the auxiliary monitoring node has deployment changes between every two adjacent moments in several moments in the target time period is the smallest; S422: Generate optimal deployment location information of the auxiliary monitoring node based on the deployment location of the auxiliary monitoring node and the deployment location of the associated topology node at each moment in the target time period in the network topology structure diagram.
9. The high-precision and low-latency node fault detection method for distributed clusters according to claim 8, characterized in that: Step S5 specifically includes: S51: extracting the deployment positions of the auxiliary monitoring nodes and the deployment positions of the associated topological nodes at each moment in the target period from the optimal deployment position information, and sending them to each node in the distributed cluster; S52: At each moment in the target time period, each node in the distributed cluster switches to the node type corresponding to the moment, and drives the auxiliary monitoring node and the corresponding associated topology node to perform node heartbeat monitoring actions to obtain the node fault detection result of the distributed cluster.
10. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: The method further comprises S6: after obtaining the node fault detection result of the distributed cluster, performing business operation content transfer and data redistribution and recovery actions on the node determined to be faulty in the node fault detection result.
Citation Information
Patent Citations
Network topology change detection method based on burst detection
CN115412443A
Intelligent optimization method and system of computing power scheduling for improving power supply reliability
CN117453398A
Intelligent auxiliary decision-making system for communication maintenance service based on topology mapping
CN118246900A
Multi-park real-time performance analysis and optimization method
CN118487945A
Method and system for determining the impact of failures in data center networks
US20130232382A1