A High-Precision and Low-Latency Node Fault Detection Method for Distributed Clusters
By analyzing the service operation information of each node in the distributed cluster and predicting communication traffic values, building a network communication traffic topology diagram, and generating the optimal deployment location of auxiliary monitoring nodes, the problem of low accuracy and real-time accuracy in the existing technology is solved, and high-precision and low-latency node fault detection is achieved.
Patent Information
- Application Number
- CN202510438166.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing fault detection methods based on heartbeat and timeout cannot accurately determine in large-scale distributed clusters whether the reason for the heartbeat monitoring timeout is caused by node failure or network delay, resulting in the accuracy and real-time accuracy of node failure detection.
By analyzing the service operation information reported by each node in the distributed cluster, estimating the predicted communication traffic value of each node at each moment in the target period, and building a network communication traffic topology chart, considering the predicted communication traffic value and the number of heartbeat monitoring nodes of the auxiliary monitoring nodes, and generating the best deployment location information of the auxiliary monitoring nodes.
The accuracy and real-timeness of fault detection of large-scale distributed cluster nodes can be improved, and the impact of network delay caused by communication traffic and communication distance of different nodes on node fault detection can be summarized, and the optimal deployment location of auxiliary monitoring nodes can be given.
Smart Images

Figure CN119945893B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed cluster node fault detection, and particularly to a high-precision and low-latency node fault detection method for distributed clusters. Background Art
[0002] A large-scale distributed cluster is a system composed of a large number of computing nodes (such as servers, virtual machines, etc.) interconnected through a network. These nodes work together to provide powerful computing power, storage capacity, and the ability to process large-scale data, and can provide business processing support for big data processing, cloud computing platforms, artificial intelligence and machine learning, and online services of large e-commerce and social media platforms.
[0003] In traditional highly available large-scale distributed clusters, due to data distribution dependencies, task assignment collaboration, and service dependencies and cascading effects among several nodes in the cluster, a single node failure may affect the availability and performance of the entire cluster. Existing fault detection methods based on heartbeat and timeout are affected by dynamic networks in practical applications. Since the business operation content executed by different nodes at different times is different, the communication traffic of the communication network link has differences at different times, which in turn causes obvious differences in the network latency of the communication network link at different times. At the same time, the network communication distance between the monitoring node and the monitored node will also affect the execution of heartbeat monitoring due to latency. When using heartbeat monitoring to judge node faults, it is impossible to determine whether the reason for the heartbeat monitoring timeout is node failure or network latency, resulting in insufficient accuracy and real-time performance of node fault detection.
[0004] Therefore, how to consider the impact of network latency brought by the business operation communication traffic of different nodes and the heartbeat monitoring network communication distance in a large-scale distributed cluster on node fault detection, and improve the accuracy and real-time performance of node fault detection in a large-scale distributed cluster, is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] The present invention provides a high-precision and low-latency node fault detection method for distributed clusters, aiming to solve at least one of the above technical problems.
[0006] To achieve the above object, the present invention provides a high-precision and low-latency node fault detection method for distributed clusters, and the method includes the following steps:
[0007] S1: Obtain the business operation information reported by each node in the distributed cluster; wherein, the business operation information includes the node identifier of each node and the business operation content in the target time period;
[0008] S2: Query the historical business operation database using the node identifier and the business operation content, and determine the predicted communication traffic value of each node at each moment within the target period;
[0009] S3: Invoke the network topology structure diagram of the distributed cluster, write the predicted communication traffic value of each node at each moment within the target period into each topology node in the network topology structure diagram, and construct the network communication traffic topology structure diagram of the distributed cluster at each moment within the target period;
[0010] S4: According to the predicted communication traffic value of each topology node recorded in the network communication traffic topology structure diagram at each moment within the target period and the heartbeat monitoring node quantity limit of each auxiliary monitoring node at each moment within the target period, generate the optimal deployment location information of the auxiliary monitoring nodes of the distributed cluster at each moment within the target period;
[0011] S5: Drive each node in the distributed cluster to perform the node heartbeat monitoring action according to the optimal deployment location information, and obtain the node fault detection result of the distributed cluster.
[0012] Optionally, step S1 specifically includes:
[0013] S11: When each node receives a business operation task, extract the business operation content and business operation time in the business operation task, determine whether the business operation time is within the target period, and if so, report the business operation content of the business operation task to the central node of the distributed cluster;
[0014] S12: The central node receives the business operation content reported by each node and extracts the node identifier of the corresponding node, and generates the business operation information of each node.
[0015] Optionally, step S2 specifically includes:
[0016] S21: Extract the node identifier in the business operation information, and query the historical business operation data of each node in the historical business operation database using the node identifier;
[0017] S22: Extract the business operation content in the business operation information, match the historical communication traffic change sequence of each node in the historical business operation data using the business operation content, and determine the predicted communication traffic value of each node at each moment within the target period according to the historical communication traffic change sequence.
[0018] Optionally, step S22 specifically includes:
[0019] S221: Extract the business operation content from the business operation information, and match several groups of historical communication traffic change sequences corresponding to the business operation content of each node type according to the historical communication traffic change sequences corresponding to different types of business operation content recorded in the historical business operation data;
[0020] S222: Calculate the average value of the historical communication traffic at each moment in several groups of historical communication traffic change sequences, and determine the predicted communication traffic value at each moment within the target period for each node.
[0021] Optionally, step S3 specifically includes:
[0022] S31: Invoke the network topology structure diagram of the distributed cluster; wherein, the network topology structure diagram includes topology nodes composed of several cluster nodes and network links connecting two topology nodes;
[0023] S32: Configure a network topology structure diagram for each moment within the target period, and write the predicted communication traffic value at each moment within the target period for each node into the topology nodes in the corresponding network topology structure diagram to construct the network communication traffic topology structure diagram of the distributed cluster at each moment within the target period.
[0024] Optionally, step S4 specifically includes:
[0025] S41: Obtain the number limit of heartbeat monitoring nodes at each moment within the target period when each topology node serves as an auxiliary monitoring node;
[0026] S42: Considering the predicted communication traffic value and the number limit of heartbeat monitoring nodes as constraint conditions, and taking the number of switching times of the auxiliary monitoring node as the optimization objective, generate the optimal deployment location information of the auxiliary monitoring node according to the predicted communication traffic value of each topology node at each moment within the target period.
[0027] Optionally, step S41 specifically includes:
[0028] S411: Query the usage status information of device hardware resources when each topology node executes the corresponding business operation content at each moment within the target period; wherein, the usage status information of the device hardware resources includes the CPU usage percentage and the memory usage percentage of the device;
[0029] S412: Estimate the number limit of heartbeat monitoring nodes at each moment within the target period when each topology node serves as an auxiliary monitoring node according to the device hardware usage status information and the upper limit value of the device hardware resource usage, considering the standard usage percentage of hardware resources obtained during each heartbeat monitoring during the test.
[0030] Optionally, step S42 specifically includes:
[0031] S421: Based on the predicted communication traffic values of each topology node at each moment within the target period, deploy a number of auxiliary monitoring nodes at the corresponding topology node positions in the network topology structure diagram at each moment within the target period, and the associated topology nodes for each auxiliary monitoring node to perform heartbeat monitoring;
[0032] Among them, the deployment of a number of auxiliary monitoring nodes and associated topology nodes corresponding to each moment within the target period needs to meet the following conditions: the total predicted communication traffic value of the network communication topology route between each auxiliary monitoring node and the associated topology node is less than the preset communication traffic value, the number of associated topology nodes for each auxiliary monitoring node to perform heartbeat monitoring is less than the heartbeat monitoring node quantity limit of each auxiliary monitoring node at each moment within the target period, and the sum of the number of topology nodes with deployment changes of auxiliary monitoring nodes between every two adjacent moments among a number of moments within the target period is minimized;
[0033] S422: Based on the deployment positions of the auxiliary monitoring nodes and the deployment positions of the associated topology nodes at each moment within the target period in the network topology structure diagram, generate the optimal deployment position information of the auxiliary monitoring nodes.
[0034] Optionally, step S5 specifically includes:
[0035] S51: Extract the deployment positions of the auxiliary monitoring nodes and the deployment positions of the associated topology nodes at each moment within the target period in the optimal deployment position information, and send them to each node in the distributed cluster;
[0036] S52: At each moment within the target period, each node in the distributed cluster switches to the corresponding node type at that moment, and drives the auxiliary monitoring nodes and the corresponding associated topology nodes to perform the node heartbeat monitoring action, and obtains the node fault detection result of the distributed cluster.
[0037] Optionally, the method further includes S6: After obtaining the node fault detection result of the distributed cluster, perform actions such as transferring the business operation content, redistributing and restoring the data on the nodes determined to be faulty in the node fault detection result.
[0038] The beneficial effects of the present invention are as follows: A high-precision and low-latency node fault detection method for distributed clusters is proposed. By analyzing the service operation information reported by each node in the distributed cluster and referring to historical service operation data, the predicted communication traffic value of each node at each moment in the target period is estimated. The predicted communication traffic value is incorporated into each topology node in the network topology structure diagram of the distributed cluster to construct a network communication traffic topology structure diagram. Considering the predicted communication traffic value and the number limit of heartbeat monitoring nodes of each auxiliary monitoring node, the optimal deployment location information of the auxiliary monitoring node is generated. Thus, by analyzing the service operation information of each node and predicting the communication traffic value, considering the number of heartbeat monitoring nodes of the predicted communication traffic value and the switching times of the auxiliary monitoring node, the optimal deployment location information of the auxiliary monitoring node is generated, which can summarize the impact of network latency caused by the communication traffic and communication distance of different nodes in a large-scale distributed cluster on node fault detection, give the optimal deployment location of the auxiliary monitoring node, and improve the accuracy and real-time performance of node fault detection in a large-scale distributed cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flowchart of an embodiment of the high-precision and low-latency node fault detection method for distributed clusters of the present invention.
[0040] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0042] An embodiment of the present invention provides a high-precision and low-latency node fault detection method for distributed clusters, referring to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the high-precision and low-latency node fault detection method for distributed clusters of the present invention.
[0043] In this embodiment, a high-precision and low-latency node fault detection method for distributed clusters, the method includes the following steps:
[0044] S1: Obtain the service operation information reported by each node in the distributed cluster; wherein, the service operation information includes the node identifier of each node and the service operation content in the target period;
[0045] S2: Query the historical service operation database by using the node identifier and the service operation content, and determine the predicted communication traffic value of each node at each moment within the target time period;
[0046] S3: Invoke the network topology structure diagram of the distributed cluster, write the predicted communication traffic value of each node at each moment within the target time period into each topology node in the network topology structure diagram, and construct the network communication traffic topology structure diagram of the distributed cluster at each moment within the target time period;
[0047] S4: Generate the optimal deployment location information of the auxiliary monitoring nodes of the distributed cluster at each moment within the target time period according to the predicted communication traffic value of each topology node recorded in the network communication traffic topology structure diagram at each moment within the target time period and the heartbeat monitoring node quantity limit of each auxiliary monitoring node at each moment within the target time period;
[0048] S5: Drive each node in the distributed cluster to execute the node heartbeat monitoring action according to the optimal deployment location information, and obtain the node fault detection result of the distributed cluster.
[0049] It should be noted that in a traditional highly available large-scale distributed cluster, due to data distribution dependencies, task assignment collaboration, and service dependencies and cascading effects among several nodes in the cluster, the failure of a single node may affect the availability and performance of the entire cluster. The existing fault detection methods based on heartbeat and timeout are affected by the dynamic network in practical applications. Since the service operation content executed by different nodes at different time periods is different, the communication traffic of the communication network link has differences at different time periods, thereby resulting in obvious differences in the network latency of the communication network link at different time periods. At the same time, the network communication distance between the monitoring node and the monitored node will also affect the execution of heartbeat monitoring due to time delay. When using heartbeat monitoring to judge node faults, it is impossible to determine whether the reason for the heartbeat monitoring timeout is caused by node failure or network delay, resulting in insufficient accuracy and real-time performance of node fault detection.
[0050] To solve the above problems, in this embodiment, by analyzing the service operation information reported by each node in the distributed cluster, the predicted communication traffic value of each node at each moment within the target time period is estimated and the network communication traffic topology structure diagram is constructed. Considering the predicted communication traffic value and the heartbeat monitoring node quantity limit of each auxiliary monitoring node, the optimal deployment location information of the auxiliary monitoring nodes is generated. Thus, by analyzing the predicted communication traffic value, heartbeat monitoring node quantity, and switching times of the auxiliary monitoring nodes of each node, the influence of network latency brought by the communication traffic and communication distance of different nodes in the large-scale distributed cluster on node fault detection can be summarized, and the optimal deployment location of the auxiliary monitoring nodes is given, improving the accuracy and real-time performance of node fault detection in the large-scale distributed cluster.
[0051] In a preferred embodiment, step S1 specifically includes:
[0052] S11: When each node receives a service operation task, it extracts the service operation content and service operation time in the service operation task, determines whether the service operation time is within the target time period. If so, it reports the service operation content of the service operation task to the central node of the distributed cluster;
[0053] S12: The central node receives the service operation content reported by each node and extracts the node identifier of the corresponding node, and generates the service operation information of each node.
[0054] In this embodiment, first, during the service operation task receiving period of the node, the service operation time of the received service operation content is considered, and the service operation content in the service operation task whose service operation time falls within the target time period is reported to the central node of the distributed cluster. Then, the central node extracts the node identifier and saves it together with the service operation content as the service operation information of each node. Thus, the central node can summarize the service operation content of all service operation tasks within the target time period, providing data support for the prediction of the subsequent communication traffic value.
[0055] In a preferred embodiment, step S2 specifically includes:
[0056] S21: Extract the node identifier in the service operation information, and use the node identifier to query the historical service operation data of each node in the historical service operation database;
[0057] S22: Extract the service operation content in the service operation information, use the service operation content to match the historical communication traffic change sequence of each node in the historical service operation data, and determine the predicted communication traffic value of each node at each moment within the target time period according to the historical communication traffic change sequence.
[0058] Furthermore, step S22 specifically includes:
[0059] S221: Extract the service operation content in the service operation information, and match several groups of historical communication traffic change sequences of the service operation content of the type to which each node belongs according to the historical communication traffic change sequences corresponding to different types of service operation content recorded in the historical service operation data;
[0060] S222: Calculate the average value of the historical communication traffic at each moment in several groups of historical communication traffic change sequences, and determine the predicted communication traffic value of each node at each moment within the target time period.
[0061] In this embodiment, the central node uses the received node identifiers to query the historical service operation data of each node in the historical service operation database, and then predicts the predicted communication traffic value of each node at each moment when executing the service operation content by querying the historical communication traffic change sequence corresponding to the service operation content of the same type and calculating the average value of the historical communication traffic.
[0062] In a preferred embodiment, step S3 specifically includes:
[0063] S31: Invoke the network topology structure diagram of the distributed cluster; wherein, the network topology structure diagram includes topology nodes composed of several cluster nodes and network links connecting two topology nodes;
[0064] S32: Configure a network topology structure diagram for each moment within the target time period, and write the predicted communication traffic value of each node at each moment within the target time period into the topology node in the corresponding network topology structure diagram, to construct the network communication traffic topology structure diagram of the distributed cluster at each moment within the target time period.
[0065] In this embodiment, after obtaining the predicted communication traffic value of each node at each moment, write it into each node of the network topology structure diagram of the distributed cluster to form a network communication traffic topology structure diagram. Use this network communication traffic topology structure diagram to reflect the network communication traffic value of the shortest network link between any two topology nodes. Through these network communication traffic values, the link congestion status caused by data forwarding and processing volume of each network link can be quantified, and it is used as an influencing factor for generating the optimal deployment location information of the subsequent auxiliary monitoring nodes.
[0066] In a preferred embodiment, step S4 specifically includes:
[0067] S41: Obtain the number limit of heartbeat monitoring nodes at each moment within the target time period when each topology node serves as an auxiliary monitoring node;
[0068] S42: According to the predicted communication traffic value of each topology node at each moment within the target time period, considering the predicted communication traffic value and the number limit of heartbeat monitoring nodes as constraint conditions, and taking the number of switching times of the auxiliary monitoring node as the optimization objective, generate the optimal deployment location information of the auxiliary monitoring node.
[0069] Furthermore, step S41 specifically includes:
[0070] S411: Query the usage status information of the device hardware resources when each topology node executes the corresponding service operation content at each moment within the target time period; wherein, the usage status information of the device hardware resources includes the CPU usage ratio and the memory usage ratio of the device;
[0071] S412: Based on the device hardware usage status information and the upper limit value of the device hardware resource usage status, considering the standard usage ratio of the hardware resources obtained in each heartbeat monitoring during the execution of the test, estimate the number limit of heartbeat monitoring nodes at each moment within the target period when each topology node serves as an auxiliary monitoring node.
[0072] Furthermore, step S42 specifically includes:
[0073] S421: Based on the predicted communication traffic values of each topology node at each moment within the target period, deploy a number of auxiliary monitoring nodes at the corresponding topology node positions in the network topology structure diagram at each moment within the target period and the associated topology nodes for each auxiliary monitoring node to perform heartbeat monitoring;
[0074] Among them, the deployment of a number of auxiliary monitoring nodes and associated topology nodes at each moment within the target period needs to meet the following requirements: the total predicted communication traffic value of the network communication topology route between each auxiliary monitoring node and the associated topology node is less than the preset communication traffic value, the number of associated topology nodes for each auxiliary monitoring node to perform heartbeat monitoring is less than the number limit of heartbeat monitoring nodes at each moment within the target period for each auxiliary monitoring node, and the sum of the number of topology nodes with deployment changes of auxiliary monitoring nodes between every two adjacent moments among a number of moments within the target period is minimized;
[0075] S422: Based on the deployment positions of the auxiliary monitoring nodes and the deployment positions of the associated topology nodes at each moment within the target period in the network topology structure diagram, generate the optimal deployment position information of the auxiliary monitoring nodes.
[0076] In this embodiment, by querying the usage status information of the device hardware resources when each topology node executes the corresponding service operation content at each moment within the target period, and considering the standard usage ratio of the hardware resources obtained from each heartbeat monitoring during the test, the number limit of heartbeat monitoring nodes at each moment within the target period is estimated when each topology node serves as an auxiliary monitoring node. After that, considering the predicted communication traffic value and the number limit of heartbeat monitoring nodes for each auxiliary monitoring node, by analyzing the service operation information of each node and predicting the communication traffic value, considering the number of heartbeat monitoring nodes and the number of switchings of the auxiliary monitoring nodes for the predicted communication traffic value, the optimal deployment location information of the auxiliary monitoring nodes is generated, which can summarize the impact of network latency caused by the communication traffic and communication distance of different nodes in a large-scale distributed cluster on node fault detection. At the same time, in order to ensure the stability of the system and the continuity of service execution, the number of switchings of the auxiliary monitoring nodes needs to be limited to the minimum to avoid the impact on the normal operation of the cluster caused by network configuration adjustment and node re-coordination due to excessive switching of node roles. Finally, the optimal deployment location of the auxiliary monitoring nodes is given to improve the accuracy and real-time performance of node fault detection in a large-scale distributed cluster.
[0077] In a preferred embodiment, step S5 specifically includes:
[0078] S51: Extract the deployment locations of the auxiliary monitoring nodes and the associated topology nodes at each moment within the target period in the optimal deployment location information, and send them to each node in the distributed cluster;
[0079] S52: At each moment within the target period, each node in the distributed cluster switches to the corresponding node type at that moment, and drives the auxiliary monitoring node and the corresponding associated topology node to perform the node heartbeat monitoring action to obtain the node fault detection result of the distributed cluster.
[0080] It should be noted that the method further includes S6: After obtaining the node fault detection result of the distributed cluster, perform service operation content transfer, data redistribution and recovery actions on the nodes determined to be faulty in the node fault detection result.
[0081] In this embodiment, the deployment positions of the auxiliary monitoring nodes and the associated topology nodes at each moment within the target time period are extracted and sent to each node in the distributed cluster, driving the distributed cluster to perform heartbeat monitoring of the associated topology nodes by the auxiliary monitoring nodes at each moment according to the role of each node. Due to the deployment positions of each auxiliary monitoring node and the corresponding associated topology node, the influence of network latency caused by the communication traffic and communication distance of different nodes in the large-scale distributed cluster on node failure detection is considered. Therefore, the network latency when each auxiliary monitoring node and the corresponding associated topology node perform heartbeat monitoring can be maintained in the optimal state of the system, improving the accuracy and real-time performance of node failure detection in the large-scale distributed cluster. After performing node heartbeat monitoring, according to the node failure detection results, operations such as business operation content transfer, data redistribution, and recovery actions are performed on the nodes determined to be faulty, thereby realizing the redeployment of the distributed cluster.
[0082] In addition, the present invention also proposes a high-precision and low-latency node failure detection device for a distributed cluster. The high-precision and low-latency node failure detection device for a distributed cluster includes: a memory, a processor, and a high-precision and low-latency node failure detection program for a distributed cluster stored on the memory and executable on the processor. When the high-precision and low-latency node failure detection program for a distributed cluster is executed by the processor, the steps of the high-precision and low-latency node failure detection method for a distributed cluster as described above are implemented.
[0083] The specific implementation manner of the high-precision and low-latency node failure detection device for a distributed cluster in this application is basically the same as that of each embodiment of the high-precision and low-latency node failure detection method for a distributed cluster described above, and will not be elaborated here.
[0084] It can be understood that in the description of this specification, the descriptions referring to terms such as "one embodiment", "another embodiment", "other embodiments", or "the first embodiment to the Nth embodiment" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0085] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such a process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, article or system including such an element.
[0086] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A high-precision and low-latency node fault detection method for distributed clusters, characterized in that: The method comprises the following steps: S1: Obtaining business operation information reported by each node in the distributed cluster; wherein the business operation information includes the node identification of each node and the business operation content in the target period; S2: using the node identifier and the service operation content, querying a historical service operation database to determine a predicted communication flow value of each node at each moment in a target period; S3: Call the network topology diagram of the distributed cluster, write the predicted communication flow value of each node at each moment in the target period into each topological node in the network topology diagram, and construct the network communication flow topology diagram of the distributed cluster at each moment in the target period; S4: Generate the optimal deployment location information of the auxiliary monitoring nodes of the distributed cluster at each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period and the limit of the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target period recorded in the network communication traffic topology structure diagram; specifically including: S41: when obtaining each topological node as an auxiliary monitoring node, the limit on the number of heartbeat monitoring nodes at each moment in the target period; S42: according to the predicted communication flow value of each topological node at each moment in the target period, taking into account the predicted communication flow value and the number of heartbeat monitoring nodes as constraints, taking the number of switching times of the auxiliary monitoring nodes as the optimization target, generating the optimal deployment location information of the auxiliary monitoring nodes; Wherein, step S41 specifically includes: S411: querying the usage status information of the device hardware resources when each topological node executes the corresponding business operation content at each moment in the target period; wherein the usage status information of the device hardware resources includes the device CPU usage ratio and memory usage ratio; S412: according to the device hardware usage status information and the upper limit of the device hardware resource usage status, considering the standard usage ratio of hardware resources for each heartbeat monitoring obtained when executing the test, estimating the limit of the number of heartbeat monitoring nodes at each moment in the target time period when each topological node is used as an auxiliary monitoring node; Wherein, step S42 specifically includes: S421: deploying a number of auxiliary monitoring nodes at corresponding topological node positions and associated topological nodes for each auxiliary monitoring node to perform heartbeat monitoring in a network topological structure diagram corresponding to each moment in the target period according to the predicted communication traffic value of each topological node at each moment in the target period; Among them, the deployment of several auxiliary monitoring nodes and associated topological nodes corresponding to each moment in the target time period needs to meet the following requirements: the sum of the predicted communication flow values of the network communication topological routes between each auxiliary monitoring node and the associated topological nodes is less than the preset communication flow value, the number of associated topological nodes for which each auxiliary monitoring node performs heartbeat monitoring is less than the limit on the number of heartbeat monitoring nodes of each auxiliary monitoring node at each moment in the target time period, and the sum of the number of topological nodes for which the auxiliary monitoring node has deployment changes between every two adjacent moments in several moments in the target time period is the smallest; S422: Generate optimal deployment location information of the auxiliary monitoring node based on the deployment location of the auxiliary monitoring node and the deployment location of the associated topology node at each moment in the target period of the network topology structure diagram; S5: driving each node in the distributed cluster to perform a node heartbeat monitoring action according to the optimal deployment location information to obtain a node fault detection result of the distributed cluster.
2. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S1 specifically includes: S11: When each node receives a business operation task, it extracts the business operation content and business operation time in the business operation task, determines whether the business operation time is within the target period, and if so, reports the business operation content of the business operation task to the central node of the distributed cluster; S12: The central node receives the business operation content reported by each node and extracts the node identifier of the corresponding node, and generates business operation information of each node.
3. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S2 specifically includes: S21: extracting the node identifier in the service operation information, and using the node identifier to query the historical service operation data of each node in the historical service operation database; S22: extracting the business operation content in the business operation information, using the business operation content to match the historical communication traffic change sequence of each node in the historical business operation data, and determining the predicted communication traffic value of each node at each moment in the target time period based on the historical communication traffic change sequence.
4. The high-precision and low-latency node fault detection method for distributed clusters according to claim 3, characterized in that: Step S22 specifically includes: S221: extracting the service operation content in the service operation information, and matching several groups of historical communication traffic change sequences of the service operation content of the type to which each node belongs according to the historical communication traffic change sequences corresponding to the different types of service operation contents recorded in the historical service operation data; S222: averaging the historical communication traffic at each moment in the plurality of groups of historical communication traffic change sequences to determine the predicted communication traffic value of each node at each moment in the target time period.
5. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S3 specifically includes: S31: calling a network topology diagram of a distributed cluster; wherein the network topology diagram includes a topology node composed of a plurality of cluster nodes and a network link connecting two topology nodes; S32: configure a network topology diagram for each moment in the target period, and write the predicted communication flow value of each node at each moment in the target period into the topology node in the corresponding network topology diagram, and construct a network communication flow topology diagram of the distributed cluster at each moment in the target period.
6. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: Step S5 specifically includes: S51: extracting the deployment positions of the auxiliary monitoring nodes and the deployment positions of the associated topological nodes at each moment in the target period from the optimal deployment position information, and sending them to each node in the distributed cluster; S52: At each moment in the target time period, each node in the distributed cluster switches to the node type corresponding to the moment, and drives the auxiliary monitoring node and the corresponding associated topology node to perform node heartbeat monitoring actions to obtain the node fault detection result of the distributed cluster.
7. The high-precision and low-latency node fault detection method for distributed clusters according to claim 1, characterized in that: The method further comprises S6: after obtaining the node fault detection result of the distributed cluster, performing business operation content transfer and data redistribution and recovery actions on the node determined to be faulty in the node fault detection result.
Citation Information
Patent Citations
Intelligent optimization method and system of computing power scheduling for improving power supply reliability
CN117453398A
Multi-park real-time performance analysis and optimization method
CN118487945A