NTP fault diagnosis method, apparatus and device, and readable storage medium

By establishing an NTP fault diagnosis multicast group between BMC devices and using an algorithm to calculate fault confidence, the problems of low efficiency and high false alarm rate in NTP fault diagnosis are solved, achieving fast and accurate fault location and resolution.

CN120602372APending Publication Date: 2025-09-05XINHUASAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510772596.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing NTP fault diagnosis techniques lack context, requiring manual troubleshooting of each machine. This is inefficient, has a high false positive rate, and makes it difficult to distinguish between different types of fault causes, such as NTP server failures, network issues, or BMC local configuration errors.

Method used

By establishing an NTP fault diagnosis multicast group between BMC devices, periodically broadcasting and receiving NTP service status, using a preset algorithm to calculate fault confidence, combined with a dynamic weight algorithm to determine the fault type and execute the corresponding handling procedure.

Benefits of technology

It achieves efficient fault location, reduces false alarm rates, speeds up fault discovery, improves operation and maintenance efficiency, and can automatically distinguish and handle different types of NTP faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602372A_ABST
    Figure CN120602372A_ABST
Patent Text Reader

Abstract

The invention provides an NTP fault diagnosis method, apparatus and device, and a readable storage medium. The method comprises the steps of adding an NTP fault diagnosis multicast group; periodically broadcasting a local NTP service state in the NTP fault diagnosis multicast group, and receiving NTP service states broadcasted by other members in the NTP fault diagnosis multicast group; and according to the NTP service state, calling a preset algorithm to calculate a fault confidence coefficient, and performing NTP fault diagnosis according to the fault confidence coefficient. Through the technical scheme of the invention, multicast communication realizes multi-node cooperation, a preset algorithm is utilized to integrate the NTP states of local and multicast reception to calculate the fault confidence, the fault type is accurately judged, efficient fault positioning is realized, the false alarm rate is reduced, the fault discovery speed is improved, and the problems of fuzzy fault positioning and high false alarm rate of a traditional scheme are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of communication technology, and in particular to an NTP fault diagnosis method, apparatus, device, and readable storage medium. Background Art

[0002] The server uses an RTC (real-time clock) to provide time services for the BMC (baseboard management controller). Even after a power outage, the CMOS continues to provide power, preserving the RTC's time information and ensuring consistent and accurate time. To accurately locate the cause of a business system failure, the BMC and operating system time must be consistent. This is typically achieved by enabling the NTP service on the OS / BMC and synchronizing time from the same NTP server.

[0003] NTP (Network Time Protocol) is a network protocol used to synchronize computer system clocks. It is designed to achieve high-precision time synchronization over the Internet or a local area network (LAN). NTP significantly impacts the precise location of faults. Therefore, when NTP synchronization fails on a server's BMC, timely reporting of the fault is crucial.

[0004] Currently, these fault messages often lack context, such as the status of other nodes, forcing operations and maintenance to manually troubleshoot each node, which is inefficient. Therefore, operations and maintenance personnel want the BMC to be able to distinguish between different types of fault causes, such as NTP server failures, network issues, or BMC local configuration errors, when a fault occurs, thereby reducing the probability of false alarms and accelerating repairs. Summary of the Invention

[0005] In view of this, this specification provides an NTP fault diagnosis method, apparatus, device, and readable storage medium to improve the above-mentioned problem of low efficiency in NTP fault diagnosis.

[0006] The specific technical solutions are as follows:

[0007] This specification provides an NTP fault diagnosis method, which is applied to a server's BMC device. The method includes: joining an NTP fault diagnosis multicast group, where members of the NTP fault diagnosis multicast group include BMC devices of several servers within a preconfigured range; periodically broadcasting the local NTP service status within the NTP fault diagnosis multicast group, and receiving the NTP service status broadcast by other members of the NTP fault diagnosis multicast group; invoking a preset algorithm to calculate a fault confidence level based on the local NTP service status and the received NTP service status broadcast by other members of the NTP fault diagnosis multicast group, and performing NTP fault diagnosis based on the fault confidence level.

[0008] As a technical solution, joining the NTP fault diagnosis multicast group, where the members of the NTP fault diagnosis multicast group include the BMC devices of several servers within a preconfigured range, includes: joining the NTP fault diagnosis multicast group associated with the TOR switch to which it belongs, where the members of the NTP fault diagnosis multicast group include the BMC devices of several servers to which the TOR switch belongs.

[0009] As a technical solution, the method of calculating the fault confidence level by calling a preset algorithm based on the local NTP service status and the NTP service status broadcast by other members in the received NTP fault diagnosis multicast group, and performing NTP fault diagnosis based on the fault confidence level includes: obtaining a weight parameter pre-configured for each member in the NTP fault diagnosis multicast group, and calculating the fault confidence level based on the abnormal status indicated in the NTP service status broadcast by other members in the received NTP fault diagnosis multicast group and the weight parameter associated with the member to which the abnormal status belongs.

[0010] As a technical solution, the NTP service status is used to indicate abnormal status and the fault type of the abnormal status according to the probe detection result; the preset algorithm is called to calculate the fault confidence based on the local NTP service status and the NTP service status broadcast by other members in the received NTP fault diagnosis multicast group, and NTP fault diagnosis is performed according to the fault confidence, including: calculating the corresponding fault confidence according to the fault type of different abnormal states, and judging whether a fault event corresponding to the fault type has occurred based on the comparison relationship between the fault confidence corresponding to the fault type of each abnormal state and the corresponding threshold.

[0011] As a technical solution, in response to the judgment result of the occurrence of a fault event, a fault handling procedure is executed according to the fault type corresponding to the fault event.

[0012] This specification also provides an NTP fault diagnosis device, which is applied to the BMC device of a server. The device includes: a first module, which is used to join an NTP fault diagnosis multicast group, where the members of the NTP fault diagnosis multicast group include the BMC devices of several servers within a preconfigured range; a second module, which is used to periodically broadcast the local NTP service status within the NTP fault diagnosis multicast group and receive the NTP service status broadcast by other members of the NTP fault diagnosis multicast group; and a third module, which is used to call a preset algorithm to calculate the fault confidence based on the local NTP service status and the received NTP service status broadcast by other members of the NTP fault diagnosis multicast group, and perform NTP fault diagnosis based on the fault confidence.

[0013] As a technical solution, joining the NTP fault diagnosis multicast group, where the members of the NTP fault diagnosis multicast group include the BMC devices of several servers within a preconfigured range, includes: joining the NTP fault diagnosis multicast group associated with the TOR switch to which it belongs, where the members of the NTP fault diagnosis multicast group include the BMC devices of several servers to which the TOR switch belongs.

[0014] As a technical solution, the method of calculating the fault confidence level by calling a preset algorithm based on the local NTP service status and the NTP service status broadcast by other members in the received NTP fault diagnosis multicast group, and performing NTP fault diagnosis based on the fault confidence level includes: obtaining a weight parameter pre-configured for each member in the NTP fault diagnosis multicast group, and calculating the fault confidence level based on the abnormal status indicated in the NTP service status broadcast by other members in the received NTP fault diagnosis multicast group and the weight parameter associated with the member to which the abnormal status belongs.

[0015] As a technical solution, the NTP service status is used to indicate abnormal status and the fault type of the abnormal status according to the probe detection result; the preset algorithm is called to calculate the fault confidence based on the local NTP service status and the NTP service status broadcast by other members in the received NTP fault diagnosis multicast group, and NTP fault diagnosis is performed according to the fault confidence, including: calculating the corresponding fault confidence according to the fault type of different abnormal states, and judging whether a fault event corresponding to the fault type has occurred based on the comparison relationship between the fault confidence corresponding to the fault type of each abnormal state and the corresponding threshold.

[0016] As a technical solution, in response to the judgment result of the occurrence of a fault event, a fault handling procedure is executed according to the fault type corresponding to the fault event.

[0017] This specification also provides an electronic device, including a processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the aforementioned NTP fault diagnosis method.

[0018] This specification also provides a readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the aforementioned NTP fault diagnosis method.

[0019] The above technical solutions provided in this specification bring at least the following beneficial effects:

[0020] Multi-node collaboration is achieved through multicast communication. A preset algorithm is used to calculate the fault confidence by combining the local and multicast received NTP status, accurately determine the fault type, achieve efficient fault location, reduce the false alarm rate, and improve the speed of fault discovery, effectively solving the problems of fuzzy fault location and high false alarm rate in traditional solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the implementation methods of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the implementation methods of this specification or the description of the prior art. Obviously, the drawings described below are only some implementation methods recorded in this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings of the implementation methods of this specification.

[0022] Figure 1 is a flow chart of an NTP fault diagnosis method in one embodiment of this specification;

[0023] Figure 2 is a structural diagram of an NTP fault diagnosis device in one embodiment of this specification;

[0024] Figure 3 is a flow chart in one embodiment of this specification;

[0025] Figure 4 This is a hardware structure diagram of an electronic device in one embodiment of this specification.

[0026] Reference numerals: first module 21 , second module 22 , third module 23 . DETAILED DESCRIPTION

[0027] The terms used in the embodiments of this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a," "the," and "the" used in this specification and claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to any or all possible combinations of one or more of the associated listed items.

[0028] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when...", "when...", or "in response to determining."

[0029] In view of this, this specification provides an NTP fault diagnosis method, apparatus, device, and readable storage medium to at least improve one of the above technical problems.

[0030] The specific technical solution is described below.

[0031] In one embodiment, this specification provides an NTP fault diagnosis method, which is applied to a BMC device of a server. The method includes: joining an NTP fault diagnosis multicast group, where members of the NTP fault diagnosis multicast group include BMC devices of several servers within a preconfigured range; periodically broadcasting the local NTP service status in the NTP fault diagnosis multicast group, and receiving the NTP service status broadcast by other members of the NTP fault diagnosis multicast group; calling a preset algorithm to calculate the fault confidence based on the local NTP service status and the received NTP service status broadcast by other members of the NTP fault diagnosis multicast group, and performing NTP fault diagnosis based on the fault confidence.

[0032] Specifically, if Figure 1 , including the following steps:

[0033] Step S11: Join an NTP fault diagnosis multicast group, where members of the NTP fault diagnosis multicast group include BMC devices of several servers within a pre-configured range.

[0034] Step S12: Periodically broadcast the local NTP service status in the NTP fault diagnosis multicast group, and receive the NTP service status broadcast by other members in the NTP fault diagnosis multicast group.

[0035] All BMC devices participating in the diagnosis join a preconfigured multicast group, such as 239.255.200.1, ensuring information sharing among all members. Within this specific multicast group, each BMC device periodically broadcasts its local NTP service status (for example, every 60 seconds) and simultaneously receives status information from other members. This design not only enables the construction of a distributed awareness network but also enables nodes to monitor each other, creating a collaborative working environment.

[0036] In a TOR switch environment with 24 ports, all servers connected to this switch are equipped with BMC devices that support this function and have automatically joined the designated multicast group. Every 60 seconds, these BMC devices will send a message containing their own NTP service status to all members in the multicast group. This message uses a structured JSON format, which includes the node IP address, NTP synchronization status, the last error code, and the results of some local tests (such as ICMP reachability, UDP port openness, NTP protocol response, etc.). At the same time, each BMC will also receive similar information sent by other members, thus forming a real-time updated network status map.

[0037] Step S13: Based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group, a preset algorithm is called to calculate the fault confidence, and NTP fault diagnosis is performed based on the fault confidence.

[0038] A method based on a dynamic weight algorithm can be used to determine the possible cause of the fault by evaluating different types of detection results (such as Ping test, port detection, NTP service response) and their corresponding weights. For example, during a specific diagnostic process, if a BMC reports that its NTP service response failed, but the Ping test and port detection are normal, then according to the preset rules, this situation will be assigned a corresponding weight value. Specifically, a successful Ping test is scored as 1 (weight), a successful port detection is scored as 2, and a failed NTP service response is scored as 3. The system then calculates a fault confidence percentage based on the ratio of the sum of the weights of all abnormal nodes to the maximum weight in the entire group. This percentage helps to distinguish between global problems (such as NTP server crashes), local problems (such as network partitions or firewall interceptions), or single-node configuration errors.

[0039] In one embodiment, if eight ports in the aforementioned 24-port TOR switch environment report NTP service response failures, then according to the formula (8×3) / 72×100%=33.33%, the system will determine that this is a global problem, likely due to an NTP server failure. Conversely, if five ports show port probe failures, that is, (5×2) / 72×100%=13.89%, it is more likely to be a localized problem caused by network partitioning or improper firewall rules. If only three ports experience ping failures, the failure confidence level is only 4.17%, indicating a single node network misconfiguration or hardware failure.

[0040] After determining the type of fault, the corresponding predefined recovery action is executed. The recovery action is determined based on the specific circumstances of the fault and is intended to restore normal service levels as quickly as possible. For example, when a global NTP server failure is detected, the system may automatically trigger the switching of the backup NTP server and notify the operation and maintenance team to check the status of the primary server; if it is a network partition or a local configuration error, it may be necessary to refresh the firewall rules or restart the BMC's network interface; for single-node configuration errors, the system can automatically roll back to the default ntp.conf configuration and restart the ntpd service. In addition, the entire fault handling process will be recorded as basic data for future optimization. Based on historical data, the effectiveness and weight distribution of each probe can also be dynamically adjusted, and even probe effectiveness reports can be automatically generated to facilitate continuous improvement of the system's diagnostic capabilities.

[0041] When there are only a small number of nodes (e.g., only one or two nodes), the threshold needs to be adjusted appropriately to avoid misjudgments. For example, when only two nodes are involved in diagnosis, the threshold for global failures should be raised to above 50% to prevent a single node failure from misjudging the entire system as a global problem. In extreme cases, if only one node is available, the specific fault type is directly mapped based on its weight without performing any confidence calculation.

[0042] The NTP fault diagnosis method based on multicast communication and dynamic weight algorithm can effectively improve the operation and maintenance efficiency and management level of data center servers without increasing additional hardware costs.

[0043] In one embodiment, joining an NTP fault diagnosis multicast group, wherein members of the NTP fault diagnosis multicast group include BMC devices of several servers within a preconfigured range, includes: joining an NTP fault diagnosis multicast group associated with a TOR switch to which the NTP fault diagnosis multicast group belongs, wherein members of the NTP fault diagnosis multicast group include BMC devices of several servers to which the TOR switch belongs.

[0044] In one embodiment, the method of calculating the fault confidence level by calling a preset algorithm based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group, and performing NTP fault diagnosis based on the fault confidence level includes: obtaining a weight parameter pre-configured for each member in the NTP fault diagnosis multicast group, and calculating the fault confidence level based on the abnormal status indicated in the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group and the weight parameter associated with the member to which the abnormal status belongs.

[0045] In one embodiment, the NTP service status is used to indicate an abnormal state and a fault type of the abnormal state based on a probe detection result; the method of calculating a fault confidence based on the local NTP service status and the received NTP service status broadcast by other members of the NTP fault diagnosis multicast group and performing NTP fault diagnosis based on the fault confidence includes: calculating corresponding fault confidences according to different fault types of abnormal states, and judging whether a fault event corresponding to the fault type has occurred based on a comparison between the fault confidences corresponding to the fault types of each abnormal state and corresponding thresholds.

[0046] In one embodiment, in response to a determination that a fault event has occurred, a fault handling procedure is executed according to a fault type corresponding to the fault event.

[0047] In one implementation, the BMC device joins a preconfigured NTP fault diagnosis multicast group. This multicast group includes the BMC devices of several servers within a preconfigured range, such as a TOR switch, EOR switch, or MOR switch. By organizing the BMC devices of multiple servers into the same multicast group, a communication framework is established for subsequent collaborative diagnosis and information sharing.

[0048] When a server boots up and initializes its BMC, it proactively sends a join request to the designated multicast address based on pre-configured parameters. In a typical data center environment, the pre-configured multicast address might be 239.255.200.1, and the servers within the pre-configured range might include all servers connected to the same TOR switch. Upon receiving the join request, the BMCs of these servers execute the appropriate network protocol operations to complete the registration of multicast group membership.

[0049] The BMC device periodically broadcasts its local NTP service status within the NTP fault diagnosis multicast group and receives NTP service status broadcasts from other members of the multicast group. The local NTP service status broadcast interval is typically set to 60 seconds, but this can be adjusted based on actual application scenarios and requirements. The broadcast data is presented in a structured format, such as JSON, which contains a rich set of information elements.

[0050] In the NTP service status broadcast in JSON format, the "node_ip" field identifies the IP address of the server corresponding to the BMC device sending the broadcast, making it easy to uniquely identify each member node in the multicast group. The "ntp_status" field indicates the current synchronization status of the local NTP service, whether it is successful ("succeeded") or failed ("failed"). The "last_error" field provides the code information of the last error encountered during the NTP synchronization process, which helps further analyze the specific cause of the failure. The "local_tests" field records the results of the detection tests performed locally on the BMC device, including network connectivity tests (ping), NTP service port (UDP / 123) reachability tests, and NTP service response tests.

[0051] By periodically broadcasting status information, each BMC device can promptly inform other members in the multicast group of its own health status and any problems it encounters, enabling all members in the multicast group to maintain real-time understanding and synchronization of each other's NTP service status.

[0052] At the same time, each BMC receives NTP service status information broadcast from other members of the multicast group. This received data is stored in a local data structure, such as a list or dictionary, for subsequent analysis and processing. By collecting and organizing this NTP status data from different member nodes, each BMC device can build a global view of the NTP status of all members in the multicast group. This global view is the key basis for subsequent fault diagnosis and confidence calculation. It allows each BMC device to not only monitor its own NTP status, but also observe and analyze the NTP synchronization status of the entire multicast group from a holistic perspective, making it possible to accurately determine the nature and scope of the fault.

[0053] After collecting NTP service status information from local and remote members, the BMC device uses a pre-set algorithm to calculate the fault confidence level. By comprehensively analyzing and weighting various status information, the BMC device derives a numerical value that quantifies the probability of a fault, known as the fault confidence level. First, the NTP status information of each member node is analyzed to determine if any anomalies exist. For any anomaly, the algorithm calculates the node's anomaly weight based on the failure of key test items during the NTP communication process. These key test items include network connectivity (ping), port reachability (port 123), and NTP service response. Each test item is assigned a different weight to reflect its importance in the NTP synchronization process and its impact on fault diagnosis. For example, the network connectivity test has a weight of 1, the port reachability test has a weight of 2, and the NTP service response test has a weight of 3. This weighting is based on a deep understanding of the NTP protocol's working principles and common failure modes. If there's a problem with network connectivity itself, NTP synchronization will inevitably fail, which is a relatively basic and serious failure scenario. If network connectivity is normal but the port is blocked or closed, this is a lower-level failure scenario. Finally, if there are no network or port issues but the NTP service itself is unresponsive, this could be a server-side failure or configuration issue, and this scenario receives the highest weighting.

[0054] After determining the anomaly weight for each abnormal node, the algorithm further calculates the fault confidence for the entire multicast group. The calculation formula is: Fault Confidence = (∑ Abnormal Node Weight / (Number of Nodes × Maximum Single Node Weight)) × 100%. The abnormal node weight is the sum of the weights of all failed test items for all abnormal nodes; the number of nodes is the total number of member nodes in the current multicast group; and the maximum single node weight is the maximum possible sum of all test item weights for a single node. For example, in a multicast group with 24 nodes, suppose three of them experience network connectivity issues (ping test failures). Each node has an abnormal weight of 1, so the total abnormal node weight is 3 × 1 = 3. Given a total number of 24 nodes, the maximum single node weight is 1 + 2 + 3 = 6. Therefore, the fault confidence is calculated as (3 / (24 × 6)) × 100% ≈ 2.08%. Based on the preset decision threshold, this confidence value is below 10%, and therefore is considered a single-node failure, likely caused by a network configuration error or other localized issue at a specific node.

[0055] This calculation method allows the fault confidence to comprehensively reflect the severity and impact of NTP failures across the entire multicast group. When the fault confidence reaches or exceeds a preset threshold, such as 30%, it can be determined to be a global NTP server failure, such as an NTP server crash or a major failure. When the confidence falls within the intermediate range, such as 10% to 30%, it may indicate a network partition or local configuration error, such as a firewall rule in a certain area mistakenly blocking NTP traffic. When the confidence is less than or equal to 10%, it is more likely to be a single-node configuration error or a local network issue. This fault confidence calculation method, based on a weighted voting mechanism, fully leverages the advantages of multi-node collaborative diagnosis. By comprehensively considering the status information and test results of each member node, it greatly improves the accuracy and reliability of fault diagnosis and avoids the false positives and false negatives that are common in traditional single-node detection methods.

[0056] In practical applications, this confidence calculation and fault diagnosis method can effectively help operations and maintenance personnel quickly locate the root cause of NTP failures. For example, in a large data center, a TOR switch connects to 24 servers, and the BMC devices of these servers are all members of the same NTP fault diagnosis multicast group. At a certain point, the BMC devices of some servers begin to experience NTP synchronization failures. Using the method of the present invention, each BMC device broadcasts its own NTP status information, including various local test results. After receiving this information, other members of the multicast group analyze and calculate it according to a preset algorithm. Suppose that over a period of time, eight nodes fail the NTP response test, while network connectivity and port reachability tests are normal. Based on the weight distribution, each node has an abnormal weight of 3 (the weight of the NTP service response test), and the total abnormal node weight is 8 × 3 = 24. The number of nodes is 24, and the maximum weight of a single node is 6. Therefore, the fault confidence is (24 / (24 × 6)) × 100% = 16.67%. Based on the decision threshold, this confidence level is between 10% and 30%, indicating a possible network partition or localized configuration error. Based on this diagnostic result, operations personnel can focus on the network area where the problematic nodes reside to check for firewall rules that are incorrectly blocking NTP traffic, or for configuration errors on specific network devices that are preventing some nodes from properly accessing the NTP server. This allows operations personnel to quickly narrow down the scope of troubleshooting, saving significant time and effort and improving troubleshooting efficiency.

[0057] To further improve the accuracy and adaptability of fault diagnosis, the impact of the number of nodes in the multicast group on confidence calculation is considered. When the number of detected nodes is small, such as only one or two nodes, the confidence calculation rules are appropriately adjusted. For a single node (N=1), due to the lack of reference information from other nodes, confidence calculation is disabled. Instead, the node's anomaly weight is directly mapped to a specific fault type. For example, if the node fails the network connectivity test (weight 1), it is directly diagnosed as a network problem; if the port reachability test fails (weight 2), it is diagnosed as a port blocking problem; if the NTP service response test fails (weight 3), it is diagnosed as an NTP server problem. For the two-node case (N=2), to prevent single-node failures from being mistakenly diagnosed as global problems, the global failure threshold is raised to 50%. A global failure is only considered when both nodes experience the same severe fault type and the total anomaly weight is sufficiently high. For three or more nodes (N≥3), the original confidence calculation rules and decision thresholds are fully applied. This flexible adjustment mechanism enables the method of the present invention to maintain stable diagnostic performance and high accuracy under different network scales and node numbers, and fully adapt to various complex and changeable practical application scenarios.

[0058] After calculating the fault confidence and making a preliminary determination of the fault cause, the BMC device can automatically trigger predefined recovery actions based on the obtained fault confidence and the corresponding fault type, achieving closed-loop self-healing of the fault. For example, if a global NTP server failure is determined, the recovery action may include triggering a switch to a backup NTP server and promptly notifying the operations and maintenance team to inspect and fix the server-side issue. In the case of network partitions or local configuration errors, the recovery action may involve automatically refreshing firewall rules, such as executing the iptables -F command to clear existing firewall rules and then reloading the correct rule configuration, or restarting the BMC device's network interface to restore network connectivity. In the case of a single-node configuration error, the default NTP.conf configuration file can be restored and the NTP service process (such as the NTPd service) can be restarted to correct the incorrect configuration and restore NTP synchronization functionality. These predefined recovery actions are carefully designed based on common failure modes and operational and maintenance practices. They can quickly and effectively resolve problems, reduce the impact of failures on system operations, and improve overall system availability.

[0059] In one embodiment, Figure 2This specification also provides an NTP fault diagnosis device, which is applied to the BMC device of a server. The device includes: a first module, which is used to join an NTP fault diagnosis multicast group, where the members of the NTP fault diagnosis multicast group include the BMC devices of several servers within a preconfigured range; a second module, which is used to periodically broadcast the local NTP service status in the NTP fault diagnosis multicast group and receive the NTP service status broadcast by other members in the NTP fault diagnosis multicast group; and a third module, which is used to call a preset algorithm to calculate the fault confidence based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group, and perform NTP fault diagnosis based on the fault confidence.

[0060] In one embodiment, joining an NTP fault diagnosis multicast group, wherein members of the NTP fault diagnosis multicast group include BMC devices of several servers within a preconfigured range, includes: joining an NTP fault diagnosis multicast group associated with a TOR switch to which the NTP fault diagnosis multicast group belongs, wherein members of the NTP fault diagnosis multicast group include BMC devices of several servers to which the TOR switch belongs.

[0061] In one embodiment, the method of calculating the fault confidence level by calling a preset algorithm based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group, and performing NTP fault diagnosis based on the fault confidence level includes: obtaining a weight parameter pre-configured for each member in the NTP fault diagnosis multicast group, and calculating the fault confidence level based on the abnormal status indicated in the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group and the weight parameter associated with the member to which the abnormal status belongs.

[0062] In one embodiment, the NTP service status is used to indicate an abnormal state and a fault type of the abnormal state based on a probe detection result; the method of calculating a fault confidence based on the local NTP service status and the received NTP service status broadcast by other members of the NTP fault diagnosis multicast group and performing NTP fault diagnosis based on the fault confidence includes: calculating corresponding fault confidences according to different fault types of abnormal states, and judging whether a fault event corresponding to the fault type has occurred based on a comparison between the fault confidences corresponding to the fault types of each abnormal state and corresponding thresholds.

[0063] In one embodiment, in response to a determination that a fault event has occurred, a fault handling procedure is executed according to a fault type corresponding to the fault event.

[0064] like Figure 3In one implementation, BMCs within the same management domain form a distributed collaborative network, pre-defining logical management groups at the TOR (Top of Rack) switch level. For example, 24 X86 servers within a cabinet are interconnected via a 1Gbps out-of-band management network. During initialization, all BMCs automatically perform a multicast group join procedure: using the Linux system's iproute add command to bind the multicast address 239.255.200.1 to the management network interface eth0, and simultaneously calling the smcroute tool to register as a multicast member. The multicast period is strictly set to 60 seconds, a value verified by laboratory stress testing. When the node scale is expanded to 50, the multicast message size is kept within 512 bytes, and the network bandwidth utilization is less than 0.1%, avoiding impact on the service network. During each broadcast period, the BMC encapsulates local status using JSON structured data. Example formats include fields such as "node_ip", "ntp_status", "last_error", "local_tests", "ping", "port_123", and "ntp_service".

[0065] The key field local_tests contains the real-time results of the three-level probes: the ping field reflects ICMP layer connectivity (executed using the fping -C 3 -q 192.168.10.1 command, with a packet loss rate of <2% considered true); the port_123 field checks UDP port reachability using the nc-uzv 192.168.10.1 123 command; and the ntp_service field verifies NTP protocol interaction capabilities using ntpdate -q.

[0066] Probe execution uses short-circuit logic: if a ping test fails (returns false), subsequent detection is immediately terminated. The same applies to port test failures. This layered detection mechanism significantly reduces resource consumption, and the CPU usage of a single detection is extremely low.

[0067] Fault confidence calculation involves two stages: dynamic weight assignment and adaptive threshold determination. Consider a sudden failure in a 24-node cluster. When BMC-101 detects a failure in its own NTP synchronization, it first executes a local probe sequence. Assuming its ping test succeeds (weight = 1) but its port probe fails (weight = 2), application-layer detection is terminated, and the accumulated local anomaly weight is 2. Simultaneously, the BMC receives status data from the other 23 nodes via multicast. Data analysis reveals that another seven nodes report port probe failures (weight = 2) and 11 nodes report NTP service failures (weight = 3). The numerator, ∑ anomaly weight, = 2(BMC-101) + 7 × 2(port failure nodes) + 11 × 3(NTP failure nodes) = 45; the denominator, = number of nodes × maximum single-node weight = 24 × 3 = 72; the fault confidence = (45 / 72) × 100% = 62.5%. Based on the preset threshold rule (≥30% is considered a global failure), the system automatically diagnoses this as an NTP server crash. This algorithm is fault-tolerant in edge scenarios: if only two nodes are online, a special threshold policy must be enabled. For example, if node A has a weight of 3 (NTP failure) and node B has a weight of 1 (ping failure), the confidence level is (3+1) / (2×3)×100% = 66.7%. However, due to the N=2 rule, the global failure threshold is raised to 50%, which still triggers an NTP service failure determination.

[0068] Confidence-based diagnostic results drive predefined self-healing actions. For the aforementioned 62.5% confidence scenario, the system automatically executes the "global_ntp_failure" policy in the recovery_actions configuration: First, it switches to a backup NTP server through DNS round-robin (by calling sed -i 's / primary.ntp / backup.ntp / g' / etc / ntp.conf to modify the configuration), then restarts the service with systemctl restart ntpd. Simultaneously, a structured alert is sent to the operations and maintenance platform via the BMC's Redfish interface: {"event":"ntp_failover","new_server":"backup.ntp.org"}.

[0069] For a local failure with a confidence level of 12% (such as port detection failures on five nodes), the "network_partition" action chain is triggered: first, the iptables rules are refreshed (execute iptables -D INPUT -p udp --dport 123 -jDROP to remove the incorrect interception), then the network interface is restarted (ifdown eth0 &&ifup eth0), and the port status is rechecked after 30 seconds.

[0070] During the system's ongoing operation, a feedback optimization mechanism improves decision-making accuracy by accumulating knowledge. After each fault event is handled, the BMC records the tuple (fault type, action taken, repair result, actual root cause) in a local SQLite database. At the end of each month, an ensemble learning model (based on Scikit-learn's RandomForestClassifier) ​​is invoked to dynamically adjust probe weights.

[0071] To accommodate heterogeneous environments, the solution provides flexible switch parameters: an "NTP Autodiagnosis" option has been added to the BIOS setup interface (enabled by default). Operations and maintenance personnel can temporarily disable this feature using the IPMI command ipmitool raw 0x320xAA0x00. In ultra-large clusters, multiple multicast groups are supported per cabinet, with group addresses allocated in the 239.255.200.[1-255] range to prevent broadcast storms.

[0072] The fault-tolerance mechanism of the fault diagnosis algorithm incorporates a specially designed node status verification process. When BMC-101 receives a broadcast message from BMC-102, it first verifies the JSON signature (the HMAC-SHA256-based signature key is rotated every 24 hours) to prevent the injection of forged data. If a logical conflict between ntp_status and last_error is detected (e.g., a success status with an error code), the node's data is marked invalid and an alarm is logged. Furthermore, a time window mechanism is implemented to address network latency: only broadcast data within the last 70 seconds is collected (10 seconds of network jitter is tolerated and configurable), and data exceeding this limit is discarded. In extreme network outage scenarios (such as a TOR switch failure), the system automatically activates local fallback mode. If no multicast packets are received within 60 seconds, the system makes a decision based on its own weights (Weight 1: Network failure; Weight 2: Port failure; Weight 3: NTP failure).

[0073] Security protections are implemented at the self-healing execution layer. All configuration changes are executed in a separate namespace via unshare-n to avoid impacting the host network stack. A snapshot is automatically created before critical configuration changes are made, and a rollback is performed if NTP fails again after a reboot. For high-frequency failure scenarios (e.g., three similar failures within 24 hours), the system automatically upgrades its handling strategy—weighting historical failure nodes by 1.2x in confidence calculations and triggering hardware-level isolation for nodes that fail to self-heal.

[0074] In one embodiment, this specification provides an electronic device, including a processor and a readable storage medium, wherein the readable storage medium stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the aforementioned NTP fault diagnosis method. From a hardware perspective, the hardware architecture diagram can be found in Figure 4 shown.

[0075] In one embodiment, this specification provides a readable storage medium storing machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the aforementioned NTP fault diagnosis method.

[0076] Here, the readable storage medium can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), solid state drive, any type of storage disk (such as optical disk, DVD, etc.), or similar storage media, or a combination thereof.

[0077] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.

[0078] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0079] Those skilled in the art will appreciate that embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of a fully hardware implementation, a fully software implementation, or an implementation combining software and hardware. Furthermore, embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0080] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0081] Furthermore, these computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0083] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0084] The foregoing is merely an embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A method for diagnosing NTP faults, characterized in that: The method is applied to a BMC device of a server and includes: Join the NTP fault diagnosis multicast group, where members include BMC devices of several servers within a pre-configured range. Periodically broadcasts the local NTP service status in the NTP fault diagnosis multicast group and receives the NTP service status broadcast by other members of the NTP fault diagnosis multicast group. Based on the local NTP service status and the NTP service status broadcast by other members in the NTP fault diagnosis multicast group, a preset algorithm is called to calculate the fault confidence level, and NTP fault diagnosis is performed based on the fault confidence level.

2. The method according to claim 1, characterized in that The joining of the NTP fault diagnosis multicast group, wherein the members of the NTP fault diagnosis multicast group include BMC devices of several servers within a pre-configured range, includes: Join the NTP fault diagnosis multicast group associated with the TOR switch to which it belongs. The members of the NTP fault diagnosis multicast group include the BMC devices of several servers belonging to the TOR switch.

3. The method according to claim 1, characterized in that The method of calculating the fault confidence level based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group and performing NTP fault diagnosis based on the fault confidence level includes: Obtain the weight parameter pre-configured for each member in the NTP fault diagnosis multicast group, and calculate the fault confidence based on the abnormal status indicated in the NTP service status broadcast by other members in the NTP fault diagnosis multicast group and the weight parameter associated with the member in the abnormal status.

4. The method according to claim 1, wherein The NTP service status is used to indicate abnormal status and fault type of the abnormal status according to the probe detection result; The method of calculating the fault confidence level based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group and performing NTP fault diagnosis based on the fault confidence level includes: According to the fault types of different abnormal states, the corresponding fault confidence is calculated respectively. According to the comparison relationship between the fault confidence corresponding to the fault type of each abnormal state and the corresponding threshold, it is determined whether the fault event corresponding to the fault type has occurred.

5. The method according to claim 4, characterized in that In response to the determination result that a fault event has occurred, a fault handling procedure is executed according to the fault type corresponding to the fault event.

6. An NTP fault diagnosis device, characterized in that: A BMC device applied to a server includes: The first module is configured to join an NTP fault diagnosis multicast group, wherein members of the NTP fault diagnosis multicast group include BMC devices of several servers within a preconfigured range; The second module is used to periodically broadcast the local NTP service status in the NTP fault diagnosis multicast group and receive the NTP service status broadcast by other members in the NTP fault diagnosis multicast group; The third module is used to call a preset algorithm to calculate the fault confidence based on the local NTP service status and the NTP service status broadcast by other members in the received NTP fault diagnosis multicast group, and perform NTP fault diagnosis based on the fault confidence.

7. The device according to claim 6, characterized in that The joining of the NTP fault diagnosis multicast group, wherein the members of the NTP fault diagnosis multicast group include BMC devices of several servers within a pre-configured range, includes: Join the NTP fault diagnosis multicast group associated with the TOR switch to which it belongs. The members of the NTP fault diagnosis multicast group include the BMC devices of several servers belonging to the TOR switch.

8. The device according to claim 6, characterized in that The method of calculating the fault confidence level based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group and performing NTP fault diagnosis based on the fault confidence level includes: Obtain the weight parameter pre-configured for each member in the NTP fault diagnosis multicast group, and calculate the fault confidence based on the abnormal status indicated in the NTP service status broadcast by other members in the NTP fault diagnosis multicast group and the weight parameter associated with the member in the abnormal status.

9. The device according to claim 6, characterized in that The NTP service status is used to indicate abnormal status and fault type of the abnormal status according to the probe detection result; The method of calculating the fault confidence level based on the local NTP service status and the received NTP service status broadcast by other members in the NTP fault diagnosis multicast group and performing NTP fault diagnosis based on the fault confidence level includes: According to the fault types of different abnormal states, the corresponding fault confidence is calculated respectively. According to the comparison relationship between the fault confidence corresponding to the fault type of each abnormal state and the corresponding threshold, it is determined whether the fault event corresponding to the fault type has occurred.

10. The device according to claim 9, characterized in that In response to the determination result that a fault event has occurred, a fault handling procedure is executed according to the fault type corresponding to the fault event.

11. An electronic device, characterized in that: include: A processor and a readable storage medium, wherein the readable storage medium stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1 to 5.

12. A readable storage medium, characterized in that: The readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the method according to any one of claims 1 to 5.