A distributed real-time control system fault diagnosis method and system
Patent Information
- Application Number
- CN202610873282.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-17
AI Technical Summary
[0005]本发明的主要目的是提供一种分布式实时控制系统故障诊断方法及系统,旨在克服现有技术触发局限、维度单一以及缺乏确定性判定的缺陷,在不依赖大量历史故障样本且不引入额外物理传感器的前提下对已检测到异常的计算节点进行快速自动分层的故障类型判定
[0017]采用该技术方案具有以下优点:
Smart Images

Figure CN122420074B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fault detection and diagnosis technology for distributed computer systems, and particularly to a fault diagnosis method and system for distributed real-time control systems. Background Technology
[0002] Distributed real-time control systems consist of multiple interconnected computing nodes. These nodes exchange control commands and status data via real-time communication protocols, requiring hard real-time performance and high reliability. When a node in the system fails, maintenance personnel need to quickly determine the fault type to take targeted recovery measures. Existing fault detection and diagnosis technologies mainly include fault detection based on periodic probing, anomaly classification based on passive data acquisition and machine learning, active detection and diagnosis based on probabilistic reasoning, and statistical diagnosis for industrial control systems. Among these, fault detection schemes based on periodic probing dynamically adjust their probing strategies to detect network device status by periodically sending probing messages and receiving in-band telemetry data. However, their probing uses a periodic polling mechanism, which is out of sync with the actual time of the fault occurrence, resulting in significant detection delays. Furthermore, their probing targets are limited to network layer reachability, only outputting a binary state of reachability or unreachability, failing to distinguish specific fault types at the process layer, functional layer, or real-time performance layer. Anomaly classification schemes based on passive data acquisition and machine learning collect multi-dimensional operational data such as CPU metrics, logs, and call chains, extract feature vectors, and input them into anomaly recognition models to output CPU anomaly detection results and anomaly types. However, this scheme relies entirely on passively collected operational data. When nodes completely lose response and cannot report data, it will completely lose its diagnostic capabilities. Furthermore, machine learning models require a large amount of labeled training data, which is difficult to build in highly reliable real-time systems with scarce fault samples. In addition, it only covers CPU-related anomalies and fails to extend to system-level fault types.
[0003] The probabilistic reasoning-based active detection and diagnosis scheme proposes a two-layer detection architecture: offline construction of detection and localization detection sets, and online selection of the optimal detector in real time based on information theory criteria and probabilistic reasoning diagnosis using Bayesian networks. However, this scheme heavily relies on the prior probability distribution of the Bayesian network, which is difficult to accurately acquire and maintain in practical deployments. Furthermore, its detector selection optimization focuses on information gain, failing to consider the hierarchical semantics of the detection action itself in dimensions such as connectivity, process, function, and real-time performance. Additionally, this scheme is geared towards general information technology distributed systems and fails to adapt to the deterministic and low-intrusion requirements of real-time control scenarios. Statistical diagnostic schemes for industrial control systems perform distributed fault diagnosis through variable block partitioning and dynamic latent variable models, or fault localization through sensor arrays and graph topology propagation path analysis. However, the former, based on statistical deviation analysis of process variables, belongs to a passive diagnostic paradigm, while the latter heavily relies on physical sensors such as vibration, temperature, and current, and targets the controlled equipment rather than the control computer system itself. Neither of these schemes includes an active hierarchical detection mechanism for computer nodes. These existing technologies result in a disconnect between the triggering mechanism and the fault event, a single detection dimension, and a reliance on probabilistic reasoning or machine learning models, lacking deterministic judgment capabilities and failing to adapt to the timing constraints of real-time control systems.
[0004] Therefore, how to overcome the limitations of existing technologies in triggering faults, single dimensions, and lack of deterministic judgment, and how to quickly and automatically classify the fault types of computing nodes that have been detected as abnormal without relying on a large number of historical fault samples or introducing additional physical sensors, has become an urgent technical problem to be solved. Summary of the Invention
[0005] The main objective of this invention is to provide a fault diagnosis method and system for distributed real-time control systems, which aims to overcome the shortcomings of existing technologies such as limited triggering, single dimension, and lack of deterministic judgment. It can quickly and automatically classify the fault types of computing nodes that have detected anomalies without relying on a large number of historical fault samples or introducing additional physical sensors.
[0006] To achieve this objective, this invention proposes a fault diagnosis method for a distributed real-time control system, which is applied to a distributed real-time control system and includes the following steps: Operational status monitoring step: Monitoring agents deployed on each computing node passively collect the operational status data of the computing nodes, including heartbeat signals, response delays, error rates, and data quality indicators; Abnormal event triggering step: When the operational status data of any computing node exceeds a preset abnormal threshold, an abnormal event is generated, and the abnormal event serves as an immediate trigger condition for initiating an active detection process; Layered detection step: In response to the abnormal event, network connectivity detection, process activity detection, business function integrity detection, and control timing real-time verification are performed. The system follows a progressive hierarchical order, performing layered active probing on the computing nodes where an anomaly occurs. If any layer of active probing fails, the progression to subsequent progressive levels is terminated, and the probing response status of each layer is recorded. A state combination step involves combining and mapping the recorded probing response statuses of each layer to generate fault feature data. The probing response statuses include a pass status indicating probe accessibility, a failure status indicating probe blockage, a timeout status indicating timing overflow, and a non-execution status indicating short-circuit control. A fault determination step involves inputting the fault feature data into a preset deterministic fault rule base for matching, and directly outputting a deterministic fault type corresponding to the state combination of the fault feature data.
[0007] Preferably, the data quality indicators include at least one of message checksum consistency, field format validity, key field integrity, and timestamp continuity.
[0008] Preferably, before the layered detection step, a strategy adaptive matching step is further included: selecting a corresponding detection strategy from a preset detection strategy table according to the abnormal behavior type in the abnormal event; the detection strategy includes the detection layer range to be executed, the detection timeout parameters for each layer, and the detection message type.
[0009] Preferably, the preset detection strategy table includes the following differentiated mapping rules: when the abnormal behavior type is heartbeat timeout, the detection level range is configured to check all levels from the first to the fourth level, and each level uses the standard timeout parameter calculated according to the chain cascading constraint relationship; when the abnormal behavior type is response delay anomaly, the detection level range is configured to check all levels from the first to the fourth level, and the real-time constraint threshold of the fourth level is tightened to 50% of the control timing round-trip time base value recorded under the no-abnormal state; when the abnormal behavior type is increased error rate, the detection level range is configured to skip the first level network connectivity detection from the second to the fourth level, and perform a preset number of repeated detections in the third level; when the abnormal behavior type is decreased data quality, the detection level range is configured to skip the first level network connectivity detection and the second level process activity detection from the third to the fourth level, and carry a data verification payload for the detection message type in the third level.
[0010] Preferably, when the abnormal behavior type is an increased error rate, the preset number of repeated probes performed at the third level is two to three.
[0011] Preferably, the types of probe messages for each level of active probing are matched and limited according to the stack semantics of each level: the network connectivity probe corresponds to the first level, using Internet Control Message Protocol (ICP-IP) echo request messages and Transmission Control Protocol (TCP) synchronization messages; the process liveness probe corresponds to the second level, using application layer port liveness request messages; the service function integrity probe corresponds to the third level, using read-only and idempotent service simulation request messages; and the control timing real-time verification corresponds to the fourth level, using lightweight round-trip time (RTT) measurement messages.
[0012] Preferably, the timeout parameters for each level of active probing satisfy the following chain-like cascading constraint relationship: the timeout time for the first-level network connectivity probing. satisfy: ,in, This is the timeout period for the first-level network connectivity probe. The system's preset node heartbeat cycle, The control coefficient is dimensionless and its value range is Timeout for second-level process activity detection satisfy: ,in, This is the timeout period for the second-level process activity detection. The first magnification factor and its value range is Timeout period for third-level business function integrity detection satisfy: ,in, Timeout period for third-level business function integrity detection It is the second magnification factor and its value range is Real-time constraint threshold for fourth-level control timing real-time performance verification satisfy: ,in, The real-time constraint threshold is used for the real-time performance verification of the fourth-level control timing. This defines the basic control cycle of a distributed real-time control system that applies the aforementioned fault diagnosis method. It is a time-tolerance factor and its value range is .
[0013] Preferably, the layered advancement detection step specifically executes the following downward atomic-level advancement control: Initiating a network layer probe to the abnormal computing node to verify its network stack reachability; if the probe fails, terminating the advancement of all subsequent layers of probes and recording a first-level failure flag; after the network layer probe passes, sending a port liveness request to the target application layer port corresponding to the computing node; if the probe fails, terminating the advancement of all subsequent layers of probes and recording a second-level failure flag; after the process liveness probe passes, sending a read-only and idempotent simulated service request to the computing node; if no response in the expected format is received within the probe timeout period corresponding to the third level, terminating the advancement of subsequent layers of probes and recording a third-level failure flag; after the service function integrity probe passes, measuring the round-trip time from sending the probe request to receiving the response; if the round-trip time exceeds a preset real-time constraint threshold, recording a fourth-level timeout flag; otherwise, recording a fourth-level pass flag.
[0014] Preferably, the deterministic fault rule base includes the following basic judgment rules based on the deterministic mapping of the execution state across the entire space: When the fault feature data is a first-level failure and the second to fourth levels are in an unexecuted state, the fault judgment step directly outputs the deterministic fault type as a first fault type, which is a node downtime and network unreachable fault; when the fault feature data is a first-level pass, a second-level failure, and the third to fourth levels are in an unexecuted state, the fault judgment step directly outputs the deterministic fault type as a second fault type, which is a control process abnormal fault; when the fault feature data is a first-level pass, a second-level pass, a third-level failure, and a fourth-level unexecuted state, the fault judgment step directly outputs the deterministic fault type as an internal service function abnormality; when the fault feature data is a first to third level pass and a fourth level timeout state, the fault judgment step directly outputs the deterministic fault type as a real-time performance degradation. When the fault feature data shows that all four levels are in a pass state, the fault determination step outputs a determination result of no persistent fault. When the fault feature data shows that the first level is in an unexecuted state, the second level has failed, and the third and fourth levels are in an unexecuted state, the fault determination step directly outputs the deterministic fault type as a control process abnormality fault. When the fault feature data shows that the first level is in an unexecuted state, the second level has passed, the third level has failed, and the fourth level is in an unexecuted state, the fault determination step directly outputs the deterministic fault type as an internal service function abnormality. When the fault feature data shows that the first and second levels are in an unexecuted state, the third level has failed, and the fourth level is in an unexecuted state, the fault determination step directly outputs the deterministic fault type as an internal service function abnormality. When the skipped level in the fault feature data is in an unexecuted state, the executed third level is in a pass state, and the fourth level is in a timeout state, the fault determination step directly outputs the deterministic fault type as a real-time performance degradation.
[0015] Preferably, after the fault determination step, a diagnostic result output step is also included: outputting a diagnostic summary report containing the node identifier of the computing node that has an anomaly, the determined deterministic fault type, the execution status details of active detection at each level, and the handling suggestions determined according to a preset handling suggestion table, wherein the preset handling suggestion table includes a fault type field, a suggested operation field, a suggested inspection object field, and a log record field.
[0016] This application also discloses a fault diagnosis system for a distributed real-time control system, including a processor, a memory, a network communication interface, and the following modules implemented by the processor calling computer program instructions stored in the memory: a running status monitoring module, used for monitoring agents deployed on each computing node to passively collect the running status data of the computing nodes, the running status data including heartbeat signals, response delays, error rates, and data quality indicators; an abnormal event triggering module, used to generate an abnormal event when the running status data of any computing node exceeds a preset abnormal threshold, and use the abnormal event as an immediate trigger condition for starting an active detection process; and a hierarchical advancement detection module, used to respond to the abnormal event by detecting network connectivity, process activity, and business function integrity. The system employs a progressive hierarchical sequence for fault detection and real-time verification of control timing. For any abnormal computing node, layered active detection is performed. If any layer of active detection fails, the progression to subsequent progressive layers is terminated, and the detection response status of each layer is recorded. A state combination module is used to combine and map the recorded detection response statuses of each layer to generate fault feature data. The detection response statuses include a pass status indicating detection accessibility, a failure status indicating detection blockage, a timeout status indicating timing overflow, and a non-execution status indicating short-circuit control. A fault determination module is used to input the fault feature data into a preset deterministic fault rule base for matching and directly output a deterministic fault type corresponding to the state combination of the fault feature data.
[0017] The advantages of adopting this technical solution are as follows: It employs an event-driven mechanism for precise triggering, detecting abnormal events generated when operational status data exceeds preset abnormal thresholds, rather than relying on periodic polling. This closely links the detection timing with the fault occurrence time, avoiding the continuous load on the system from periodic detection and significantly reducing the startup delay for fault diagnosis. By constructing a four-layer progressive detection architecture—network connectivity detection, process activity detection, business function integrity detection, and control sequence real-time verification—it can meticulously distinguish various system-level fault types, such as node downtime, process crashes, internal service function anomalies, and real-time performance degradation, completely overcoming the limitation of traditional health checks that can only output a binary state of reachability or unreachability. During the judgment process, it makes decisions based on a combination of deterministic rules from the detection results at each layer, without relying on Bayesian prior probabilities, machine learning models, or large language models. This results in highly consistent, reproducible, and auditable judgment results, perfectly suited for high-reliability real-time control systems with stringent reliability requirements. All probe actions are standard network and application layer operations, initiated only against nodes that have exhibited anomalies. Normal nodes remain completely unaffected. Furthermore, probe messages employ read-only, idempotent, lightweight requests, neither modifying the target node's state nor interfering with normal real-time control operations, achieving non-intrusiveness to the real-time control system. In addition, the deterministic fault rule base is predefined based on system architecture knowledge and fault mode analysis, completely independent of training with a large number of historical fault samples, making it highly suitable for high-reliability systems with few fault cases. During execution, the system implements a downward atomic-level progression control logic. If any layer of active probe fails, the progression to subsequent progressive layers is immediately terminated. Combined with differentiated mapping rules in the pre-set probe strategy table, different anomaly types trigger probe strategies of varying depths and focuses, avoiding a one-size-fits-all approach of performing full-level probes for all anomaly scenarios. This significantly reduces unnecessary probe interference while maintaining diagnostic accuracy. Attached Figure Description
[0018] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a system architecture block diagram of a distributed real-time control system fault diagnosis system provided in an embodiment of the present invention.
[0019] Figure 2 The flowchart illustrates the overall process steps of a fault diagnosis method for a distributed real-time control system provided in this embodiment of the invention. Detailed Implementation
[0020] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. The following examples are used to illustrate the technical solutions of the present invention, so as to enable those skilled in the art to implement the corresponding fault diagnosis process according to the modules, steps, parameters and rules described.
[0021] In Example 1, this embodiment provides an adaptive fault type detection and diagnosis method for distributed real-time control computer systems, based on anomaly event-driven and hierarchical progressive detection. Distributed real-time control systems, such as plasma discharge control systems, motion control systems, or flight control systems, typically consist of multiple interconnected computing nodes. These nodes exchange control commands and status data via real-time communication protocols, exhibiting millisecond-level hard real-time control cycles and extremely high reliability requirements. When a computing node in the system fails, traditional solutions relying solely on periodic heartbeat polling not only suffer from severe detection delays but also only output a binary state of reachability, failing to accurately distinguish specific fault types such as control process collapse, internal functional abnormalities, or real-time performance degradation. Furthermore, relying entirely on passive machine learning models makes it difficult to construct effective classification models in high-reliability control scenarios with scarce fault samples, and completely loses diagnostic capabilities when nodes become unresponsive. To address these shortcomings in the background technology, this embodiment proposes an automated diagnostic closed-loop method triggered by anomaly signals, with active hierarchical verification and automatic fault type output, specifically executed by the distributed real-time control system 100. Figure 1 As shown, the distributed real-time control system 100 consists of multiple interconnected computing nodes 101, each of which is equipped with a corresponding software module to collaboratively achieve a diagnostic closed loop.
[0022] First, the operational status monitoring step, i.e., step S1, is executed. An operational status monitoring module 201 is deployed on each computing node 101 of the distributed real-time control system 100. The operational status monitoring module 201 specifically includes a monitoring agent 102, which continuously and passively collects operational status data from the computing nodes 101. This operational status data specifically includes heartbeat signals, response latency, error rate, and data quality indicators. The monitoring agent 102 obtains these status indicators by passively monitoring network protocol stack data and reading local control logs. During normal system operation, it does not actively send probe traffic to the network, thus ensuring zero intrusion and low overhead on control operations and ensuring the determinism of the control cycle. Specifically, the heartbeat signal serves as the first status indicator, representing the basic liveness status of the computing node 101; the response latency serves as the second status indicator, quantifying the timeliness of data interaction; the error rate serves as the third status indicator, statistically analyzing the probability of business execution failure; and the data quality indicator serves as the fourth status indicator, verifying the integrity and correctness of the control data. The data quality indicators include at least one of the following: message checksum consistency, field format validity, key field integrity, and timestamp continuity.
[0023] Next, the abnormal event triggering step, i.e., step S2, is executed. The abnormal event triggering module 202 generates an abnormal event when the running status data of any of the computing nodes 101 exceeds a preset abnormal threshold, and uses the abnormal event as an immediate trigger condition to start the active detection process. The preset abnormal threshold is stored in the system configuration file and includes a heartbeat timeout threshold, a response delay threshold, an error rate threshold, and a data quality threshold. The heartbeat timeout threshold is determined by the number of consecutive missing heartbeat cycles, the response delay threshold is determined by the maximum response time allowed by the system, the error rate threshold is determined by the ratio of the number of failed requests within a unit statistical window to the total number of requests, and the data quality threshold is determined by the ratio of the number of data verification failures within a unit statistical window to the total number of data frames. The abnormal event includes at least an abnormal node identifier and an abnormal behavior type, which includes at least one of heartbeat timeout, abnormal response delay, increased error rate, or decreased data quality. In this embodiment, the generated abnormal event is used as the start signal for the diagnostic process, thereby transforming the traditional diagnostic paradigm of passively waiting for data or blindly polling into an active diagnostic paradigm driven by precise abnormal signals. For example, in the actual operation scenario of a distributed real-time control system 100, the basic control cycle of the system is set to 1ms, and the normal heartbeat transmission interval of the nodes is 10ms. The monitoring agent 102 sets the heartbeat timeout threshold to 30ms, i.e., three consecutive missing heartbeat cycles; the set response delay threshold to 500μs; and the set error rate threshold to 5%. When a computing node 101 stops reporting its heartbeat signal due to a sudden cause at a certain moment, and the continuous missing time reaches 30ms, the monitoring agent 102 determines that the computing node 101 has experienced a heartbeat timeout anomaly, and generates an anomaly event containing the node's identifier and the heartbeat timeout behavior type, which serves as the start signal for the diagnostic process.
[0024] Before the layered detection, this embodiment also includes a strategy adaptive matching step, namely step S3, where the diagnostic system automatically selects the corresponding differentiated detection strategy from the preset detection strategy table 104 based on the specific abnormal behavior type in the abnormal event. The detection strategy includes the detection level range to be executed, the timeout parameters for each level of detection, and the detection message type used. The preset detection strategy table 104 pre-constructs multiple differentiated mapping rules. The first mapping rule is that when the abnormal behavior type is heartbeat timeout, the detection level range is configured to be a full-level investigation from the first to the fourth level, and each level uses standard timeout parameters; the standard timeout parameters are the detection timeout parameters for each level calculated according to the chain cascading constraint relationship. The second mapping rule is that when the anomaly type is response delay anomaly, the detection level range is configured to check all levels from the first to the fourth level, but with emphasis on the real-time verification of control timing at the fourth level. The real-time constraint threshold of the fourth level needs to be tightened to 50% of the normal timing baseline value. The normal timing baseline value is the control timing round-trip time baseline value recorded by the system in an anomaly-free state. The third mapping rule is that when the anomaly type is increased error rate, the detection level range is configured to skip the first level connectivity detection from the second to the fourth level, focusing on the integrity of business functions, and performing a preset number of repeated detections at the third level, with the preset number of repeated detections being two to three. The fourth mapping rule is that when the anomaly type is decreased data quality, the detection level range is configured to skip the first level connectivity detection and the second level process activity detection from the third to the fourth level, directly targeting the integrity of business functions and the real-time performance of control timing, and additionally carrying data verification payload for the detection message type at the third level. The data verification payload includes a preset verification field, a payload sequence number, a timestamp field, and a checksum field, used to request the target computing node 101 to return a verification result corresponding to the data verification payload. This adaptive matching mechanism enables different anomaly types to trigger detection strategies with varying depths and focuses, avoiding a one-size-fits-all approach of performing full-level detection on all anomaly scenarios. This significantly reduces unnecessary detection actions that interfere with the system while ensuring diagnostic accuracy.
[0025] To ensure that the detection process meets the strong timing constraints of the distributed real-time control system, the detection timeout parameters for each level of active detection in the preset detection strategy table 104 strictly satisfy a set of chain-like cascading constraints. Specifically, the timeout for the first-level network connectivity detection can be expressed by the formula: ,in, This is the timeout period for the first-level network connectivity probe. The system's preset node heartbeat cycle, The control coefficient is dimensionless and its range is specified as follows: For example, when the system's heartbeat cycle is 10ms and the control coefficient is set to 20, the timeout time for the first level is... The calculated timeout is 200ms. The timeout for the second-level process activity detection can be expressed by the formula: ,in, This is the timeout period for the second-level process activity detection. The first magnification factor is defined as having a value range of 1. For example, when the first magnification factor is selected as 2.5, the timeout time of the second level... The calculated timeout is 500ms. The timeout for the third-level business function integrity probe can be expressed by the formula: ,in, Timeout period for third-level business function integrity detection The second magnification factor is defined as follows: For example, when the second magnification factor is selected as 4, the timeout time of the third level... The calculated time is 2000ms. The real-time constraint threshold for the fourth-level control timing real-time performance verification can be expressed by the formula: ,in, The real-time constraint threshold is used for the real-time performance verification of the fourth-level control timing. This defines the basic control cycle of a distributed real-time control system that applies the aforementioned fault diagnosis method. The time tolerance factor is defined as follows: For example, when the system's basic control cycle is 1ms and the timing tolerance factor is set to 1, the real-time constraint threshold of the fourth level... The specific timeout value is set to 1ms. These timeout parameters and cascading constraints can be configured and adjusted according to the actual communication latency and control cycle of the specific system being diagnosed.
[0026] Subsequently, a layered detection step, namely step S4, is executed. In response to the abnormal event, the layered detection module 203 initiates precise layered active detection on the abnormal computing node 101 according to the progressive order of network connectivity detection, process activity detection, service function integrity detection, and control timing real-time verification. The detection message types for each layer of active detection are strictly matched and limited according to the stack semantics of each layer. Specifically, the network connectivity detection corresponds to the first layer, using Internet Control Message Protocol (ICP-A) echo request messages and Transmission Control Protocol (TCP) synchronization messages; the process activity detection corresponds to the second layer, using application layer port activation request messages; the service function integrity detection corresponds to the third layer, using read-only and idempotent service simulation request messages; and the control timing real-time verification corresponds to the fourth layer, using lightweight round-trip time (RTT) measurement messages. The service simulation request message includes a node identifier field, a request sequence number field, a read-only service type field, a verification payload field, and a desired response format field; the lightweight RTT measurement message includes a send timestamp field, a receive timestamp field, and a response sequence number field. The expected response format includes at least a response sequence number consistent with the request sequence number, an execution result code, a response data field, and a response verification field. During the execution process, the system implements a downward atomic-level progression control logic. That is, if any layer's active probe fails, the progression to subsequent progressive levels is terminated, and the probe response status of each layer's active probe is recorded in detail. The specific progression control process is as follows: The system initiates a first-level network layer probe to the abnormal computing node 101 to verify the reachability of its operating system network stack. If no correct response message is received within the set timeout period, the first-level probe is deemed a failure, and the system terminates all subsequent layer probes and records the first-level failure flag. If the first-level network layer probe passes, the system sends a port activation request to the target application layer port corresponding to the computing node 101 after the network layer probe passes, performing a second-level process activity probe. If the port does not respond within the corresponding timeout period, the second-level probe is deemed a failure, and the system terminates all subsequent layer probes and records the second-level failure flag. After the process liveness detection passes, the system then initiates a third-level business function integrity detection to the computing node 101, sending a read-only and idempotent simulated business request. If no response in the expected format is received within the corresponding timeout period, or if abnormal data is returned, the third-level detection is deemed a failure, and the subsequent level detection is terminated, with a third-level failure flag recorded. Only after the business function integrity detection passes will the system finally proceed to the fourth-level control timing real-time verification, measuring the round-trip time from sending the detection request to receiving the response. If the round-trip time exceeds a preset real-time constraint threshold, the system records a fourth-level timeout flag; otherwise, the system records a fourth-level pass flag.
[0027] After all permitted detection actions are completed, the state combination step, i.e., step S5, is executed. The state combination module 204 combines and maps the detection response states of each level of active detection recorded earlier to generate fault feature data. The detection response states include a pass state representing detection accessibility, a failure state representing detection blocking, a timeout state representing timing overflow, and a non-execution state representing short-circuit control. Through this state combination, discrete hierarchical detection results can be summarized into a structured fault feature data, providing clear data input for subsequent deterministic logic matching.
[0028] Next, the fault determination step, namely step S6, is executed. The fault determination module 205 inputs the generated fault feature data into the preset deterministic fault rule base 108 for matching, thereby directly outputting the deterministic fault type corresponding to the state combination of the fault feature data. The deterministic fault rule base 108 is constructed based on the deterministic mapping of the execution state global space. It takes the deterministic combination of the detection results of each layer as a premise and the fault type as a conclusion. It does not rely on Bayesian prior probability reasoning or complex machine learning training models at all. It has deterministic judgment capability, and the judgment results are highly consistent, reproducible, and auditable. The basic rules of the deterministic fault rule base 108 specifically include the following deterministic mapping rules: The first basic rule is that when the fault feature data is a first-level failure and the second to fourth levels are in an unexecuted state, it indicates that the network stack of the target computing node 101 is unreachable at the physical layer or network layer. The fault determination module 205 directly outputs the deterministic fault type as the first fault type, which is a node downtime and network unreachability fault. The second basic rule states that when the fault characteristic data shows that the first level is passed, the second level fails, and the third and fourth levels are in an unexecuted state, it indicates that the operating system and network protocol stack of the target computing node 101 are running properly, but the port where the target control service is located is inaccessible. The fault determination module 205 directly outputs the deterministic fault type as the second fault type, which is a control process abnormality fault. The third basic rule states that when the fault characteristic data shows that the first level is passed, the second level is passed, the third level fails, and the fourth level is in an unexecuted state, it indicates that although the control process of the target computing node 101 is online and alive, it cannot correctly execute business logic, respond to simulated requests, or return an incorrect response format. The fault determination module 205 directly outputs the deterministic fault type as an internal service function abnormality. The fourth basic rule states that when the fault characteristic data shows that all levels from the first to the third are in a passable state and the fourth level is in a timeout state, it indicates that although the various control functions of the target computing node 101 are sound, its response round-trip time has exceeded the control cycle constraint of hard real-time. The fault determination module 205 directly outputs the deterministic fault type as real-time performance degradation. The fifth basic rule states that when the fault characteristic data shows that all levels from the first to the fourth are in a passable state, it indicates that no persistent blockage or delay was found after the target computing node 101 underwent layered progressive verification. The fault determination module 205 outputs the determination result as no persistent fault. The sixth basic rule states that when the detection strategy corresponding to the increased error rate skips the first-level network connectivity detection, and the fault characteristic data shows that the first level was not executed, the second level failed, and the third to fourth levels were not executed, the fault determination module 205 directly outputs the deterministic fault type as a control process abnormality fault.The 7th basic rule is that when the detection strategy corresponding to an increased error rate skips the first-level network connectivity detection, and the fault characteristic data shows that the first level was not executed, the second level passed, the third level failed, and the fourth level was not executed, the fault determination module 205 directly outputs the deterministic fault type as "service internal function anomaly". The 8th basic rule is that when the detection strategy corresponding to a decrease in data quality skips the first-level network connectivity detection and the second-level process activity detection, and the fault characteristic data shows that the first and second levels were not executed, the third level failed, and the fourth level was not executed, the fault determination module 205 directly outputs the deterministic fault type as "service internal function anomaly". The 9th basic rule is that when the skipped level in the fault characteristic data is in an "not executed" state, the executed third level is in a "passed" state, and the fourth level is in a "timeout" state, the fault determination module 205 directly outputs the deterministic fault type as "real-time performance degradation". The deterministic fault rule base 108 also includes a mapping relationship between rule number, state combination condition, deterministic fault type and handling suggestion number, which is used to match fault feature data with handling suggestions in the diagnostic summary report.
[0029] Finally, the diagnostic result output step, i.e., step S7, is executed. After the fault determination step, the distributed real-time control system 100 executes the diagnostic result output step through the diagnostic result output module 206, directly outputting a diagnostic summary report containing the node identifier of the abnormal computing node 101, the determined deterministic fault type, the execution status details of active detection at each level, and the handling suggestions determined according to the preset handling suggestion table. The preset handling suggestion table includes a fault type field, a suggested operation field, a suggested inspection object field, and a log record field; wherein, the first fault type corresponds to checking power supply, network link, and node reachability; the second fault type corresponds to checking control process status and service startup status; service internal function abnormality corresponds to checking business logic logs and simulated request response format; and real-time performance degradation corresponds to checking processor load, real-time thread scheduling status, and network latency status. This automated diagnostic closed loop, triggered by an abnormal signal, undergoes active layered verification, and finally automatically outputs the fault type, can be completed within milliseconds to seconds. The detection actions are all standard network and application layer operations, initiated to the nodes that have experienced abnormalities, while normal nodes are unaffected. Meanwhile, the probe messages are read-only, idempotent lightweight requests that do not modify the target node state or interfere with normal real-time control operations, thus improving the autonomous operation and maintenance efficiency of the distributed real-time control system 100 in complex industrial control scenarios.
[0030] In Example 2, based on Example 1, this example further expands and supplements the strategy adaptive matching mechanism, data indicator system, and differentiated detection process under different anomaly types in the fault detection and diagnosis method. In the distributed real-time control system 100, the operating status data continuously collected by the monitoring agent 102 is refined into four operating status indicators: a first status indicator, a second status indicator, a third status indicator, and a fourth status indicator. Specifically, the first status indicator is a heartbeat signal, the second status indicator is response latency, the third status indicator is error rate, and the fourth status indicator is a data quality indicator. When any status indicator of the computing node 101 exceeds a preset anomaly threshold, the process is triggered. Depending on the anomaly type, the strategy adaptive matching step dynamically extracts the detection strategy that best suits the current network and application status from the preset detection strategy table 104 to avoid the additional load caused by blind, full-level detection.
[0031] Specifically, when the abnormal behavior type is heartbeat timeout, it means that the target computing node 101 may have experienced a failure at any level, ranging from the lowest-level physical power outage and network interruption to the upper-level process deadlock. In this case, the preset detection strategy table 104 guides the system to configure the detection level range to a full-level investigation from the first to the fourth level, with each level using standard timeout parameters. If an actual failure occurs, such as a control center node process crash, the monitoring agent 102 detects that the heartbeat signal of computing node 101 stops reporting at a certain moment, and after three consecutive missing heartbeat cycles (30ms), it determines that the heartbeat is abnormal and generates an abnormal event. The system adaptively matches the full-level detection strategy and then executes layered detection sequentially. First, an Internet Control Message Protocol (ICP-IP) echo request message is sent to the node. If its operating system network stack responds normally, the first level is considered successful. Then, a liveness request message is sent to the node's control process port. If the port does not respond, the second level is considered a failure. At this point, the system triggers the short-circuit mechanism in the downward atomic level propagation control, terminates the detection propagation of all subsequent levels, records the second level failure flag, and finally inputs the fault feature data into the rule base to determine the second fault type.
[0032] When the abnormal behavior type is response latency anomaly, it indicates that the basic network layer and process layer of the target computing node 101 still maintain a basic accessibility state; otherwise, the system would not be able to collect quantified latency values. However, high latency indicates a higher risk of timing timeout at the fourth level. Therefore, although the system's matching detection strategy still configures the detection level range to check all levels from the first to the fourth level, it introduces a timing-focused mechanism, tightening the real-time constraint threshold in the control timing real-time performance verification of the fourth level to 50% of the normal timing baseline value. Through this tightening, the system can assess the degree of latency degradation of the node with higher sensitivity. For example, in a real-world scenario where the data acquisition node experiences real-time performance degradation, the monitoring agent 102 detects that the response latency of the data acquisition node to control commands gradually increases from the normal 50μs to 800μs, exceeding the preset 500μs latency threshold, thereby generating a latency anomaly event. After the system passes the first-level network detection, the second-level process liveness detection, and the third-level simulated data query, it performs a real-world test at the fourth level using lightweight round-trip time measurement messages. If the measured round-trip time is 1.2ms, exceeding the real-time constraint threshold of 1ms (corresponding to one control cycle), a fourth-level timeout flag is directly recorded. Finally, the rule base outputs a deterministic fault type of real-time performance degradation and prompts operations personnel to check the CPU load, real-time thread scheduling status, or whether there is any business thread pause caused by garbage collection.
[0033] When the abnormal behavior is characterized by an increased error rate, it indicates that the process of the target computing node 101 is online and alive, and its network reachability has been indirectly verified by the existence of passively received data. Specifically, a preset number of repeated probes are performed at the third level, namely two to three. By continuously initiating two to three read-only and idempotent business simulation requests, the system can distinguish whether the current error of the computing node 101 is an occasional error caused by minor network disturbances, or a persistent failure caused by code defects or internal resource exhaustion, thereby improving the robustness of the diagnostic conclusions.
[0034] When the anomaly is described as data quality degradation, it means that while the underlying communication of computing node 101 is not only unimpeded and its control process can respond to routine liveness probes, the content of the business control data it sends has been corrupted, failed verification, or is formatted incorrectly. To verify the correctness of its data processing pipeline and memory data, the system carries a data verification payload in the business simulation request message initiated at the third layer. This requires the target computing node 101 to perform algorithmic calculations on this specific payload and return the verification result, thereby verifying the integrity of its internal business logic functions.
[0035] In Example 3, this example provides a more quantitative and formulaic in-depth expansion of the cascading constraints of detection semantics, message matching, and timeout parameters mentioned in Examples 1 and 2. In the distributed real-time control system 100, if the timeout time for each layer of detection is set in an isolated static manner, it is highly susceptible to false alarms caused by network transient fluctuations or uneven load on computing nodes 101, or the timeout may be too long, violating the hard real-time timing constraints of the distributed control system. Therefore, in this example, the timeout parameters for each layer of active detection strictly satisfy a set of chain-like cascading constraints, ensuring that the detection window exhibits a deterministic, step-like expansion on the time axis.
[0036] Timeout for Level 1 network connectivity probe It must be in sync with the heartbeat cycle of the underlying nodes in the system. Strong binding, their cascading relationship satisfies the formula ,in, This is the timeout period for the first-level network connectivity probe. The system's preset node heartbeat cycle, The control coefficient is dimensionless, and its value range is strictly limited to... Between. Control coefficient Its purpose is to reserve sufficient statistical safety margin for instantaneous congestion at the network level. This is based on the system's preset node heartbeat cycle. Specifically, at 10ms, if the control coefficient is selected... Specifically, it is 20, which is the timeout period for the first level. The value is set at 200ms.
[0037] Timeout for second-level process activity detection Cascade amplification is performed based on the first-level timeout time, and the cascade relationship satisfies the formula. ,in, This is the timeout period for the second-level process activity detection. This is the timeout period for the first-level network connectivity probe. This is the first magnification factor, and its value range is strictly limited to... Between. When the application layer port responds to a probe request, it needs to switch from the operating system kernel network stack to the application layer user-mode process, which inevitably takes longer than the response time of a pure network stack. If the first amplification factor If we select 2.5 specifically, then the timeout time for the second level... The calculation yielded a specific time of 500ms.
[0038] Timeout period for third-level business function integrity detection Based on the second-level timeout, a deeper level of cascaded amplification is performed, and the cascade relationship satisfies the formula. ,in, Timeout period for third-level business function integrity detection This is the timeout period for the second-level process activity detection. This is the second magnification factor, and its value range is strictly limited to... Between. Because the third layer uses read-only and idempotent business simulation request messages, after receiving the message, the target computing node 101's control process must actually call the internal business logic code, and may even need to read local registers, memory, or execute lightweight algorithm processing, thus lengthening its processing cycle axially. If the second amplification factor If we select 4 specifically, then the timeout time for the third level... The calculation yielded a specific time of 2000ms.
[0039] Real-time constraint threshold for fourth-level control timing real-time performance verification Based directly on the basic control cycle of the distributed control system The relationship is set to satisfy the formula. ,in, The real-time constraint threshold is used for the real-time performance verification of the fourth-level control timing. This is the basic control cycle of the system. It is a time-tolerance factor and its value range is strictly limited to 1 / 2. In hard real-time control scenarios, if the round-trip time of a node's service exceeds several control cycles, it will directly cause instability in the control loop. When the system's basic control cycle... Specifically, at 1ms, if the timing tolerance factor If we select 1 specifically, then the real-time constraint threshold for the fourth level... It is strictly limited to 1ms.
[0040] Through the formulaic chain-cascade constraints, the entire hierarchical propulsion detection structure has clear basic logic support for each time it advances to the next atomic level, thus ensuring the timing determinism of the fault diagnosis system at the algorithm level and adapting to hard real-time control environments with microsecond to millisecond control cycles.
[0041] In Example 4, this example provides a distributed real-time control system fault diagnosis system that completely corresponds to Examples 1 to 3, used to implement the adaptive detection and diagnosis process. For example... Figure 1As shown, the fault diagnosis system of the distributed real-time control system specifically includes an operation status monitoring module 201, an abnormal event triggering module 202, a hierarchical advancement detection module 203, a status combination module 204, a fault determination module 205, and a diagnosis result output module 206.
[0042] The operation status monitoring module 201 specifically includes a monitoring agent 102 deployed on each computing node 101, which is used to passively collect the operation status data of the computing node 101. The operation status data specifically includes heartbeat signals, response latency, error rate and data quality indicators.
[0043] The abnormal event triggering module 202 is connected to the operation status monitoring module 201. It is used to generate an abnormal event immediately when the operation status data of any of the computing nodes 101 exceeds a preset abnormal threshold, and to use the abnormal event as an immediate trigger condition for starting the active detection process. The abnormal event triggering module 202 is internally configured with a dedicated policy adaptive matching unit, which is used to select the corresponding detection policy from the preset detection policy table 104 according to the abnormal behavior type in the abnormal event, and to determine the detection level range to be executed, the detection timeout parameters of each level, and the detection message type.
[0044] The hierarchical advancement detection module 203 is connected to the abnormal event triggering module 202. In response to the abnormal event, it performs hierarchical active detection on the abnormal computing node 101 according to a progressive hierarchical order: network connectivity detection, process activity detection, service function integrity detection, and control timing real-time verification. The hierarchical advancement detection module 203 is internally equipped with an atomic advancement control unit. If any layer of active detection fails, it executes short-circuit control, immediately terminating the advancement to subsequent progressive layers, and completely records the detection response status of each layer of active detection.
[0045] The state combination module 204 is connected to the hierarchical advancement detection module 203 and is used to combine and map the detection response states of each level of active detection to generate fault feature data that includes the pass state representing detection accessibility, the failure state representing detection blockage, the timeout state representing timing overflow, and the non-execution state representing short-circuit control.
[0046] The fault determination module 205 is connected to the state combination module 204 and is used to input the fault feature data into a preset deterministic fault rule base 108 for matching, and directly output the deterministic fault type corresponding to the state combination of the fault feature data. The deterministic basic mapping rules are internally embedded in the deterministic fault rule base 108.
[0047] After outputting the deterministic fault type, the fault determination module 205 will also call the internal diagnostic result output module 206 to execute the diagnostic result output step. Finally, it will output a diagnostic summary report containing the node identifier of the abnormal computing node 101, the determined deterministic fault type, details of the execution status of active detection at each level, and handling suggestions determined according to a preset handling suggestion table. The diagnostic summary report is sent to the main control and maintenance terminal via the network or written to the system's non-volatile security log memory for post-event auditing and review by maintenance personnel. Physically, each module is driven by the microprocessor inside the computing node 101 calling computer program instructions stored in memory to implement the hardware communication interface, thereby constructing a high-performance, highly deterministic automated fault diagnosis hardware closed loop at the hardware level.
[0048] In this embodiment, the various modules, method steps, and parameter ranges are combined according to the system configuration of the distributed real-time control scenario. The microprocessor, memory, network communication interface, monitoring agent 102, preset detection strategy table 104, and deterministic fault rule base 108 of computing node 101 jointly complete the generation of abnormal events, selection of detection strategies, hierarchical active detection, combination of fault feature data, determination of deterministic faults, and output of diagnostic summary reports.
Claims
1. A fault diagnosis method for a distributed real-time control system, which is applied to a distributed real-time control system, characterized in that, Includes the following steps: Operational status monitoring steps: The monitoring agent deployed on each computing node passively collects the operational status data of the computing node, including heartbeat signals, response latency, error rate and data quality indicators; Abnormal event triggering steps: When the operational status data of any computing node exceeds a preset abnormal threshold, an abnormal event is generated, and the abnormal event is used as an immediate triggering condition to start the active detection process. Layered Probing Steps: In response to the abnormal event, layered active probing is performed on the abnormal computing node in a progressive order of network connectivity detection, process activity detection, service function integrity detection, and control timing real-time verification. If any layer of active probing fails, the progression to subsequent progressive layers is terminated, and the detection response status of each layer is recorded. State Combination Steps: The recorded detection response statuses of each layer are combined and mapped to generate fault characteristic data. The detection response statuses include a pass status indicating probe accessibility, a failure status indicating probe blockage, a timeout status indicating timing overflow, and a non-execution status indicating short-circuit control. Fault determination steps: Input the fault feature data into a preset deterministic fault rule base for matching, and directly output the deterministic fault type corresponding to the state combination of the fault feature data.
2. The fault diagnosis method for a distributed real-time control system according to claim 1, characterized in that, The data quality indicators include at least one of the following: message checksum consistency, field format validity, key field integrity, and timestamp continuity.
3. The fault diagnosis method for a distributed real-time control system according to claim 1, characterized in that, Before the layered detection step, a strategy adaptive matching step is also included: selecting the corresponding detection strategy from the preset detection strategy table according to the abnormal behavior type in the abnormal event; the detection strategy includes the detection layer range to be executed, the detection timeout parameters of each layer, and the detection message type.
4. The fault diagnosis method for a distributed real-time control system according to claim 3, characterized in that, The preset detection strategy table includes the following differentiated mapping rules: When the abnormal behavior type is heartbeat timeout, the detection level range is configured to check all levels from the first to the fourth level, and each level uses the standard timeout parameter calculated according to the chain cascading constraint relationship; when the abnormal behavior type is response delay anomaly, the detection level range is configured to check all levels from the first to the fourth level, and the real-time constraint threshold of the fourth level is tightened to 50% of the control timing round-trip time base value recorded under the no-abnormal state; when the abnormal behavior type is increased error rate, the detection level range is configured to skip the first level network connectivity detection from the second to the fourth level, and perform a preset number of repeated detections in the third level; when the abnormal behavior type is data quality degradation, the detection level range is configured to skip the first level network connectivity detection and the second level process activity detection from the third to the fourth level, and carry data verification payload for the detection message type in the third level.
5. The fault diagnosis method for a distributed real-time control system according to claim 4, characterized in that, When the abnormal behavior type is an increase in error rate, the preset number of repeated probes performed at the third level is two to three.
6. The fault diagnosis method for a distributed real-time control system according to claim 1, characterized in that, The types of probe messages for each active probing layer are matched and limited according to the stack semantics of each layer: the network connectivity probe corresponds to the first layer, using Internet Control Message Protocol (ICP-IP) echo request messages and Transmission Control Protocol (TCP) synchronization messages; the process liveness probe corresponds to the second layer, using application layer port liveness request messages; the service function integrity probe corresponds to the third layer, using read-only and idempotent service simulation request messages; and the control timing real-time verification corresponds to the fourth layer, using lightweight round-trip time (RTT) measurement messages.
7. The fault diagnosis method for a distributed real-time control system according to claim 1, characterized in that, The timeout parameters for active probing at each level satisfy the following chain-like cascading constraint: the timeout period for the first-level network connectivity probe. satisfy: ,in, This is the timeout period for the first-level network connectivity probe. The system's preset node heartbeat cycle, The control coefficient is dimensionless and its value range is Timeout for second-level process activity detection satisfy: ,in, This is the timeout period for the second-level process activity detection. The first magnification factor and its value range is Timeout period for third-level business function integrity detection satisfy: ,in, Timeout period for third-level business function integrity detection It is the second magnification factor and its value range is Real-time constraint threshold for fourth-level control timing real-time performance verification satisfy: ,in, The real-time constraint threshold is used for the real-time performance verification of the fourth-level control timing. This defines the basic control cycle of a distributed real-time control system that applies the aforementioned fault diagnosis method. It is a time-tolerance factor and its value range is .
8. The fault diagnosis method for a distributed real-time control system according to claim 1, characterized in that, The layered advancement detection step specifically executes the following downward atomic level advancement control: initiate network layer detection to the computing node that has an anomaly to verify its network stack reachability; if the detection fails, terminate the advancement of detection at all subsequent levels and record the first level failure flag. After the network layer probe is successful, a port liveness probe request is sent to the target application layer port corresponding to the computing node. If the probe fails, the probe progress of all subsequent layers is terminated and a second-level failure flag is recorded. After the process liveness probe is successful, a read-only and idempotent simulated service request is sent to the computing node. If no response in the expected format is received within the probe timeout period corresponding to the third level, the probe progress of subsequent layers is terminated and a third-level failure flag is recorded. After the business function integrity detection is passed, the round-trip time from sending the detection request to receiving the response is measured. If the round-trip time exceeds the preset real-time constraint threshold, a fourth-level timeout flag is recorded; otherwise, a fourth-level pass flag is recorded.
9. The fault diagnosis method for a distributed real-time control system according to claim 1, characterized in that, The deterministic fault rule base includes the following basic judgment rules based on the deterministic mapping of the execution state across the entire space: When the fault feature data is a first-level failure and the second to fourth levels are in an unexecuted state, the fault judgment step directly outputs the deterministic fault type as the first fault type, which is a node crash and network unreachable fault; when the fault feature data is a first-level pass, a second-level failure, and the third to fourth levels are in an unexecuted state, the fault judgment step directly outputs the deterministic fault type as the second fault type, which is a control process abnormal fault; when the fault feature data is a first-level pass, a second-level pass, a third-level failure, and a fourth-level unexecuted state, the fault judgment step directly outputs the deterministic fault type as an internal service function abnormality. When the fault characteristic data shows that all three levels are in a pass state and the fourth level is in a timeout state, the fault determination step directly outputs the deterministic fault type as real-time performance degradation; when the fault characteristic data shows that all three levels are in a pass state, the fault determination step outputs the determination result as no persistent fault; when the fault characteristic data shows that the first level is in an inactive state, the second level has failed, and the third and fourth levels are in an inactive state, the fault determination step directly outputs the deterministic fault type as control process abnormality fault; when the fault characteristic data shows that the first level is in an inactive state, the second level is pass, the third level has failed, and the fourth level is in an inactive state, the fault determination step directly outputs the deterministic fault type as service internal function abnormality. When the first and second levels of the fault feature data are in an unexecuted state, the third level is in a failed state, and the fourth level is in an unexecuted state, the fault determination step directly outputs the deterministic fault type as service internal function abnormality; when the skipped level in the fault feature data is in an unexecuted state, the executed third level is in a passed state, and the fourth level is in a timeout state, the fault determination step directly outputs the deterministic fault type as real-time performance degradation.
10. The fault diagnosis method for a distributed real-time control system according to claim 1, characterized in that, Following the fault determination step, a diagnostic result output step is also included: outputting a diagnostic summary report containing the node identifier of the computing node that experienced the anomaly, the determined deterministic fault type, the execution status details of active detection at each level, and the handling suggestions determined according to a preset handling suggestion table. The preset handling suggestion table includes a fault type field, a suggested operation field, a suggested inspection object field, and a log record field.
11. A fault diagnosis system for a distributed real-time control system, characterized in that, The system includes a processor, a memory, a network communication interface, and the following modules implemented by the processor calling computer program instructions stored in the memory: a running status monitoring module, used for monitoring agents deployed on each computing node to passively collect the running status data of the computing nodes, the running status data including heartbeat signals, response latency, error rate, and data quality indicators; and an abnormal event triggering module, used to generate an abnormal event when the running status data of any computing node exceeds a preset abnormal threshold, and to use the abnormal event as an immediate trigger condition for starting an active detection process. The hierarchical detection module is used to respond to the abnormal event by performing hierarchical active detection on the abnormal computing node in a progressive order of network connectivity detection, process activity detection, service function integrity detection, and control timing real-time verification. If any level of active detection fails, the progression to subsequent progressive levels is terminated, and the detection response status of each level is recorded. The state combination module is used to combine and map the recorded detection response statuses of each level to generate fault characteristic data. The detection response status includes a pass status indicating detection accessibility, a failure status indicating detection blockage, a timeout status indicating timing overflow, and a non-execution status indicating short-circuit control. The fault determination module is used to input the fault feature data into a preset deterministic fault rule base for matching, and directly output the deterministic fault type corresponding to the state combination of the fault feature data.
Citation Information
Patent Citations
HBase client main and standby switching method and system based on fault perception
CN119537484A
Full-link dynamic monitoring and abnormity simulation test system and method
CN120743680A