Maintenance troubleshooting method for complex system

By combining Bayesian networks and double-Q learning, a causal relationship model is constructed and maintenance decisions are optimized, solving the problems of low decision-making efficiency and high cost in traditional methods, and realizing intelligent fault handling of complex systems.

CN120894007APending Publication Date: 2025-11-04BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511000049.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Traditional methods for diagnosing and repairing complex system faults rely on human experience, resulting in low decision-making efficiency, unstable strategies, and high maintenance costs, making it difficult to meet the intelligent requirements of modern high-reliability equipment.

Method used

An intelligent maintenance and troubleshooting method combining Bayesian networks and double-Q learning is proposed. By constructing a causal relationship model, introducing prior probability and action masking mechanisms, and using Monte Carlo simulation to generate training data, a double-Q learning framework is built to optimize maintenance decisions.

Benefits of technology

It enables efficient, targeted, and cost-optimized maintenance decisions, enhances the intelligent fault handling capabilities of complex systems, reduces reliance on expert experience, and improves strategy stability and learning effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894007A_ABST
    Figure CN120894007A_ABST
Patent Text Reader

Abstract

The invention provides a complex system-oriented intelligent maintenance troubleshooting method combining a Bayesian network and double-Q learning, and aims to solve the problems of low decision-making efficiency, poor path and high maintenance cost of a traditional heuristic strategy. The method comprises the following steps: firstly, constructing a Bayesian network model according to a system structure and a causal dependency relationship, and setting a prior probability of each node in combination with an FMECA analysis result; and then, a troubleshooting sample is generated by using a Monte Carlo method, a double-Q learning model is trained offline, observation and maintenance costs are fused, and a maintenance action strategy is optimized. In order to improve strategy convergence and stability, an observation action screening mechanism and an action shielding technology are introduced, and dynamic decision optimization in the multi-step troubleshooting process is achieved. The method has a causal modeling capability and reinforcement learning adaptivity, has relatively high robustness and real-time performance, and is suitable for intelligent maintenance and fault diagnosis scenes of key equipment such as aircrafts and automobiles.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent maintenance and fault diagnosis, in particular to a complex system intelligent maintenance and fault diagnosis method combining Bayesian network and double Q learning, and belongs to the cross technical field of artificial intelligence, equipment health management and intelligent decision support system. BACKGROUND

[0002] With the increasing complexity of the structure of complex systems (such as automotive electronic systems, aerospace vehicles, industrial equipment, etc.), the coupling relationship between the components in the system is enhanced, which leads to a significant increase in the difficulty of fault diagnosis and maintenance. The traditional troubleshooting methods mostly rely on manual experience or heuristic rules, which have problems such as low decision-making efficiency, poor strategy adaptability, and high overall maintenance cost, and have been difficult to meet the requirements of modern high-reliability equipment for efficient intelligent fault handling.

[0003] In recent years, Bayesian networks have been widely used in fault diagnosis and causal modeling, which can describe the probabilistic dependency relationship between components in the system and realize reasoning under incomplete information. However, Bayesian networks themselves do not have decision optimization capabilities and are difficult to deal with the selection and sequencing of maintenance actions in complex troubleshooting processes.

[0004] Reinforcement learning, especially Q-learning, has become a research hotspot in the field of intelligent decision-making because it can gradually optimize strategies through interaction with the environment. However, in traditional Q-learning, the action value estimation is prone to be high, which leads to unstable strategy convergence and poor robustness, especially in troubleshooting scenarios that may cause high-cost path selection.

[0005] Therefore, there is an urgent need for a troubleshooting decision-making method that combines the causal modeling capabilities of Bayesian networks and improved reinforcement learning algorithms to improve troubleshooting efficiency, reduce overall maintenance costs, and achieve intelligent and autonomous fault handling of complex systems. SUMMARY

[0006] To overcome the shortcomings of existing heuristic troubleshooting methods in terms of low decision-making efficiency, unstable strategies, and high troubleshooting costs, the present application proposes a complex system intelligent troubleshooting method combining Bayesian network and double Q learning, which realizes efficient, targeted, and cost-optimized maintenance decision-making.

[0007] The method of the present application mainly includes the following steps:

[0008] Firstly, based on the structure and functional dependency relationship of the system to be maintained, a Bayesian network model is constructed. The network expresses the causal relationship between the internal components of the system in the form of a directed acyclic graph, and the nodes cover measurable components, non-measurable components, indirect observation indicators, problem definition nodes, and implicit factor nodes. Each type of node has a clear functional division to support efficient fault reasoning and maintenance operation decision-making.

[0009] Secondly, according to the failure mode, effect and criticality analysis (FMECA) results of the system, the original failure probability of the component is extracted, and by normalization and mapping processing, the prior probability value of the Bayesian network node is set, and the engineering credibility and applicability of the model are enhanced.

[0010] Then, the troubleshooting strategy is trained and modeled by using reinforcement learning technology. In this framework, the combination of the states of each component of the system is defined as a state space, and the observation and maintenance operation are defined as an action space. By designing observation-maintenance joint action pairs, combined with the action mask mechanism, invalid or repeated operations are effectively shielded, and the complexity of the action space is significantly reduced. At the same time, the operation costs of various types are introduced into the reward function, so that the learning process aims to minimize the total maintenance cost.

[0011] In terms of training sample generation, the Monte Carlo simulation method is used to construct a large number of "state-action-reward" sequences containing different fault injection conditions, which truly reflect the fault propagation path and the behavior mode of the agent, and provide data support for reinforcement learning training.

[0012] Then, the double Q learning mechanism is introduced to construct the training framework, and two parallel updated action value functions (Q1 and Q2) are maintained. By randomly selecting the network to perform action selection and update, the problems of strategy bias and action value overestimation existing in the traditional Q-learning algorithm are effectively alleviated, and the learning stability is improved.

[0013] Finally, in the actual decision-making stage, the Q1 and Q2 strategy models trained are applied to the actual troubleshooting task. According to the initial state information generated in the fault diagnosis stage and the real-time observation results, under the action mask restriction, the optimal action sequence is dynamically selected by the value maximization strategy until the system fault is eliminated, so as to realize the maintenance troubleshooting decision-making process with intelligent judgment, adaptive ability and cost sensitivity.

[0014] The method of the application combines the causal modeling capability of Bayesian network and the strategy optimization capability of double Q learning, has the advantages of model-driven and data-driven, and is suitable for the fields of intelligent fault diagnosis and maintenance troubleshooting of key complex systems such as aerospace, aviation and automobile equipment, and has wide application prospect and actual engineering value. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 The overall flowchart of the maintenance troubleshooting method of the application;

[0016] Figure 2 The schematic diagram of the Bayesian network modeling structure in the system of the application;

[0017] Figure 3 The flowchart of the double Q learning training stage described in the application;

[0018] Figure 4 Flowchart of the online troubleshooting decision-making phase of the present application. Specific embodiments

[0019] S1: According to the structure and function dependence of the troubleshooting system to be maintained, a Bayesian network model is established to represent the causal relationship between the components and functional units in the system. The Bayesian network is a directed acyclic graph (DAG) structure, in which each node represents a key element in the system, and each edge represents the conditional dependence or causal transmission path between the nodes.

[0020] Specifically, the node types in the Bayesian network include the following categories: measurable component nodes: representing components in the system whose health status can be directly measured by sensors, test instruments or other means, which can be observed and repaired; unmeasurable component nodes: representing components whose state cannot be directly obtained by observation means, and can only be inferred indirectly through system behavior, and can only be excluded by maintenance operation; indirect observation nodes: representing the health status that can be obtained by observation, but not directly corresponding to a specific component, but reflecting some comprehensive system characteristics or fault signs, which are often used to support fault reasoning; problem definition nodes: used to represent whether the system is in a fault state, which is the target judgment basis for troubleshooting process, and the state value change is used to judge whether the fault is successfully excluded; implicit nodes: used to assist in modeling potential factors in the system that cannot be directly observed but affect other nodes, to enhance the causal modeling capability.

[0021] In addition, in the model construction stage, economic information related to each node needs to be obtained in advance. Specifically, it includes the maintenance cost of all component nodes (measurable and unmeasurable), and the observation cost of measurable component nodes and indirect observation nodes. The above cost parameters will be part of the reward function in the reinforcement learning algorithm, used to measure the resource consumption of each step in the troubleshooting process, so as to guide the policy model to optimize the maintenance path with lower cost and higher efficiency.

[0022] S2: After completing the Bayesian network structure modeling, reasonable prior probabilities need to be set for each node in the network to reflect the preliminary estimate of the fault state of each component under the condition of no observation information. In order to ensure that the prior probability has engineering interpretability and actual reliability, combined with the results of failure mode, effects and criticality analysis (FMECA), the fault probability of each component is normalized and mapped as the prior probability input of the corresponding node. The normalization formula is as follows, where is the prior fault probability of the i-th node; is the fault probability of component c i defined in the FMECA table.

[0023]

[0024] S3: Based on the preliminary estimation of component health status obtained by FMECA analysis and normalization processing in S2, further construct a reinforcement learning model suitable for intelligent maintenance and troubleshooting of complex systems. For this purpose, five core elements need to be defined, including state space, action space, reward function, state transition probability and policy function, to form a complete learning and decision-making framework.

[0025] In reinforcement learning modeling, the state of the environment should be consistent with the actual state of the equipment in the maintenance and troubleshooting task. Therefore, the running state of each component in the equipment can be combined to form the state space of reinforcement learning. Since the state of each component is a discrete variable ("normal", "fault" or "unknown"), the overall state of the system is also composed of a limited number of discrete states, so that the reinforcement learning problem has a discrete state space structure.

[0026] The action space of the reinforcement learning agent corresponds to the set of actions that the decision maker can perform in the maintenance and troubleshooting process. Specifically, the actions on the nodes of the Bayesian network in the maintenance and troubleshooting task include two types: one is the observation action on the observable component nodes and indirect observation nodes; the second is the repair action on all component nodes. Therefore, the transformed reinforcement learning action space is also a discrete space. In view of the fact that repeated execution of the same action by the agent cannot obtain new information or repair the equipment, but will increase the system cost, in order to avoid such inefficient behavior, a hard limit mechanism, action mask (A mask ), is introduced in the action space design to mask illegal actions and prevent the agent from selecting them. When the action mask value is 1, the action is selectable, and when the value is 0, the action is not selectable, and all action masks are initialized to 1).

[0027] To further reduce the action size of double Q learning, the observation-repair action pair is jointly defined for the component nodes, that is, first perform the observation action on the component node, if the state is determined to be normal, end the current action; if it is determined to be faulty, then perform the repair action immediately. Based on this design, the action space only contains observation actions on indirect observation nodes and observation-repair composite actions on component nodes, and at this time the component node state can only be in two states, fuzzy (-1) or normal (0), which significantly simplifies the complexity of the state and action space.

[0028] In terms of reinforcement learning strategy, in the offline training phase, the agent selects an action from the action space according to the current state and the policy function, and then updates the state according to the action and the state transition probability. The reward function is used to evaluate the performance of the agent, and the policy function is updated based on the reward function and the Q value of the action. maskThe legal action space with a value of 1 is selected by adopting an epsilon-greedy strategy. The strategy randomly explores the action space with a probability of epsilon to prevent falling into a local optimum, and selects the currently estimated optimal action with a probability of 1-epsilon to achieve a balance between exploration and utilization. In the actual decision-making stage online, the agent selects the action with the highest estimated value in the legal action space constrained by the action mask to maximize the immediate return, thereby realizing optimal maintenance and fault diagnosis decision-making.

[0029] In the reinforcement learning model of the application, the perfect action assumption is based on the assumption that all maintenance actions and observation actions do not introduce new device faults during execution, that is, the influence of such actions on the system state is safe and has no side effects. The state transition probability is simplified as a deterministic transition, that is, under the condition of a given current state s and action a, the system must transition to a unique next state s', which is mathematically expressed as formula (3).

[0030] P(s'|s,a) = 1 (3)

[0031] Specifically, the state transition rule is as follows:

[0032] (1) Observation action: after performing observation on the node in the fuzzy state (-1), the node state must change to its real state (normal 0 or fault 1), and the state transition is a deterministic process.

[0033] (2) When the observation action acts on the node in the determined state (normal state 0 or fault state 1), the node state remains unchanged, and the state transition exhibits an identity mapping.

[0034] (3) If the maintenance action acts on the faulty node, the node state must change from the fault state (1) to the normal state (0), realizing deterministic state transition.

[0035] (4) If the maintenance action acts on the non-faulty node (normal or fuzzy state), the node state remains unchanged, and the state transition is an identity mapping.

[0036] The goal of reinforcement learning is to minimize the total cost generated in the process of maintenance decision-making, and the agent learns the policy that maximizes the cumulative reward. To this end, the reward value R is negatively correlated with the action cost C, that is, the reward value is proportional to the inverse of the action cost, so as to encourage the agent to select low-cost observation and maintenance actions. In the maintenance troubleshooting problem, the next state s' of the environment is uniquely determined by the current state s and the executed action a, and the selection of the action depends on the current state and the policy function. Since the action cost is independent of the node state, the reward function is designed as shown in equation (4), where C represents the observation cost or maintenance cost corresponding to the action. This design effectively guides the agent to complete the maintenance troubleshooting task by minimizing the total cost. If the action is the problem definition node that returns to normal, a larger positive reward value is returned.

[0037] R(s,a,s')=R(s,a)=R(s)=-C (4)

[0038] S4: The application generates training samples by the Monte Carlo simulation method for offline training of the reinforcement learning model. The specific steps are as follows: first, a component node is randomly selected and a fault is injected into it, and the state of the node is set to fault (1); at the same time, the state of the problem definition node is also set to fault (1) to represent that the system as a whole has been identified as a fault. Subsequently, by means of the inference mechanism of the Bayesian network, all nodes that have causal relationship with the known fault node and whose fault probability is inferred to be 1 under the current condition are set to fault (1) at the same time. The states of the remaining nodes whose fault probability is not 1 are set to normal (0), thereby constructing a "real fault state vector" containing the state information of all component nodes.

[0039] On this basis, combined with the structural logic of the system and the executability of the observation and maintenance actions, the observation and maintenance decision-making process that can be executed by the agent under this fault state is simulated step by step, the state transition process and the reward value corresponding to each action are recorded, and finally a complete "state-action-reward" sequence is formed. This sequence can be used as a high-quality sample required for training of the reinforcement learning model, improving the effectiveness and accuracy of policy learning.

[0040] S5: Based on the large number of "state-action-reward" sequence samples generated by Monte Carlo simulation in S4, the application further designs and trains a reinforcement learning strategy model. Considering that the traditional Q-learning method is prone to maximum value deviation (overestimation bias) in action value estimation, in order to improve the stability and learning accuracy of the strategy, a Double Q-learning framework is constructed, that is, two independent action value functions Q networks (Q1 and Q2) are set up. In each step of action selection and update, one of the networks is randomly selected for action selection, and the other is used for action value estimation. Each set of training samples contains the current state, the selected action, the immediate cost (as negative reward) corresponding to the action, and the next state information after executing the action, forming a complete state-action-reward-next state tuple (s, a, r, s').

[0041] S6: In the training phase, for each episode, the following operations are performed:

[0042] (1) Initialize the state: set the problem definition node state to the fault state (1), and set the remaining node states to the fuzzy state (-1) to form the initial state vector s0.

[0043] Parameter setting: set the learning rate a (0 < a < 1), the discount factor g (0 < g < 1), and the exploration rate e (0 < e < 1); construct the state space S and the action space A, and initialize the value table of the two action value functions Q1 and Q2, assigning an initial value of 0 to all state-action pairs.

[0044] (3) Action selection strategy: in the current action mask A mask , select an action a using the e-greedy strategy, that is: with a probability of e, randomly select an action from the actions with mask 1; with a probability of 1-e, select the action with the highest value from the actions with mask 1 based on the average of Q1 and Q2. Set the mask of the executed action to 0 to prevent repeated selection.

[0045]

[0046] (4) State update and Q value update: after executing the selected action, obtain the immediate reward r and the next state s'. Update Q1 with a probability of 0.5 and update Q2 with a probability of 0.5. The corresponding Q value update formula is as follows:

[0047] Q i (s,a)←Q i (s,a)+α(r+γQ j (s′,argmaxQ(s′,a′))-Q i (s,a))(6)

[0048] wherein i≠j, i, j∈{1, 2}.

[0049] (5) Fault judgment: if the state of the problem definition node is changed from fault (1) to normal (0), it is considered that the maintenance troubleshooting task is completed this time; otherwise, the next action is selected from the actions with mask value 1 according to the ε-greedy strategy.

[0050] S7: After the training is completed, in the decision-making stage of actual application, the following process is executed:

[0051] (1) Initialize the action mask A mask , so that all actions are optional, that is, A mask (a) = 1,

[0052] (2) Under the premise that the troubleshooting task is not completed, the average value of each action is calculated from the trained Q1 and Q2, and the action with the maximum average value is selected according to formula (5) in the action set with mask value 1 to perform.

[0053] (3) Record the immediate reward r and the new state s' after the action is executed, and set the mask value of the executed action to 0;

[0054] (4) If the state of the problem definition node returns to normal (0), it is determined that the troubleshooting task is completed, and the decision-making process is ended; otherwise, step 2 is continued.

[0055] The maintenance troubleshooting method for complex systems provided by the application combines Bayesian network modeling and double Q learning strategy optimization. By introducing FMECA prior knowledge to construct a causal fault model, training data is generated using Monte Carlo simulation, and the action mask mechanism is used to improve troubleshooting efficiency and accuracy. Compared with traditional heuristic methods, the application has stronger systematicity and interpretability, reducing the dependence on expert experience; compared with standard Q-learning, the combination of double Q structure and Bayesian reasoning improves the stability of the strategy and the learning effect. The method realizes the fusion of model-driven and data-driven, and is suitable for efficient intelligent troubleshooting of complex systems.

Claims

1. A method for maintenance troubleshooting of a complex system, characterized in that, The method comprises the following steps: According to the structure information of the complex system and the Bayesian network model of the causal dependence relationship of the components, the nodes represent the device components and their states, and the edges represent the causal relationship between the components; Based on the failure mode, effects and criticality analysis (FMECA) results, the failure probabilities of various types are obtained, and through normalization and mapping processing, they are used as the prior probabilities of the nodes of the Bayesian network; Based on the observation and maintenance actions of the system, a reinforcement learning environment model is constructed, which includes states, actions, rewards, and transitions; Monte Carlo simulation is used to generate troubleshooting samples, and by injecting faults, Bayesian network reasoning and strategy simulation, a state-action-reward sequence is constructed for training; In the offline reinforcement learning phase, a double Q-learning algorithm is used to train the strategy, and by alternately selecting and evaluating actions through two independent Q-value networks, the action value overestimation problem of traditional Q-learning is alleviated; During the training process, an observation action screening mechanism and an action masking strategy are introduced to dynamically eliminate invalid or redundant actions and improve the stability of the strategy; The trained troubleshooting strategy is deployed in the actual system, and based on the current real-time observation state and the fuzzy fault set generated during the fault diagnosis phase, an optimal observation and maintenance operation sequence is dynamically generated to achieve the goal of low cost and high efficiency troubleshooting.

2. The method of claim 1, wherein: The nodes of the Bayesian network include indirect observation nodes, measurable component nodes, unmeasurable component nodes, hidden nodes, and problem definition nodes, and the causal relationships between them are described by conditional probability tables.

3. The method of claim 1, wherein: Based on the failure mode, effects and criticality analysis (FMECA) results, the failure probabilities of various types are obtained, and through normalization and mapping processing, they are used as the prior probabilities of the nodes of the Bayesian network, which reflect the initial occurrence probability of each type of fault under the condition of no observation.

4. The method of claim 1, wherein: The action masking strategy is used to mask the executed maintenance and observation actions in each decision step.

5. The method of claim 1, wherein: During the Monte Carlo sample generation process, faults are randomly injected into system component nodes, and Bayesian network is used for fault reasoning to generate a state-action-reward triple sequence for training.

6. The method of claim 1, wherein: In the double Q-learning framework, two Q networks are independently initialized and alternately used in the training process, one network selects actions and the other network evaluates action values to reduce the overestimation error in traditional Q-learning.

7. The method of claim 1: after the training of the troubleshooting strategy is completed, the optimal observation and maintenance action sequence can be dynamically calculated and executed based on real-time observation input and fuzzy fault set to achieve optimal fault location and exclusion efficiency.