Intelligent fault detection method for speed reducer

Through domain-adaptive transfer learning and multi-agent causal reasoning technology, the problems of sample scarcity and system-level collaborative analysis in reducer fault detection are solved, high-accuracy and real-time fault diagnosis is achieved, and a comprehensive intelligent fault detection solution is built.

CN120671045APending Publication Date: 2025-09-19SHAANXI LINKEZHI MASCH EQUIP CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510767266.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing reducer fault detection methods have low diagnostic accuracy under conditions of sample scarcity, find it difficult to capture the system-level collaborative relationship of multiple reducers, lack in-depth exploration of the root causes of faults, and have difficulty in knowledge accumulation, making it difficult to formulate accurate maintenance and prevention strategies.

Method used

Domain-adaptive transfer learning, multi-agent causal reasoning, knowledge graph evolution, and reinforcement learning techniques are used to migrate the source domain reducer fault feature knowledge to the target domain, build a multi-agent causal model, dynamically update the knowledge graph, combine Bayesian uncertainty estimation to generate maintenance decisions, and deploy an edge-cloud collaborative architecture for real-time detection and optimization.

Benefits of technology

In the case of scarce samples, the diagnostic accuracy rate reaches 88%, the system-level fault identification accuracy rate is increased to 92%, the unplanned downtime is reduced by 75%, maintenance costs are reduced, equipment availability is improved, system adaptability is improved by 65%, and the response time is less than 100 milliseconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671045A_ABST
    Figure CN120671045A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of equipment fault diagnosis, and discloses an intelligent fault detection method for a speed reducer, and the method comprises the steps: extracting fault feature knowledge from a source domain based on a domain adaptive transfer learning algorithm, migrating the fault feature knowledge to a target domain, and generating a synthetic fault sample based on physical model constraints; constructing interaction influence among the multi-subject causal capture equipment; executing causal intervention and anti-factual reasoning, and determining a fault root cause; constructing a fault knowledge graph and continuously optimizing the fault knowledge graph; and constructing a preventive maintenance decision system to generate an optimal maintenance strategy. According to the method, high-accuracy fault diagnosis can be realized under the condition of sample scarcity, interaction influence among multiple devices is analyzed, fault root causes are traced, and reliable preventive maintenance decision support is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment fault diagnosis, and more particularly to an intelligent fault detection method for a reducer. Background Art

[0002] In large-scale industrial facilities, speed reducers are critical transmission devices. Failure can cause the entire production line to shut down, resulting in significant economic losses. With the development of Industry 4.0 and smart manufacturing, intelligent detection and prevention of speed reducer failures have become crucial for ensuring efficient and stable operation of production systems.

[0003] Existing methods for detecting reducer faults primarily include physical model-based methods and data-driven methods. Physical model-based methods build mathematical models of reducer components and analyze their vibration, noise, and other characteristics to determine if a fault exists. Data-driven methods utilize machine learning or deep learning techniques to learn fault characteristics and patterns from large amounts of historical data to detect and classify faults.

[0004] However, existing technologies have the following problems: First, in actual industrial environments, especially for new reducers or rare fault types, it is often difficult to obtain sufficient labeled samples, resulting in a decline in the performance of fault diagnosis methods based on supervised learning; second, traditional fault detection systems mainly perform independent analysis on a single device, ignoring the system-level correlation when multiple reducers work together, and it is difficult to capture complex interactive relationships; third, existing technologies mainly focus on the identification and classification of fault symptoms, lack the ability to deeply explore the root causes of faults, and it is difficult to formulate accurate maintenance and prevention strategies; finally, traditional systems find it difficult to effectively accumulate and utilize historical case knowledge. Every time they face a new device or a new fault mode, they need to be retrained or undergo a lot of manual intervention and adjustment.

[0005] Therefore, how to achieve highly accurate diagnosis under conditions of scarce samples, have system-level analysis capabilities, trace the root cause of the fault and continuously self-optimize intelligent fault detection methods has become an urgent problem to be solved. Summary of the Invention

[0006] The present invention provides an intelligent fault detection method for a reducer, which solves the technical problem in the prior art of how to achieve highly accurate diagnosis under conditions of scarce samples, have system-level analysis capabilities, trace the root cause of the fault and continuously self-optimize the intelligent fault detection method.

[0007] The present invention provides an intelligent fault detection method for a reducer, comprising the following steps:

[0008] Transfer the knowledge of fault characteristics extracted from the source domain reducer to the target domain reducer; and constrain the generation of fault samples to expand the training data set;

[0009] Based on the expanded training dataset, multiple reducers are regarded as interrelated entities, and a multi-agent causal model is constructed and dynamically updated to capture the interactive impact between devices.

[0010] Through the constructed multi-agent causal detection system, abnormal states are detected, causal intervention is performed on abnormal variables, counterfactual conditions are calculated, and the causal effects of various factors on the failure are quantified based on sensitivity analysis. The key root causes are determined, and the root cause analysis results are used as an important input for knowledge graph construction.

[0011] Combining root cause analysis results with historical failure cases, we build a fault knowledge graph to represent the relationships between fault types, symptoms, causes, and solutions. This graph is continuously optimized as new cases accumulate.

[0012] By integrating the results of multi-agent causal analysis, root cause diagnosis information and fault knowledge graph, a preventive maintenance decision-making system is constructed to predict the reducer system status, define the maintenance decision space, train the decision strategy, and quantify the decision reliability by combining Bayesian uncertainty estimation to generate the optimal maintenance strategy.

[0013] Furthermore, the steps of extracting fault feature knowledge from the source domain reducer and migrating it to the target domain reducer include:

[0014] Apply convolutional neural network to the historical data of reducer in the source domain to extract fault feature representation;

[0015] Construct a domain adaptation layer to map the source domain feature space to a feature space compatible with the target domain. Its optimization goal is to minimize the JS divergence between the source and target domain feature distributions.

[0016] Generate synthetic fault samples based on the reducer physical model and constraints;

[0017] Combine real samples from the target domain and synthetic samples to train the fault diagnosis model.

[0018] Furthermore, considering multiple reducers as interrelated entities, the steps of constructing and dynamically updating multi-agent causal relationships include:

[0019] Construct initial multi-agent causal relationships based on device topology and domain knowledge;

[0020] Using causal discovery algorithms, we learn the structure and parameters of multi-agent causality based on historical data.

[0021] Based on the learned graph structure, calculate the conditional probability table of each node;

[0022] As new fault cases appear, the likelihood under the current graph structure is calculated. If the likelihood is lower than the threshold, the structure update process is triggered.

[0023] Furthermore, the steps of detecting abnormal conditions based on multi-agent causality and determining the key root causes include:

[0024] For each monitoring variable of each device, the deviation between its observed value and the expected value calculated based on the parent node value is calculated. If the deviation exceeds the threshold, the variable is marked as an abnormal variable;

[0025] Perform causal intervention on each variable in the set of variables marked as abnormal, and analyze the change of system state after the intervention;

[0026] Perform counterfactual calculations on key variables to assess their impact on failures under specific scenarios;

[0027] Based on sensitivity analysis, the causal effect of each factor on the failure is quantified, the potential root causes are ranked according to the size of the causal effect, and the set of key root cause variables is determined.

[0028] Furthermore, the root cause ranking not only considers the size of the causal effect, but also the intervention cost and feasibility. The root cause priority is calculated by dividing the causal effect value by the intervention cost and multiplying it by the intervention feasibility.

[0029] Furthermore, the steps of constructing a fault knowledge graph to represent the relationship between fault types, symptoms, causes, and solutions include:

[0030] Construct a knowledge graph of reducer faults and represent the relationships between entities through triples;

[0031] Reasoning based on knowledge graphs supports complex fault diagnosis queries;

[0032] Build a knowledge graph dynamic update algorithm to continuously enrich and optimize the knowledge graph as new fault cases emerge;

[0033] Distill and integrate the knowledge in the knowledge graph to extract more refined fault diagnosis rules.

[0034] Furthermore, the steps of constructing a preventive maintenance decision system based on the analysis results include:

[0035] Based on multi-agent causality and knowledge graph, the reducer system status is represented and predicted;

[0036] Define a maintenance decision space containing various possible maintenance actions, with each decision associated with a cost function;

[0037] Build a Markov decision process model and train reinforcement learning strategies;

[0038] Combined with Bayesian uncertainty estimation, the reliability of decisions is quantified and explainable decision reports are generated.

[0039] Furthermore, the cost function considers various factors, including direct repair costs, downtime losses, spare parts inventory status, and maintenance personnel availability.

[0040] Furthermore, by deploying an edge-cloud collaborative architecture, lightweight fault detection and early warning models are deployed on the edge side, responsible for real-time data processing and preliminary anomaly detection; complex causal reasoning, knowledge graphs and decision generation models are deployed on the cloud side, responsible for in-depth analysis and long-term optimization.

[0041] Furthermore, a computer-readable storage medium is provided, characterized in that it is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, an intelligent fault detection method for a reducer can be executed.

[0042] The beneficial effects of the present invention are as follows: the present invention organically combines domain adaptive transfer learning, multi-agent causal reasoning, knowledge graph evolution and reinforcement learning technology, effectively solving technical problems in the prior art such as sample scarcity, single-machine independent analysis limitations, separation of fault symptoms and root causes, and difficulty in knowledge precipitation, and has the following significant beneficial effects: First, through domain adaptive transfer learning and physical constraint sample synthesis technology, the sample scarcity problem is overcome. In the case of only 10 labeled samples, the diagnostic accuracy can still reach 88%, which is 45% higher than the traditional method, and the adaptation time for new reducers is shortened from the traditional month level to the day level; secondly, the introduction of multi-agent causal modeling of the relationship between devices breaks through the limitations of single-machine independent analysis, can capture system-level faults, and realize multi-device collaborative analysis, which is particularly suitable for complex scenarios where multiple reducers work in series; thirdly, the use of causal intervention and counterfactual reasoning The fault diagnosis technology solves the problem of separating the fault symptoms from the root cause, and the accuracy of fault root cause identification is increased to 92%, which reduces the false alarm rate and missed alarm rate, providing a reliable basis for accurate maintenance. In addition, the fault knowledge graph is constructed and continuously optimized to achieve effective sedimentation and reuse of knowledge, and the system adaptability is improved by 65%. Each diagnosis not only solves the current problem, but also strengthens the overall capability of the system. At the same time, based on the preventive maintenance decision system, the average unplanned downtime is reduced by 75%, and the fault prediction lead time is extended from hours to days. The system's decision recommendation adoption rate is increased from the initial 46% to 92%, which greatly reduces maintenance costs and improves equipment availability. Finally, the edge-cloud collaborative deployment architecture is adopted to achieve the dual advantages of real-time response and continuous learning. The system response time is less than 100 milliseconds, which meets the real-time monitoring needs of industrial sites and has powerful analysis and learning capabilities. Overall, the present invention constructs a comprehensive reducer intelligent fault detection solution, which significantly improves the accuracy, timeliness and reliability of fault diagnosis and provides a new technical path for intelligent maintenance of industrial equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flow chart of an intelligent fault detection method for a reducer provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0044] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0045] At least one embodiment of the present invention discloses an intelligent fault detection method for a reducer, comprising:

[0046] like Figure 1 As shown, a method for intelligent fault detection of a reducer includes the following steps:

[0047] Step 1: Transfer the fault feature knowledge extracted from the source domain reducer to the target domain reducer; constrain the generation of fault samples and expand the training data set;

[0048] The step provided in this application uses a domain-adaptive transfer learning algorithm to extract knowledge from the source domain (reducer types with sufficient samples) and migrate it to the target domain (reducer types with scarce samples). At the same time, it generates synthetic samples based on physical model constraints to expand the training data set.

[0049] Step 1-1: Source domain feature extraction.

[0050] According to one embodiment of the present application, a convolutional neural network is applied to the source domain reducer historical data to extract fault feature representation. It should be noted that the source domain data includes a variety of sensor signals (vibration, temperature, sound, current, etc.), which are pre-processed and then input into the feature extraction network. The source domain dataset D is s for:

[0051]

[0052] in, is the feature vector of the i-th sample, is the corresponding fault label, N s is the number of source domain samples. The feature extraction network can be expressed as a function where θ s is the network parameter, and the feature obtained after training is expressed as

[0053] In some embodiments, the source domain feature extraction network can adopt different structures. For example, for vibration signals, a one-dimensional convolutional network (1D-CNN) can be used, whose structure includes multiple convolutional layers, pooling layers, and fully connected layers; for sound signals, a combination of mel-frequency cepstral coefficient (MFCC) feature extraction and recurrent neural network (RNN) can be used; for time series signals such as temperature and current, a long short-term memory network (LSTM) or a gated recurrent unit (GRU) can be used. Optionally, the features of these different modes can be integrated through a multimodal fusion algorithm to form a more comprehensive fault feature representation.

[0054] Step 1-2: Construction of domain adaptation layer.

[0055] In this step, a domain adaptation layer is constructed to map the source domain feature space to a feature space compatible with the target domain. It should be understood that the target domain dataset is recorded as:

[0056]

[0057] in, is the feature vector of the j-th target domain sample, is the corresponding fault label, N t <<N s Indicates that the number of samples in the target domain is much smaller than that in the source domain. The domain adaptation layer is represented by the function g φ , the parameter is φ.

[0058] The optimization goal of this layer is to minimize the JS divergence of the feature distribution between the source domain and the target domain:

[0059]

[0060] Among them, D JS represents JS divergence (Jensen-Shannon divergence), P s and P t denote the feature distribution of source domain and target domain respectively, g φ represents the domain adaptation layer function, represents the source domain feature extraction network function, x s and x t denote the input samples of the source domain and the target domain respectively, φ denotes the parameters of the domain adaptation layer, and θ s Represents the parameters of the source domain feature extraction network.

[0061] According to another embodiment of the present application, the domain adaptation layer can be implemented by adversarial training. Specifically, a domain discriminator D is introduced. disc , whose goal is to distinguish whether the features come from the source domain or the target domain, while the goal of the feature extraction network and the domain adaptation layer is to generate features that the discriminator cannot distinguish.

[0062] Optionally, in order to further improve the migration effect, the idea of ​​conditional adversarial network (CGAN) can be introduced so that the domain adaptation process also considers category information to ensure that the features after migration still have good discriminability.

[0063] In a specific application example, the source domain is a rolling mill reducer, and the target domain is a wind turbine gearbox. Although the two work in similar principles, due to differences in load characteristics, working environment, etc., directly applying the fault diagnosis model of the rolling mill reducer to the wind turbine gearbox will result in a significant performance degradation. Through the domain adaptation method provided in this application, the diagnostic accuracy rate is improved from 32% of direct migration to 86%, which is close to the performance of the model trained with sufficient target domain samples (92%).

[0064] Steps 1-3: Physically constrained sample synthesis.

[0065] According to the embodiments of the present application, synthetic fault samples are generated based on the reducer's physical model and constraints. First, a physical model M of the reducer is established, including dynamic and thermodynamic equations, describing the physical relationships between the various components of the reducer. Then, normal operating data is used to determine the model parameters, and synthetic fault data is generated by introducing changes in the physical parameters corresponding to different fault modes.

[0066] The synthetic sample generation function is denoted as G(z, c, M), where z is the random noise vector, c is the fault type condition vector, and M is the physical model parameter. It should be noted that the generated synthetic samples must satisfy the physical constraints:

[0067] C(G(z,c,M))≤ε

[0068] Among them, C is the physical constraint function and ε (epsilon) is the constraint threshold.

[0069] The synthetic sample set generated by this method is recorded as:

[0070]

[0071] in, is the feature vector of the kth synthetic sample, is the corresponding fault label, N syn is the number of synthetic samples.

[0072] In some embodiments, physical constraints can take various forms. For example, for gear faults, a meshing stiffness variation model can be used to represent the fault as a periodic change in meshing stiffness; for bearing faults, an impulse response model can be used to represent the fault as a sequence of impulses of a specific frequency. Optionally, these physical models can be combined with deep generative models (such as variational autoencoders (VAEs) or generative adversarial networks (GANs)) to achieve more flexible sample generation. Specifically, the deep generative model is responsible for generating the initial samples, which are then post-processed using physical constraints to ensure that the samples satisfy the laws of physics.

[0073] According to a preferred embodiment of the present application, physical constraints include the following aspects:

[0074] 1. Frequency domain constraints: The fault signal should exhibit specific characteristic frequencies in the frequency domain, such as the fault frequencies of the inner and outer races of bearings, the gear meshing frequency, etc.

[0075] 2. Energy conservation constraint: The total energy of the generated signal should be close to the actual signal;

[0076] 3. Time-varying characteristic constraints: The time-varying characteristics of the fault signal should conform to the actual fault development law;

[0077] 4. Signal-to-noise ratio constraint: The signal-to-noise ratio of the synthetic signal should match the actual industrial environment.

[0078] In a specific application example, traditional methods struggle to obtain sufficient samples for crack faults in reducer gears, especially samples of early-stage, tiny cracks. Using the physically constrained sample synthesis method presented in this application, samples of varying crack sizes, locations, and operating conditions were generated, enabling the model to identify early-stage crack faults. The detection rate increased from 45% with traditional methods to 87%, detecting potential faults an average of 15 days in advance.

[0079] Steps 1-4: Small sample model training.

[0080] The method provided in this application combines real samples and synthetic samples from the target domain to train a fault diagnosis model. The diagnosis model consists of a feature extraction network, a domain adaptation layer, and a classification layer. The overall parameters are:

[0081] Θ=θ s ,φ,ψ

[0082] where θ s is the source domain feature extraction network parameter, φ is the domain adaptation layer parameter, and ψ is the classification layer parameter.

[0083] The model training objective function is:

[0084]

[0085] Among them, Θ represents the overall model parameter set, F Θ represents the overall model, L is the classification loss function, R is the regularization term, λ is the regularization coefficient, (x, y) represents the input sample and the corresponding label, D t represents the target domain dataset, D syn represents a synthetic sample dataset, and ∪ represents a set union operation.

[0086] Optionally, to address sample imbalance, the loss function can use weighted cross-entropy or focal loss, giving higher weight to rare fault types. In some implementations, meta-learning can be introduced to improve the model's generalization ability under sample-scarce conditions by simulating small-sample learning tasks. For example, methods such as model-agnostic meta-learning (MAML) or prototypical networks can be used to enable the model to "rapidly adapt."

[0087] According to another embodiment of the present application, during the training process, the ratio of real samples to synthetic samples can be dynamically adjusted, with synthetic samples being the primary focus in the early stages and gradually increasing the proportion of real samples as real samples accumulate. This curriculum learning strategy can smooth the model's learning process and improve final performance.

[0088] Step 2: Based on the expanded training dataset, multiple reducers are regarded as interrelated entities, and a multi-agent causal model is constructed and dynamically updated to capture the interactive effects between devices.

[0089] In view of the limitations of independent analysis of a single machine, this step of the present application regards multiple reducers as interrelated entities, constructs and dynamically updates multi-agent causality, and captures the interactive impact between devices.

[0090] Step 2-1: Multi-agent causal initialization.

[0091] According to one embodiment of the present application, an initial multi-agent causal relationship is constructed based on the device topology and domain knowledge. The multi-agent causal relationship is denoted as:

[0092] G=G1,G2,...,G n ,R

[0093] Among them, G n =(V i , E i ) represents the internal causal relationship of device i, V i is a node set (including sensor indicators and state variables), E i is the edge set (representing the causal relationship between variables), n is the total number of devices, i is the device number; R is the relationship set between devices.

[0094] Step 2-2: Application of causal discovery algorithm.

[0095] In this step, the causal discovery algorithm is used to learn the structure and parameters of multi-agent causality based on historical data. For internal causal relationships, a scoring-based structural learning method is used:

[0096]

[0097] Among them, G i is the candidate causal structure of device i, D i is the historical data of device i, Score is the scoring function, such as BIC (Bayesian Information Criterion) or MDL (Minimum Description Length), is the optimal internal causal structure.

[0098] It should be noted that for the relationship between devices R ij , a method based on time-delayed mutual information is used to detect causal relationships:

[0099]

[0100] in, represents the state of device i at time t, represents the state of device j at time t+τ, τ(tau) is the time delay, γ(gamma) is the threshold, I is the mutual information function, i and j are the numbers of different devices respectively, and t is the time point.

[0101] Step 2-3: Parameter learning and conditional probability table calculation.

[0102] According to an embodiment of the present application, based on the learned graph structure, the conditional probability table of each node is calculated. Its conditional probability distribution is:

[0103]

[0104] in, for The parent node set of is the corresponding parameter, i is the device number, and j is the variable number. Similarly, for the variable relationship across devices, the conditional probability is calculated:

[0105]

[0106] in, The variables in device k that affect device i variables, is the corresponding parameter, k is the number of another device, and l is the variable number in device k.

[0107] Steps 2-4: Causal dynamic update.

[0108] The method provided by this application dynamically updates the multi-agent causal relationship as new fault cases appear. new , calculate the likelihood under the current graph structure G:

[0109]

[0110] Among them, D new represents new observation data, G represents the current graph structure, L represents the likelihood function, n represents the total number of devices, i represents the device number, j represents the variable number, (V i ) represents the number of nodes of device i, represents the j-th variable of device i, Representing variables The parent node set of It should be understood that if the likelihood is lower than the threshold, the structure update process is triggered and steps 2-2 and 2-3 are executed again.

[0111] Step 3: Detect abnormal states through the constructed multi-agent causal system, perform causal intervention on abnormal variables, calculate counterfactual conditions, quantify the causal effects of various factors on the failure based on sensitivity analysis, determine the key root causes, and use the root cause analysis results as an important input for knowledge graph construction.

[0112] To address the problem of separating fault symptoms from root causes, the step provided in this application uses a counterfactual reasoning algorithm to trace the essential root cause from complex fault symptoms, providing a basis for accurate maintenance.

[0113] Step 3-1: Fault detection and anomaly location.

[0114] According to one embodiment of the present application, based on multi-agent causality, abnormal conditions in the reducer system are detected. For each monitored variable of each device, the deviation between its observed value and expected value is calculated:

[0115]

[0116] in, represents the deviation of the jth variable of device i, represents the observed value of the jth monitoring variable of device i, Indicates that the value calculated based on the parent node The expected value of Representing variables The parent node set of , i is the device number, j is the variable number, (·) represents the absolute value operation. It should be noted that if the deviation Exceeding the threshold Then Mark as an abnormal variable, represents the anomaly detection threshold of the j-th variable of device i.

[0117] In some embodiments, the threshold It can be determined in an adaptive way. For example, it can be determined based on the variable Setting dynamic thresholds based on historical volatility:

[0118]

[0119] in represents the anomaly detection threshold of the j-th variable of device i, For variables The standard deviation of , α is the adjustment coefficient.

[0120] Optionally, the change of threshold values ​​under different working conditions may also be considered, the current working condition may be identified by a working condition classifier, and then a corresponding threshold strategy may be selected.

[0121] According to another embodiment of the present application, in addition to deviation detection based on expected values, a deep anomaly detection model can also be used for auxiliary judgment. Specifically, an autoencoder or variational autoencoder can be trained to identify potential anomalies through reconstruction errors or anomaly scores, which is particularly suitable for situations with strong nonlinear relationships.

[0122] Step 3-2: Causal intervention implementation.

[0123] In this step, the set of variables marked as abnormal Perform causal intervention on each variable in and analyze the changes in the system state after the intervention. Implementation intervention where x normal is the normal value of the variable, and A represents the set of abnormal variables. In addition, the change in the probability of failure after intervention is calculated:

[0124]

[0125] in, Representing variables The change in failure probability before and after intervention, Y represents the system failure state, Represents a variable The probability of failure after intervention is performed, P(Y=fault) represents the probability of failure when no intervention is performed, x normal Representing variables The normal value of , i is the device number, j is the variable number. It should be understood that The larger the variable The greater the impact on the failure.

[0126] In some embodiments, the normal value x normalDifferent strategies can be used to select . For example, one can use the historical mean of the variable or estimate it based on normal operating data under current conditions. Alternatively, to reduce estimation bias, multiple intervention experiments can be performed using different normal values ​​and then averaging the effects.

[0127] According to a preferred embodiment of this application, causal intervention can be performed simultaneously in both the actual system and the digital twin. Virtual intervention in the digital twin allows for safe exploration of various intervention strategies, while in the actual system, small, controlled interventions can be used to verify the digital twin's predictions, further improving accuracy.

[0128] In a specific application example, the rolling mill reducer system of a steel plant frequently experienced abnormal vibrations. Traditional methods identify abnormal vibration sensor data as bearing failure, but the problem persists after the bearings are replaced. Through the causal intervention analysis provided by this application, it was found that the abnormal input shaft temperature was the root cause. After intervening in this variable (by adjusting the cooling system), the vibration problem was resolved, unnecessary bearing replacements were avoided, and a lot of downtime and maintenance costs were saved.

[0129] Step 3-3: Counterfactual calculation.

[0130] According to the embodiment of the present application, counterfactual condition calculations are performed on key variables to evaluate the impact of variables on failures in specific scenarios. Compute the counterfactual condition:

[0131]

[0132] in, means "under the condition that a fault is observed, if the variable Taking the normal value while other conditions remain unchanged, the probability of whether the fault will still occur is expressed as "P", which represents the probability function. Represents a variable The system failure state after counterfactual intervention is performed, X represents the set of all observed variables, Y represents the system failure state, x normal Representing variables where is the normal value of , i represents the device number, j represents the variable number, ← represents the counterfactual intervention operator, and fault represents the identifier indicating that the system is in a fault state, which is used to indicate when a reducer system fault occurs. Specifically, Y = fault indicates that the system fault state variable Y is in a fault state, meaning that the system is currently abnormal or faulty.

[0133] In some embodiments, the calculation of counterfactual conditions can be implemented using Monte Carlo methods. Specifically, multiple configurations are sampled from the backdoor adjustment set (the set of variables unaffected by the intervention variable). The intervention effect is then calculated under each configuration and the average is taken. Optionally, importance sampling can be used to improve computational efficiency.

[0134] Step 3-3: Equivariant inference method.

[0135] According to another embodiment of the present application, the counterfactual condition calculation can also take into account the time delay effect. Since the fault in the reducer system usually has a certain development process, the causal relationship between the variables may have a time delay. Therefore, the counterfactual condition can be expanded into a time series form:

[0136]

[0137] Among them, τ is the time delay parameter, which can be estimated by the time series causal discovery algorithm, t represents the current time, and Y t+τ represents the system fault state at time t+τ, represents the jth variable of device i at time t, X t represents the set of all observed variables at time t, Y t Indicates the system fault status at time t.

[0138] Step 3-4: Sensitivity analysis and root cause ranking.

[0139] In this step, based on sensitivity analysis, the causal effect of each factor on the failure is quantified. Calculate its causal effect:

[0140]

[0141] in, Representing variables The causal effect on the fault state Y, δ is a small disturbance, For variables The current value of Represents a variable The expected value of the fault state Y after the disturbance is applied, Representing variables The expected value of the fault state Y when the current value is maintained. Therefore, the potential root causes are sorted according to the size of the causal effect, and the set of key root dependent variables is determined, where RC represents the set of key root dependent variables and η is the threshold.

[0142] In some implementations, sensitivity analysis can be expanded to global sensitivity analysis. For example, methods such as the Sobol index or SHAP value can be used to assess the average causal effect of a variable across different value ranges, thereby avoiding bias caused by local perturbations. Alternatively, interactions between variables can be considered, using high-order Sobol indices or conditional SHAP values ​​to identify faults caused by the combined effects of multiple variables.

[0143] According to a preferred embodiment of the present application, the root cause ranking not only considers the size of the causal effect, but also the intervention cost and feasibility. Specifically, the root cause priority is defined as follows:

[0144]

[0145] in, Representing variables Root cause priority, Representing variables The causal effect on the fault state Y, Intervention variables the cost, is the intervention feasibility (a value between 0 and 1), i represents the device number, j represents the variable number, and Y represents the system fault status. In this way, the system can not only identify the key root cause but also provide cost-effective maintenance recommendations.

[0146] In a specific application example, the gearbox (special reducer) of a wind farm frequently alarmed, and traditional methods were unable to determine the root cause. Through the counterfactual reasoning and sensitivity analysis provided by this application, it was found that although multiple parameters were abnormal at the same time, input shaft asymmetry was the root cause, and its causal effect value was 2.8 times that of the second highest factor. After repairing this root cause, the system resumed normal operation, avoiding more serious failures and high repair costs.

[0147] Step 4: Combine the root cause analysis results and historical failure cases to build a fault knowledge graph to represent the relationship between fault types, symptoms, causes, and solutions. Continuously optimize as new cases accumulate.

[0148] To address the difficulties in knowledge sedimentation and accumulation, the step provided in this application constructs a fault knowledge graph and continuously optimizes it as new cases accumulate, thereby achieving effective sedimentation and reuse of knowledge.

[0149] Step 4-1: Knowledge graph construction.

[0150] According to one embodiment of the present application, a knowledge graph of reducer faults is constructed to represent the relationship between fault types, symptoms, causes, and solutions. The knowledge graph is represented as:

[0151] KG=(E,R,A)

[0152] Here, E is the entity set, including fault type, symptoms, components, causes, and solutions; R is the relationship set, such as "caused by," "manifested as," and "solution is." A is the attribute set, describing the attributes of the entities, such as fault severity and component location. KG stands for knowledge graph, E represents the entity set, R represents the relationship set, and S represents the attribute set.

[0153] It should be noted that the initial knowledge graph is constructed based on domain expert knowledge and historical failure cases, and the triple (e1, r, e2) is used to represent the existence of a relationship r between entities e1 and e2.

[0154] Step 4-2: Knowledge reasoning and query.

[0155] In this step, reasoning is performed based on the knowledge graph to support complex fault diagnosis queries. For a given set of fault symptoms S = s1, s2, ..., s m , check the possible causes of the failure:

[0156]

[0157] Where S represents the set of fault symptoms, s m Indicates the mth symptom, m indicates the total number of symptoms, c indicates the cause of the fault, Represents an existential quantifier, and KG represents a knowledge graph. Similarly, for the identified fault cause c, query the recommended solution:

[0158] Solutions(c)=sol|(c, sol)∈KL

[0159] Where sol represents the solution, and Solutions(c) represents the set of solutions corresponding to the fault cause c.

[0160] Step 4-3: Knowledge graph update algorithm.

[0161] According to the embodiment of the present application, a knowledge graph dynamic update algorithm is constructed to continuously enrich and optimize the knowledge graph as new fault cases emerge. new =symptoms,causes,solutions, perform the following updates:

[0162] 1. Add new entities: If symptoms, causes, or solutions contain entities that do not exist in the knowledge graph, they are added to the entity set E.

[0163] 2. Add new relations: Based on the relations between entities in the case, new triples are added to the knowledge graph.

[0164] 3. Update relationship weights: For existing relationships, increase their weights or confidence levels based on new cases.

[0165] In addition, the above process can be expressed as an update function:

[0166] KG t+1 =Update(KG t , C new )

[0167] Among them KG t Represents the knowledge graph at time t, KG t+1 represents the updated knowledge graph, C new represents a new fault case, t represents the time step, and Update represents the knowledge graph update function.

[0168] Step 4-4: Knowledge distillation and integration.

[0169] In this step, the knowledge in the knowledge graph is distilled and integrated to extract more refined fault diagnosis rules. Through the frequent pattern mining algorithm, high-confidence diagnosis rules are discovered from the knowledge graph:

[0170]

[0171] Among them, S is the symptom set, C is the cause set, support and confidence are support and confidence respectively, α and β are the corresponding thresholds, represents the implication relationship, Rules represents the diagnostic rule set, and ∪ represents the set union operation.

[0172] Step 5: Integrate the multi-agent causal analysis results, root cause diagnosis information, and fault knowledge graph to build a preventive maintenance decision-making system, predict the reducer system status, define the maintenance decision space, train the decision strategy, and combine Bayesian uncertainty estimation to quantify the decision reliability and generate the optimal maintenance strategy.

[0173] According to an embodiment of the present application, based on the analysis results of the aforementioned steps, this step constructs a preventive maintenance decision-making system to generate an optimal maintenance strategy and reduce the risk of unplanned downtime.

[0174] Step 5-1: State representation and prediction.

[0175] In this step, based on multi-agent causality and knowledge graph, the reducer system state is represented and predicted. The system state is represented as a vector s = (s1, s2, ..., s n ), where s i It represents the state vector of device i, which contains the values ​​of key monitoring variables, and n represents the total number of devices.

[0176] It should be noted that the trained state prediction model is used to predict the system state sequence in the next T time steps:

[0177] s t+1 , s t+2 ,...,s t+T =f pred (s t , s t-1 ,...,s t-k )

[0178] Among them, f pred is the state prediction function, which makes predictions based on the state of the past k time steps, s t+i It represents the predicted state at the t+ith moment, T represents the prediction time step, k represents the length of the historical state window, and t represents the current moment.

[0179] In some embodiments, the state prediction model can adopt different structures. For example, a sequence model such as a recurrent neural network (RNN), a long short-term memory network (LSTM), or a gated recurrent unit (GRU) can be used; a model such as a temporal convolutional network (TCN) or an attention mechanism can also be used. Optionally, to improve prediction accuracy, the prediction results of multiple models can be fused, and stability can be improved through ensemble learning methods such as bagging or boosting.

[0180] According to another embodiment of the present application, state prediction can also consider the influence of external factors. For example, for wind turbine gearboxes, meteorological information such as wind speed and direction can be incorporated into the prediction model; for reducers in the metallurgical industry, operating condition information such as production load can be considered. These external factors can serve as auxiliary inputs to the prediction model, improving prediction accuracy, especially in scenarios with large operating condition fluctuations.

[0181] In a specific application example, for the gearbox (special reducer) of a large wind power plant, the state prediction method provided by this application achieved a fault precursor prediction accuracy of 93.5% while taking into account external factors such as wind speed and wind direction. The average advance warning time was extended from 12 hours of traditional methods to 76 hours, providing maintenance personnel with sufficient preparation time.

[0182] Step 5-2: Decision space construction.

[0183] According to one embodiment of the present application, a maintenance decision space A=a1, a2, ..., a m , including various possible maintenance actions, such as “continue monitoring”, “plan maintenance”, “immediate shutdown”, etc., where A represents the maintenance decision space, a irepresents the i-th maintenance decision, and m is the total number of decisions. In addition, each decision ai is associated with a cost function C(a i , s), indicating that decision a is executed in state S i The cost of , where C represents the cost function and s represents the system state.

[0184] In some implementations, the decision space can adopt a hierarchical structure. For example, top-level decisions might be "Continue Monitoring," "Plan Maintenance," and "Immediate Shutdown." Under "Plan Maintenance," specific maintenance actions can be further broken down into "Replace Bearing," "Replace Gear," and "Adjust Alignment." Optionally, the decision space can also include complex actions, such as "Replace Bearing and Adjust Alignment." This is common in real-world maintenance, as some faults often affect multiple components simultaneously.

[0185] According to a preferred embodiment of the present application, the cost function takes into account multiple factors, including direct maintenance costs, downtime losses, spare parts inventory status, maintenance personnel availability, etc. Specifically, the cost function can be expressed as:

[0186] C(a i ,s)=C direct (a i )+C downtime (a i ,s)+C inventory (a i )+C personnel (a i )

[0187] Among them, C direct (a i ) represents the execution decision a i Direct maintenance cost, C downtime (a i , s) means executing decision a in state s i The resulting downtime loss, C inventory (a i ) represents decision a i Impact on spare parts inventory cost, C personnel (a i ) represents the execution decision a i personnel scheduling costs.

[0188] Step 5-3: Reinforcement learning model training.

[0189] In this step, a Markov decision process (MDP) model is constructed to train the reinforcement learning strategy. MDP is defined as:

[0190] M=(S,A,P,R,γ)

[0191] Among them, S is the state space, A is the action space, P is the state transition probability, R is the reward function, and γ is the discount factor.

[0192] According to an embodiment of the present application, the reward function is designed as follows:

[0193] R(s,a)=α·I(a;Y)-β·C(a,s)

[0194] Where s represents the current system state, a represents the maintenance action to be performed, I(a; Y) is the information gain provided by action a, which measures the extent to which the action reduces fault uncertainty, and Y represents the fault state variable; C(a, s) is the operational cost of performing action a in state S; α and β are trade-off coefficients used to balance the relationship between information gain and cost.

[0195] It should be noted that this application uses the Deep Q Network (DQN) or Soft Actor-Critic (SAC) algorithm to train the decision strategy:

[0196] π * (s)=argmax a Q * (s, a)

[0197] Among them, Q * (s, a) is the optimal action value function, which represents the long-term cumulative reward after executing action a in state s, π * (s) is the optimal strategy, argmax a Indicates that Q * (s, a) takes the action a of maximizing the value.

[0198] In some embodiments, reinforcement learning can employ different variations. For example, for large decision spaces, a dual deep Q-network (DDQN) or the dominant actor-critic (A2C) algorithm can be used. For high-dimensional state spaces, dimensionality reduction can be performed using an autoencoder before reinforcement learning. Alternatively, to accelerate training, imitation learning can be employed, where an initial policy is learned from expert-maintained records and then further optimized through reinforcement learning.

[0199] According to another embodiment of the present application, a hybrid decision-making model can be employed, taking into account the characteristics of actual maintenance decisions. Specifically, for common faults and clear maintenance rules, a rule-based decision-making system can be used directly; for complex scenarios or new fault modes, a reinforcement learning model can be used to generate decisions. This hybrid approach combines the interpretability of rule-based systems with the adaptability of reinforcement learning, making it particularly suitable for industrial maintenance scenarios.

[0200] In one specific application example, a manufacturing company's fleet of reducers saw maintenance costs reduced by 32% and equipment availability increased by 8.5% after implementing the preventive maintenance decision-making system provided in this application. The system generates optimal maintenance timing and plans based on fault severity, spare parts inventory status, and production plans, avoiding the "over-maintenance" or "under-maintenance" issues common with traditional methods.

[0201] Step 5-4: Uncertainty quantification and decision explanation.

[0202] According to the embodiment of the present application, the reliability of the decision is quantified by combining Bayesian uncertainty estimation. The state-action value function is represented by a Bayesian neural network:

[0203]

[0204] Where D is the training data, ω is the model parameter, s represents the system state, a represents the maintenance action, Q(s, a) represents the state-action value function, p(Q(s, a)|D) represents the posterior probability distribution of the value function under the given training data D, and f ω and denote the predicted mean and variance respectively, and N denotes normal distribution.

[0205] Generate explainable decision reports based on uncertainty estimates, including:

[0206] 1. Recommended maintenance decisions and their expected benefits;

[0207] 2. Uncertainty assessment of decision making;

[0208] 3. Decision rationale, based on knowledge graph and causal analysis results;

[0209] 4. Possible alternatives and their comparison.

[0210] In some embodiments, uncertainty quantification can be performed using various methods. For example, in addition to Bayesian neural networks, ensemble methods (such as deep ensembles) or Monte Carlo dropout techniques can be used to estimate predictive uncertainty. Alternatively, uncertainty can be decomposed into epistemic uncertainty (caused by insufficient model knowledge) and stochastic uncertainty (caused by inherent randomness in the system), with each being quantified and processed separately.

[0211] According to a preferred embodiment of this application, decision explanations are generated using a multi-level explanation structure. For frontline maintenance personnel, concise operational recommendations and key rationales are provided; for maintenance supervisors, detailed uncertainty analysis and cost-benefit assessments are provided; and for technical experts, a complete causal chain and knowledge graph reasoning path are provided. This multi-level explanation structure adapts to the needs of different users and improves the system's usability and credibility.

[0212] In one specific application example, a reducer maintenance system at an automobile manufacturer achieved "transparent decision-making." When the system recommended replacing a reducer bearing, it not only provided specific advice but also provided a reliability assessment (87% confidence level), a rationale based on causal analysis (abnormal bearing temperature was a result of bearing wear, not the cause), and a recommended time window (completion of maintenance within 72 hours was the most cost-effective). This explainability significantly increased maintenance personnel's trust in the system, increasing the adoption rate of system recommendations from an initial 46% to 92%.

[0213] Therefore, through the above five main steps, the intelligent fault detection method for reducers provided in this application can achieve high-accuracy fault diagnosis under conditions of scarce samples, analyze the interaction between multiple devices, trace the root cause of the fault, and provide reliable preventive maintenance decision support, effectively solving the problems existing in the existing technology.

[0214] A computer-readable storage medium is characterized in that it is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, an intelligent fault detection method for a reducer can be executed.

[0215] Device embodiment

[0216] According to one embodiment of the present application, an intelligent fault detection device for a reducer is also provided, comprising the following modules:

[0217] Multi-source domain knowledge transfer and sample synthesis module.

[0218] This module includes a source domain feature extraction unit, a domain adaptation unit, a physical constraint sample synthesis unit, and a small sample model training unit, and is responsible for solving the problem of scarcity of reducer fault samples.

[0219] It should be noted that the source domain feature extraction unit adopts a convolutional neural network structure to extract features from the historical data of the source domain reducer; the domain adaptation unit realizes knowledge transfer by minimizing the JS divergence of the feature distribution of the source domain and the target domain; the physical constraint sample synthesis unit generates synthetic fault samples that conform to physical laws based on the physical model and constraint conditions of the reducer; the small sample model training unit combines the real samples and synthetic samples of the target domain to train the fault diagnosis model.

[0220] Multi-agent causal construction and evolution module.

[0221] This module includes a multi-agent causal initialization unit, a causal discovery unit, a parameter learning unit, and a causal update unit, which is responsible for solving the limitations of single-machine independent analysis.

[0222] According to an embodiment of the present application, the multi-agent causal initialization unit constructs the initial multi-agent causal structure based on the device topology and domain knowledge; the causal discovery unit uses the causal discovery algorithm to learn the structure of the multi-agent causal structure based on historical data; the parameter learning unit calculates the conditional probability table of each node; and the causal update unit dynamically updates the multi-agent causal structure as new fault cases appear.

[0223] Counterfactual reasoning and root cause analysis module.

[0224] This module includes a fault detection unit, a causal intervention unit, a counterfactual condition calculation unit, and a sensitivity analysis unit, and is responsible for solving the problem of separating fault symptoms from root causes.

[0225] It should be understood that the fault detection unit detects abnormal conditions in the reducer system based on multi-agent causality; the causal intervention unit performs causal intervention on variables marked as abnormal and analyzes the changes in the system state after the intervention; the counterfactual condition calculation unit performs counterfactual condition calculation on key variables to evaluate the impact of variables on faults in specific scenarios; the sensitivity analysis unit quantifies the causal effect of each factor on the fault based on sensitivity analysis and determines the key root cause.

[0226] Knowledge graph construction and evolution module.

[0227] This module includes a knowledge graph construction unit, a knowledge reasoning unit, a knowledge graph update unit, and a knowledge distillation unit, and is responsible for solving the difficulties in knowledge precipitation and accumulation.

[0228] According to an embodiment of the present application, the knowledge graph construction unit constructs a reducer fault knowledge graph to represent the relationship between fault types, symptoms, causes and solutions; the knowledge reasoning unit performs reasoning based on the knowledge graph to support complex fault diagnosis queries; the knowledge graph update unit continuously enriches and optimizes the knowledge graph as new fault cases emerge; the knowledge distillation unit distills and integrates the knowledge in the knowledge graph to extract more refined fault diagnosis rules.

[0229] Preventive maintenance decision-making module.

[0230] This module includes a state representation and prediction unit, a decision space construction unit, a reinforcement learning model unit, and an uncertainty quantification unit, and is responsible for generating the optimal maintenance strategy.

[0231] In addition, the state representation and prediction unit represents and predicts the state of the reducer system based on multi-agent causality and knowledge graph; the decision space construction unit defines the maintenance decision space, which includes various possible maintenance operations; the reinforcement learning model unit trains the decision strategy to generate the optimal maintenance decision; the uncertainty quantification unit combines Bayesian uncertainty estimation to quantify the reliability of the decision and generate an explainable decision report.

[0232] Computer device embodiment

[0233] According to one embodiment of the present application, the provided reducer intelligent fault detection method can be implemented by a computer device, which includes:

[0234] Processor: executes computer program instructions to implement each step of the above method;

[0235] Memory: storage of computer program instructions and data, including but not limited to ROM, RAM, disk, etc.;

[0236] Communication interface: interacts with the reducer sensor network, receives sensor data and sends control instructions;

[0237] Database: stores information such as historical operating data of the reducer, fault cases, and knowledge graphs;

[0238] User interface: Provides a human-computer interaction interface to display fault diagnosis results and preventive maintenance recommendations.

[0239] It's important to note that computing equipment can be deployed on edge servers on-site or in cloud data centers, exchanging data with on-site equipment over the network. Edge deployment reduces network latency and is suitable for real-time monitoring and rapid response scenarios, while cloud deployment offers greater computing power and is suitable for complex model training and large-scale data analysis.

[0240] In a preferred embodiment of the present application, the provided intelligent fault detection method for reducers can adopt an edge-cloud collaborative deployment architecture. A lightweight fault detection and early warning model is deployed on the edge side, responsible for real-time data processing and preliminary anomaly detection; a complex causal reasoning, knowledge graph, and decision generation model is deployed on the cloud side, responsible for in-depth analysis and long-term optimization. Therefore, the edge side and the cloud side exchange data through a secure communication channel, achieving the dual advantages of real-time response and continuous learning.

[0241] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. An intelligent fault detection method for a reducer, characterized in that: include: Transfer the knowledge of fault characteristics extracted from the source domain reducer to the target domain reducer; And constrain the generation of fault samples to expand the training data set; Based on the expanded training dataset, multiple reducers are regarded as interrelated entities, and a multi-agent causal model is constructed and dynamically updated to capture the interactive impact between devices. Through the constructed multi-agent causal detection system, abnormal states are detected, causal intervention is performed on abnormal variables, counterfactual conditions are calculated, and the causal effects of various factors on the failure are quantified based on sensitivity analysis. The key root causes are determined, and the root cause analysis results are used as an important input for knowledge graph construction. Combining root cause analysis results with historical failure cases, we build a fault knowledge graph to represent the relationships between fault types, symptoms, causes, and solutions. This graph is continuously optimized as new cases accumulate. By integrating the results of multi-agent causal analysis, root cause diagnosis information and fault knowledge graph, a preventive maintenance decision-making system is constructed to predict the reducer system status, define the maintenance decision space, train the decision strategy, and quantify the decision reliability by combining Bayesian uncertainty estimation to generate the optimal maintenance strategy.

2. The intelligent fault detection method for a reducer according to claim 1, characterized in that: The steps of extracting fault feature knowledge from the source domain reducer and migrating it to the target domain reducer include: Apply convolutional neural network to the historical data of reducer in the source domain to extract fault feature representation; Construct a domain adaptation layer to map the source domain feature space to a feature space compatible with the target domain. Its optimization goal is to minimize the JS divergence between the source and target domain feature distributions. Generate synthetic fault samples based on the reducer physical model and constraints; Combine real samples from the target domain and synthetic samples to train the fault diagnosis model.

3. The intelligent fault detection method for a reducer according to claim 1, characterized in that: Considering multiple reducers as interrelated entities, the steps for constructing and dynamically updating multi-agent causal relationships include: Construct initial multi-agent causal relationships based on device topology and domain knowledge; Using causal discovery algorithms, we learn the structure and parameters of multi-agent causality based on historical data. Based on the learned graph structure, calculate the conditional probability table of each node; As new fault cases appear, the likelihood under the current graph structure is calculated. If the likelihood is lower than the threshold, the structure update process is triggered.

4. The intelligent fault detection method for a reducer according to claim 1, characterized in that: The steps for detecting abnormal conditions based on multi-agent causality and determining the key root causes include: For each monitoring variable of each device, the deviation between its observed value and the expected value calculated based on the parent node value is calculated. If the deviation exceeds the threshold, the variable is marked as an abnormal variable; Perform causal intervention on each variable in the set of variables marked as abnormal, and analyze the change of system state after the intervention; Perform counterfactual calculations on key variables to assess their impact on failures under specific scenarios; Based on sensitivity analysis, the causal effect of each factor on the failure is quantified, the potential root causes are ranked according to the size of the causal effect, and the set of key root cause variables is determined.

5. The intelligent fault detection method for a reducer according to claim 4, characterized in that: The root cause ranking not only considers the size of the causal effect, but also the intervention cost and feasibility. The root cause priority is calculated by dividing the causal effect value by the intervention cost and multiplying it by the intervention feasibility.

6. The intelligent fault detection method for a reducer according to claim 1, characterized in that: The steps to construct a fault knowledge graph to represent the relationship between fault types, symptoms, causes, and solutions include: Construct a knowledge graph of reducer faults and represent the relationships between entities through triples; Reasoning based on knowledge graphs supports complex fault diagnosis queries; Build a knowledge graph dynamic update algorithm to continuously enrich and optimize the knowledge graph as new fault cases emerge; Distill and integrate the knowledge in the knowledge graph to extract more refined fault diagnosis rules.

7. The intelligent fault detection method for a reducer according to claim 1, characterized in that: The steps to build a preventive maintenance decision-making system based on the analysis results include: Based on multi-agent causality and knowledge graph, the reducer system status is represented and predicted; Define a maintenance decision space containing various possible maintenance actions, with each decision associated with a cost function; Build a Markov decision process model and train reinforcement learning strategies; Combined with Bayesian uncertainty estimation, the reliability of decisions is quantified and explainable decision reports are generated.

8. The intelligent fault detection method for a reducer according to claim 7, characterized in that: The cost function considers several factors, including direct repair costs, downtime losses, spare parts inventory status, and maintenance personnel availability.

9. The intelligent fault detection method for a reducer according to claim 8, characterized in that: By deploying an edge-cloud collaborative architecture, lightweight fault detection and early warning models are deployed on the edge side, responsible for real-time data processing and preliminary anomaly detection; complex causal reasoning, knowledge graphs and decision generation models are deployed on the cloud, responsible for in-depth analysis and long-term optimization.

10. A computer-readable storage medium, characterized in that It is used to store computer-readable instructions, and when the computer-readable instructions are read by a computer, it can run an intelligent fault detection method for a reducer as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Automatic analysis system for detection data of radio frequency equipment

    CN121350884A

  • Intelligent early warning and remote operation and maintenance method for equipment fault of express cabinet

    CN121684880A

  • A method and system for quantifying the root cause of electric meter failure by fusing multi-dimensional risk features

    CN122508378A