Dynamic Perception and Adaptive Decision-making Fault Diagnosis System Based on Deep Reinforcement Learning
Through a dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning, the problem that traditional fault diagnosis methods are difficult to adapt to equipment degradation and complex fault modes is solved, and high-precision fault diagnosis and resource optimization are achieved.
Patent Information
- Application Number
- CN202510340126.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-21
AI Technical Summary
Traditional fault diagnosis methods are difficult to adapt to equipment degradation, environmental changes and complex fault modes, resulting in insufficient diagnostic accuracy and low resource utilization.
The dynamic perception and adaptive decision fault diagnosis system based on deep reinforcement learning is adopted, and data is collected through the multi-source sensor module, and the real-time state encoder dynamically adjusts the sampling rate and feature extraction. The deep Q network decision module designs the action space based on the state space and outputs the decision action through the ε-greedy strategy. The dynamic execution controller adjusts the hardware parameters and schedules the diagnostic algorithm, and the feedback learning loop performs online incremental learning and uncertainty processing.
It realizes high-precision identification of complex fault modes, improves diagnostic accuracy, optimizes computing resource utilization, adapts to equipment degradation and environmental changes, and reduces the risk of misjudgment and maintenance costs.
Smart Images

Figure CN119884980B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of decision-making fault diagnosis, and in particular to a dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning. Background Art
[0002] In industrial production, the fault diagnosis of mechanical equipment is a key link in preventing accidents, ensuring personal and equipment safety, and improving economic benefits. Fault diagnosis technology can reduce the probability of major sudden accidents, and at the same time, through prediction technology, it can achieve the purpose of extending the service life of mechanical equipment and reducing the probability and time of downtime. However, traditional fault diagnosis methods usually rely on fixed threshold judgments or static rule guidance, and it is difficult to meet the detection requirements of equipment degradation, environmental changes, and complex fault modes. In the prior art, the selection of sampling rate, feature extraction method, and diagnostic model is often fixed and cannot be dynamically adjusted according to the real-time state of the equipment, resulting in low utilization rate of available resources and insufficient diagnostic accuracy. In addition, the multi-source sensor data fusion and real-time processing capabilities in complex industrial scenarios are limited, further restricting the accuracy and real-time performance of fault diagnosis. Therefore, an intelligent fault diagnosis method that can dynamically perceive the equipment state and adaptively adjust decisions is needed to improve the accuracy of fault diagnosis, rationally utilize resources, and meet the real-time requirements of industrial scenarios. Summary of the Invention
[0003] The object of the present invention is to propose a dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning for the problem that traditional fault diagnosis methods in the background art usually rely on fixed threshold judgments or static rule guidance and are difficult to meet the detection requirements of equipment degradation, environmental changes, and complex fault modes.
[0004] The technical solution of the present invention: A dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning, including: a multi-source sensor module for collecting vibration, temperature, and acoustic signals of the equipment and performing preprocessing;
[0005] A real-time state encoder that dynamically adjusts the data sampling rate and sliding window length, extracts time-domain, frequency-domain, and time-frequency domain features, fuses the environmental noise level and CPU occupancy rate, and generates a low-dimensional state vector;
[0006] A deep Q-network (DQN) decision module that designs a hierarchical action space and a reward function based on the state space. The state space includes root mean square (RMS), kurtosis, peak frequency, energy entropy, noise level, and CPU occupancy rate. It balances exploration and exploitation through the ε-greedy strategy and outputs the optimal decision-making action;
[0007] The dynamic execution controller adjusts the hardware parameters according to the optimal decision actions and schedules the diagnostic algorithm to perform action combination;
[0008] The feedback learning loop dynamically updates the network parameters of the Deep Q-Network (DQN) and evaluates the decision confidence through online incremental learning and uncertainty processing, forming a closed-loop optimization system.
[0009] Optionally, the multi-source sensor module includes:
[0010] A three-axis acceleration sensor for collecting vibration signals;
[0011] A K-type thermocouple sensor for collecting temperature signals;
[0012] A microphone sensor for collecting acoustic signals;
[0013] The preprocessing includes synchronizing multi-modal data signals using hardware timestamps, denoising and filtering using wavelet thresholding, and normalization processing.
[0014] Optionally, the operations of the real-time state encoder include:
[0015] (a) Dynamic sampling rate adjustment: low-frequency sampling is used in the stable state of the device, and high-frequency sampling is switched to in the abnormal state to capture transient features;
[0016] (b) Sliding window length optimization: the window length is adaptively adjusted according to the signal period and real-time requirements;
[0017] (c) Multi-dimensional feature extraction, specifically including:
[0018] Time-domain features: Root Mean Square (RMS), kurtosis;
[0019] Frequency-domain features: Fast Fourier Transform (FFT) peak, wavelet packet energy;
[0020] Time-frequency domain features: wavelet packet node energy.
[0021] Optionally, the hierarchical action space specifically includes:
[0022] Sampling frequency: including the sampling frequencies of the temperature sensor F0, the acoustic sensor F1, and the vibration sensor F2;
[0023] Feature set: time-domain feature set SetA, frequency-domain feature set SetB, time-frequency domain feature set SetC;
[0024] Diagnostic model: threshold alarm M0, Support Vector Machine (SVM) M1, One-dimensional Convolutional Neural Network (1D-CNN) M2.
[0025] Optionally, the reward function is:
[0026] Among them, is the basic reward item, is the penalty item, is the accuracy weight, is the resource consumption penalty, is the false alarm penalty, is the fault detection accuracy, is the computing resource occupancy rate, is the number of false alarms.
[0027] Optionally, the action combinations of the dynamic execution controller include:
[0028] Hardware adjustment: Dynamically switch sensor channels and adjust the sampling rate of the analog-to-digital converter (ADC);
[0029] Algorithm scheduling: Generate action options through the cross-combination of sampling rate, feature set, and diagnostic model.
[0030] Optionally, the action options include:
[0031] {F0_SetA_M0}: Low-frequency sampling, time-domain features, threshold alarm;
[0032] {F1_SetB_M1}: Medium-frequency sampling, frequency-domain features, SVM classification;
[0033] {F2_SetC_M2}: High-frequency sampling, time-frequency domain features, 1D-CNN model.
[0034] Optionally, the feedback learning loop specifically includes:
[0035] Online incremental learning: Store instant reward data through the experience replay buffer and randomly sample and update DQN parameters regularly;
[0036] Uncertainty handling: When the decision confidence is lower than the preset threshold, trigger redundant detection or manual intervention.
[0037] Compared with the prior art, the present application includes at least one of the following beneficial technical effects:
[0038] By collaborating with multi-modal sensors such as vibration, temperature, and acoustics to collect data, combined with hardware timestamp alignment and wavelet noise reduction technology, the synchronization and accuracy of data are ensured, and the operating state of the device is comprehensively captured; Dynamically adjust the sampling rate and sliding window length, adopt low-frequency sampling when the device is stable to save resources, and switch to high-frequency sampling when abnormal to capture transient fault features, ensuring the timeliness and accuracy of signal analysis. Combining multi-dimensional feature extraction methods in the time domain, frequency domain, and time-frequency domain can effectively identify complex fault patterns and improve the diagnostic accuracy.
[0039] By using a reward function to balance the diagnostic accuracy, resource consumption, and false alarm rate, the optimal balance between diagnostic efficiency and cost is achieved. Real-time data is continuously stored in an experience replay buffer to dynamically update the parameters of the DQN model, adapting to equipment degradation, environmental noise changes, and new fault modes to avoid model obsolescence. Through confidence-based evaluation, when the decision confidence is low, redundant detection or manual intervention is triggered to reduce the risk of misjudgment and ensure the high reliability of the diagnostic results.
[0040] The present invention proposes a dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning, which can effectively solve the problems of dynamic data changes and complex fault mode recognition in mechanical equipment fault diagnosis, and at the same time significantly optimize the utilization rate of computing resources, providing an efficient and accurate solution for intelligent fault diagnosis in actual industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a structural diagram of the dynamic perception and adaptive decision-making system.
[0042] Figure 2 It is a distribution diagram of the execution actions proposed by the present invention.
[0043] Figure 3 It is a flowchart for updating the dynamic perception and adaptive decision-making system. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0045] Embodiment
[0046] As Figure 1 shown, the dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning proposed by the present invention includes a multi-source sensor module, a real-time state encoder, a deep Q-network (DQN) decision module, a dynamic execution controller, and a feedback learning loop. Each module will be described in detail below.
[0047] In this embodiment, the multi-source sensor module preprocesses the collected vibration, temperature, and acoustic signals through hardware timestamp alignment, wavelet threshold denoising, and normalization. The real-time state encoder dynamically adjusts the sampling rate and sliding window length, extracts time-domain (RMS, kurtosis), frequency-domain, and time-frequency-domain features, and generates a low-dimensional state vector containing the equipment health score and environmental noise level.
[0048] The calculation formula for the root mean square value is:
[0049] Where represents the th sampling value of the signal, is the total number of sampling points; serves to normalize the total energy (or total power) of the signal to obtain the average energy of the signal;
[0050] The formula for kurtosis is:
[0051] where: N is the total number of samples, is the i-th data value, represents the mean of the data, and -3 is an adjustment term used to define the kurtosis of the normal distribution as 0.
[0052] Among them, the decision-making module defines the state space and the hierarchical action space, adopts the ε-greedy strategy to balance exploration and exploitation. At the beginning of training, a random number between 0 and 1 is generated. If the random number is less than the exploration rate, a random action is executed; otherwise, the best action value function is selected according to the learned information, that is , where , is the long-term cumulative reward; the reward function design is divided into a basic reward and a penalty term:
[0053] In the above formula: where, is the basic reward term, is the penalty term, is the accuracy weight, is the resource consumption penalty, is the false alarm penalty, is the fault detection accuracy, is the computing resource occupancy rate, is the number of false alarms.
[0054] The designed reward function combines the diagnostic accuracy and the computing cost to achieve intelligent decision-making. During the training process, the mean squared error is used as the loss function to find the optimal parameter values, and its formula is:
[0055] In the above formula: is the optimal parameter, N is the total number of samples, r is the immediate reward, is the discount factor, Q(s,a) is the action value function at the current moment, is the action value function at the next moment.
[0056] After calculating the loss function, gradient descent is performed, and its formula is:
[0057] In the above formula: is the parameter vector, is the learning rate, is the gradient of the loss function with respect to the parameter.
[0058] As Figure 2 and Figure 3 shown, the dynamic execution controller executes the output actions of the decision-making module through hardware adjustment and algorithm scheduling. There are 27 selectable actions in total. For example, {F0_SetA_M0}, {F1_SetB_M1}, and {F2_SetC_M2} respectively represent three actions: the sampling rate is F0, time-domain feature extraction, and the diagnostic model is the threshold method; the sampling rate is F1, frequency-domain feature extraction, and the diagnostic model is SVM; the sampling rate is F2, time-frequency domain feature extraction, and the diagnostic model is 1D-CNN. The three-level linkage of the sampling rate, feature extractor, and diagnostic model, combined with their respective advantages, is applicable to scenarios with different complexities, computing speeds, and requirements. Through dynamic adjustment and optimization, while ensuring the diagnostic accuracy, the system significantly reduces the consumption of computing resources, is applicable to the fault diagnosis of mechanical equipment in complex industrial environments, can significantly reduce the maintenance cost, extend the equipment life, and improve the reliability and safety of the production line. The feedback learning loop realizes the online incremental learning and uncertainty processing mechanism, dynamically updates the model parameters and evaluates the decision confidence, forming a closed-loop optimization system. Through dynamic adjustment and optimization, while ensuring the diagnostic accuracy, the system significantly reduces the consumption of computing resources, is applicable to the fault diagnosis of mechanical equipment in complex industrial environments, can significantly reduce the maintenance cost, extend the equipment life, and improve the reliability and safety of the production line.
[0059] In this implementation, through the collaborative acquisition of multi-modal sensors such as vibration, temperature, and acoustics, combined with hardware timestamp alignment and wavelet denoising, high-precision data synchronization and redundancy removal are achieved, and the device status information is comprehensively captured. The sampling rate and sliding window length are dynamically adjusted to save resources when the device is stable and perform high-frequency sampling to capture transient features when abnormal, ensuring the timeliness and accuracy of signal analysis. The multi-dimensional feature fusion of the time domain (RMS, kurtosis), frequency domain (FFT peak, wavelet packet energy), and time-frequency domain (wavelet packet node energy) effectively identifies complex fault modes (such as nonlinear faults and multi-fault superposition). Through the dynamic perception and adaptive decision-making architecture of deep reinforcement learning, the limitations of traditional static diagnostic methods are broken through, and high-precision, low-latency, and low-cost fault diagnosis is achieved in complex industrial environments.
[0060] It should be noted that by continuously storing real-time data in the experience replay buffer and dynamically updating the DQN model parameters, it can adapt to equipment degradation, environmental noise changes, and new fault modes, avoiding model obsolescence. Based on confidence evaluation, when the decision confidence is low, redundant detection or manual intervention is triggered to reduce the risk of misjudgment and ensure the high reliability of the diagnostic results. By dynamically adjusting the sampling rate and diagnostic model, computing power can be saved during the stable period of the equipment, and resources can be concentrated during the abnormal period for rapid response, significantly reducing the consumption of computing resources and reducing the average power consumption by 20%-30%. It supports millisecond-level fault detection and decision-making, meeting the real-time monitoring requirements of the production line; the modular design facilitates the expansion of new sensors or diagnostic algorithms to adapt to different industrial scenarios. Through accurate prediction and early intervention, the unplanned downtime can be reduced by more than 50%, the service life of the equipment can be extended, the maintenance cost can be reduced, and the production efficiency can be improved.
[0061] The above specific embodiments are merely several alternative embodiments of the present invention. Based on the technical solution of the present invention and the relevant inspirations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning, characterized in that: include: Multi-source sensor module, used to collect vibration, temperature and acoustic signals of equipment and perform pre-processing; Real-time state encoder dynamically adjusts data sampling rate and sliding window length, extracts time domain, frequency domain and time-frequency domain features, integrates environmental noise level and CPU occupancy rate, and generates low-dimensional state vector; The deep Q network decision module designs a hierarchical action space and reward function based on the state space. The state space includes root mean square, kurtosis, peak frequency, energy entropy, noise level, and CPU occupancy. It balances exploration and utilization through the ε-greedy strategy and outputs the optimal decision action. A dynamic execution controller adjusts hardware parameters and schedules diagnostic algorithms according to the optimal decision action to perform action combination; The feedback learning loop dynamically updates the network parameters of the deep Q network and evaluates the decision confidence through online incremental learning and uncertainty processing, forming a closed-loop optimization system; The hierarchical action space specifically includes: Sampling frequency: including the sampling frequency of temperature sensor F0, acoustic sensor F1, and vibration sensor F2; Feature set: time domain feature set SetA, frequency domain feature set SetB, time-frequency domain feature set SetC; Diagnostic model: threshold alarm M0, support vector machine M1, one-dimensional convolutional neural network M2; The reward function is: R1=A-B=α×P accuracy -β×C occupancy_level -γ×FP Among them, A is the basic reward item, B is the penalty item, α is the accuracy weight, β is the resource consumption penalty, γ is the false alarm penalty, P accuracy is the fault detection accuracy, C occupancy_level To calculate resource occupancy, FP is the number of false positives; The action combination of the dynamic execution controller includes: Hardware adjustment: dynamically switch sensor channels and adjust analog-to-digital converter sampling rate; Algorithm scheduling: Generate action options by cross-combining sampling rate, feature set and diagnosis model; The action options include: {F0_SetA_M0}: low frequency sampling, time domain characteristics, threshold alarm; {F1_SetB_M1}: intermediate frequency sampling, frequency domain features, SVM classification; {F2_SetC_M2}: high-frequency sampling, time-frequency domain features, and 1D-CNN model.
2. According to claim 1, a dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning is characterized in that: The multi-source sensor module comprises: Three-axis acceleration sensor, used to collect vibration signals; K-type thermocouple sensor, used to collect temperature signals; A microphone sensor for collecting acoustic signals; The preprocessing includes aligning the multimodal data signal synchronization with hardware timestamps and performing noise reduction filtering and normalization processing using a wavelet threshold method.
3. According to claim 1, a dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning is characterized in that: The operation of the real-time state encoder includes: (a) Dynamic sampling rate adjustment: low-frequency sampling is used when the device is in a stable state, and high-frequency sampling is switched to capture transient features when the device is in an abnormal state; (b) Sliding window length optimization: adaptively adjust the window length according to the signal period and real-time requirements; (c) Multi-dimensional feature extraction, including: Time domain characteristics: root mean square, kurtosis; Frequency domain features: Fast Fourier transform peak, wavelet packet energy; Time-frequency domain features: wavelet packet node energy.
4. According to claim 1, a dynamic perception and adaptive decision-making fault diagnosis system based on deep reinforcement learning is characterized in that: The feedback learning loop specifically includes: Online incremental learning: store immediate reward data through the experience replay buffer and regularly update DQN parameters through random sampling; Uncertainty handling: When the decision confidence is lower than the preset threshold, redundant detection or manual intervention is triggered.
Citation Information
Patent Citations
Planetary gear box fault diagnosis method based on deep reinforcement learning model
CN112633245A
Single intersection signal lamp control method and device, terminal and storage medium
CN117649776A
Photovoltaic module fault diagnosis system and method based on deep learning
CN119474671A