Equipment fault diagnosis and prediction method based on deep learning

By installing multiple sensors on the equipment to collect multimodal data and building a hybrid deep learning model for data fusion, the problem of the difficulty in effectively fusing multiple modal data in existing technologies is solved, and high-accuracy diagnosis and prediction of equipment failures are achieved, thereby improving the safety and efficiency of equipment operation.

CN120632777APending Publication Date: 2025-09-12SHENZHEN JITON INTELLIGENT TECH CO LTD
View PDF 0 Cites 48 Cited by

Patent Information

Application Number
CN202510736099.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively integrate equipment operation data from multiple modalities, resulting in insufficient accuracy in equipment fault diagnosis. Especially in complex industrial environments, the diversity and complexity of equipment operation data place higher demands on the accuracy of fault diagnosis.

Method used

By installing a variety of sensors to collect multimodal data in real time, including vibration signals, temperature data, sound signals and image data, and preprocessing the data, a hybrid deep learning model is constructed. The convolutional neural network is used to process image data, and the recurrent neural network is used to process time series data. The features of different modal data are dynamically weighted and fused through the attention mechanism, and finally a comprehensive feature representation is generated.

Benefits of technology

It achieves high-accuracy diagnosis of complex equipment failures, improves the accuracy of fault diagnosis to over 95%, and supports dynamic optimization of equipment maintenance decisions by generating potential fault warning signals, thereby reducing maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632777A_ABST
    Figure CN120632777A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of equipment fault diagnosis, and discloses an equipment fault diagnosis and prediction method based on deep learning, and the method comprises the following steps: S1, collecting multi-modal data in real time through a plurality of sensors installed on equipment; s2, preprocessing the collected data; s3, constructing a hybrid deep learning model; s4, dynamic weighted fusion is performed on the features of different modal data by using an attention mechanism, and comprehensive feature representation is generated; s5, using the marked fault data and normal data to supervise and train the model; s6, inputting equipment operation data acquired in real time into the trained model, and judging the state of the equipment; and S7, generating a potential fault early warning signal based on a prediction result of the model. A piezoelectric vibration sensor and a thermal infrared imager are arranged on a motor bearing through vibration, temperature and sound sensors, vibration waveforms, thermal imaging slices and time-frequency diagrams are synchronously captured, and composite state characteristics such as mechanical wear and temperature anomaly of equipment are comprehensively reflected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment fault diagnosis, and specifically to a method for equipment fault diagnosis and prediction based on deep learning. Background Art

[0002] In the industrial field, timely diagnosis and prediction of equipment failures are core links to ensure production safety and improve equipment operating efficiency. Its technological development is closely related to the innovation of the Internet of Things, sensor technology and data analysis methods.

[0003] In the industrial sector, timely diagnosis and prediction of equipment failures are crucial for ensuring production safety, reducing maintenance costs, and improving equipment operating efficiency. Traditional fault diagnosis methods mainly rely on a single type of data (such as vibration signals or temperature data), which makes it difficult to fully reflect the complex operating status of the equipment. With the development of the Internet of Things, big data, and artificial intelligence technologies, multimodal data fusion has become a new trend. However, the existing technology lacks a method that can effectively fuse multiple modal data and utilize deep learning models for efficient processing. Especially in complex industrial environments, the diversity and complexity of equipment operating data place higher demands on the accuracy of fault diagnosis. Summary of the Invention

[0004] In response to the shortcomings of the existing technology, the present invention provides a method for equipment fault diagnosis and prediction based on deep learning, which solves the problem that traditional fault diagnosis methods mainly rely on a single type of data and are difficult to fully reflect the complex operating status of the equipment.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for diagnosing and predicting equipment faults based on deep learning, comprising the following steps: S1. Collect multimodal data in real time through various sensors installed on the equipment, including vibration signals, temperature data, sound signals and image data; S2. Preprocess the collected data, including data cleaning, normalization, noise reduction, and downsampling or interpolation of time series data; S3. Build a hybrid deep learning model that processes image or spatial data using a convolutional neural network (CNN) to extract spatial features, and processes time series data using a recurrent neural network (RNN) or a long short-term memory network (LSTM) to extract temporal features. S4. Use the attention mechanism to dynamically weight the features of different modal data to generate a comprehensive feature representation; S5. Use labeled fault data and normal data to supervise the model training, and optimize the model parameters through cross-validation to improve the generalization ability; S6. Input the real-time collected equipment operation data into the trained model, extract features and output fault diagnosis results to determine the equipment status and fault type; S7. Generate potential fault warning signals based on the model's prediction results and optimize equipment maintenance plans in combination with historical data.

[0006] By employing this technical solution, equipment operating data is collected synchronously through vibration, temperature, and sound sensors and cameras. This constructs a cross-dimensional dataset encompassing both spatial features (such as surface cracks and hot spot distribution) and temporal features (such as vibration amplitude fluctuations and temperature trends), overcoming the information limitations of traditional single-modal data. Taking motor bearings as an example, the simultaneous capture of vibration waveforms (reflecting mechanical wear), thermal imaging slices (reflecting temperature anomalies), and time-frequency maps (reflecting acoustic abnormalities) comprehensively characterizes the combined fault characteristics of bearing wear, temperature rise, and noise, addressing the problem of missed diagnoses caused by a single data dimension and increasing the recognition rate of combined faults to over 95%.

[0007] Preferably, the multimodal data fusion includes: synchronously collecting equipment operation data through a vibration sensor, a temperature sensor, a sound sensor and a camera, and aligning the timestamps of each modal data to ensure data synchronization.

[0008] Preferably, in the hybrid deep learning model: Input the 224×224 pixel RGB image or thermal image into the convolutional neural network of ResNet-50 architecture through industrial camera or infrared thermal imager, extract features through 7×7 convolution kernel, 3×3 convolution kernel and void convolution layer, and follow batch normalization and ReLU activation function after each layer: ReLU(x)=max(0,x), where x is a scalar value, which is the output result of the linear transformation of the previous layer of neural network. The 1024-dimensional spatial feature vector is output through maximum pooling. At the same time, the vibration signal of ≥20kHz is converted by STFT, and the temperature data of 30℃~90℃ is normalized. Differential processing is performed on 20Hz-20kHz sound signals to extract 20-dimensional MFCCs. Time series data is input into a 3-layer, 128-node bidirectional LSTM network with a 5-second time window. The tanh activation function and peephole connection structure are used to output a 256-dimensional time series feature vector. Each modality is synchronized with the IEEE1588 protocol through FPGA pulses to achieve a timestamp of ≤10ms. Linear interpolation is used to align data at different sampling rates to form a time-space matrix. Each row corresponds to a multimodal data block at the same time. After being input into CNN and LSTM respectively, the feature vectors are fused through a gated attention mechanism.

[0009] Preferably, the attention mechanism is implemented as follows: attention weights are calculated for the feature vectors output by the convolutional neural network and the long short-term memory network respectively, the weight distribution is dynamically adjusted according to the importance of the features to fault diagnosis, and a weighted fusion feature vector is generated.

[0010] Preferably, the normalization processing of the data preprocessing includes: mapping sensor data of different dimensions to the interval [0,1], and segmenting the time series data through a sliding window to meet the model input format requirements.

[0011] Preferably, the output of the fault diagnosis result includes: calculating the probability of the device being in a normal state or different fault types based on the fused feature vector, and determining the final diagnosis conclusion according to a preset threshold.

[0012] Preferably, the potential fault warning signal is generated by combining the failure probability predicted by the model with the historical operation trend of the equipment to establish a fault risk level assessment rule, and triggering a warning when the risk level exceeds a preset threshold.

[0013] Preferably, during the training process of the model: a labeled multimodal dataset is used, the loss function is optimized by a back-propagation algorithm, and a Dropout layer and regularization technology are introduced to prevent overfitting.

[0014] Preferably, the method further comprises: in equipment maintenance decision support, generating a maintenance priority list according to the early warning result.

[0015] Preferably, the method is applied to industrial equipment monitoring scenarios, including rotating machinery, power equipment, and production line devices, to achieve a fault diagnosis accuracy rate of ≥95% and a prediction lead time of ≥24 hours.

[0016] The present invention provides a method for diagnosing and predicting equipment faults based on deep learning. It has the following beneficial effects: 1. The present invention uses vibration, temperature, sound sensors and cameras to synchronously collect equipment operation data, performs microsecond-level alignment based on a unified timestamp, constructs a time-space matrix to store cross-modal correlation data, deploys piezoelectric vibration sensors and infrared thermal imagers on motor bearings, and synchronously captures vibration waveforms, thermal imaging slices and time-frequency graphs, thus overcoming the limitations of a single data dimension and comprehensively reflecting complex state characteristics such as mechanical wear and temperature anomalies of the equipment.

[0017] 2. This invention utilizes a layered architecture. The spatial feature extraction module processes image data using a ResNet-50 convolutional neural network to extract spatial features such as cracks and hot spots on the equipment surface. The temporal feature module analyzes vibration and temperature series using a bidirectional LSTM network to capture dynamic patterns such as amplitude fluctuations and trend changes. For example, for gearbox data, the CNN identifies gear wear images, while the LSTM analyzes periodic impacts in the vibration signal. The outputs of these two are then combined to form a joint spatiotemporal feature representation.

[0018] 3. This invention introduces a multi-head attention module to calculate attention weights for the spatial feature vectors output by the CNN and the temporal feature vectors output by the LSTM. Using a learnable matrix, the contribution of each modal feature to the fault is evaluated. For example, in bearing fault diagnosis, the weights of abnormal impact features in the vibration signal and high-temperature areas in the thermal imaging are dynamically increased, suppressing environmental noise interference and generating a fused feature vector focused on key information, thereby improving fault diagnosis accuracy by over 15%. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flow chart of the deep learning-based device fault diagnosis and prediction method of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solution of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0021] Please see the attached Figure 1 , an embodiment of the present invention provides a method for diagnosing and predicting device faults based on deep learning, comprising the following steps: S1. Collect multimodal data in real time through various sensors installed on the equipment, including vibration signals, temperature data, sound signals and image data; S2. Preprocess the collected data, including data cleaning, normalization, noise reduction, and downsampling or interpolation of time series data; S3. Build a hybrid deep learning model that processes image or spatial data using a convolutional neural network (CNN) to extract spatial features, and processes time series data using a recurrent neural network (RNN) or a long short-term memory network (LSTM) to extract temporal features. S4. Use the attention mechanism to dynamically weight the features of different modal data to generate a comprehensive feature representation; S5. Use labeled fault data and normal data to supervise the model training, and optimize the model parameters through cross-validation to improve the generalization ability; S6. Input the real-time collected equipment operation data into the trained model, extract features and output fault diagnosis results to determine the equipment status and fault type; S7. Generate potential fault warning signals based on the model's prediction results and optimize equipment maintenance plans in combination with historical data.

[0022] Specifically, for data collection and preprocessing, multiple types of sensors are deployed at key operating locations of the equipment, such as bearings, motors, and gearboxes. These sensors include vibration sensors for collecting high-frequency vibration signals, infrared temperature sensors for monitoring the surface temperature distribution of the equipment, acoustic sensors for capturing the noise spectrum of equipment operation, and industrial cameras for acquiring images of the equipment's appearance or thermal imaging data. Through an embedded data acquisition module, the multimodal data output by each sensor is synchronously collected in real time, and data alignment is performed based on a unified timestamp to ensure temporal matching of vibration signals, temperature curves, acoustic waveforms, and image frames. The collected raw data is then transmitted to the edge computing node for preprocessing: a sliding window filtering algorithm is used to remove high-frequency noise from the vibration signal, a median filter is used to smooth the temperature data, and a wavelet transform is used to reduce noise on the sound signal. For image data, grayscale normalization and contrast enhancement are used to optimize feature recognizability. In addition, sensor data of different dimensions are normalized, vibration acceleration values ​​are mapped to the [0,1] interval, temperature data are converted to relative percentages, and sound pressure level data are standardized to a decibel scale to meet the input requirements of the deep learning model. For time series data, linear interpolation is used to align time steps based on sampling frequency differences to ensure temporal consistency of multimodal data.

[0023] Hybrid deep learning model construction and training: A layered fusion deep learning architecture includes a spatial feature extraction module, a temporal feature extraction module, and a cross-modal fusion module. The spatial feature extraction module uses the ResNet-50 convolutional neural network to perform multi-layer convolution and pooling operations on the equipment image data to extract the spatial features of cracks, wear, or thermal anomalies on the equipment surface; the temporal feature extraction module uses a bidirectional LSTM network to encode the temporal data of vibration, temperature, and sound signals to capture the dynamic evolution of the equipment's operating status over time (such as periodic fluctuations in vibration amplitude and the trend of abnormal temperature increases). In the feature fusion stage, a multi-head attention mechanism is introduced: the spatial feature vector output by the convolutional neural network (CNN) and the temporal feature vector output by the long short-term memory network (LSTM) are respectively input into the attention weight calculation layer. The contribution of each modal feature to fault diagnosis is evaluated through a learnable weight matrix. After dynamically assigning weights, the weighted feature vectors are spliced ​​into a comprehensive feature representation. During the model training phase, a labeled dataset containing normal states and six typical faults (such as bearing wear, motor overheating, gear tooth breakage, etc.) was used. The cross-entropy loss function was used as the target, and the parameters were updated through the Adam optimizer. A five-fold cross-validation strategy was used to adjust the hyperparameters. At the same time, the Dropout layer (dropout rate 0.3) and L2 regularization (λ=0.001) were introduced to suppress overfitting, and finally a fault diagnosis model with excellent generalization performance was obtained.

[0024] For fault diagnosis and prediction, a convolutional neural network processes 224×224 pixel RGB images or thermal images (covering the surface appearance of the equipment, hot spot distribution, etc.), extracting the spatial features of surface defects or structural anomalies using the ResNet-34 / 50 architecture. A long short-term memory network processes time series data of vibration signals with a sampling rate of ≥20kHz (converted to time-frequency plots via a short-time Fourier transform (STFT, window length 1024)), temperature data normalized to the [0,1] range (converted to relative percentages based on a range of 30°C to 90°C), and 20Hz-20kHz sound signals (extracting 20-dimensional Mel-frequency cepstral coefficients (MFCCs)). A three-layer, 128-node bidirectional LSTM network captures the dynamic changes in the equipment's operating status. Each modal data is synchronized with a timestamp error of ≤10ms through hardware synchronization triggering (FPGA pulse signals) and linear interpolation. The data is input into the model in the form of a time-space matrix. The model generates a fused feature vector through the feature extraction layer, and the fully connected classification layer outputs the fault probability distribution. The diagnostic module determines the equipment status based on preset thresholds (such as failure probability ≥ 85%). If an anomaly is detected, it will further match the fault type library. For example, the harmonic components in the vibration spectrum indicate bearing imbalance, and abnormal temperature gradients indicate lubrication failure. At the prediction level, the model analyzes the fault evolution patterns in historical data, such as the continuous increase in vibration energy entropy and the temperature change rate exceeding the safety threshold, and combines the current diagnostic results to calculate the risk level (low / medium / high) of serious failures in the next 24 hours. When the risk level reaches "high", the early warning system automatically triggers a three-level response mechanism: pushing alarm information to the operation and maintenance terminal, generating a work order containing the fault location and maintenance suggestions, and linking the equipment control system to perform load reduction or shutdown protection. At the same time, the maintenance decision module dynamically optimizes the maintenance plan based on the equipment's historical maintenance records and real-time health status, achieving a reduction of more than 30% in predictive maintenance costs.

[0025] Multimodal data fusion includes: synchronously collecting equipment operation data through vibration sensors, temperature sensors, sound sensors, and cameras, and aligning the timestamps of each modal data to ensure data synchronization; Specifically, a high-precision sensor array is deployed at key monitoring points on the equipment (such as motor bearing end caps, gearbox housings, and high-voltage cable connectors). These sensors include piezoelectric vibration sensors (sampling rate ≥ 20kHz, range ±50g), infrared thermal imagers (resolution 640×480, temperature range -20°C to 500°C), broadband acoustic sensors (frequency response 20Hz-20kHz, sensitivity 50mV / Pa), and industrial-grade high-speed cameras (frame rate 120fps, 2 million pixels). Each sensor is connected to the data acquisition system via a unified hardware trigger interface. A field-programmable logic controller (FPGA) generates synchronized trigger pulse signals, ensuring that all sensors initiate data acquisition synchronously within microseconds. During the acquisition process, vibration signals are recorded as waveform streams, temperature data is stored as a two-dimensional thermal map sequence, acoustic signals are converted into time-frequency plots, and image data is stored as consecutive frames according to timestamps.

[0026] To address the data timeline offset caused by differences in sensor sampling rates, a three-level time alignment mechanism was designed: Hardware-level synchronization: A nanosecond-level time reference is provided to all data acquisition nodes via GPS modules or the IEEE 1588 Precision Time Protocol (PTP), eliminating system clock bias. Data-level alignment: A sliding window interpolation method is used for high-frequency vibration signals (20kHz) and low-frequency temperature data (1Hz), resampling the low-frequency data according to the vibration signal's timestamp to ensure one-to-one correspondence between data points. Software-level correction: A timestamp verification algorithm is run on edge computing nodes to detect and remove anomalous timestamps (e.g., time jumps > 10ms) caused by network latency or buffer overflows, and missing data segments are repaired through linear interpolation. The aligned multimodal data is stored in a time-space matrix, where each row corresponds to a vibration waveform segment, thermal imaging slice, acoustic spectrum block, and device appearance image at the same moment, forming a cross-modal correlation dataset.

[0027] In the hybrid deep learning model: Input the 224×224 pixel RGB image or thermal image into the convolutional neural network of ResNet-50 architecture through industrial camera or infrared thermal imager, extract features through 7×7 convolution kernel, 3×3 convolution kernel and void convolution layer, and follow batch normalization and ReLU activation function after each layer: ReLU(x)=max(0,x), where x is a scalar value, which is the output result of the linear transformation of the previous layer of neural network. The 1024-dimensional spatial feature vector is output through maximum pooling. At the same time, the vibration signal of ≥20kHz is converted by STFT, and the temperature data of 30℃~90℃ is normalized. Differential processing is performed on 20Hz-20kHz sound signals to extract 20-dimensional MFCCs. Time series data is input into a 3-layer, 128-node bidirectional LSTM network with a 5-second time window. The tanh activation function and peephole connection structure are used to output a 256-dimensional time series feature vector. Each modality is synchronized with the IEEE1588 protocol through FPGA pulses to achieve a timestamp of ≤10ms. Linear interpolation is used to align data at different sampling rates to form a time-space matrix. Each row corresponds to a multimodal data block at the same time. After being input into CNN and LSTM respectively, the feature vectors are fused through a gated attention mechanism.

[0028] Specifically, the convolutional neural network (CNN) uses the ResNet-34 architecture as the basic network for processing device image data. The input data is RGB images with a resolution of 224×224, covering the surface appearance of the device, thermal images and close-ups of key components. A multi-scale convolution kernel group is set at the front end of the network: the first layer uses a 7×7 convolution kernel (step size 2) to extract global contour features, and then a 3×3 convolution kernel stack (step size 1) is used to capture local details (such as bearing cracks, gear tooth wear, and cable insulation damage). A hollow convolution layer (expansion rate 2) is introduced in the deep layer of the network to expand the receptive field to identify large-scale structural anomalies (such as motor winding deformation). Each convolution layer is followed by batch normalization and ReLU activation function, where x is a scalar value (i.e., a single real number), which is the output result of the previous neural network layer (such as a convolution layer or a fully connected layer) after linear transformation, and the feature dimension is compressed through maximum pooling (2×2 window). Finally, the spatial feature extraction module outputs a 1024-dimensional feature vector, which contains the geometric distribution of surface defects of the equipment, the spatial gradient of the thermal anomaly area, and the texture change information.

[0029] A bidirectional long short-term memory (LSTM) network is constructed for time series analysis of vibration, temperature, and acoustic signals. The input data consists of preprocessed multi-channel time series signals with a time window length of 5 seconds (corresponding to 50,000 samples for vibration signals, 50 samples for temperature data, and 10,000 samples for acoustic signals). The network consists of three layers of LSTM units, each with 128 hidden nodes. It uses a tanh activation function and a peephole connection structure to enhance the modeling of long-term dependencies. Before the vibration signal is input, it is first converted into a time-frequency plot using a short-time Fourier transform (STFT, window length 1024) to extract frequency domain energy distribution features. Temperature data is differentially processed to generate a rate-of-change curve to highlight abnormal temperature increases. For the acoustic signal, 20-dimensional acoustic features are extracted using Mel-Frequency Cepstral Coefficients (MFCC). The LSTM network sequentially encodes three types of time series data: the vibration signal's time-frequency energy matrix is ​​passed through the LSTM layer to extract the time-varying patterns of harmonic components (such as the energy fluctuations in the 2kHz sideband caused by a bearing failure); the temperature rate curve captures the effects of thermal inertia and heat dissipation delay (such as the exponential temperature rise caused by lubrication failure); and the sound MFCC features analyze the transient characteristics of the noise spectrum (such as the impact sound pulse sequence caused by a broken gear tooth). Ultimately, the time series feature extraction module outputs a 256-dimensional feature vector that encodes the dynamic evolution of the equipment's operating status.

[0030] Cross-modal feature fusion and optimization: The 1024-dimensional spatial features output by the CNN and the 256-dimensional temporal features output by the LSTM are fed into the cross-modal fusion layer. A gated attention mechanism is used for dynamic weighting: First, a fully connected layer maps both feature types to the same dimension (512). The similarity matrix between the spatial and temporal features is then calculated to generate attention weights (ranging from 0 to 1). Features with higher weights are considered to contribute more to the current fault mode. The FocalLoss loss function is used during training to address sample imbalance (normal data accounts for 70%). Class weight coefficients (fault class weight is 3.0) are set, and a cosine annealing learning rate schedule (initial value 1e-3, period 10 epochs) is used. Model convergence is achieved within 100 epochs.

[0031] The implementation method of the attention mechanism is as follows: the attention weights are calculated for the feature vectors output by the convolutional neural network and the long short-term memory network respectively, the weight distribution is dynamically adjusted according to the importance of the features to fault diagnosis, and a weighted fusion feature vector is generated.

[0032] Specifically, during the feature fusion stage, the 1024-dimensional spatial feature vector extracted by the convolutional neural network (CNN) and the 256-dimensional temporal feature vector output by the long short-term memory network (LSTM) are first input into independent feature mapping layers. Their dimensions are then unified to 512 dimensions through a fully connected network to eliminate the impact of modal differences on weight calculation. Subsequently, a dual-channel attention scoring module is constructed: for spatial features, a channel attention mechanism is used to compress the spatial dimensions of the feature map through global average pooling (GAP) to generate a channel importance vector. This is then activated by a sigmoid function to obtain channel weights (ranging from 0 to 1) to enhance the feature response of key areas. For temporal features, a temporal attention mechanism is used to calculate the autocorrelation of the LSTM hidden state sequence and generate time step weights using the Softmax function to highlight key time nodes in fault evolution (such as the onset of the vibration impact signal).

[0033] The attention weights of the two types of modalities are cross-modally interacted: a cross-attention gating unit is designed, with the spatial feature weights as the query vector and the temporal feature weights as the key-value vector. The similarity matrix is ​​calculated through dot product to generate a dependency graph between the modalities.

[0034] During model training, the parameters of the attention module are jointly optimized through backpropagation and classification loss. An auxiliary supervisory signal is introduced to impose a sparsity constraint on the attention weight (L1 regularization, λ=0.01), forcing the model to focus on a few key features and avoid irrelevant noise interference. Field tests have shown that this mechanism increases the feature weight of the harmonic component in bearing fault diagnosis by 47%, and reduces the image background noise weight to below 0.05. In industrial pump unit tests, the attention fusion model achieved an accuracy of 97.2% for identifying seal leakage faults, an increase of 12.8% compared to the baseline model without the attention mechanism, and reduced the false alarm rate to 1.3%, significantly enhancing the model's robustness to complex working conditions.

[0035] The normalization process of data preprocessing includes mapping sensor data of different dimensions to the interval [0,1] and segmenting the time series data through a sliding window to meet the model input format requirements.

[0036] Specifically, for the normalization of multi-source sensor data, a hierarchical normalization strategy is adopted for sensor data of different physical dimensions: Vibration signal: Taking the output of a piezoelectric accelerometer as an example, the original data is the acceleration value (unit: g, range ±50g), which is normalized by minimum-maximum Linearly map the data to the interval [0,1], where X min and X max Determined based on historical data statistics during normal equipment operation; Temperature data: The two-dimensional temperature matrix (unit: °C) collected by the infrared thermal imager is normalized pixel by pixel. The temperature value T is converted into a relative percentage based on the rated operating temperature range of the equipment (such as 30°C to 90°C). At the same time, data exceeding the threshold (e.g. >100°C) are marked as abnormal areas; Sound signal: Sound pressure level data (dB) were converted to a linear scale using logarithmic compression. A piecewise normalization strategy was then used: steady-state background noise (<80 dB) was mapped to [0, 0.3], and transient impulse noise (≥80 dB) was mapped to [0.3, 1] to enhance the model's sensitivity to sudden noise. Image data: RGB images captured by industrial cameras are subjected to channel separation and grayscale normalization. Each pixel value is divided by 255 to achieve [0, 1] standardization, and the dynamic range of the radiation value of the thermal image is compressed (14-bit raw data is converted to 8-bit).

[0037] To adapt to the fixed-length input requirements of deep learning models, an adaptive sliding window mechanism is designed: Window parameter definition: Set the window length and step size according to the time scale of the equipment fault characteristics. For example, for the microsecond vibration shock of the early bearing failure, the window length is set to 200ms (covering 5 fault cycles) with a step size of 50ms; for the temperature slow-changing failure, the window length is extended to 10 minutes with a step size of 1 minute; Multi-rate data alignment: For sensor data with inconsistent sampling rates (such as vibration signal 20kHz, temperature 1Hz), linear interpolation is used within the window to unify the time resolution. Based on the vibration signal, the low-frequency temperature data is interpolated to a 20kHz timestamp to generate a synchronized multimodal data block; Edge filling and exception handling: When the window slides to the end of the data sequence, the mirror filling method is used to expand the data to avoid information truncation. If the missing data ratio in the window is greater than 10% (such as a sensor momentary disconnection), the window is discarded and the data retransmission mechanism is triggered; Feature normalization enhancement: The data in each window is individually Z-score normalized, and the mean μ and standard deviation σ in the window are calculated. Eliminate baseline drift caused by operating condition fluctuations, especially suitable for motor vibration signals with load changes.

[0038] The output of the fault diagnosis results includes: calculating the probability of the device being in a normal state or different fault types based on the fused feature vector, and determining the final diagnosis conclusion based on the preset threshold.

[0039] Specifically, after the hybrid deep learning model completes the multimodal data feature extraction and weighted fusion of the attention mechanism, the generated fusion feature vector will be input into the fully connected classification layer, and the probability distribution value of the equipment being in normal operation or various typical fault types (such as bearing wear, motor overheating, gear tooth breakage, etc.) will be calculated through the softmax activation function. The system presets a differentiated diagnostic threshold. When the probability of a certain type of fault exceeds the corresponding threshold, a three-level diagnostic conclusion generation mechanism is triggered: first, the current state of the equipment is determined to be abnormal. Secondly, based on the cosine similarity of the feature vector and the historical fault pattern matching based on the fault type library, the specific fault type is locked. Finally, a diagnostic report containing the fault probability value, type label and feature matching confidence is generated. The report is pushed to the operation and maintenance terminal in real time through a visual interface, showing the risk level of each fault type in the form of a heat map, and annotating key feature parameters to provide a precise positioning basis for on-site maintenance.

[0040] The softmax activation function is a normalized exponential function in a multi-classification scenario, which converts the original score output by the fully connected layer into a probability distribution. The formula is: Make the sum of the state probabilities equal to 1, where Zi: the raw score of the i-th category output by the fully connected layer, corresponding to the equipment state (such as normal, bearing wear, etc.), e: the base of the natural logarithm (about 2.71828), used to exponentially amplify the raw score, Sum the index values ​​of all n categories (including normal state and fault type) to achieve normalization, pi: the probability value of category i, reflecting the possibility of the device being in this state, all p i The sum is 1. In fault diagnosis, the output device is normal or has the probability of each fault type. It is often trained in conjunction with the cross entropy loss function to achieve accurate multi-category classification by amplifying the score difference, providing a quantitative basis for threshold determination and fault type matching.

[0041] The potential fault warning signal is generated by combining the fault probability predicted by the model with the historical operation trend of the equipment to establish a fault risk level assessment rule, and triggering a warning when the risk level exceeds the preset threshold.

[0042] Specifically, the system uses the real-time failure probability output by the model, combined with the evolutionary characteristics of similar failures in the equipment's historical operation database (such as a 15% increase in the vibration root mean square value for three consecutive days and an expansion of the temperature gradient standard deviation to 1.8 times the threshold), to extract trend characteristics of key parameters using a sliding time window. The system then fits the parameter change rate using an exponentially weighted moving average model. Based on this, a three-tiered risk assessment rule is established: when the failure probability is less than 40% and the trend is flat, it is judged as "low risk"; when the probability is 40% ≤ and the probability is less than 70%, or a single parameter trend exceeds the warning threshold (such as a temperature change rate of more than 2°C / h), it is marked as "medium risk"; when the probability is ≥70% and multiple parameter trends show a coordinated deterioration (such as a simultaneous increase in vibration energy entropy and an excessive sound signal kurtosis value), it is designated as "high risk." The system presets a risk level threshold (such as the high risk threshold is 70%). When the real-time calculated risk level exceeds the corresponding threshold, the graded warning mechanism is immediately triggered: low risk is prompted by a pop-up window on the operation and maintenance platform, medium risk is pushed through a warning SMS containing a trend curve, and high risk is automatically generated with a maintenance work order with fault location (such as motor bearing position), and the equipment control system is synchronously linked to execute reduced speed operation, and the abnormal area is highlighted on the digital twin interface.

[0043] During the model training process: a labeled multimodal dataset is used, the loss function is optimized through the back-propagation algorithm, and the Dropout layer and regularization technology are introduced to prevent overfitting.

[0044] Specifically, during the model training process, a labeled dataset containing multimodal data such as vibration signals, temperature curves, and sound spectra is used. The data is first cleaned, normalized, and enhanced, such as sliding window interception of time series data, random rotation and scaling of image data, and GAN is used to generate simulated fault data to balance the sample ratio. Then, a hybrid deep learning model is constructed (CNN processes spatial features, LSTM captures temporal dependencies, and Transformer extracts cross-modal associations). Dropout layers are inserted between fully connected layers, and L2 regularization is introduced to constrain model complexity. During training, the backpropagation algorithm is used to minimize the weighted cross entropy loss function using the Adam optimizer. At the same time, the early stopping mechanism and the addition of Gaussian noise to the input layer to simulate sensor errors are combined to prevent overfitting. During this period, the training indicators are monitored in real time through TensorBoard, the optimal model checkpoints are saved regularly, and Bayesian optimization is used to search for hyperparameters, ultimately improving the model generalization ability.

[0045] The method also includes: in equipment maintenance decision support, generating a maintenance priority list based on the early warning results, and associating the equipment operation log to recommend a maintenance plan.

[0046] Specifically, in the equipment maintenance decision support link, the system builds a dynamic maintenance strategy system based on the early warning results: first, based on the fault risk level (low / medium / high), the equipment downtime loss cost (such as the weight coefficient of key equipment on the production line × 1.5) and the maintenance history record (the number of days since the last maintenance), the maintenance priority index is calculated through the hierarchical analysis method (AHP) to generate a visual maintenance priority Gantt chart.

[0047] The method is applied to industrial equipment monitoring scenarios, including rotating machinery, power equipment, and production line devices, achieving a fault diagnosis accuracy rate of ≥95% and a prediction lead time of ≥24 hours.

[0048] Specifically, for rotating machinery (such as fans, motors, and pumps), high-frequency vibration sensors (sampling rate 20kHz) and infrared temperature sensors are deployed in key locations such as bearing seats and couplings. A CNN-LSTM hybrid model is used to analyze the envelope spectrum entropy and temperature field gradient changes of the vibration signal, successfully achieving a 96.8% diagnostic accuracy for bearing outer ring faults and a 48-hour lead time for early gearbox wear prediction. For power equipment (such as transformers and high-voltage switchgear), acoustic sensors are used to collect partial discharge noise and ultraviolet imagers to capture corona discharge images. Multi-source features are fused through an attention mechanism, achieving a 95.3% diagnostic accuracy for 110kV transformer oil chromatogram anomaly warning and a 36-hour lead time for bushing overheating failures. For production line equipment (such as automated production line conveyor belts and CNC machine tools), industrial cameras are used to identify abnormalities in the robot arm's motion trajectory, current sensors are used to monitor spindle load fluctuations, and combined with trend analysis of time series data, the diagnostic accuracy for tool wear in gear processing equipment on a certain automotive production line reached 97.1%, with a stable lead time of 24-30 hours.

[0049] Example 1 Rotating machinery (fan / motor / pump group) bearing fault diagnosis Application Scenario It is suitable for bearing health monitoring of rotating mechanical equipment such as fans, motors, and pumps, and for fault diagnosis and prediction such as bearing wear and imbalance.

[0050] Technical Solution Data collection Sensor deployment: Install high-frequency vibration sensors (sampling rate ≥ 20kHz, range ±50g) and infrared temperature sensors (temperature measurement range -20℃~500℃) at key locations such as bearing seats and couplings.

[0051] Multimodal data: vibration signal (time domain waveform, frequency domain envelope spectrum), temperature data (surface temperature distribution), sound signal (noise spectrum).

[0052] Data preprocessing of vibration signals: noise reduction is performed through sliding window filtering, short-time Fourier transform (STFT) is used to convert the signals into time-frequency diagrams, and frequency domain energy features are extracted.

[0053] Temperature data: Normalized to the [0,1] interval, calculate the temperature gradient change rate, and highlight abnormal warming trends.

[0054] Time alignment: Ensure the consistency of multi-modal data timestamps through hardware synchronization triggering (FPGA pulse signal) and linear interpolation.

[0055] Model building Hybrid Deep Learning Model: CNN module: ResNet-34 architecture, processes bearing thermal images, extracts spatial features such as surface cracks and wear (outputs 1024-dimensional vectors).

[0056] LSTM module: A bidirectional LSTM network that analyzes vibration time-frequency graphs and temperature change rates to capture dynamic patterns (outputs a 256-dimensional vector).

[0057] Attention mechanism: Dynamically weighted fusion of CNN and LSTM output features to enhance the weight of bearing fault features (such as 2kHz sideband energy).

[0058] Diagnosis and prediction of outcomes Diagnostic accuracy: The diagnostic accuracy of bearing outer ring faults reaches 96.8%.

[0059] Prediction advance time: The prediction of early gearbox wear can reach 48 hours in advance.

[0060] Early warning mechanism: When the vibration energy entropy continues to rise and the temperature gradient exceeds the standard, a high-risk warning is triggered and the equipment is linked to slow down.

[0061] Example 2 Power equipment (transformer / high-voltage switchgear) fault warning Application Scenario It is suitable for transformers, high-voltage switchgear and other equipment in power systems to monitor faults such as partial discharge and bushing overheating.

[0062] Technical Solution Data collection Sensor deployment: acoustic sensor (captures partial discharge noise, frequency response 20Hz-20kHz), ultraviolet imager (monitors corona discharge images), infrared thermal imager (resolution 640×480).

[0063] Multimodal data: sound signals (time-frequency diagram), UV discharge images, and temperature field thermograms.

[0064] Data preprocessing of sound signals: 20-dimensional acoustic features are extracted through Mel-Frequency Cepstral Coefficients (MFCC), and segmented normalization is performed to enhance sensitivity to impact noise.

[0065] Image data: UV image grayscale normalization, thermal image dynamic range compression (14-bit → 8-bit), highlighting discharge areas and temperature anomalies.

[0066] Model building Hybrid Deep Learning Model: CNN module: ResNet-50 architecture, analyzes the spatial features of UV discharge images (such as discharge spot location and morphology).

[0067] LSTM module: Processes sound MFCC features and time-series temperature data to capture the transient characteristics of discharge noise and the slow temperature change trend.

[0068] Attention mechanism: Channel attention is used to enhance the characteristic response of the discharge area in the image, and temporal attention is used to highlight the key time points of the noise pulse.

[0069] Diagnosis and prediction of outcomes Diagnostic accuracy: The diagnostic accuracy of 110kV transformer oil chromatographic abnormalities is 95.3%.

[0070] Prediction advance time: The prediction time for casing overheating failure is up to 36 hours.

[0071] Maintenance decision-making: Generates a maintenance priority list based on risk level (low / medium / high) and recommends maintenance plans (such as oil sampling and insulation layer replacement) based on historical data.

[0072] Example 3 Tool wear prediction for production line equipment (automatic conveyor belts / CNC machine tools) Application Scenario It is suitable for automated production line equipment, such as CNC machine tools and conveyor belt robotic arms, to monitor tool wear, abnormal robotic arm trajectory and other faults.

[0073] Technical Solution Data collection Sensor deployment: industrial camera (frame rate 120fps, 2 million pixels, monitoring the motion trajectory of the robotic arm), current sensor (monitoring spindle load fluctuations), and vibration sensor (sampling rate 20kHz).

[0074] Multimodal data: visual images (robot arm posture, tool appearance), current timing data (load curve), vibration signals (tool cutting vibration).

[0075] Data preprocessing visual data: Identify the tool wear area through target detection algorithm (such as YOLO), and normalize the image to the [0,1] range.

[0076] Current data: Differential processing is used to generate load change rate, and sliding window segmentation is used (window length 5 seconds, step length 1 second).

[0077] Vibration data: Wavelet transform is used to reduce noise and the envelope spectrum entropy value is extracted to reflect the degree of wear.

[0078] Model building Hybrid Deep Learning Model: CNN module: Processes tool images and extracts spatial features such as wear edge sharpness and hot spot coordinates.

[0079] LSTM module: Analyzes the current load curve and vibration envelope spectrum to capture load mutations and vibration energy trends.

[0080] Cross-modal fusion: Dynamically fuse visual and temporal features through a gated attention mechanism to suppress interference from conveyor belt background noise.

[0081] Diagnosis and prediction of outcomes Diagnostic accuracy: The diagnostic accuracy of tool wear for gear processing equipment on automotive production lines reached 97.1%.

[0082] Prediction lead time: The prediction lead time is stable at 24-30 hours, triggering the tool change work order in advance to avoid abnormal processing quality.

[0083] Optimization effect: Dynamically adjust maintenance plans based on prediction results, reducing maintenance costs by more than 30% and production line downtime by 40%.

[0084] The above examples verify the effectiveness of multimodal data fusion and deep learning in industrial fault diagnosis. Through differentiated sensor configurations and model architectures, they adapt to the physical characteristics of different equipment and achieve the transition from post-repair to predictive maintenance, which has good engineering application value and economic benefits.

[0085] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for diagnosing and predicting equipment faults based on deep learning, characterized in that: The following steps are involved: S1. Collect multimodal data in real time through various sensors installed on the equipment, including vibration signals, temperature data, sound signals and image data; S2. Preprocess the collected data, including data cleaning, normalization, noise reduction, and downsampling or interpolation of time series data; S3. Build a hybrid deep learning model that processes image or spatial data using convolutional neural networks to extract spatial features, and processes time series data using recurrent neural networks or long short-term memory networks to extract temporal features. S4. Use the attention mechanism to dynamically weight the features of different modal data to generate a comprehensive feature representation; S5. Use labeled fault data and normal data to supervise the model training, and optimize the model parameters through cross-validation to improve the generalization ability; S6. Input the real-time collected equipment operation data into the trained model, extract features and output fault diagnosis results to determine the equipment status and fault type; S7. Generate potential fault warning signals based on the model's prediction results and optimize equipment maintenance plans in combination with historical data.

2. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: The multimodal data fusion includes: synchronously collecting equipment operation data through vibration sensors, temperature sensors, sound sensors and cameras, and aligning the timestamps of each modal data to ensure data synchronization.

3. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: In the hybrid deep learning model: Input the 224×224 pixel RGB image or thermal image into the convolutional neural network of ResNet-50 architecture through industrial camera or infrared thermal imager, extract features through 7×7 convolution kernel, 3×3 convolution kernel and void convolution layer, and follow batch normalization and ReLU activation function after each layer: ReLU(x)=max(0,x), where x is a scalar value, which is the output result of the linear transformation of the previous layer of neural network. The 1024-dimensional spatial feature vector is output through maximum pooling. At the same time, the vibration signal of ≥20kHz is converted by STFT, and the temperature data of 30℃~90℃ is normalized. Differential processing is performed on 20Hz-20kHz sound signals to extract 20-dimensional MFCCs. Time series data is input into a 3-layer, 128-node bidirectional LSTM network with a 5-second time window. The tanh activation function and peephole connection structure are used to output a 256-dimensional time series feature vector. Each modality is synchronized with the IEEE1588 protocol through FPGA pulses to achieve a timestamp of ≤10ms. Linear interpolation is used to align data at different sampling rates to form a time-space matrix. Each row corresponds to a multimodal data block at the same time. After being input into CNN and LSTM respectively, the feature vectors are fused through a gated attention mechanism.

4. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: The attention mechanism is implemented as follows: attention weights are calculated for the feature vectors output by the convolutional neural network and the long short-term memory network respectively, the weight distribution is dynamically adjusted according to the importance of the features to fault diagnosis, and a weighted fusion feature vector is generated.

5. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: The normalization processing of the data preprocessing includes: mapping sensor data of different dimensions to the interval [0, 1], and segmenting the time series data through a sliding window to meet the model input format requirements.

6. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: The output of the fault diagnosis result includes: calculating the probability of the device being in a normal state or different fault types based on the fused feature vector, and determining the final diagnosis conclusion according to a preset threshold.

7. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: The potential fault warning signal is generated by combining the fault probability predicted by the model with the historical operation trend of the equipment to establish a fault risk level assessment rule, and triggering a warning when the risk level exceeds a preset threshold.

8. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: During the training process of the model: a labeled multimodal dataset is used, the loss function is optimized through the back-propagation algorithm, and the Dropout layer and regularization technology are introduced to prevent overfitting.

9. The method for diagnosing and predicting equipment faults based on deep learning according to claim 1, characterized in that: The method further includes: in equipment maintenance decision support, generating a maintenance priority list according to the early warning result.

10. The method for diagnosing and predicting equipment faults based on deep learning according to any one of claims 1 to 9, characterized in that: The method is applied to industrial equipment monitoring scenarios, including rotating machinery, power equipment, and production line devices, achieving a fault diagnosis accuracy rate of ≥95% and a prediction lead time of ≥24 hours.

Citation Information

Cited By

  • Cable joint abnormity early warning method and system

    CN120822162A

  • Vehicle steering system abnormal sound real-time detection method

    CN120869339A

  • Intelligent interconnection switch system based on deep learning and working method

    CN120879963A

  • Intelligent interconnected switch system based on deep learning and working method

    CN120879963B

  • Automobile chassis abnormal state detection system based on multi-modal fusion

    CN120907861A