Equipment fault prediction method and device based on deep learning, equipment and medium

Through the equipment failure prediction method based on deep learning, the equipment historical operation data is used to extract features and build a prediction model, which solves the problems of equipment downtime and resource waste in traditional maintenance methods, and realizes intelligent management and efficient maintenance of equipment failures.

CN120336813APending Publication Date: 2025-07-18JIANGYIN ACREL ELECTRICAL APPLIANCE MFGCO +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424487.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In traditional equipment maintenance methods, after-maintenance causes equipment to be shut down, and regular maintenance causes excessive repairs or insufficient repairs, resulting in waste of resources or hidden dangers in equipment.

Method used

The equipment failure prediction method based on deep learning is adopted, and the equipment historical operation data is collected, the time domain, frequency domain and time frequency domain characteristics are extracted, the deep learning model is constructed, the equipment failure probability is predicted, and maintenance measures are taken in advance.

Benefits of technology

It realizes intelligent prediction of equipment failures, avoids equipment downtime, reduces unnecessary regular maintenance, improves equipment reliability and management efficiency, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention relates to an equipment fault prediction method and device based on deep learning, equipment and a medium. The equipment fault prediction method based on deep learning comprises the following steps: collecting historical operation data of equipment, and marking an occurring equipment fault in the historical operation data; extracting a time domain feature, a frequency domain feature and a time-frequency domain feature related to the equipment fault from the historical operation data, and combining the time domain feature, the frequency domain feature and the time-frequency domain feature related to the equipment fault with the historical operation data to form training data; constructing a deep learning model taking the equipment operation data as input and the equipment fault probability as output; training a deep learning model by using the training data to obtain an equipment fault prediction model; and collecting current operation data of the equipment, inputting the current operation data into the equipment fault prediction model, and obtaining a current equipment fault probability output by the equipment fault prediction model. According to the method, the equipment failure probability is predicted, measures are taken in advance for maintenance, equipment shutdown is avoided, and unnecessary periodic maintenance is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of fault repair technology, and in particular to a method, device, equipment and medium for predicting equipment faults based on deep learning. Background Art

[0002] Traditional equipment maintenance methods are mainly divided into two types: post-maintenance and regular maintenance. Post-maintenance is to repair the equipment after it fails, which will cause equipment downtime, affect production efficiency, and even cause safety accidents. Regular maintenance is to repair and maintain the equipment at fixed time intervals. Although this method can prevent some failures, there are problems of over-maintenance or under-maintenance, resulting in waste of resources or equipment hidden dangers. Summary of the invention

[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, device, equipment and medium for predicting equipment failure based on deep learning.

[0004] In a first aspect, the present disclosure provides a method for predicting equipment failures based on deep learning, comprising: Collect historical operation data of the equipment and mark the equipment failures that have occurred in the historical operation data. The historical operation data includes at least one of the vibration frequency, temperature, pressure, current and voltage of the equipment during operation; Extracting time domain features, frequency domain features, and time-frequency domain features related to equipment failure from historical operation data, and combining the time domain features, frequency domain features, and time-frequency domain features related to equipment failure with the historical operation data as training data; Build a deep learning model with equipment operation data as input and equipment failure probability as output; Use the training data to train the deep learning model to obtain the equipment failure prediction model; The current operation data of the equipment is collected and input into the equipment failure prediction model to obtain the current equipment failure probability output by the equipment failure prediction model.

[0005] Optionally, before extracting the time domain features, frequency domain features and time-frequency domain features related to the equipment failure from the historical operation data, the method further includes: Preprocess the historical operation data, including data cleaning, denoising and normalization; Extract time domain features, frequency domain features, and time-frequency domain features related to equipment failure from historical operation data, and combine the time domain features, frequency domain features, and time-frequency domain features related to equipment failure with historical operation data as training data, including: Extract time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from the preprocessed historical operation data, and combine the time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures with the preprocessed historical operation data as training data.

[0006] Optionally, extract time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from historical operation data, including: According to the occurrence time point of the equipment failure, extract the target operation data within a preset time range before and after the occurrence time point; Calculate the mean, variance, peak-to-peak value, kurtosis, waveform factor, and impulse factor of the target operation data as time-domain features related to equipment failures; Perform a fast Fourier transform on the target operation data to generate a spectrum, and extract the spectral energy, main frequency component, and frequency band entropy in the spectrum as frequency-domain features related to equipment failures; Perform wavelet packet decomposition on the target operation data to extract the energy entropy of each node in the target operation data, perform short-time Fourier transform on the target operation data to generate a time-frequency diagram, extract the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram, and use the energy entropy of each node in the target operation data and the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram as frequency-domain features related to equipment failures.

[0007] Optionally, the deep learning model includes an input layer, a CNN layer, an Attention layer, an LSTM layer, and a fully connected layer; Among them, the training data is used as the input of the input layer, the output of the input layer is used as the input of the CNN layer, the output of the CNN layer is used as the input of the Attention layer, the output of the Attention layer is used as the input of the LSTM layer, the output of the LSTM layer is used as the input of the fully connected layer, and the fully connected layer outputs the failure probability of the equipment.

[0008] Optionally, the CNN layer includes a 1D convolutional layer and a max pooling layer, and the Attention layer is a Multi-HeadAttention layer.

[0009] Optionally, use the training data to train the deep learning model to obtain an equipment failure prediction model, including: Divide the training data into a training set, a validation set, and a test set; Input the training set into the deep learning model for forward propagation, and output the failure prediction probability of the equipment; Calculate the loss between the failure prediction probability and the real equipment failure in the training set; Update the network weights of the deep learning model according to the loss to complete one training round of the deep learning model; After each training epoch of the deep learning model, use the validation set to evaluate the precision, recall, and F1-Score of the deep learning model, and obtain the validation metrics after each training epoch of the deep learning model; If the validation metrics obtained by the deep learning model within a continuous preset number of training epochs do not improve, terminate the training of the deep learning model to obtain the target deep learning model after training in the preset number of training epochs; Use the test set to evaluate each target deep learning model, and determine the target deep learning model as the device fault prediction model according to the evaluation performance of each target deep learning model.

[0010] Optionally, input the training set into the deep learning model for forward propagation to output the device fault prediction probability, including: Input the training data into the input layer, and transmit the training data to the CNN layer through the input layer; Extract the local spatio-temporal features in the training data through the CNN layer and compress them to output the first feature sequence; Input the first feature sequence into the Attention layer, and perform dynamic weighting on the first feature sequence through the Attention layer to input the second feature sequence with attention weighting; Input the second feature sequence into the LSTM layer, and model the evolution law of the device state represented by the second feature sequence through the LSTM layer to output the third feature sequence representing the device state trend; Input the third feature sequence into the fully connected layer, and map the third feature sequence to the device fault prediction probability through the fully connected layer to output the fault prediction probability.

[0011] In a second aspect, the present disclosure provides a device fault prediction device based on deep learning, including: An acquisition module, configured to acquire the historical operation data of the device and mark the device faults that have occurred in the historical operation data. The historical operation data includes at least one of the vibration frequency, temperature, pressure, current, and voltage during device operation; A data processing module, configured to extract the time-domain features, frequency-domain features, and time-frequency domain features related to the device fault from the historical operation data, and combine the time-domain features, frequency-domain features, and time-frequency domain features related to the device fault with the historical operation data into training data; A model construction module, configured to construct a deep learning model with the device operation data as the input and the device fault probability as the output; A model training module, configured to train the deep learning model using the training data to obtain a device fault prediction model; A fault prediction module, configured to acquire the current operation data of the device and input it into the device fault prediction model to obtain the current device fault probability output by the device fault prediction model.

[0012] Optionally, before extracting time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from historical operation data, the data processing module is further configured to: Preprocess the historical operation data, where the preprocessing includes data cleaning, denoising, and normalization processing; When the data processing module extracts time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from historical operation data and combines the time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures with the historical operation data into training data, it is specifically configured to: Extract time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from the preprocessed historical operation data, and combine the time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures with the preprocessed historical operation data into training data.

[0013] Optionally, when the data processing module extracts time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from historical operation data, it is specifically configured to: According to the occurrence time point of the equipment failure, extract the target operation data within a preset time range before and after the occurrence time point; Calculate the mean, variance, peak-to-peak value, kurtosis, waveform factor, and impulse factor of the target operation data as time-domain features related to equipment failures; Perform a fast Fourier transform on the target operation data to generate a spectrum, and extract the spectral energy, main frequency component, and frequency band entropy in the spectrum as frequency-domain features related to equipment failures; Perform wavelet packet decomposition on the target operation data to extract the energy entropy of each node in the target operation data, perform short-time Fourier transform on the target operation data to generate a time-frequency diagram, extract the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram, and use the energy entropy of each node in the target operation data and the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram as frequency-domain features related to equipment failures.

[0014] Optionally, the deep learning model includes an input layer, a CNN layer, an Attention layer, an LSTM layer, and a fully connected layer; Among them, the training data is used as the input of the input layer, the output of the input layer is used as the input of the CNN layer, the output of the CNN layer is used as the input of the Attention layer, the output of the Attention layer is used as the input of the LSTM layer, the output of the LSTM layer is used as the input of the fully connected layer, and the fully connected layer outputs the failure probability of the equipment.

[0015] Optionally, the CNN layer includes a 1D convolutional layer and a max pooling layer, and the Attention layer is a Multi-HeadAttention layer.

[0016] Optionally, when the model training module trains a deep learning model using training data to obtain a device fault prediction model, it is specifically used for: Dividing the training data into a training set, a validation set, and a test set; Inputting the training set into the deep learning model for forward propagation, and outputting the fault prediction probability of the device; Calculating the loss between the fault prediction probability and the true device faults in the training set; Updating the network weights of the deep learning model according to the loss, and completing one training round of the deep learning model; After each training round of the deep learning model is completed, using the validation set to evaluate the precision, recall rate, and F1-Score of the deep learning model, and obtaining the validation metrics after each training round of the deep learning model; If the validation metrics obtained by the deep learning model within a continuous preset number of training rounds are not improved, terminate the training of the deep learning model, and obtain the target deep learning model after training in the preset number of training rounds; Using the test set to evaluate each target deep learning model, and determining the target deep learning model as the device fault prediction model according to the evaluation performance of each target deep learning model.

[0017] Optionally, when the model training module inputs the training set into the deep learning model for forward propagation and outputs the fault prediction probability of the device, it is specifically used for: Inputting the training data into the input layer, and transmitting the training data to the CNN layer through the input layer; Extracting and compressing the local spatio-temporal features in the training data through the CNN layer, and outputting the first feature sequence; Inputting the first feature sequence into the Attention layer, dynamically weighting the first feature sequence through the Attention layer, and inputting the second feature sequence with attention weighting; Inputting the second feature sequence into the LSTM layer, modeling the evolution law of the device state represented by the second feature sequence through the LSTM layer, and outputting the third feature sequence representing the device state trend; Inputting the third feature sequence into the fully connected layer, mapping the third feature sequence to the fault prediction probability of the device through the fully connected layer, and outputting the fault prediction probability.

[0018] In a third aspect, the present disclosure provides an electronic device, including a memory and a processor. Among them, a computer program is stored in the memory. When the computer program is executed by the processor, the method according to any item of the first aspect is implemented.

[0019] In a fourth aspect, the present disclosure provides a computer-readable storage medium, on which program instructions are stored. When the program instructions are executed, the method according to any item of the first aspect is implemented.

[0020] The technical solution provided by the present disclosure has the following advantages compared with the prior art: The device fault prediction method, device, equipment and medium based on deep learning provided by the present disclosure first collect the historical operation data of the device, mark the device faults that have occurred in the historical operation data, and then extract the time-domain features, frequency-domain features and time-frequency domain features related to the device faults from the historical operation data. The time-domain features, frequency-domain features and time-frequency domain features related to the device faults are combined with the historical operation data to form training data. Then, a deep learning model with the device operation data as the input and the device fault probability as the output is constructed, and the deep learning model is trained using the training data to obtain a device fault prediction model. Finally, the current operation data of the device is collected and input into the device fault prediction model, and the current device fault probability is obtained through the device fault prediction model. Through data analysis and deep learning algorithms, the present disclosure can predict the probability of device faults, so that measures can be taken in advance for maintenance, avoiding device downtime, improving device reliability, reducing unnecessary regular maintenance, avoiding over-maintenance, reducing maintenance costs, realizing intelligent device management, and improving management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a flowchart of the device fault prediction method based on deep learning provided by the embodiment of the present disclosure; Figure 2 It is a schematic diagram of the model architecture of the deep learning model provided by the embodiment of the present disclosure; Figure 3 It is a schematic diagram of the structure of the device fault prediction device based on deep learning provided by the embodiment of the present disclosure; Figure 4 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In order to be able to more clearly understand the above objects, features and advantages of the present disclosure, the following will further describe the solution of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0025] A lot of specific details are set forth in the following description to facilitate a full understanding of the present disclosure. However, the present disclosure may also be implemented in other ways different from those described herein. Obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.

[0026] Figure 1 The following is a flowchart of a device fault prediction method based on deep learning provided for an embodiment of the present disclosure. This method can be executed by a device fault prediction device based on deep learning, and the device fault prediction device based on deep learning can be configured in an electronic device, such as a server or a terminal. As Figure 1 shown, the device fault prediction method based on deep learning includes the following steps: S101. Collect historical operation data of the device, and mark the device faults that have occurred in the historical operation data. The historical operation data includes at least one of vibration frequency, temperature, pressure, current, and voltage during device operation.

[0027] Exemplarily, a three-axis acceleration sensor can be used to collect the vibration frequency during device operation, an infrared thermocouple temperature sensor can be used to collect the temperature during device operation, a piezoelectric pressure sensor can be used to collect the pressure during device operation, and a Hall effect sensor can be used to collect the current and voltage during device operation.

[0028] These data can indicate the state of the device during operation. The data can be transmitted to the cloud (such as AWS IoT Core) in real time through an industrial Internet of Things gateway, stored in a time series database, and form the historical operation data of the device. At the same time, according to the maintenance reports of the device by relevant staff or the maintenance logs saved for the device, determine the device faults that have occurred during the period corresponding to the historical operation data of the device, and mark these occurred device faults and the time nodes when the device faults occurred in the historical operation data of the device.

[0029] S102. Extract time domain features, frequency domain features, and time-frequency domain features related to device faults from the historical operation data, and combine the time domain features, frequency domain features, and time-frequency domain features related to device faults with the historical operation data to form training data.

[0030] According to the device faults marked in the historical operation data and the time nodes when they occurred, the historical operation data near the time nodes when the device faults occurred can be extracted as operation parameters related to the device faults, and the time domain features, frequency domain features, and time-frequency domain features in the device operation parameters related to the device faults can be extracted through calculation.

[0031] Time-domain features can be used to capture short-term fluctuations or long-term trends of device operating parameters. For example, a slow increase in motor current may indicate bearing wear, and a sudden peak in vibration signals may indicate mechanical shock. Frequency-domain features can identify periodic faults during device operation. For example, specific frequency vibrations caused by spalling on the tooth surface of a gearbox can be identified. Time-frequency domain features can identify non-steady signals during device operation. For example, transient vibrations during the start-up and shutdown phases of a device can be identified.

[0032] In some embodiments, before extracting time-domain features, frequency-domain features, and time-frequency domain features related to device faults from historical operation data, it further includes: preprocessing the historical operation data, and the preprocessing includes data cleaning, denoising, and normalization.

[0033] Preprocessing the historical operation data can improve data quality. Data cleaning includes missing value processing, outlier detection, etc. Missing value processing is performed by linear interpolation or forward filling, and outlier detection can use the Isolation Forest algorithm to remove outliers. Denoising includes time-domain denoising and frequency-domain denoising. Time-domain denoising can use a time-domain filter, and frequency-domain denoising can use the wavelet threshold denoising algorithm. Normalization includes data size alignment and standardization processing. Data size alignment can use the Min-Max normalization algorithm to scale the data into a preset data range interval, and standardization processing is used for non-uniformly distributed data. For example, the Z-Score statistical standardization method can be used.

[0034] Correspondingly, extracting time-domain features, frequency-domain features, and time-frequency domain features related to device faults from historical operation data, and combining the time-domain features, frequency-domain features, and time-frequency domain features related to device faults with the historical operation data into training data includes: extracting time-domain features, frequency-domain features, and time-frequency domain features related to device faults from the preprocessed historical operation data, and combining the time-domain features, frequency-domain features, and time-frequency domain features related to device faults with the preprocessed historical operation data into training data.

[0035] Embodiments of the present disclosure improve data quality by preprocessing historical operation data, which is beneficial for subsequent feature extraction and model training.

[0036] In some embodiments, time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures are extracted from historical operation data, including: according to the occurrence time point of the equipment failure, extracting the target operation data within a preset time range before and after the occurrence time point; calculating the mean, variance, peak-to-peak value, kurtosis, waveform factor, and impulse factor of the target operation data as time-domain features related to equipment failures; performing a fast Fourier transform on the target operation data to generate a spectrum, and extracting the spectral energy, main frequency component, and frequency band entropy in the spectrum as frequency-domain features related to equipment failures; performing wavelet packet decomposition on the target operation data to extract the energy entropy of each node in the target operation data, performing a short-time Fourier transform on the target operation data to generate a time-frequency diagram, and extracting the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram, and taking the energy entropy of each node in the target operation data and the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram as frequency-domain features related to equipment failures.

[0037] Centering on the time point of fault occurrence, expand a fixed time window forward and backward (for example: 2 hours before the fault to 10 minutes after the fault), and extract the historical operation data within this window, which can capture the complete state evolution before and after the fault occurrence, such as fault precursors, fault outbreaks, and post-fault states.

[0038] Calculate the following statistics on the historical operation data within the fixed time window as time-domain features related to equipment failures, including mean, variance, peak-to-peak value, kurtosis, waveform factor, and impulse factor. The mean can reflect the average intensity of the signal. The variance can characterize the degree of signal fluctuation. The peak-to-peak value is the difference between the maximum value and the minimum value, which can detect abnormal amplitudes. Kurtosis can measure the sharpness of the signal distribution and can capture instantaneous impacts (such as bearing cracks). The waveform factor is the ratio of the root mean square value to the absolute average value, which can distinguish between stationary and non-stationary signals. The impulse factor is the ratio of the peak value to the absolute average value, which can identify sudden impacts. For example, when the vibration peak-to-peak value of a certain bearing of the equipment continuously increases and the kurtosis suddenly increases 1 hour before the fault, it can be determined as a sign of early wear.

[0039] Perform a fast Fourier transform on the historical operation data within the fixed time window, and extract features such as spectral energy, main frequency component, and frequency band entropy after generating the spectrum. The spectral energy can represent the energy distribution of each frequency band and is used to detect resonance or harmonic anomalies. The main frequency component is the frequency component with the highest energy, such as the motor fundamental frequency and the gear meshing frequency. The frequency band entropy can measure the complexity of the spectrum, and a sudden increase in the entropy value may indicate a random fault, such as gear looseness.

[0040] Wavelet packet decomposition of historical operation data within a fixed time window can obtain multiple sub-band nodes in the historical operation data. Calculating the energy entropy of each node can measure the complexity of the historical operation data within the fixed time window in the time-frequency domain. For example, a crack in the fan blade will cause a sudden increase in the energy entropy of the high-frequency sub-band. Performing short-time Fourier transform on the historical operation data within the fixed time window generates a time-frequency diagram. The energy concentration and instantaneous frequency change rate in a specific time-frequency region of the time-frequency diagram indicate that the device may have certain faults. For example, when the motor is blocked, the low-frequency energy surges, and when the bearing lacks oil, it causes periodic fluctuations in the high-frequency components.

[0041] In the embodiments of the present disclosure, by extracting the time-domain features, frequency-domain features, and time-frequency domain features of the historical operation data within a preset range near the fault occurrence time point, the multi-dimensional feature representation of the device state during the fault occurrence is strengthened. When the subsequent model training is performed using the training data containing these features, the model can learn these multi-dimensional features with stronger fault correlation, improving the accuracy and robustness of the model's fault prediction.

[0042] S103. Construct a deep learning model with the device operation data as the input and the device fault probability as the output.

[0043] Figure 2 is a schematic diagram of the model architecture of the deep learning model provided by the embodiments of the present disclosure. As Figure 2 shown, the deep learning model includes an input layer, a CNN layer, an Attention layer, an LSTM layer, and a fully connected layer. Among them, the training data is used as the input of the input layer, the output of the input layer is used as the input of the CNN layer, the output of the CNN layer is used as the input of the Attention layer, the output of the Attention layer is used as the input of the LSTM layer, the output of the LSTM layer is used as the input of the fully connected layer, and the fully connected layer outputs the fault probability of the device.

[0044] In some embodiments, the CNN layer includes a 1D convolutional layer and a max pooling layer, and the Attention layer is a Multi-Head Attention layer.

[0045] The input layer receives the input device historical operation data, can define the data format, clarify the dimension and type of the data (such as floating-point numbers), and ensure that the subsequent neural network layers can correctly parse the data.

[0046] The CNN layer is a convolutional neural network layer, including a 1D convolutional layer and a max pooling layer. The 1D convolutional layer slides along the time dimension of the input historical operation data to extract local spatio-temporal features of the historical operation data. The max pooling layer can compress the length of the feature data extracted by the 1D convolutional layer, retain key features, reduce the amount of feature processing, and output a feature sequence containing key features. The deep learning model can automatically learn local features in the time domain through the CNN layer. At the same time, because the filter response of the convolutional kernel in the CNN layer can also implicitly capture frequency domain features.

[0047] The Attention layer is connected after the CNN layer and dynamically weights the feature map output by the CNN. The Attention layer in the embodiments of the present disclosure selects a Multi-Head Attention (multi-head self-attention) layer. The Attention layer can calculate the importance weights of different time steps and channels in the feature sequence, perform weighted summation on the feature sequence, enhance key features, and generate an attention-weighted feature sequence. The deep learning model can focus on key time periods related to faults through the Attention layer and solve the long-distance dependence problem, such as the association between early minor anomalies and later faults.

[0048] The LSTM layer is connected after the Attention layer and processes the weighted feature sequence. When processing the input sequence, the LSTM layer will default to generate a hidden state for each time step. The hidden state is a dynamic encoding of the device operation history, integrating short-term events and long-term trend information of the device operation state. The output of the LSTM layer is a feature sequence representing the hidden states of all time steps, and the sequence output can be compressed into a fixed-length feature vector by means such as pooling, weighting, or taking the last hidden state for the fully connected layer to map to probability values. The deep learning model can model the long-term evolution law of the device state through the LSTM layer and integrate the information of the device operation state in the historical operation data.

[0049] The fully connected layer is connected after the LSTM layer and receives the hidden state feature sequence output by the LSTM. In the embodiments of the present disclosure, the activation function of the fully connected layer is set to the sigmoid function. The deep learning model maps the feature sequence output by the LSTM to a fault probability through the fully connected layer.

[0050] The embodiments of the present disclosure construct a hybrid deep learning model including a CNN layer, an Attention layer, and an LSTM layer, and realize high-precision prediction of complex device faults by multi-level collaborative learning of fault-related features in historical operation data.

[0051] S104. Use the training data to train the deep learning model to obtain a device fault prediction model.

[0052] In specific implementation, a deep learning model is trained using training data to obtain a device fault prediction model, including: dividing the training data into a training set, a validation set, and a test set; inputting the training set into the deep learning model for forward propagation to output the fault prediction probability of the device; calculating the loss between the fault prediction probability and the actual device faults in the training set; updating the network weights of the deep learning model according to the loss to complete one training round of the deep learning model; after each training round of the deep learning model is completed, the precision, recall rate, and F1-Score of the deep learning model are evaluated using the validation set to obtain the validation metrics after each training round of the deep learning model; if the validation metrics obtained by the deep learning model within a continuous preset number of training rounds are not improved, the training of the deep learning model is terminated to obtain the target deep learning model trained in the preset number of training rounds; the test set is used to evaluate each target deep learning model, and the target deep learning model serving as the device fault prediction model is determined according to the evaluation performance of each target deep learning model.

[0053] The training data can be divided into a training set (70% of the training data), a validation set (20% of the training data), and a test set (10% of the training data) in chronological order to avoid future information leakage. The training set data is input into the deep learning model for forward propagation. The training set data passes through CNN → Attention → LSTM → fully connected layer in sequence to output the fault prediction probability, and then the loss between the fault prediction probability and the actual fault labels in the training set data is calculated. The loss function can be set as the Focal Loss function, and then the network weights of the deep learning model are updated according to the calculated loss through the gradient descent algorithm, thus completing one training round (Epoch) of the deep learning model. When inputting the training set data into the deep learning model for training, the training set can be divided into multiple small batches for training to improve the efficiency and memory utilization rate during model training.

[0054] The deep learning model is trained repeatedly for multiple rounds. After each training round is completed, the precision, recall rate, and F1-Score of the trained deep learning model are evaluated using the validation set as the validation metrics after each training round of the deep learning model to judge the generalization ability of the model. In addition, Bayesian optimization can be used to search for the best learning rate, the best convolution kernel size, and other parameters for hyperparameter optimization of the model. Precision is the proportion of samples predicted as positive classes that are actually positive classes, reflecting the accuracy of the model prediction. Recall rate is the proportion of samples that are actually positive classes and are correctly predicted, reflecting the ability of the model to retrieve all relevant samples. F1-Score is used to balance precision and recall rate. For example, when the cost of missed detection of device faults is much higher than false detection (such as high downtime losses), the recall rate of the model is monitored preferentially. When it is necessary to balance false alarms and missed reports, the F1-Score of the model is monitored preferentially.

[0055] After completing a preset number of training rounds or when the validation metrics obtained by the deep learning model within a continuous preset number of training rounds do not improve, terminate the model training. Select the trained deep learning models obtained in multiple training rounds where the validation metrics do not improve as the target deep learning models. These multiple target deep learning models are the ones with the highest validation metrics, that is, the deep learning models with the best performance obtained through training in multiple training rounds. Evaluate each target deep learning model using the test set, and based on the evaluation performance of each target deep learning model, select the target deep learning model with the highest evaluation performance as the equipment fault prediction model.

[0056] In some embodiments, input the training set into the deep learning model for forward propagation to output the fault prediction probability of the equipment, including: input the training data into the input layer, and transmit the training data to the CNN layer through the input layer; extract the local spatio-temporal features in the training data through the CNN layer and compress them to output the first feature sequence; input the first feature sequence into the Attention layer, dynamically weight the first feature sequence through the Attention layer, and input the second feature sequence with attention weighting; input the second feature sequence into the LSTM layer, model the evolution law of the equipment state represented by the second feature sequence through the LSTM layer, and output the third feature sequence representing the equipment state trend; input the third feature sequence into the fully connected layer, map the third feature sequence to the fault prediction probability of the equipment through the fully connected layer, and output the fault prediction probability.

[0057] The embodiments of the present disclosure train the constructed deep learning model through the above model training steps, enabling the deep learning model to efficiently learn the multi-dimensional features of the equipment operation state in historical operation data and improving the accuracy of the deep learning model in predicting the fault probability.

[0058] S105: Collect the current operation data of the equipment and input it into the equipment fault prediction model to obtain the current equipment fault probability output by the equipment fault prediction model.

[0059] Deploy the fault prediction model on the device side, collect the current operation data of the equipment and input it into the equipment fault prediction model to obtain the current equipment fault probability of the device output by the model. When deploying the model fault prediction model, the model can be deployed in the cloud, and the currently collected operation data of the equipment is input into the fault prediction model deployed in the cloud through the industrial Internet of Things.

[0060] In addition, a fault probability threshold can be set. When the currently output equipment fault probability by the model is greater than this fault probability threshold, an alarm is issued to the equipment maintenance personnel to enable the equipment maintenance personnel to perform early repairs.

[0061] The device fault prediction method, device, equipment and medium based on deep learning provided by the embodiments of the present disclosure first collect the historical operation data of the device, mark the occurred device faults in the historical operation data, then extract the time-domain features, frequency-domain features and time-frequency domain features related to the device faults from the historical operation data, combine the time-domain features, frequency-domain features and time-frequency domain features related to the device faults with the historical operation data into training data, then construct a deep learning model with the device operation data as the input and the device fault probability as the output, use the training data to train the deep learning model to obtain a device fault prediction model, and finally collect the current operation data of the device and input it into the device fault prediction model to obtain the current device fault probability through the device fault prediction model. Through data analysis and deep learning algorithms, the present disclosure can predict the probability of device faults, so that measures can be taken in advance for maintenance, avoiding device downtime, improving device reliability, reducing unnecessary regular maintenance, avoiding over-maintenance, reducing maintenance costs, realizing intelligent device management, and improving management efficiency.

[0062] Figure 3 FIG. is a schematic structural diagram of a device fault prediction device based on deep learning provided by an embodiment of the present disclosure. The device fault prediction device based on deep learning provided by the embodiments of the present disclosure can execute the processing flow provided by the embodiment of the device fault prediction method based on deep learning, as Figure 3 shown, the device fault prediction device 300 based on deep learning includes: A collection module 301, configured to collect the historical operation data of the device and mark the occurred device faults in the historical operation data, and the historical operation data includes at least one of the vibration frequency, temperature, pressure, current, and voltage during device operation; A data processing module 302, configured to extract the time-domain features, frequency-domain features and time-frequency domain features related to the device faults from the historical operation data, and combine the time-domain features, frequency-domain features and time-frequency domain features related to the device faults with the historical operation data into training data; A model construction module 303, configured to construct a deep learning model with the device operation data as the input and the device fault probability as the output; A model training module 304, configured to use the training data to train the deep learning model to obtain a device fault prediction model; A fault prediction module 305, configured to collect the current operation data of the device and input it into the device fault prediction model to obtain the current device fault probability output by the device fault prediction model.

[0063] In some embodiments, before the data processing module 302 extracts the time-domain features, frequency-domain features and time-frequency domain features related to the device faults from the historical operation data, it is further configured to: perform preprocessing on the historical operation data, and the preprocessing includes data cleaning, denoising and normalization processing; When the data processing module 302 extracts time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from historical operation data and combines the time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures with the historical operation data into training data, it is specifically used for: extracting time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from the preprocessed historical operation data, and combining the time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures with the preprocessed historical operation data into training data.

[0064] In some embodiments, when the data processing module 302 extracts time-domain features, frequency-domain features, and time-frequency domain features related to equipment failures from historical operation data, it is specifically used for: according to the occurrence time point of the equipment failure, extracting the target operation data within a preset time range before and after the occurrence time point; calculating the mean, variance, peak-to-peak value, kurtosis, waveform factor, and impulse factor of the target operation data as the time-domain features related to equipment failures; performing fast Fourier transform on the target operation data to generate a spectrum, and extracting the spectral energy, main frequency component, and frequency band entropy in the spectrum as the frequency-domain features related to equipment failures; performing wavelet packet decomposition on the target operation data to extract the energy entropy of each node in the target operation data, performing short-time Fourier transform on the target operation data to generate a time-frequency diagram, and extracting the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram, and taking the energy entropy of each node in the target operation data and the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram as the frequency-domain features related to equipment failures.

[0065] In some embodiments, the deep learning model includes an input layer, a CNN layer, an Attention layer, an LSTM layer, and a fully connected layer; Among them, the training data is used as the input of the input layer, the output of the input layer is used as the input of the CNN layer, the output of the CNN layer is used as the input of the Attention layer, the output of the Attention layer is used as the input of the LSTM layer, the output of the LSTM layer is used as the input of the fully connected layer, and the fully connected layer outputs the failure probability of the equipment.

[0066] In some embodiments, the CNN layer includes a 1D convolutional layer and a max pooling layer, and the Attention layer is a Multi-HeadAttention layer.

[0067] In some embodiments, when the model training module 304 trains a deep learning model using training data to obtain a device fault prediction model, it is specifically configured to: divide the training data into a training set, a validation set, and a test set; input the training set into the deep learning model for forward propagation to output the fault prediction probability of the device; calculate the loss between the fault prediction probability and the actual device faults in the training set; update the network weights of the deep learning model according to the loss to complete one training round of the deep learning model; after each training round of the deep learning model, use the validation set to evaluate the precision, recall, and F1-Score of the deep learning model to obtain the validation metrics after each training round of the deep learning model; if the validation metrics obtained by the deep learning model within a continuous preset number of training rounds do not improve, terminate the training of the deep learning model to obtain the target deep learning model trained in the preset number of training rounds; use the test set to evaluate each target deep learning model, and determine the target deep learning model as the device fault prediction model according to the evaluation performance of each target deep learning model.

[0068] In some embodiments, when the model training module 304 inputs the training set into the deep learning model for forward propagation to output the fault prediction probability of the device, it is specifically configured to: input the training data into the input layer, and transmit the training data to the CNN layer through the input layer; extract local spatio-temporal features in the training data through the CNN layer and compress them to output a first feature sequence; input the first feature sequence into the Attention layer, dynamically weight the first feature sequence through the Attention layer, and input the second feature sequence with attention weighting; input the second feature sequence into the LSTM layer, model the evolution law of the device state represented by the second feature sequence through the LSTM layer, and output a third feature sequence representing the device state trend; input the third feature sequence into the fully connected layer, map the third feature sequence to the fault prediction probability of the device through the fully connected layer, and output the fault prediction probability.

[0069] Figure 3 The device fault prediction device based on deep learning in the illustrated embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0070] Figure 4 The following is a schematic structural diagram of an electronic device in an embodiment of the present disclosure. Specifically refer to Figure 4 which shows a schematic structural diagram of the electronic device 400 suitable for implementing the present disclosure. Figure 4 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0071] As Figure 4As shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 402 or the program loaded from the storage device 408 into the random access memory (RAM) 403 to implement the deep learning-based device fault prediction method of the embodiments described in the present disclosure. In the RAM 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0072] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0073] In particular, according to the embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts, so as to implement the deep learning-based device fault prediction method as described above. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above functions defined in the methods of the embodiments of the present disclosure are executed.

[0074] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0075] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0076] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device.

[0077] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: Collect the historical operation data of the acquisition device, and mark the occurred device failures in the historical operation data. The historical operation data includes at least one of the vibration frequency, temperature, pressure, current, and voltage during device operation; Extract the time-domain features, frequency-domain features, and time-frequency domain features related to device failures from the historical operation data, and combine the time-domain features, frequency-domain features, and time-frequency domain features related to device failures with the historical operation data into training data; Construct a deep learning model with the device operation data as the input and the device failure probability as the output; Use the training data to train the deep learning model to obtain a device failure prediction model; Collect the current operation data of the device and input it into the device failure prediction model to obtain the current device failure probability output by the device failure prediction model.

[0078] Optionally, when one or more of the above programs are executed by the electronic device, the electronic device can also execute the other steps described in the above embodiments.

[0079] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0081] The units involved in the embodiments described in the present disclosure can be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.

[0082] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and the like.

[0083] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0084] The embodiments of the present disclosure also provide a computer-readable storage medium. A computer program is stored in the storage medium. When the computer program is executed by a processor, the methods of any of the above embodiments can be implemented. Their execution manners and beneficial effects are similar and will not be elaborated herein.

[0085] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features. It should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0086] In addition, although the operations are depicted in a specific order, this should not be construed as requiring that the operations be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0087] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

[0088] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0089] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A device fault prediction method based on deep learning, characterized in that, Including: Collecting historical operation data of the acquisition device and marking the occurred device failures in the historical operation data, where the historical operation data includes at least one of vibration frequency, temperature, pressure, current, and voltage during device operation; Extracting time-domain features, frequency-domain features, and time-frequency domain features related to the device failure from the historical operation data, and combining the time-domain features, frequency-domain features, and time-frequency domain features related to the device failure with the historical operation data as training data; Constructing a deep learning model with device operation data as input and device failure probability as output; Training the deep learning model with the training data to obtain a device failure prediction model; Collecting the current operation data of the device and inputting it into the device failure prediction model to obtain the current device failure probability output by the device failure prediction model.

2. The device fault prediction method based on deep learning according to claim 1, wherein Before extracting the time-domain features, frequency-domain features, and time-frequency domain features related to the device failure from the historical operation data, it further includes: Performing preprocessing on the historical operation data, where the preprocessing includes data cleaning, denoising, and normalization processing; Extracting the time-domain features, frequency-domain features, and time-frequency domain features related to the device failure from the historical operation data, and combining the time-domain features, frequency-domain features, and time-frequency domain features related to the device failure with the historical operation data as training data, including: Extracting the time-domain features, frequency-domain features, and time-frequency domain features related to the device failure from the preprocessed historical operation data, and combining the time-domain features, frequency-domain features, and time-frequency domain features related to the device failure with the preprocessed historical operation data as training data.

3. The device fault prediction method based on deep learning according to claim 1, characterized in that Extracting the time-domain features, frequency-domain features, and time-frequency domain features related to the device failure from the historical operation data, including: According to the occurrence time point of the device failure, extracting the target operation data within a preset time range before and after the occurrence time point; Calculating the mean, variance, peak-to-peak value, kurtosis, waveform factor, and impulse factor of the target operation data as the time-domain features related to the device failure; Performing fast Fourier transform on the target operation data to generate a spectrum, and extracting the spectral energy, main frequency component, and frequency band entropy in the spectrum as the frequency-domain features related to the device failure; Performing wavelet packet decomposition on the target operation data to extract the energy entropy of each node in the target operation data, performing short-time Fourier transform on the target operation data to generate a time-frequency diagram, and extracting the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram, and taking the energy entropy of each node in the target operation data and the energy concentration degree and instantaneous frequency change rate in the time-frequency diagram as the frequency-domain features related to the device failure.

4. The device fault prediction method based on deep learning according to claim 1, wherein The deep learning model includes an input layer, a CNN layer, an Attention layer, an LSTM layer, and a fully connected layer; Among them, the training data serves as the input of the input layer, the output of the input layer serves as the input of the CNN layer, the output of the CNN layer serves as the input of the Attention layer, the output of the Attention layer serves as the input of the LSTM layer, the output of the LSTM layer serves as the input of the fully connected layer, and the fully connected layer outputs the failure probability of the device.

5. The device fault prediction method based on deep learning according to claim 4, wherein The CNN layer includes a 1D convolutional layer and a max pooling layer, and the Attention layer is a Multi-Head Attention layer.

6. The method for predicting device failures based on deep learning according to claim 4, wherein, Training the deep learning model using the training data to obtain a device failure prediction model includes: Dividing the training data into a training set, a validation set, and a test set; Inputting the training set into the deep learning model for forward propagation to output the failure prediction probability of the device; Calculating the loss between the failure prediction probability and the actual device failures in the training set; Updating the network weights of the deep learning model according to the loss to complete one training round of the deep learning model; After each completion of one training round of the deep learning model, using the validation set to evaluate the precision, recall, and F1-Score of the deep learning model to obtain the validation metrics after each training round of the deep learning model; If the validation metrics obtained by the deep learning model within a continuous preset number of training rounds have not improved, terminate the training of the deep learning model to obtain the target deep learning model trained in the preset number of training rounds; Using the test set to evaluate each target deep learning model and determining the target deep learning model as the device failure prediction model according to the evaluation performance of each target deep learning model.

7. The method for predicting equipment faults based on deep learning according to claim 6, wherein The inputting the training set into the deep learning model for forward propagation to output the failure prediction probability of the device includes: Inputting the training data into the input layer and transmitting the training data to the CNN layer through the input layer; Extracting and compressing the local spatio-temporal features in the training data through the CNN layer to output a first feature sequence; Inputting the first feature sequence into the Attention layer and dynamically weighting the first feature sequence through the Attention layer to input a second feature sequence with attention weighting; Inputting the second feature sequence into the LSTM layer and modeling the evolution law of the device state represented by the second feature sequence through the LSTM layer to output a third feature sequence representing the device state trend; Inputting the third feature sequence into the fully connected layer and mapping the third feature sequence to the failure prediction probability of the device through the fully connected layer to output the failure prediction probability.

8. A device fault prediction device based on deep learning, characterized in that Including: An acquisition module for acquiring the historical operation data of the device and marking the device failures that have occurred in the historical operation data. The historical operation data includes at least one of the vibration frequency, temperature, pressure, current, and voltage during device operation. A data processing module, configured to extract time-domain features, frequency-domain features, and time-frequency domain features related to the equipment failure from the historical operation data, and combine the time-domain features, frequency-domain features, and time-frequency domain features related to the equipment failure with the historical operation data to form training data; A model construction module, configured to construct a deep learning model with equipment operation data as input and equipment failure probability as output; A model training module, configured to train the deep learning model using the training data to obtain an equipment failure prediction model; A failure prediction module, configured to collect the current operation data of the equipment and input it into the equipment failure prediction model to obtain the current equipment failure probability output by the equipment failure prediction model.

9. An electronic device, characterized in that, It includes a memory and a processor. Among them, a computer program is stored in the memory, and when the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that, Program instructions are stored thereon, and when the program instructions are executed, the method according to any one of claims 1-7 is implemented.