Wavelet enhancement and lstm-based equipment rul prediction method and system
Patent Information
- Application Number
- CN202610995639.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-07-06
AI Technical Summary
然而,工业现场采集的传感器数据通常包含大量噪声,在处理高度噪声、非平稳且具有多尺度特征的时间序列数据时,其鲁棒性仍有提升空间
[0066]1.本发明技术方案提出了一种全新的小波增强LSTM与注意力机制融合模型,通过对多尺度退化特征的强化提取及长时序依赖的自适应建模,实现设备剩余使用寿命的高精度估计。即先通过小波增强操作生成时序特征序列得到表征设备健康状态的时序特征序列,再通过LSTM神经网络的特征提取生成隐藏状态序列,用以表示设备的潜在退化趋势,再经分段注意力增强机制得到注意力增强的隐藏状态,进而最终基于注意力增强的隐藏状态得到预测结果。
Smart Images

Figure CN122508550B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equipment remaining useful life prediction, specifically to a method and system for predicting equipment RUL based on wavelet enhancement and LSTM. Background Technology
[0002] Remaining Useful Life (RUL) is of profound significance for ensuring the continuity of industrial production, reducing maintenance costs, and improving equipment reliability. Existing RUL prediction methods are mainly divided into three categories: model-driven, data-driven, and hybrid-driven. Model-driven methods predict RUL by establishing physical degradation failure mechanism models of the equipment. These methods rely on a deep understanding of the equipment's operating principles and failure modes. However, for equipment with complex structures and unclear degradation mechanisms, it is difficult to establish accurate physical models, and model parameters are also difficult to obtain. Data-driven methods utilize machine learning to analyze historical operating data of the equipment to learn degradation patterns and predict RUL.
[0003] With the rapid development of deep learning technology, its powerful capabilities in processing complex data and automatically extracting features have led to its widespread application. Recurrent Neural Networks (RNNs) and their variants, such as Long Short-Term Memory (LSTM) networks, can effectively model temporal dependencies in time series data, demonstrating great potential in RUL prediction. LSTMs effectively solve the problems of gradient vanishing or exploding during traditional RNN training by introducing gating units to control the updating and propagation of time series states, enabling them to capture degradation trends and patterns during long-term equipment operation. However, sensor data collected in industrial fields often contains a large amount of noise, and its robustness in processing highly noisy, non-stationary time series data with multi-scale characteristics still has room for improvement. LSTMs are limited by their insufficient ability to focus on data quality and key information; therefore, noise and redundant information in time series data may obscure key degradation features, affecting the learning efficiency and prediction accuracy of LSTMs for core degradation patterns.
[0004] Therefore, how to effectively utilize LSTM networks to reduce or eliminate the impact of noise and redundant information in time series data, and achieve more accurate and reliable device RUL prediction, remains a hot topic and an urgent problem to be solved in this field. Summary of the Invention
[0005] This invention aims to overcome the limitations of existing LSTM methods and achieve more accurate and reliable device RUL prediction. It provides a device RUL prediction method and system based on wavelet enhancement and LSTM. The method innovatively proposes a wavelet-enhanced LSTM segmented attention enhancement mechanism fusion model, which aims to improve the accuracy of device RUL prediction.
[0006] Therefore, the present invention provides the following technical solution:
[0007] On the one hand, the present invention provides a device RUL prediction method based on wavelet enhancement and LSTM, comprising the following steps:
[0008] Step S1: Obtain monitoring data of the device to be predicted throughout its entire lifecycle or during a specified operating phase;
[0009] Step S2: Perform wavelet enhancement on the monitoring data. The wavelet enhancement operation involves sequentially performing multi-scale wavelet decomposition, thresholding, and signal reconstruction to obtain a time-series feature sequence that characterizes the health status of the equipment.
[0010] Step S3: Input the time-series feature sequence into the LSTM neural network for feature extraction to generate a hidden state sequence, representing the potential degradation trend of the device;
[0011] Step S4: The hidden state sequence is used as the input of the segmented attention enhancement mechanism to obtain the attention-enhanced hidden state, and then nonlinear mapping is performed through the regression output layer to obtain the device RUL prediction result;
[0012] The segmented attention enhancement mechanism involves sequentially performing sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping.
[0013] This invention introduces segmented attention enhancement (local attention aggregation) and global context modeling (global attention aggregation) on the hidden state sequence output by LSTM to achieve hierarchical feature aggregation of degraded sequences. On the one hand, segmented attention weights and selects key degradation information within each sub-sequence, alleviating the problem of insufficient feature representation caused by information decay and difficulty in capturing long-distance dependencies in long-sequence scenarios. On the other hand, global attention models cross-segment associations between sub-sequences, which can more effectively capture degradation trends over long periods. Furthermore, by concatenating the global context vector with the key degradation features output by LSTM and inputting it into the regression output layer, RUL prediction is finally achieved, simultaneously fusing two complementary types of information: "current degradation state" and "long-term degradation trend".
[0014] Preferably, the threshold processing in the wavelet enhancement operation adopts a segmented enhancement threshold processing rule that is aware of degradation features, and the corresponding mathematical model is:
[0015] ;
[0016] In the formula: The lower threshold used for noise reduction. The upper threshold used to enhance the interval is selected based on the baseline threshold. ; These are the wavelet coefficients after thresholding. These are the wavelet coefficients obtained from multi-scale wavelet decomposition; Here, J represents the wavelet decomposition scale index, and J represents the total number of scale types. ; k is the wavelet coefficient location index at this scale. j Let j be the number of wavelet coefficients at scale j. ; This is the enhancement coefficient.
[0017] Among them, when Set the wavelet coefficients to zero to suppress pure noise components; when ,according to The wavelet coefficients are amplified to highlight the weak degradation features near the noise level; when The original coefficients are kept unchanged to avoid distortion of strong degradation features and normal signal components.
[0018] Preferably, the lower threshold and upper threshold Selection and benchmark threshold Relevant, specifically:
[0019] First, regarding scale wavelet coefficient set Candidate thresholds are constructed based on Stein unbiased risk estimation. Risk function:
[0020] ;
[0021] In the formula: For scale Corresponding candidate threshold The risk function; This is an estimate of the noise standard deviation; The candidate threshold; As an indicator function, when the absolute value of the wavelet coefficients... Less than or equal to candidate threshold hour, The item takes a value of 1, otherwise The value of the item is 0; For scale The first, second, and kth wavelet coefficients in the set of wavelet coefficients j Wavelet coefficients;
[0022] Then, at scale The upper search finds the baseline threshold that minimizes the risk function: And based on the signal at scale Energy distribution on and kurtosis indicating impact Constructing degradation sensitivity factors :
[0023] ;
[0024] In the formula: , These are adjustable weighting coefficients used to balance the contributions of energy features and kurtosis features to the degradation sensitivity factor. ;
[0025] Finally, based on the benchmark threshold and degradation sensitivity factor Adaptive construction of dual thresholds:
[0026] ;
[0027] ;
[0028] In the formula: , This is the threshold adjustment coefficient.
[0029] Preferably, the process in step S4 of using the hidden state sequence as input to the segmented attention enhancement mechanism to obtain the attention-enhanced hidden state is as follows:
[0030] Step S41: Sequence segmentation, dividing the hidden state sequence into segments of equal length L. A sequence of jokes;
[0031] Step S42: Local attention aggregation. For each subsequence generated in step S41, perform local attention aggregation to generate segment-level hidden state sequences. , These are the segment-level hidden vectors for the 1st, 2nd, and Mth subsequences, respectively;
[0032] Step S43: Global attention aggregation, processing the segment-level hidden state sequence generated in step S42. Perform global attention aggregation to obtain a global context vector representing device degradation features;
[0033] Step S44: Feature concatenation and mapping, converting the global context vector from step S43... Key degradation features of the LSTM output in step S3 Feature concatenation is performed to form a comprehensive feature vector. The comprehensive feature vector As a hidden state that enhances attention.
[0034] Preferably, the last hidden state of each subsequence is used as the local query vector, and then the 1st... The subsequence is denoted as , These are all position markers for each hidden state in the subsequence, and the hidden state is... As a local query vector The implementation process of local attention aggregation in step S42 is as follows:
[0035] First, perform local attention aggregation to compute each hidden state within the subsequence. Local attention score ;
[0036] Then, the local attention score for the subsequence. Softmax normalization is performed to obtain local attention weights, and the hidden states within the subsequence are weighted and summed to obtain the segment-level hidden vector of the subsequence.
[0037] ;
[0038] In the formula, For the first The segment-level hidden vectors of each subsequence are used to generate the segment-level hidden state sequence. Preferably, the implementation process of global attention aggregation in step S43 is as follows:
[0039] First, use the sequence The last segment-level hidden vector As a global query vector Perform global attention aggregation on each segment-level hidden vector. Calculate the global attention score ;
[0040] Then process the segment-level hidden vectors The global attention weights are obtained by softmax normalization of the global attention scores. ;
[0041] Finally, the segment-level hidden vectors of all subsequences are weighted and summed to obtain the global context vector z:
[0042] .
[0043] Preferably, the wavelet enhancement operation is implemented as follows:
[0044] First, multi-scale wavelet decomposition is performed on the discrete sequence data to obtain different scales. and location wavelet coefficients ;
[0045] Then, for the wavelet coefficients Perform threshold processing;
[0046] Finally, the wavelet coefficients after thresholding are used. The reconstructed signal yields a time-series feature sequence:
[0047]
[0048]
[0049] In the formula: For the first Temporal characteristics of discrete-time sampling points; It is a wavelet function; , , They represent the 1st, 2nd, and 3rd respectively. Temporal characteristics of discrete-time sampling points.
[0050] Preferably, the feature extraction process of the LSTM neural network in step S3 is as follows:
[0051] Step S31: After the time-series feature sequence is input into the LSTM neural network, the corresponding output value is obtained through the forget gate, input gate, and output gate of the gating unit. , and ;
[0052] Step S32: Obtain the corresponding output value using the forget gate, input gate, and output gate. , and And update the hidden state of the LSTM neural network.
[0053] Preferably, the monitoring data of the equipment is a vibration signal, a temperature signal, or a current signal.
[0054] Secondly, the present invention provides a prediction system based on the above method, comprising:
[0055] The data acquisition module is used to acquire monitoring data of the device to be predicted throughout its entire life cycle or during a specified operating phase.
[0056] The wavelet enhancement module is used to perform wavelet enhancement operations on monitoring data. The wavelet enhancement operation involves sequentially performing multi-scale wavelet decomposition, threshold processing, and signal reconstruction to obtain a time-series feature sequence that characterizes the health status of the device.
[0057] The LSTM module is used to input the temporal feature sequence into the LSTM neural network for feature extraction and generate a hidden state sequence to represent the potential degradation trend of the device.
[0058] The segmented attention enhancement module is used to take the hidden state sequence as input to the segmented attention enhancement mechanism to obtain the attention-enhanced hidden state, and then perform nonlinear mapping through the regression output layer to obtain the device RUL prediction result.
[0059] The segmented attention enhancement mechanism involves sequentially performing sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping.
[0060] In three aspects, the present invention provides a computer device comprising: one or more processors and a memory storing one or more computer programs;
[0061] The processor invokes a computer program to achieve the following:
[0062] The steps of the above-described device RUL prediction method based on wavelet enhancement and LSTM.
[0063] In three aspects, the present invention provides a computer-readable storage medium storing a computer program, which is invoked by a processor to implement:
[0064] The steps of the above-described device RUL prediction method based on wavelet enhancement and LSTM.
[0065] Compared with the prior art, the present invention achieves the following effects:
[0066] 1. This invention proposes a novel fusion model of wavelet-enhanced LSTM and attention mechanism. Through enhanced extraction of multi-scale degradation features and adaptive modeling of long-term temporal dependencies, it achieves high-precision estimation of the remaining lifespan of equipment. Specifically, it first generates a temporal feature sequence representing the health status of the equipment through wavelet enhancement, then generates a hidden state sequence through feature extraction using an LSTM neural network to represent the potential degradation trend of the equipment, and finally obtains attention-enhanced hidden states through a segmented attention enhancement mechanism. Finally, the prediction result is obtained based on the attention-enhanced hidden states.
[0067] 2. This invention further optimizes the threshold processing in wavelet enhancement operations. By constructing a lower and upper threshold and introducing an enhancement coefficient, it achieves noise suppression while specifically amplifying weakly degenerate features, avoiding the problem of neglecting degenerate features caused by traditional single-threshold denoising. Furthermore, it proposes a scale-adaptive threshold selection based on the Stein unbiased risk estimation principle, enabling the threshold to dynamically adjust with different scale features. This improves feature extraction capabilities in complex scenarios such as early degradation and low signal-to-noise ratio, providing a more stable and discriminative input representation for subsequent RUL prediction, thereby reducing prediction errors and improving robustness.
[0068] 3. This invention designs a segmented attention enhancement mechanism based on LSTM temporal modeling, performing hierarchical feature aggregation on the hidden state sequence. This enables adaptive weighted selection of key degradation information, alleviating the limitations of LSTM's feature representation capabilities caused by information decay and difficulty in capturing long-distance dependencies in long-sequence scenarios. Furthermore, cross-segment association modeling is performed, which can more effectively capture degradation trends over long time spans. By further concatenating the global context vector with the key degradation features output by LSTM and inputting this concatenation into the regression output layer, the model simultaneously considers both the current degradation state and the long-term degradation trend, improving adaptability to operating condition fluctuations and noise disturbances, and achieving high-precision estimation of the remaining service life of equipment and better generalization performance. Attached Figure Description
[0069] Figure 1 This is a schematic diagram of the construction method of the method in an embodiment of the present invention;
[0070] Figure 2 This is a schematic diagram of the wavelet enhancement processing flow in the method of this embodiment of the invention;
[0071] Figure 3 This is a schematic diagram of the segmented enhanced threshold processing rules used in the method of this embodiment of the invention;
[0072] Figure 4 This is a schematic diagram of the unit structure of the LSTM neural network in the method of this embodiment of the invention; Detailed Implementation
[0073] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention. The technical features involved in the various embodiments of the invention described below can be combined with each other as long as they do not conflict with each other.
[0074] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0075] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0076] The adjustable parameters such as weights and biases involved in this invention can be divided into learning parameters and preset parameters: learning parameters can be adaptively determined based on historical monitoring data with degradation labels through offline training or intelligent optimization processes; preset parameters can be preset to fixed values according to device type and engineering experience, or determined through parameter selection methods such as grid search and Bayesian optimization. The above adjustable parameters can be adjusted according to the application scenario without constituting a limitation of this invention.
[0077] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.
[0078] This invention provides a method and system for predicting the remaining useful life (RUL) of equipment based on wavelet enhancement and LSTM. It aims to offer a robust RUL prediction technique for degradation monitoring signals, with enhanced noise resistance and temporal modeling capabilities. This significantly improves upon the limitations of traditional methods relying solely on a single LSTM network, which suffer from insufficient degradation feature representation, limited long-sequence dependency capture, and inadequate prediction accuracy and learning efficiency. To this end, this invention creatively proposes a fusion model of wavelet-enhanced LSTM and an attention mechanism. Through enhanced extraction of multi-scale degradation features and adaptive modeling of long-sequence dependencies, it achieves high-precision estimation of the remaining useful life of equipment. The following section uses a permanent magnet synchronous motor as an example to further illustrate the technical solution of this invention. However, it should be understood that the technical solution of this invention is not limited to this device. Without departing from the spirit of this invention, it can also be applied to life prediction scenarios for key power equipment such as transformers, circuit breakers, instrument transformers, rotating machinery, or other industrial equipment with degradation processes.
[0079] Example 1
[0080] like Figure 1 As shown, this embodiment of the invention takes a permanent magnet synchronous motor as an example and provides the following technical approach for a device RUL prediction method based on wavelet enhancement and LSTM:
[0081] Step S1: Acquire monitoring data of the permanent magnet synchronous motor throughout its entire lifecycle or a specified operating phase. In this embodiment, the monitoring data includes, but is not limited to, vibration signals, temperature signals, or current signals collected by sensors. Assume the sensor sampling period is... The sampled length is Discrete sequence ,in This is the index for discrete-time sampling points.
[0082] Step S2: For discrete sequences Perform wavelet enhancement, which involves sequentially performing multi-scale wavelet decomposition, thresholding, and signal reconstruction to obtain a time-series feature sequence. , , , They represent the 1st, 2nd, and 3rd respectively. The temporal characteristics of each discrete time sampling point are used to characterize the health status of the device at the corresponding time sampling point.
[0083] Among them, such as Figure 2 As shown, the wavelet enhancement operation process is as follows:
[0084] Step S21: For discrete sequences Multi-scale wavelet decomposition is performed to obtain wavelet coefficients at different scales j and positions k. The corresponding data model in this embodiment is as follows:
[0085] ;
[0086] ;
[0087] In the formula: Here, J represents the wavelet decomposition scale index, and J represents the total number of scale types. ; k is the wavelet coefficient location index at this scale. j Let j be the number of wavelet coefficients at scale j. ; These are wavelet coefficients; Index of time sampling points; This represents summing over all sample points; It is a wavelet function; for The conjugate form of .
[0088] Considering the sensitivity of wavelet functions to non-stationary signals and the greater energy concentration during signal reconstruction after denoising, this embodiment preferably uses the Daubechies wavelet as the wavelet function; other feasible embodiments may also select other types of wavelet functions; in practical applications, the sampling period can be adjusted according to the prediction accuracy requirements. and sampling length .
[0089] Step S22: Adjust wavelet coefficients Perform threshold processing.
[0090] like Figure 3 As shown, this embodiment preferably employs a segmented enhancement thresholding rule that is aware of degradation features to suppress noise and enhance weak degradation features in wavelet coefficients at each scale. The specific formula is as follows:
[0091] ;
[0092] In the formula: The lower threshold used for noise reduction. The upper threshold used to enhance the interval is selected in relation to the baseline threshold. related; These are the wavelet coefficients after thresholding. These are the wavelet coefficients obtained from multi-scale wavelet decomposition. This is the enhancement coefficient.
[0093] Based on dual thresholds and enhancement coefficients, piecewise enhanced thresholding with degradation feature awareness is performed on wavelet coefficients at each scale to obtain the processed wavelet coefficients. . Specifically, when Set the wavelet coefficients to zero to suppress pure noise components; when ,according to The wavelet coefficients are amplified to highlight the weak degradation features near the noise level; when The original coefficients are kept unchanged to avoid distortion of strong degradation features and normal signal components. Through the above segmented enhanced thresholding process, noise can be suppressed while enhancing multi-scale details related to equipment degradation, thereby providing a feature sequence with a higher signal-to-noise ratio and more prominent degradation features for subsequent LSTM-based remaining lifetime prediction.
[0094] Step S23: Use the thresholded wavelet coefficients The reconstructed signal yields a time-series feature sequence:
[0095] ;
[0096] ;
[0097] In the formula: For the first Temporal characteristics of discrete-time sampling points; It is a wavelet function; , , They represent the 1st, 2nd, and 3rd respectively. Temporal characteristics of discrete-time sampling points.
[0098] It should be noted that in some embodiments, the threshold processing rule in step S22 is preferably optimized. After obtaining the wavelet coefficients through multi-scale wavelet decomposition, risk functions for the corresponding scales are constructed respectively, and the optimal benchmark threshold is independently calculated at each scale to achieve scale-adaptive denoising. Traditional unbiased risk estimation methods only consider noise statistics, which are insufficient for preserving and enhancing the weak early degradation features of equipment. To enhance the early degradation features, this invention specifically amplifies the weak degradation features in the middle interval of the threshold function. The specific implementation process is as follows:
[0099] First, regarding scale wavelet coefficient set Candidate thresholds are constructed based on Stein unbiased risk estimation. Risk function:
[0100] ;
[0101] In the formula: For scale Number of wavelet coefficients; This is an estimate of the noise standard deviation; The candidate threshold; As an indicator function, when the absolute value of the wavelet coefficients... Less than or equal to candidate threshold hour, The item takes a value of 1, otherwise The value of the item is 0; For scale The first, second, and kth wavelet coefficients in the set of wavelet coefficients j Wavelet coefficients.
[0102] Then, at scale The search above finds the baseline threshold that minimizes the risk function. :
[0103]
[0104] Noise approximately follows a zero-mean Gaussian distribution in the wavelet transform domain. Therefore, in some embodiments, the noise standard deviation is estimated using the median absolute deviation method based on the wavelet coefficients of the highest frequency scale and is shared across the risk functions at each scale. For example, the noise standard deviation estimate... :
[0105] ;
[0106] In the formula: The noise standard deviation is estimated by the median absolute deviation of the detail coefficients of the first-level decomposition (scale 1), which represents the highest frequency band among all scales. is the median operator; 0.6745 is the calibration coefficient between the absolute deviation of the median and the standard deviation when the noise approximately follows a zero-mean Gaussian distribution; when the noise distribution deviates from the Gaussian distribution, this calibration coefficient can also be calibrated using historical data or replaced with an empirical coefficient without affecting the implementation of the present invention.
[0107] Next, we define and normalize the statistics strongly correlated with degradation. Energy at each scale:
[0108] ;
[0109] ;
[0110] ;
[0111] In the formula: The energy at scale j; Let be the number of wavelet coefficients in scale j.
[0112] Cliff (impact level):
[0113] ;
[0114] ;
[0115] In the formula: A preset smoothing constant is used to avoid the denominator being too small and to normalize the kurtosis to a stable range; and respectively scale The energy proportion (energy distribution) and the mean of wavelet coefficients; For scale The normalized kurtosis.
[0116] Degradation sensitivity factors are constructed based on the energy distribution and impulsivity of the signal at various scales, enabling targeted weighting of wavelet scales highly correlated with equipment degradation characteristics. Degradation sensitivity factors are defined. :
[0117] ;
[0118] In the formula: , These are adjustable weighting coefficients used to balance the contributions of energy features and kurtosis features to the degradation sensitivity factor. In some embodiments, , The parameters to be learned can be jointly determined through an optimization process using historical data labeled with remaining lifetime, in order to minimize the RUL prediction error; in addition, non-negativity and normalization constraints are imposed on the weighting coefficients, so that... and .
[0119] Finally, based on the baseline threshold and degradation sensitivity factor An adaptive dual threshold and enhancement coefficients are constructed to obtain a parameter set for noise suppression and weak degradation feature enhancement. Specifically, the baseline threshold is... Corrected to the lower threshold according to the preset functional relationship. and upper threshold :
[0120] ;
[0121] ;
[0122] In the formula: , This is the threshold adjustment coefficient, used to control the threshold shrinkage and enhancement magnitude for degradation-sensitive scales.
[0123] In some embodiments, enhancement coefficients at each scale are also calculated based on the degradation sensitivity factor:
[0124] ;
[0125] In the formula: To enhance the weighting coefficients, which are used to control the amplification magnitude at degradation-sensitive scales.
[0126] It should be understood that the above , , As preset parameters, they can be pre-set according to the equipment type, noise level and operating condition fluctuation range, and the parameter values can be determined by combining grid search to balance the noise suppression effect and the weak degradation feature enhancement effect. The specific value selection process is not limited in this invention.
[0127] Step S3: Transform the time-series feature sequence The input is used to extract features from an LSTM neural network, generating a sequence of hidden states.
[0128] like Figure 4 As shown, the feature extraction process of the LSTM neural network in this embodiment is as follows:
[0129] Step S31: After the time-series feature sequence is input into the LSTM neural network, the corresponding output value is obtained through the forget gate, input gate, and output gate of the gating unit. , and .
[0130] ;
[0131] ;
[0132] ;
[0133] In the formula: , and These represent the output values of the input gate, forget gate, and output gate, respectively. For the first Temporal characteristics of discrete-time sampling points; For the first The hidden state corresponds to each discrete-time sampling point; W and b are the weight matrix and bias to be learned. Inputs The weights of the input gate, forget gate, and output gate. The hidden states corresponding to the previous discrete-time sampling point are respectively The weights of the input gate, forget gate, and output gate. These are the biases for the input gate, forget gate, and output gate, respectively. This is the Sigmoid function.
[0134] Step S32: Obtain the corresponding output value using the forget gate, input gate, and output gate. , and Update the hidden state of the LSTM neural network.
[0135] ;
[0136] ;
[0137] In the formula: This is element-wise multiplication; For the updated memory; For the first Candidate memory elements corresponding to discrete-time sampling points; Inputs The hidden state corresponding to the previous discrete-time sampling point The weight matrix of candidate memory elements; For the first The hidden state corresponding to each discrete-time sampling point; These are the bias parameters. The final hidden state sequence is generated. and output key degradation features .
[0138] The aforementioned LSTM neural network selectively retains and updates information through forget gates, input gates, and output gates, thereby learning the complex nonlinear dynamic patterns and time dependencies in the device degradation process.
[0139] In some embodiments, to balance computational efficiency in long sequence scenarios with the ability to model multi-scale degradation features, a segmented attention enhancement module is designed after LSTM.
[0140] Step S4: Construct a segmented attention enhancement module for the hidden state sequence. Deep feature mining is performed. The segmented attention enhancement module includes sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping. Specifically: First, a sequence segmentation strategy is executed; then, local attention aggregation is performed, calculating local attention weights for the hidden states within each sub-sequence and performing weighted fusion to obtain the segment-level hidden vectors of the sub-sequences; next, global attention aggregation is performed, using all segment-level hidden vectors as input, and leveraging the global attention mechanism to capture long-distance dependencies between segments to generate a global context vector; finally, through feature concatenation and mapping, the global context vector is combined with key degenerate features (generally selected from the final generated hidden state sequence). The hidden state of the last time step in the model is concatenated and then input into the regression output layer, thereby significantly improving the model's ability to capture long-sequence degradation trends while achieving high-precision prediction of device RUL.
[0141] Specifically, the implementation process of the segmented attention enhancement module is as follows:
[0142] Step S41: Input the hidden state sequence Divided into Segment sequence, each subsequence having a length of Preferably, the subsequence length can be selected based on the device sampling period or experience. and satisfy If the last paragraph is insufficient, zero padding or truncation can be used. For example, the first... The subsequence is denoted as The last hidden state of the subsequence As a local query vector .
[0143] Step S42: For each subsequence generated in step S41, perform local attention aggregation. First, calculate the hidden state within each subsequence. Local attention score:
[0144] ;
[0145] In the formula: This represents the segment number of the corresponding subsequence; , These are the local attention weight parameters for the local query vector and the hidden state, respectively. This is a bias term for local attention; Let be the local attention score vector to be learned; This is the matrix transpose symbol; For the first Each discrete-time sampling point corresponds to a hidden state. The local attention score.
[0146] For the The local attention scores of each subsequence are softmax normalized to obtain the local attention weights. Based on this, the hidden states within the subsequence are weighted and summed to obtain the th... Segment-level hidden vectors of subsequences .
[0147] ;
[0148] ;
[0149] Step S43: Process the segment-level hidden state sequence generated in step S42. Global attention aggregation is performed to obtain a global context vector representing device degradation features. Preferably, the last segment-level hidden vector in the sequence is used. As a global query vector For each segment-level hidden vector Calculate the global attention score:
[0150] ;
[0151] In the formula: , These are the global attention weight parameters for the global query vector and the segment-level hidden vector, respectively. This is a bias term for global attention; Let be the global attention score vector to be learned; This is the matrix transpose symbol; Indicates the corresponding to the first Segment-level hidden vectors of subsequences The global attention score.
[0152] Segment-level hidden vectors The global attention scores are then softmax normalized to obtain the global attention weights.
[0153] ;
[0154] Based on this, the segment-level hidden vectors of all subsequences are weighted and summed to obtain the global context vector:
[0155] ;
[0156] Step S44: Transfer the global context vector from step S43. Key degradation features of the LSTM output in step S3 Feature concatenation is performed to form a comprehensive feature vector. The comprehensive feature vector is then input into the regression output layer for nonlinear mapping to obtain the predicted remaining useful life (RUL) of the current device. Preferably, this embodiment uses a fully connected network for nonlinear mapping, and the final output is the predicted RUL of the device:
[0157] ;
[0158] In the formula: Key degradation features of LSTM output; For global context vectors; This represents the vector concatenation operation; , These are the weight parameters for the fully connected layer and the output layer, respectively. , These are the biases for the fully connected layer and the output layer, respectively. The activation function is an element-wise nonlinear activation function, preferably a ReLU function or a hyperbolic tangent function; This represents the predicted remaining useful life.
[0159] In summary, this invention introduces segmented attention enhancement and global context modeling on the hidden state sequence output by LSTM to achieve hierarchical feature aggregation of degraded sequences. On the one hand, segmented attention weights and selects key degradation information within each sub-sequence, alleviating the problem of insufficient feature representation caused by information decay and difficulty in capturing long-distance dependencies in long-sequence scenarios. On the other hand, global attention models cross-segment associations between sub-sequences, which can more effectively capture degradation trends over long periods. Furthermore, by concatenating the global context vector with the key degradation features output by LSTM and inputting it into the regression output layer to finally achieve RUL prediction, this invention simultaneously integrates two complementary types of information: "current degradation state" and "long-term degradation trend," thereby improving robustness to noise disturbances and operating condition fluctuations, reducing RUL estimation bias, and achieving high-precision prediction of the remaining service life of equipment.
[0160] Example 2
[0161] The present invention also provides a prediction system based on the above method, including a data acquisition module, a wavelet enhancement module, an LSTM module, and a segmented attention enhancement module that are sequentially linked or interconnected.
[0162] The system comprises the following modules: a data acquisition module for acquiring monitoring data of the device under prediction throughout its entire lifecycle or at a specified operational stage; a wavelet enhancement module for performing wavelet enhancement operations on the monitoring data, which involves sequentially performing multi-scale wavelet decomposition, thresholding, and signal reconstruction to obtain a temporal feature sequence; an LSTM module for inputting the temporal feature sequence into an LSTM neural network for feature extraction, learning the temporal dependencies of the degradation process, and generating a hidden state sequence; and a segmented attention enhancement module for using the hidden state sequence as input to the segmented attention enhancement mechanism to obtain attention-enhanced hidden states, which are then nonlinearly mapped through a regression output layer to obtain the device's RUL prediction result. The segmented attention enhancement mechanism involves sequentially performing sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping.
[0163] It should also be understood that the specific implementation process of each module is described in the above method. This invention will not repeat it here. The above division of functional modules is only for illustrative purposes. In some embodiments, some functional modules can be combined and some functional modules can be separated. Each functional module can be implemented in software, hardware, or a combination of software and hardware. The software and hardware devices include, but are not limited to, general-purpose computer equipment, programmable gate arrays, digital signal processors, microprocessors and their corresponding programming or burning software.
[0164] Example 3
[0165] An embodiment of the present invention also provides a computer device, comprising: one or more processors and a memory storing one or more computer programs; the processor invokes the computer programs to implement:
[0166] The steps of the above-described device RUL prediction method based on wavelet enhancement and LSTM.
[0167] Specific implementation:
[0168] Step S1: Obtain monitoring data of the device to be predicted throughout its entire lifecycle or during a specified operating phase;
[0169] Step S2: Perform wavelet enhancement on the monitoring data. The wavelet enhancement operation involves sequentially performing multi-scale wavelet decomposition, thresholding, and signal reconstruction to obtain a time-series feature sequence.
[0170] Step S3: Input the temporal feature sequence into the LSTM neural network for feature extraction to generate the hidden state sequence;
[0171] Step S4: The hidden state sequence is used as the input of the attention mechanism to obtain the attention-enhanced hidden state, and then the device RUL prediction result is obtained by nonlinear mapping through the regression output layer.
[0172] The segmented attention enhancement mechanism involves sequentially performing sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping.
[0173] For details on the implementation of each step, please refer to the aforementioned embodiment of the device RUL prediction method based on wavelet enhancement and LSTM.
[0174] In some embodiments, the electronic components of a computer device include:
[0175] The processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present invention.
[0176] The memory can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory and the processor calls and executes the algorithm program of the device RUL prediction method based on wavelet enhancement and LSTM in the embodiments of this invention.
[0177] Input / output interfaces are used to implement information input and output.
[0178] The communication interface is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0179] A bus is used to transfer information between various components of a device, such as processors, memory, input / output interfaces, and communication interfaces.
[0180] The processor, memory, input / output interfaces, and communication interfaces communicate with each other within the device via a bus.
[0181] Example 4
[0182] The present invention also provides a computer-readable storage medium storing a computer program that is called by a processor to implement the steps of the above-described device RUL prediction method based on wavelet enhancement and LSTM.
[0183] Specific implementation:
[0184] Step S1: Obtain monitoring data of the device to be predicted throughout its entire lifecycle or during a specified operating phase;
[0185] Step S2: Perform wavelet enhancement on the monitoring data. The wavelet enhancement operation involves sequentially performing multi-scale wavelet decomposition, thresholding, and signal reconstruction to obtain a time-series feature sequence.
[0186] Step S3: Input the temporal feature sequence into the LSTM neural network for feature extraction to generate the hidden state sequence;
[0187] Step S4: The hidden state sequence is used as the input of the attention mechanism to obtain the attention-enhanced hidden state, and then the device RUL prediction result is obtained by nonlinear mapping through the regression output layer.
[0188] The segmented attention enhancement mechanism involves sequentially performing sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping.
[0189] For details on the implementation of each step, please refer to the description of the aforementioned prediction method embodiment.
[0190] The readable storage medium is a computer-readable storage medium, which can be an internal storage unit of the hardware or software device in any of the foregoing embodiments, such as the hard drive or memory of the controller. The readable storage medium can also be an external storage device of the controller, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the controller. Furthermore, the readable storage medium can include both internal storage units and external storage devices of the controller. The readable storage medium is used to store computer programs and other programs and data required by the controller. The readable storage medium can also be used to temporarily store data that has been output or will be output.
[0191] Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0192] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application refers to flowchart illustrations and / or instructions executed by a processor of a method, apparatus (system), and computer program product according to embodiments of this application to create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams. These computer program instructions may also be stored in a computer-readable storage medium capable of directing a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowchart illustrations and / or one or more block diagrams. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.
[0193] It should be emphasized that the examples described in this invention are illustrative rather than limiting. Therefore, this invention is not limited to the examples described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solutions of this invention, without departing from the spirit and scope of this invention, whether modifications or substitutions, are also within the protection scope of this invention.
Claims
1. A device RUL prediction method based on wavelet enhancement and LSTM, characterized in that: Includes the following steps: Step S1: Obtain monitoring data of the device to be predicted throughout its entire lifecycle or during a specified operating phase; Step S2: Perform wavelet enhancement on the monitoring data. The wavelet enhancement operation involves sequentially performing multi-scale wavelet decomposition, thresholding, and signal reconstruction to obtain a time-series feature sequence that characterizes the health status of the equipment. Step S3: Input the time-series feature sequence into the LSTM neural network for feature extraction to generate a hidden state sequence, representing the potential degradation trend of the device; Step S4: The hidden state sequence is used as the input of the segmented attention enhancement mechanism to obtain the attention-enhanced hidden state, and then nonlinear mapping is performed through the regression output layer to obtain the device RUL prediction result; The segmented attention enhancement mechanism involves sequentially performing sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping. In the wavelet enhancement operation, the threshold processing adopts a segmented enhancement threshold processing rule that is aware of degradation features, and the corresponding mathematical model is as follows: ; In the formula: The lower threshold used for noise reduction. This is the upper threshold used to enhance the interval; These are the wavelet coefficients after thresholding. These are the wavelet coefficients obtained from multi-scale wavelet decomposition; Here, J represents the wavelet decomposition scale index, and J represents the total number of scale types. ; k is the wavelet coefficient location index at this scale. j Let j be the number of wavelet coefficients at scale j. ; Enhancement coefficient; lower threshold and upper threshold Selection and benchmark threshold Relevant, specifically: First, regarding scale wavelet coefficient set Candidate thresholds are constructed based on Stein unbiased risk estimation. Risk function: ; In the formula: For scale Corresponding candidate threshold The risk function; This is an estimate of the noise standard deviation; The candidate threshold; As an indicator function, when the absolute value of the wavelet coefficients... Less than or equal to candidate threshold hour, The item takes a value of 1, otherwise The value of the item is 0; For scale The first, second, and kth wavelet coefficients in the set of wavelet coefficients j Wavelet coefficients; Then, at scale The upper search finds the baseline threshold that minimizes the risk function: And based on the signal at scale Energy distribution on and kurtosis indicating impact Constructing degradation sensitivity factors : ; In the formula: , These are adjustable weighting coefficients used to balance the contributions of energy features and kurtosis features to the degradation sensitivity factor. ; Finally, based on the benchmark threshold and degradation sensitivity factor Adaptive construction of dual thresholds: ; ; In the formula: , This is the threshold adjustment coefficient.
2. The method according to claim 1, characterized in that: The process of using the hidden state sequence as input to the segmented attention enhancement mechanism in step S4 to obtain the attention-enhanced hidden state is as follows: Step S41: Sequence segmentation, dividing the hidden state sequence into segments of equal length L. A sequence of jokes; Step S42: Local attention aggregation. For each subsequence generated in step S41, perform local attention aggregation to generate segment-level hidden state sequences. , These are the segment-level hidden vectors for the 1st, 2nd, and Mth subsequences, respectively; Step S43: Global attention aggregation, processing the segment-level hidden state sequence generated in step S42. Perform global attention aggregation to obtain a global context vector representing device degradation features. ; Step S44: Feature concatenation and mapping, converting the global context vector from step S43... Key degradation features of the LSTM output in step S3 Feature concatenation is performed to form a comprehensive feature vector. The comprehensive feature vector The key degenerate feature serves as an attention-enhanced hidden state. This refers to the hidden state at the last time step in the hidden state sequence of step S3.
3. The method according to claim 2, characterized in that: Using the last hidden state of each subsequence as a local query vector, we then define the... The subsequence is denoted as , These are all position markers for each hidden state in the subsequence, and the hidden state is... As a local query vector The implementation process of local attention aggregation in step S42 is as follows: First, perform local attention aggregation to compute each hidden state within the subsequence. Local attention score ; Then, the local attention score for the subsequence. Softmax normalization is performed to obtain local attention weights, and the hidden states within the subsequence are weighted and summed to obtain the segment-level hidden vector of the subsequence. ; In the formula, For the first The segment-level hidden vectors of each subsequence are used to generate the segment-level hidden state sequence. , For local attention weights.
4. The method according to claim 2, characterized in that: The implementation process of global attention aggregation in step S43 is as follows: First, use the sequence The last segment-level hidden vector As a global query vector Perform global attention aggregation on each segment-level hidden vector. Calculate the global attention score ; Then process the segment-level hidden vectors The global attention weights are obtained by softmax normalization of the global attention scores. ; Finally, the segment-level hidden vectors of all subsequences are weighted and summed to obtain the global context vector z: 。 5. The method according to claim 1, characterized in that: The monitoring data of the device to be predicted during its entire life cycle or a specified operating phase are vibration signals, temperature signals, or current signals.
6. A prediction system based on the method of any one of claims 1-5, characterized in that: include: The data acquisition module is used to acquire monitoring data of the device to be predicted throughout its entire life cycle or during a specified operating phase. The wavelet enhancement module is used to perform wavelet enhancement operations on monitoring data. The wavelet enhancement operation involves sequentially performing multi-scale wavelet decomposition, threshold processing, and signal reconstruction to obtain a time-series feature sequence that characterizes the health status of the device. The LSTM module is used to input the temporal feature sequence into the LSTM neural network for feature extraction and generate a hidden state sequence to represent the potential degradation trend of the device. The segmented attention enhancement module is used to take the hidden state sequence as input to the segmented attention enhancement mechanism to obtain the attention-enhanced hidden state, and then perform nonlinear mapping through the regression output layer to obtain the device RUL prediction result. The segmented attention enhancement mechanism involves sequentially performing sequence segmentation, local attention aggregation, global attention aggregation, and feature concatenation and mapping.
7. A computer device, characterized in that: include: One or more processors; And a memory that stores one or more computer programs; The processor invokes a computer program to achieve the following: The steps of the device RUL prediction method based on wavelet enhancement and LSTM as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: The computer program is stored and is invoked by the processor to implement: The steps of the device RUL prediction method based on wavelet enhancement and LSTM as described in any one of claims 1-5.
Citation Information
Patent Citations
Heating load ultra-short-term prediction method and system based on deep residual network
CN121809734A
Metal tool wear prediction and maintenance suggestion generation method and system
CN122134319A