Isolator remaining useful life prediction method based on deep learning
By constructing a multi-dimensional sensor sequence and a deep learning model, and combining time-frequency analysis of vibration and temperature signals with a bidirectional recurrent neural network, the accuracy and stability issues of predicting the remaining service life of isolators were solved, and the accurate quantification and dynamic adaptation of the isolator's health status were achieved.
Patent Information
- Application Number
- CN202511142309.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies struggle to accurately predict the remaining lifespan of isolators. Traditional methods have limited accuracy when processing high-dimensional, nonlinear time-series data and lack adaptive capabilities, failing to fully reflect the health status of isolators.
By collecting vibration and temperature signals from the isolator, a multidimensional sensor sequence is constructed, time-frequency analysis is performed, and sub-sequence units are divided. Features are extracted using a deep convolutional network, and a bidirectional recurrent neural network model with an attention mechanism is used to predict the remaining service life and dynamically correct the predicted values.
It achieves comprehensive quantification and accurate prediction of isolator health status, can adapt to dynamic changes, improves prediction accuracy and stability, and reduces deviations caused by fixed model parameters.
Smart Images

Figure CN120744767B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of isolator life prediction, in particular to an isolator residual service life prediction method based on deep learning. BACKGROUND
[0002] As a key component in mechanical and electronic devices, isolators are widely used in aerospace, rail transportation, industrial manufacturing and other fields. Their operating status is directly related to the safety and stability of the entire system. During long-term operation of the device, the isolator will gradually degrade due to factors such as material aging, mechanical wear, and environmental erosion. If the residual service life cannot be grasped in time, it may lead to equipment failure, production interruption, safety accidents and other serious consequences.
[0003] Currently, the methods for predicting the residual service life of isolators mainly include physical model-based methods and data-driven methods. The physical model-based method establishes a mathematical equation for isolator degradation, and combines physical information such as material properties and structural parameters to deduce the life. However, the degradation process of the isolator is influenced by the coupling of multiple complex factors, and its physical mechanism is often difficult to model accurately, resulting in large errors in the prediction results.
[0004] The data-driven method relies on the operating data collected by sensors, and extracts the degradation rules hidden in the data through statistical analysis or machine learning algorithms. Traditional data-driven methods mostly use shallow machine learning models such as support vector machines and random forests. These models have difficulty in effectively extracting deep features when dealing with high-dimensional and nonlinear time series data, resulting in limited prediction accuracy of the residual service life of the isolator. At the same time, existing methods often rely only on a single type of sensor signal, ignoring the relevance between different signals, and cannot fully reflect the health status of the isolator. In addition, when facing dynamic changes in the operating environment, traditional models lack the ability to adaptively adjust, making it difficult to ensure the stability of long-term prediction. SUMMARY
[0005] The purpose of the present application is to provide an isolator residual service life prediction method based on deep learning to solve the problems raised in the background.
[0006] To achieve the above purpose, the present application provides an isolator residual service life prediction method based on deep learning, which comprises:
[0007] Collecting vibration signals and temperature signals during the operation of the isolator to construct a multi-dimensional sensor sequence; performing time-frequency analysis on the multi-dimensional sensor sequence to divide it into multiple subsequence units containing fixed time windows; extracting time domain features and frequency domain features of each subsequence unit through a deep convolutional network to generate a fusion feature vector;
[0008] Calculate the statistical distribution characteristics and trend characteristics of the key indicators in the fusion feature vector, and output the isolator health state index;
[0009] A bidirectional recurrent neural network model containing an attention mechanism is constructed, the isolator health state index and historical degradation data are input, the time sequence dependence is processed through multi-layer residual connection, and a remaining useful life prediction value is generated.
[0010] The multi-dimensional sensor sequence is updated according to the real-time collected sensor signals, and the remaining useful life prediction value is dynamically corrected.
[0011] The calculation method of the isolator health state index includes: calculating the skewness and kurtosis of each dimension data in the fusion feature vector to generate a statistical distribution index; calculating the dynamic time warping distance between the fusion feature vectors of adjacent subsequence units to generate a trend change index; and weighting and aggregating the statistical distribution index and the trend change index, and outputting the isolator health state index after normalization.
[0012] Preferably, the subsequence unit division method includes: wavelet packet decomposition of the vibration signal to obtain different frequency band energy distributions; detecting the extreme point positions of each frequency band energy distribution to segment the vibration subsequence with adjacent extreme points as boundaries; taking the mean and variance of the temperature signal within the corresponding time window as the temperature subsequence; and combining the vibration subsequence and its associated temperature subsequence to form the subsequence unit.
[0013] Preferably, the fusion feature vector generation method includes: inputting the vibration subsequence into a one-dimensional convolution layer to extract time domain waveform features; inputting the frequency band energy distribution after Fourier transform into a two-dimensional convolution layer to extract frequency domain texture features; inputting the temperature subsequence into a fully connected layer to extract thermodynamic features; splicing the time domain waveform features, frequency domain texture features and thermodynamic features, and outputting the fusion feature vector after dimension reduction by a feature compression layer.
[0014] Preferably, the construction method of the bidirectional recurrent neural network model includes: setting a gated recurrent unit as a basic unit to receive an isolator health state index sequence arranged in time sequence; embedding an attention weight module in the input layer to calculate the contribution weight of each time step feature; capturing forward and reverse time sequence dependence through a bidirectional hidden layer; adding a residual connection path between the hidden layers to fuse shallow and deep features; and connecting a fully connected network in the output layer to map to a remaining useful life prediction value.
[0015] Preferably, the dynamic correction method comprises: appending the newly collected sensor signal to the end of the multi-dimensional sensor sequence to generate a new sub-sequence unit; extracting the fusion feature vector of the new sub-sequence unit and calculating the isolator health state index thereof; inputting the isolator health state index into the trained bidirectional recurrent neural network model to replace the input feature of the earliest time step; re-executing the time sequence calculation process to update the remaining useful life prediction value.
[0016] Preferably, the model training step of the bidirectional recurrent neural network model comprises: collecting historical isolator full-life-cycle sensor data and labeling the real label of the remaining useful life; dividing the sensor data into a training set and a validation set; initializing the parameters of the bidirectional recurrent neural network model; minimizing the mean square error of the prediction value and the real label by using the stochastic gradient descent method; and stopping training when the validation set loss function converges.
[0017] Preferably, the method further comprises an abnormal condition processing: if the continuous decline amplitude of the isolator health state index exceeds the degradation threshold, triggering a high-priority sampling instruction; increasing the acquisition frequency of the multi-dimensional sensor sequence; inputting the high-frequency sampling data into the bidirectional recurrent neural network model to perform real-time life prediction compensation calculation.
[0018] Preferably, the real-time life prediction compensation calculation comprises: using a sliding time window to intercept high-frequency sampling data; extracting the fusion feature vector of the multi-dimensional sensor sequence in the window; calculating the current window health state index and the decay rate of the previous window; adjusting the forget gate parameter of the bidirectional recurrent neural network model according to the decay rate; and performing prediction value recalculation.
[0019] Preferably, the operation of the feature compression layer comprises: calculating the variance contribution rate of the spliced feature; screening the feature dimensions with a variance contribution rate greater than a set threshold; mapping the screened features to a low-dimensional space through a linear projection matrix; and outputting the reduced fusion feature vector.
[0020] The implementation of the residual connection path comprises: adding the output feature of the previous hidden layer to the input feature of the next hidden layer; adjusting the scale of the added feature through a learnable weight matrix; inputting the adjusted feature into the gated recurrent unit; and transmitting the final residual output to the next time sequence calculation node.
[0021] Compared with the prior art, the present application has the following beneficial effects:
[0022] The multi-dimensional sensor sequence is constructed by collecting vibration signals and temperature signals during the operation of the isolator, breaking through the limitation of traditional methods relying on single signals. The vibration signals can reflect the state changes of the isolator mechanical structure such as wear and looseness, and the temperature signals can reflect the heat dissipation performance and the aging degree of the internal components. The combination of the two can more comprehensively capture the health state information of the isolator, providing more abundant basis for subsequent life prediction.
[0023] The multi-dimensional sensor sequence is subjected to time-frequency analysis and divided into multiple sub-sequence units containing fixed time windows, so that the originally continuous time series data is transformed into a unit set with local characteristics. This processing method not only retains the time sequence characteristics of the data, but also facilitates targeted analysis of the state characteristics of different time periods, avoiding feature dilution caused by excessively long data, and helping to more accurately extract the degradation characteristics of the isolator in different operating stages.
[0024] The time domain features and frequency domain features of each sub-sequence unit are extracted by a deep convolutional network to generate a fusion feature vector, which fully utilizes the advantages of deep convolutional networks in processing grid-like data and extracting local features. The time domain features can reflect the trend of signal change over time, and the frequency domain features can reveal the changes in frequency components hidden in the signal. The fusion of the two realizes multi-dimensional representation of the state characteristics of the isolator, improving the richness and discriminability of the features.
[0025] The statistical distribution characteristics and change trend characteristics of the key indicators in the fusion feature vector are calculated, and the isolator health state index is output, which converts the high-dimensional feature vector into a quantitative indicator that can intuitively reflect the health state. The statistical distribution characteristics reflect the overall distribution of the key indicators, and the change trend characteristics show the evolution of the key indicators over time. The combination of the two makes the health state of the isolator quantifiable and traceable, and can more clearly present its degradation process.
[0026] A bidirectional recurrent neural network model containing an attention mechanism is constructed, taking the health state index and historical degradation data as input, processing the time sequence dependence relationship through multiple layers of residual connection, effectively solving the gradient disappearance or gradient explosion problem of traditional recurrent neural networks when processing long time series data. The attention mechanism can automatically focus on the historical data segments that are more critical to life prediction, enhancing the model's ability to capture important information. The bidirectional recurrent neural network can utilize both past and future time sequence information, improving the modeling ability of time sequence dependence relationship. The multiple layers of residual connection alleviate the training difficulty of deep network through skip connection, ensuring the fitting effect of the model on complex degradation rules.
[0027] The multi-dimensional sensor sequence is updated according to the real-time collected sensor signals, and the residual service life prediction value is dynamically corrected, so that the model can adapt to the dynamic changes in the operation process of the isolator. With the continuous input of new data, the model can continuously learn the latest degradation features and timely adjust the prediction results, avoiding the prediction deviation caused by fixed model parameters, and ensuring that the prediction results can be consistent with the actual state of the isolator in the long-term operation process. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 The working principle diagram of the deep learning-based isolator residual service life prediction method is provided.
[0029] Figure 2 The flow chart for generating the fusion feature vector is provided.
[0030] Figure 3 The flow chart for constructing the bidirectional recurrent neural network model is provided.
[0031] Figure 4 The flow chart for dynamic correction is provided.
[0032] Figure 5 The flow chart for real-time life prediction compensation calculation is provided. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application.
[0034] Please refer to Figure 1 The present application provides a deep learning-based isolator residual service life prediction method, which comprises:
[0035] During the operation of the isolator, the vibration acceleration sensor and the temperature sensor synchronously collect signals to form a multi-dimensional sensor sequence containing time stamps. After preprocessing, the original signals are divided into subsequence units of fixed length by using a sliding time window, and each unit contains a vibration signal segment and the temperature statistics of the corresponding time window. The deep convolutional network performs parallel feature extraction on the subsequence: the one-dimensional convolutional layer processes the vibration waveform, the two-dimensional convolutional layer analyzes the frequency energy distribution, and the fully connected layer processes the temperature features. After the extracted features are fused by the compression layer, the statistical distribution characteristics and trend change characteristics are calculated, and the health state index in the range of 0-1 is output. A bidirectional recurrent neural network model is constructed, the input layer of which is embedded with an attention mechanism, the hidden layer uses GRU units with residual connections, and the output layer maps to the remaining useful life value. The model continuously receives new sensor data during operation, dynamically updates the prediction results, and triggers the abnormal working condition processing mechanism.
[0036] Example 1: see Figure 2 The vibration signal is collected by a three-axis acceleration sensor, and the sampling frequency is set to 10 kHz. After anti-aliasing filtering, the signal enters the preprocessing link. The preprocessing includes DC component elimination and amplitude normalization, and the normalization range is limited within the ±5V sensor range. The wavelet packet decomposition selects the Daubechies4 wavelet basis function, and the decomposition layer is 3 layers, generating an energy distribution of 8 frequency bands. The signal of each frequency band is extracted by Hilbert transform to obtain the envelope line, and the local extreme point detection of the envelope signal uses the sliding window comparison method, and the window width is set to 50 sampling points. When the amplitude of a certain sampling point is greater than the amplitudes of the previous 25 points and the next 25 points within the window range, it is determined as a local extreme point. The signal segment between adjacent extreme points is intercepted as a vibration subsequence, and the minimum length constraint is about 100 sampling points. The segments with insufficient length are merged with the previous valid subsequence.
[0037] The temperature signal is collected by a PT100 platinum resistance temperature sensor, and the sampling frequency is synchronized with the vibration signal, which is set to 1 kHz. Within the time window corresponding to the vibration subsequence, the temperature signal is first subjected to median filtering to remove pulse interference, and then the 60-second moving average and standard deviation are calculated. The temperature statistics calculation uses an overlapping window method, and the window sliding step is 1 second, ensuring time alignment with the vibration subsequence. The temperature feature vector contains four components: the temperature mean value, the temperature standard deviation, the temperature change rate, and the temperature difference value with the previous window.
[0038] The temporal feature extraction network employs a three-layer one-dimensional convolutional structure. The input layer receives the raw waveform data of the vibrational subsequence. The first convolutional layer contains 16 kernels of length 5 with a stride of 1, using symmetric padding to preserve the feature map size. The activation function is LeakyReLU, with a negative slope coefficient set to 0.01. The second layer increases the number of kernels to 32, and the third layer to 64, with batch normalization following each layer. Max pooling layers are placed between every two convolutional layers, with a pooling window size of 2 and a stride of 2. In the frequency domain feature extraction path, the energy distribution of each frequency band is first subjected to a 256-point Fast Fourier Transform, and then combined into a time-frequency matrix, which is input into the two-dimensional convolutional network. The first two-dimensional convolutional layer uses 32 kernels of 3×3, and the number increases to 64 in the second layer, with a stride of 2×2 for both layers. A spatial pyramid pooling layer is added after each convolutional layer in the frequency domain network, outputting a fixed-dimensional feature vector.
[0039] The temperature feature processing path employs a fully connected network structure. The first layer contains 128 neurons with the ReLU activation function, and the second layer is compressed to 64 dimensions. The network input includes the temperature statistics for the current window and the temperature change trends for the previous three historical windows. A 20% Dropout layer is added after each fully connected layer to prevent overfitting. In the feature fusion stage, time-domain features, frequency-domain features, and temperature features are all normalized to L2 norm before concatenation. The fused high-dimensional features are then dimensionality-reduced using principal component analysis (PCA), retaining 95% of the cumulative variance contribution. PCA calculations utilize an incremental learning method to adapt to the needs of online updates.
[0040] The feature compression layer employs a two-stage processing strategy. The first stage calculates the covariance matrix between feature dimensions to determine the variance contribution of each dimension. The second stage constructs an orthogonal projection matrix, linearly mapping the high-dimensional features to a 256-dimensional latent space. The projection matrix update mechanism uses a sliding window approach, recalculating the PCA parameters after processing every 1000 subsequences. The dimensionality-reduced fused feature vectors are appended with timestamp information and stored in a circular buffer for subsequent processing. The buffer management adopts a least recently used strategy, with a maximum capacity of 10,000 feature vectors.
[0041] During vibration subsequence segmentation, when three consecutive extreme points are detected with an interval exceeding 200ms, a virtual segmentation point is automatically inserted to avoid generating excessively long subsequences. The segmentation algorithm maintains a priority queue and dynamically adjusts the sensitivity threshold for extreme point detection. In temperature signal processing, when a sample value jump exceeding 5℃ is detected, an outlier removal procedure is initiated, replacing the outlier with the median of the preceding and following 10 sample points. All convolutional layers in the feature extraction network are initialized using the He normal distribution method, with the bias term initialized to zero. The learnable parameters of the batch normalization layer are initialized with unit scaling and zero offset.
[0042] A quality monitoring mechanism is set for the time-frequency analysis link. When the energy proportion of a certain frequency band is less than 0.1% of the total energy for 5 consecutive subsequences, the feature extraction path of the frequency band is automatically shielded. The frequency domain convolution network adopts a separable convolution structure, and the ratio of the depth convolution kernel to the point convolution kernel is set to 3:1. The temperature feature network adopts a residual connection structure, and the output of the first layer is added to the output of the second layer before being activated. The normalization processing before feature fusion includes a learnable scaling factor, and the initial value is set to the reciprocal of the standard deviation of the output of each feature path.
[0043] During online processing, the system maintains a feature extraction pipeline, including three parallel threads: vibration signal processing thread, temperature signal processing thread and feature fusion thread. Each thread adopts a double buffering mechanism to ensure the continuity of real-time data processing. The vibration feature extraction path sets a dynamic computation graph, which automatically adjusts the padding parameters of the convolution layer according to the length of the input subsequence. The frequency domain analysis path adopts a multi-phase filter group to realize the combination of wavelet packet decomposition and Fourier transform into a single calculation process. The temperature feature network is implemented as a lightweight structure, using 8-bit integer quantization to accelerate calculation.
[0044] The feature storage adopts a columnar compression format, and each feature vector is attached with a timestamp, a device ID and a signal quality flag. The signal quality evaluation is based on three indicators: the signal-to-noise ratio of the vibration signal, the stability of the temperature signal and the reconstruction error of the feature vector. When any of the indicators exceeds the preset threshold, the corresponding feature vector is marked as a low-quality sample. The feature buffer realizes an automatic cleaning function, removing historical data more than 7 days every 24 hours. The intermediate results generated during feature extraction are saved in a memory-mapped file, supporting breakpoint resume processing function.
[0045] The entire processing flow adopts an event-driven architecture, which triggers the feature extraction pipeline when new sensor data arrives. The events generated by the vibration signal processing thread include: subsequence segmentation completion, time domain feature extraction completion, frequency domain analysis completion, etc. The events generated by the temperature signal processing thread include: temperature statistical quantity calculation completion, temperature feature extraction completion, etc. The feature fusion thread listens to the above events and executes feature splicing and dimensionality reduction operations when all the prerequisites are met. The system state monitoring module tracks the delay and resource usage of each processing stage in real time, dynamically adjusts the thread priority and resource allocation.
[0046] Example 2: see Figure 3The health state index is calculated and the recurrent neural network is constructed, and the degradation state quantitative evaluation and life prediction are realized through multi-dimensional feature analysis and time series modeling. After the fusion of the feature vector input analysis module, the system first selects the top 20 feature dimensions with the highest variance contribution rate as the core analysis object. The skewness calculation uses the standard definition method, and the ratio of the third central moment to the standard deviation cube of each dimension data distribution is calculated. The kurtosis calculation analyzes the relationship between the fourth central moment and the square of the variance, and quantitatively represents the long-tailed distribution. The abnormal value processing mechanism adopts the Huber loss function principle, and sets the conversion point to 3 times the standard deviation. Data exceeding this point is rescaled. The calculation result forms a 20-dimensional statistical distribution index matrix.
[0047] The trend change detection module starts the dynamic time warping algorithm, and configures the curved path window constraint to be 15% of the current sequence length. The feature vector of the current subsequence is compared with the latest 5 subsequences in the history buffer one by one, and the buffer uses a circular queue structure. The distance calculation uses the Euclidean distance metric, and the deformed distance after path constraint is used as the similarity criterion. After the algorithm outputs 5 distance values, the minimum value is recorded as the trend change indicator. The history buffer management strategy is set to a fixed capacity of 200 groups of data, and the new calculation result is updated and stored according to the first-in first-out principle.
[0048] The health state index synthesis stage presets the weighting coefficient, and the statistical distribution index accounts for 60% of the total weight, and the trend change indicator accounts for 40%. The weighting operation uses the element-by-element multiplication and linear addition method, and the synthesis value domain is the intermediate variable in the range [-1, 1]. The normalization module uses the modified form of the hyperbolic tangent function to map the intermediate variable to the [0, 1] interval. The system adds a timestamp and device identification information, encapsulates it as a health state record package and writes it into the time series database.
[0049] The bidirectional recurrent neural network is configured with two layers, and each layer uses 128 gated recurrent units. The attention module deployed in the input layer includes three calculation processes: the feature vector generates a 128-dimensional hidden state representation through a fully connected layer; the dot product score of the hidden state and the trainable query vector is calculated; and the softmax function with a temperature parameter of 0.5 is applied to obtain the standardized weight. The weighted input sequence is processed in two directions: the original time series data is input into the forward recurrent layer, and the data arranged in reverse order is input into the reverse recurrent layer. The output state of each recurrent layer is a 128-dimensional vector, and the outputs of the same time step in two directions are concatenated in the channel dimension to form a 256-dimensional representation.
[0050] The residual connection acts between the network hidden layers. The output of the first layer of the recurrent network is adjusted by a 1x1 convolution kernel to adjust the number of channels, and the convolution kernel is initialized as a unit matrix. The adjusted features and the original features input to the second layer perform element-level addition operations, and the vector is directly input to the second layer of the recurrent unit. The internal calculation of the gated recurrent unit is divided into three branches: the update gate uses Sigmoid activation to determine the proportion of historical state to be retained; the reset gate controls the degree of historical state participating in the calculation of the current candidate state; and the candidate state is generated by a hyperbolic tangent function. The three gate outputs are combined to form a new hidden state.
[0051] The output layer structure includes two fully connected modules. The first module performs linear transformation on the spliced features to a 64-dimensional space, and the activation function uses a modified linear unit. The second module maps the 64-dimensional features to a scalar output, and the activation function is a smooth variant of ReLU. The network prediction output is the estimated value of the remaining useful life, and the unit is consistent with the unit of the training data label.
[0052] The system maintains the hidden state transmission mechanism during the running period. The current hidden state is retained after each time step calculation for initialization in the next iteration. The sequence processing length is set to 200 time steps, and the input sequence that is too long is processed by using a sliding window. The input features are disturbed by Gaussian noise, and the noise standard deviation is set to 5% of the feature standard deviation. The gradient calculation uses the truncation method, and the norm threshold is set to 1.0 to prevent numerical instability.
[0053] The network parameter initialization strategy uses the orthogonal initial method, and the connection weight between units is initialized using a Gaussian distribution. The bias term is configured in three parts: the update gate bias is initialized as a full 1 vector, the reset gate bias is set to -1, and the state transition bias is initialized to zero. The network running environment is configured for double-precision floating-point operation, and the derivative calculation of the activation function is realized by using the analytical expression. Each layer output is subjected to layer normalization processing before internal loop calculation to reduce the influence of covariant bias.
[0054] The time series data processing pipeline sets a feature cache mechanism. The health state calculation module outputs the results to a ring buffer with a capacity of 1000 records. The recurrent neural network reads data from the buffer in chronological order in batches, and the batch size is dynamically adjusted to 80% of the device processing capacity. The long-term dependence modeling adopts a segmented recurrent strategy, and the hidden state is stored in the long-term memory library after processing 50 time steps. The memory library uses key-value pair storage, with the key value being the device state hash value and the value being the corresponding hidden state tensor.
[0055] The parameters of the residual connection are regularized, and an L2 penalty term is added to the loss function with a coefficient of 0.001. The gradient transmission in the multi-layer structure uses a straight-through estimator design, which skips the complex gate unit derivative calculation during backpropagation. The state diagnosis function is implemented inside the network to monitor the unit activation value distribution in real time, and automatically insert the reset instruction when saturation occurs. The model service layer maintains multiple computation graph instances, dynamically loads the corresponding pre-trained parameters according to the device type, and supports parallel prediction tasks for heterogeneous device groups.
[0056] The health index calculation module sets an output verification link. The statistical indicators and trend indicators are set with effective range thresholds, and the calculation results exceeding the limits trigger data re-sampling commands. The recurrent neural network checks the continuity of the time stamp when receiving input data, and automatically linearly interpolates the missing data when a breakpoint occurs. The model version management adopts a blue-green deployment scheme, with new and old models running in parallel, and the final switching time is determined by correlation detection of the prediction results. The entire prediction system maintains device resource usage monitoring during operation, and starts the intermediate data transfer mechanism when memory usage reaches 80%.
[0057] Example 3: refer to Figure 4 , the dynamic correction mechanism and the model training process, through real-time data update and optimization algorithm to realize the continuous improvement of the prediction system. During system operation, newly collected sensor data triggers a multi-dimensional sequence update process, which uses a sliding window mechanism to maintain fixed-length time series data. The window size is set to 200 sub-sequence units, and when new data arrives, the system performs a first-in, first-out strategy, removing the earliest time step data and appending the latest collected information. The feature extraction path of the new sub-sequence unit maintains structural consistency with the offline training phase, and all convolution kernel weights and fully connected parameters are frozen to the trained values, ensuring consistent expression of the feature space. The health status index calculation module loads the pre-defined statistical weight vector, which is determined by principal component analysis in the training phase to determine the contribution of each feature dimension.
[0058] The core of dynamic correction is the incremental inference mechanism of the time series model. The bidirectional recurrent neural network model maintains hidden state memory, which is matched with the network structure. When a new health status index is input, the model performs local forward calculation, only processing the newly added time series node. The calculation process involves dynamic adjustment of attention weights, and the importance score of the current time step is calculated as follows:
[0059]
[0060] where, represents the attention weight of the th time step, is a trainable parameter vector, and transformation matrix corresponding to hidden state and input feature respectively, is the hidden state of the previous time, is the current input feature, represents the total length of the sequence, is the index variable of time step, used to traverse all time points of the entire time sequence (from step 1 to step , where is the total length of the sequence), : is the index variable in the summation process, used to traverse each time step from 1 to , : is the hidden state of the th time, : is the input feature of the th time step, : exponential function, used to map the calculation results to the positive value range, facilitating subsequent normalization processing, : hyperbolic tangent function, as an activation function, used for nonlinear processing of transformed hidden state and input feature, : summation symbol, used to accumulate the calculation results of the exponential function from 1 to each time step to realize the normalization of attention weight. The parameter symbol in this formula has uniqueness in this embodiment and does not repeat with the mathematical expressions of other embodiments.
[0061] The model training phase uses a full life cycle dataset to construct samples, each sample containing all health state sequences from device activation to retirement. The time alignment operation is performed in the data preprocessing link to unify the sensor data of different sampling frequencies to a common time reference. The division of training set and validation set is based on the device serial number hash value, ensuring that the data of the same device does not appear in both sets at the same time. The network parameter initialization adopts the orthogonal matrix method, and the weight matrix of the cyclic connection is obtained by QR decomposition to obtain the orthogonal basis. The optimization algorithm selects a variant version of Adam, the momentum parameter β1 is set to 0.9, β2 is adjusted to 0.999, the initial value of the learning rate is 0.001 and the cosine annealing strategy is used for adjustment.
[0062] The loss function is designed as the weighted sum of smooth L1 loss and prediction bias penalty term. The former measures the difference between the predicted value and the true remaining life, and the latter constrains the mutation amplitude between consecutive prediction results. The early stopping strategy monitors the trend of validation loss, and terminates the training process when there is no decrease for 15 consecutive training cycles. The data enhancement module injects two types of disturbances during training: Gaussian process noise in the time dimension, with a strength of 10% of the feature standard deviation; random mask in the feature dimension, with a discard probability of 5%. The model saving adopts the checkpoint mechanism, recording a best parameter snapshot once every 5 training cycles.
[0063] During online updating, the system maintains two independent memory regions: real-time inference buffer and historical reference database. The former stores the last 200 time step's feature vectors and model states, supporting multi-process access with shared memory design. The latter saves the complete device running history, using columnar storage format for compression and archiving. When new data triggers prediction updating, the system first checks the timestamp continuity, filling in the missing data with linear interpolation of adjacent values. The model inference results are processed with moving average filtering, with a window width of 5 time steps and a dynamic adjustment of the smoothing coefficient to adapt to the prediction fluctuation level.
[0064] The abnormal recovery mechanism designs three layers of protection measures. The primary protection detects the reasonable range of input features, and data beyond the physical range is discarded and requested to be retransmitted. The intermediate protection analyzes the self-consistency of the feature vector, and the reconstruction error is calculated by the pre-trained autoencoder. The threshold value of the sample triggers the feature recalculation process. The advanced protection monitors the time sequence consistency of the prediction results. When the life degradation rate of the continuous 3 predictions exceeds the maximum degradation rate of the device, the model state rollback operation is started, and the parameters of the previous stable version are loaded to continue serving.
[0065] The model version management adopts a differential update strategy. Each parameter adjustment only transmits the difference from the previous version, and the update package is checked for integrity by cyclic redundancy check. When deploying, the blue-green release mode is used, and the new model runs in shadow mode. Its prediction results are correlated with the output of the actual service model, and the quality is confirmed to meet the standard before switching traffic. The fallback mechanism keeps two stable versions available, which can quickly recover to the previous version when the new version appears abnormal.
[0066] The resource management subsystem dynamically allocates computing resources. CPU-intensive tasks such as feature extraction are deployed on dedicated computing nodes equipped with AVX instruction set optimization. Memory-sensitive operations such as recurrent network inference run in large memory containers equipped with NUMA core binding strategies. GPU resources are shared through time slicing, with each task getting a fixed amount of computing quota. Network bandwidth reserves 20% of the margin to meet sudden data transmission needs. All computing tasks are tagged with priority labels, and when the system load exceeds 70%, the quality of service degradation mechanism is started, and non-critical background tasks are suspended.
[0067] The state monitoring component collects three types of running indicators: time series feature extraction delay, model inference throughput, and prediction result confidence. These indicators are smoothed by the exponential weighted moving average algorithm, and input into the resource scheduling decision maker. When any indicator exceeds the preset range, the system automatically adjusts the parallelism or computing accuracy of the related module. For example, when the feature extraction delay increases, the floating-point operation may be temporarily reduced from double precision to single precision mode. All adjustment operations are recorded in the audit log for subsequent performance analysis.
[0068] The data persistence scheme adopts a tiered storage architecture. The latest 200 records are saved in an in-memory database to ensure real-time access performance. The medium-term historical data is written to a solid-state drive with a compression ratio set to 30% of the original size. The long-term archived data is stored in a distributed file system with error correction code encoding to provide fault tolerance. The data migration process is completely transparent. Applications access data at different levels through a unified interface, and the system automatically handles location resolution and format conversion. The backup strategy performs daily incremental backups and weekly full backups, with a retention period set to a three-month rolling window.
[0069] Example 4: Refer to Figure 5 The abnormal condition handling and real-time compensation mechanism ensures the stable operation of the prediction system during the accelerated degradation period of the equipment through multi-level monitoring and dynamic adjustment. Taking the centrifuge set isolator monitoring in a chemical plant as an example, the system triggers the abnormal handling process after 8760 hours of continuous operation. The health status index monitoring module records a decrease in the value of the last three sampling periods, with the specific data changes shown in Table 1.
[0070] Table 1: The specific data changes of the last three sampling periods are as follows.
[0071] Time stamp Health status index Single change amount Cumulative change amount Sampling frequency flag 2024-03-0514:00 0.82 -0.03 -0.03 Standard (1 Hz) 2024-03-0515:00 0.78 -0.04 -0.07 Standard (1 Hz) 2024-03-0516:00 0.71 -0.07 -0.14 Standard (1 Hz) 2024-03-0516:05 0.69 -0.02 -0.16 Enhanced (4 Hz) 2024-03-0516:10 0.67 -0.02 -0.18 Enhanced (4 Hz)
[0072] When the cumulative change exceeds the preset threshold of 0.15, the system automatically activates the high-priority sampling mode. The vibration sensor sampling rate is increased from 1 Hz to 4 Hz, and the temperature sensor is increased from 0.2 Hz to 1 Hz. The data acquisition card switches to direct memory access mode, bypassing the system buffer and directly writing the measurement values to the predetermined memory area. The high-frequency sampling data is temporarily stored in a circular buffer, which is divided into 16 data blocks, each block containing 5 minutes of raw waveform data. The real-time processing thread extracts the latest data block from the buffer and performs fast feature extraction calculations.
[0073] The decay rate analysis module uses a sliding window comparison method. The current window selects the last 10 minutes of high-frequency data, and the reference window takes the normal sampling data 1 hour before the abnormal trigger. After aligning the feature vectors of the two windows through dynamic time warping, the relative change rate of each dimension feature is calculated. Among the vibration features, the energy change in the 2-4 kHz frequency band is the most significant, reaching 2.3 times the reference value. In the temperature features, the gradient change rate increases to 1.8 times the normal operating condition. These indicators are input into the decision tree model, and the output forgetting gate adjustment coefficient is 0.65, indicating that the model needs to strengthen the memory weight of recent data.
[0074] After the real-time compensation calculation is started, the system creates a dedicated prediction instance. This instance loads the same network parameters as the main model, but modifies three key configurations: the time step is compressed to 1 / 4 of the original to accommodate high-frequency data, the forget gate parameter is multiplied by an adjustment coefficient, and the output layer is added with a moving average filter. The instance runs on a separate computing unit to avoid interfering with the main prediction process. Each calculation consumes 4 consecutive high-frequency sampling points, outputs an intermediate prediction result, and marks it as a compensation value. The main prediction system continuously receives standard frequency data, but will refer to the compensation value to correct the final output.
[0075] The sampling frequency management adopts a gradual adjustment strategy. The initial enhancement phase maintains 4Hz sampling for 30 minutes, and if the decay rate continues to be higher than the threshold, it will be further increased to 8Hz. The recovery condition is set to a fluctuation range of less than 0.01 in the health index of 6 consecutive sampling periods. During the frequency switching process, the system performs sensor recalibration procedures, including zero correction and sensitivity verification. The data acquisition unit automatically records the switching timestamp and parameter snapshot at the mode conversion time for subsequent analysis.
[0076] The abnormal event report generates a structured log, including five parts: trigger condition record, data quality assessment, processing measure details, resource consumption statistics, and prediction difference comparison. Log entries are sent to the central monitoring platform through the message queue, and a rolling cache is retained locally. The analysis service on the platform side automatically extracts key fields, generates device health trend charts and processing effect heat maps. Operation and maintenance personnel can view the superimposed display of the original waveform and feature changes to assist in judging the nature of the anomaly.
[0077] The resource guarantee mechanism is dynamically adjusted during abnormal processing. The compensation calculation priority is raised to real-time level by the computing task scheduler, and CPU core binding avoids context switching. Memory allocation uses a pre-reservation strategy to ensure that high-frequency sampling buffers are not recycled due to system pressure. Network transmission enables a dedicated channel, and the transmission interval of sampling data packets is shortened from the regular 10 seconds to 1 second. When the overall system load exceeds 85%, non-essential background services such as data archiving and regular diagnostics are automatically suspended.
[0078] The quality control module implements a triple verification mechanism. The original data verification checks the physical reasonableness of the sampling value, and the segment with a vibration amplitude exceeding 80% of the range will be marked. The feature verification compares the homologous features extracted at different frequencies, and the difference exceeding 10% triggers feature recalculation. The prediction verification monitors the output difference between the main model and the compensation model, and if the deviation is greater than 15% for 3 consecutive times, the model consistency check is started. All abnormal events will generate a verification report, recording detailed environmental variables and system status.
[0079] The fault recovery process is designed as a step-by-step rollback. The first stage attempts to recalculate the data for the last 5 minutes using the backup feature extraction parameters. If the problem persists, the second stage rolls back to the last stable model version and maintains high-frequency sampling to continue observation. The final stage switches to safe mode when two consecutive recovery attempts fail, providing only raw sensor data without prediction. Each recovery operation generates an audit trail record, including timestamp, operation type, impact range, and execution result.
[0080] Version compatibility management maintains three interface adaptation layers. The data format adapter handles data structure conversion at different sampling frequencies to ensure that new and old modules can correctly parse the input. The feature space mapper aligns the feature dimension differences of different version models by linear projection to maintain compatibility. The prediction result converter unifies the output scale and unit to eliminate display differences caused by version iteration. The adaptation layer configuration information is stored in a separate metadata database, supporting dynamic loading and hot updating.
[0081] During an actual processing process, the system detected an abnormal decrease in health index at 14:15 and confirmed the entry into the accelerated degradation phase at 16:20. High-frequency sampling continued until 02:00 the next day, during which 47 real-time compensation calculations were performed. The main model predicted the remaining life from 650 hours to 320 hours, and the compensation model output was 290 hours. Finally, the weighted average result of 305 hours was used as the official prediction value, with an error of 312 hours from the actual failure time, which was within a reasonable range. The entire event processing consumed 2.3 times the normal operating condition, and did not affect the data acquisition tasks of other monitoring points. The root cause analysis generated by the event report indicated that the sudden change in vibration characteristics caused by insufficient lubrication of the bearing was the main cause, consistent with the subsequent disassembly and inspection conclusion. Among them, bearing lubrication refers to the lubrication state of the core rotating parts (such as bearings) inside the isolator. As a key component in mechanical systems, the bearings of the isolator reduce friction loss through lubricating media (such as lubricating oil, grease) to maintain stable operation. When lubrication is insufficient, the friction coefficient of the bearing contact surface increases, which can cause the vibration amplitude to rise (especially the energy significantly increases in the 2-4 kHz high frequency band), and the temperature signal gradient rises due to friction heat. These changes will be captured by vibration and temperature sensors and reflected in multi-dimensional sensor sequences.
[0082] Example 5: Detailed implementation of feature compression and residual connections, enhancing the expressive power of deep neural networks through dimensionality reduction optimization and information preservation strategies. The feature compression layer operation begins with the calculation of the covariance matrix of the concatenated features, using a randomized singular value decomposition algorithm to process the high-dimensional matrix. This method first projects the original features onto a Gaussian random matrix to generate a low-dimensional approximate matrix, and then performs precise singular value decomposition on this matrix. During the decomposition process, a convergence threshold of 1e-6 (in scientific notation, corresponding to a value of 0.000001, or one in a million) is set, and the maximum number of iterations is limited to within 100. Variance contribution rate analysis selects the right singular vectors corresponding to the top k singular values, with the k value dynamically determined based on a preset cumulative energy percentage. In the feature selection stage, a dimensionality importance score is established, comprehensively considering the variance contribution of each dimension, its correlation with other dimensions, and its business interpretability.
[0083] The projection matrix is constructed using truncated singular vectors as basis vectors, and the orthogonality of the basis is ensured through Gram-Schmidt orthogonalization. Matrix elements are stored as single-precision floating-point numbers with compressed encoding of the sign and exponent bits. The dimensionality reduction mapping operation consists of two steps: first, the original features are centered by subtracting the mean vector, and then multiplied right by the projection matrix to obtain the low-dimensional representation. The compressed feature vector is appended with two metadata fields: the original variance proportions for each dimension and the upper bound of the reconstruction error, for reference by subsequent modules. The feature cache uses memory-mapped file storage, supporting multi-process shared access and fast serialization.
[0084] The implementation of the residual connection path includes the initialization and optimization of the learnable weight matrix. The matrix dimension is automatically calculated based on the feature dimensions of the preceding and following layers. When the dimensions match, it is initialized as an identity matrix; otherwise, the Xavier initialization method is used. During forward propagation, the residual branch performs an identity mapping or linear projection, adding element-wise to the output of the main path. The summed features are adjusted for magnitude using a trainable scaling factor, initialized to 0.5 and constrained within the range [0.1, 2.0]. During backpropagation, the gradients flow to both the main path and the residual path, with the gradient ratio between the two paths controlled by adaptive weights.
[0085] The process of candidate state generation changes after the introduction of residual connections in the internal state update of the gated recurrent unit. Historical hidden states, after being filtered by a reset gate, are concatenated with the current input features and then dimensionality reduced using a projection matrix. The dimensionality-reduced features are added to the skipped features of the residual connections and then input into the hyperbolic tangent activation function. When updating the mixing ratio of old and new states by the gate control, the confidence score of the residual path is additionally considered. This score is calculated based on the reconstruction error of the residual features; the smaller the error, the higher the weight of the residual's contribution.
[0086] The network training phase implements special regularization strategies for the residual connections. A path difference penalty term is added to the loss function to encourage complementary representations learned by the main path and the residual path. Gradient clipping is used to limit the maximum norm when the gradient flows through the residual connection, preventing the amplitude from being too large to interfere with the learning of the main path. The learning rate of the residual path is set to 1.2 times that of the main path during parameter update, accelerating the formation of feature reuse capability. After each training batch, the activation sparsity of each residual connection is counted, and the connection path with low usage is subjected to dropout enhancement.
[0087] The online update mechanism of the feature compression layer uses incremental learning. The newly arrived sample features are first standardized, and then the running mean and variance estimates are updated. The maintenance of the covariance matrix is realized through rank 1 update, and singular value decomposition recalculation is triggered every 100 new samples. The version management of the projection matrix uses generation markers, retaining the last three valid versions for rollback. When feature distribution drift is detected, the matrix retraining process is automatically started, and the update is completed in the background thread without affecting real-time prediction.
[0088] The dynamic adjustment module of the residual connection monitors the feature fusion effect. Evaluation indicators include path correlation coefficient, contribution balance factor and information entropy ratio. When the output correlation of the main path and the residual path exceeds 0.7, a feature decorrelation operation is automatically inserted. When the contribution imbalance reaches a 3:1 ratio, the learning rate of the scaling coefficient is adjusted for correction. The information entropy ratio reflects the diversity of the two paths, and a value between 1.5 and 2.0 is considered ideal. All monitoring data is visualized to help understand the information flow within the network.
[0089] The computational graph optimization is specifically accelerated for the residual structure. The adjacent linear projection and addition operations are fused into a single operator, reducing the number of memory accesses. Parallel processing of multiple residual branch operations is achieved using CUDA streams for asynchronous execution. Static memory allocation reserves a fixed space for residual features, avoiding the overhead of dynamic allocation. Operator-level optimizations include loop unrolling, vector instruction usage, and shared memory caching, specifically for matrix-vector multiplication and addition operations.
[0090] The quantization process in the deployment phase is configured separately for the residual connection path. The main path uses 8-bit integer quantization, and the residual path retains 16-bit floating-point precision. This mixed precision strategy balances computational efficiency and numerical stability. The quantization parameter calibration uses a moving average method to track the minimum / maximum value range of each connection. Dynamic range adjustment sets a safety margin to prevent feature distortion caused by overflow. When deployed on edge devices, the scaling coefficients of the residual connection are converted to fixed-point representations, and the bit width is configured according to device capabilities.
[0091] Error recovery mechanism is designed for specific checkpoints of feature compression and residual operation. Feature compression layer records the norm ratio of input and output, and triggers data re-normalization when abnormal fluctuation is detected. Residual connection checks the sign of gradient explosion, and switches to safe mode automatically when abnormality is detected, temporarily disabling the residual path. Memory error handling includes buffer overflow protection and wild pointer detection, with boundary check for all array access. Computational consistency verification compares the difference of results from different precision modes, and falls back to high precision computation when the difference exceeds the threshold.
[0092] Performance analysis tools are integrated into key links. Feature compression layer records the information loss rate before and after dimension reduction, and draws the energy distribution spectrum. Residual connection analyzer statistics path activation frequency, gradient propagation intensity and parameter update amplitude. Real-time monitoring interface shows the feature evolution in the depth direction of the network, and presents the residual compensation effect with a heat map. Performance counters measure the latency and throughput of each module, and identify computational bottlenecks. All analysis data are time-stamped, supporting cross-module correlation analysis.
[0093] Version compatibility is guaranteed through interface abstraction layer. Feature compression algorithm is encapsulated as a unified interface, and the internal implementation can be replaced with different dimension reduction methods. Residual connection standardizes the input and output tensor format, and hides the specific implementation details. Module registry manages available implementation versions, and dynamically loads at runtime according to configuration. Data format converter handles transparent conversion between different precisions, ensuring seamless integration between modules. Compatibility test suite covers boundary conditions and abnormal scenarios, ensuring that upgrades do not affect existing functions.
[0094] It should be noted that the relational terms, such as first and second, and the like, are used solely to distinguish one from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0095] Although embodiments of the present application have been shown and described, it is to be understood that various modifications, substitutions, replacements and changes can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalents.
Claims
1. A deep learning-based isolator residual useful life prediction method, characterized by, The application relates to a vibration signal and a temperature signal collected during the operation of an isolator, a multi-dimensional sensor sequence is constructed, time-frequency analysis is performed on the multi-dimensional sensor sequence, a plurality of subsequence units containing fixed time windows are divided, time domain features and frequency domain features of each subsequence unit are extracted through a deep convolution network to generate a fusion feature vector, statistical distribution characteristics and variation trend characteristics of key indicators in the fusion feature vector are calculated, and an isolator health state index is output, a bidirectional recurrent neural network model containing an attention mechanism is constructed, the isolator health state index is taken as input, time sequence dependence is processed through multi-layer residual connection, and a remaining service life prediction value is generated, the multi-dimensional sensor sequence is updated according to real-time collected sensor signals, and the remaining service life prediction value is dynamically corrected, the calculation method of the isolator health state index comprises the following steps: the skewness and kurtosis of each dimension data in the fusion feature vector are calculated to generate a statistical distribution index; a dynamic time warping distance between fusion feature vectors of adjacent subsequence units is calculated to generate a trend change index; and the statistical distribution index and the trend change index are weighted and aggregated, and the isolator health state index is output after normalization, abnormal working condition processing is further included, if the continuous descending amplitude of the isolator health state index exceeds a degradation threshold, a high-priority sampling instruction is triggered, the collection frequency of the multi-dimensional sensor sequence is increased, high-frequency sampling data is input into the bidirectional recurrent neural network model, and real-time life prediction compensation calculation is performed, the real-time life prediction compensation calculation comprises the following steps: high-frequency sampling data is intercepted by using a sliding time window; the fusion feature vector of the multi-dimensional sensor sequence in the window is extracted; the health state index of the current window and the decay rate of the previous window are calculated; the forgetting gate parameter of the bidirectional recurrent neural network model is adjusted according to the decay rate; and the prediction value is recalculated. The subsequence unit division method comprises the following steps: wavelet packet decomposition is performed on the vibration signal to obtain different frequency band energy distributions; extreme point positions of the frequency band energy distributions are detected, and the vibration subsequences are segmented with adjacent extreme points as boundaries; the mean and variance of the temperature signal in the corresponding time window are taken as temperature subsequences; and the vibration subsequences and the temperature subsequences associated with the vibration subsequences are combined to form the subsequence units. The fusion feature vector generation method comprises the following steps: the vibration subsequence is input into a one-dimensional convolution layer to extract time domain waveform features; the frequency band energy distribution is input into a two-dimensional convolution layer after Fourier transform to extract frequency domain texture features; the temperature subsequence is input into a fully connected layer to extract thermodynamic features; the time domain waveform features, the frequency domain texture features and the thermodynamic features are spliced, the fusion feature vector is output after dimension reduction through a feature compression layer. 2.The deep learning-based isolator remaining useful life prediction method of claim 1, wherein, 3.The deep learning-based isolator remaining useful life prediction method of claim 2, wherein, 4.The deep learning-based isolator remaining useful life prediction method of claim 3, wherein, The construction method of the bidirectional recurrent neural network model comprises: setting a gated recurrent unit as a basic unit, and receiving an isolator health state index sequence arranged in time sequence; embedding an attention weight module in an input layer to calculate contribution weight of each time step feature; capturing forward and reverse time sequence dependency through a bidirectional hidden layer; adding a residual connection path between hidden layers to fuse shallow and deep features; and connecting a full connection network in an output layer to map a remaining useful life prediction value. 5.The deep learning-based isolator remaining useful life prediction method of claim 4, wherein, The method for dynamically correcting comprises: appending a newly collected sensor signal to an end of the multi-dimensional sensor sequence to generate a new sub-sequence unit; extracting a fusion feature vector of the new sub-sequence unit and calculating an isolator health state index thereof; inputting the isolator health state index into the trained bidirectional recurrent neural network model to replace an input feature of an earliest time step; re-executing a time sequence calculation process to update a remaining useful life prediction value. 6.The deep learning-based isolator remaining useful life prediction method of claim 5, wherein, The model training step of the bidirectional recurrent neural network model comprises: collecting historical isolator full-life-cycle sensor data and labeling a real label of a remaining useful life; dividing the sensor data into a training set and a verification set; initializing parameters of the bidirectional recurrent neural network model; minimizing mean square error of a prediction value and a real label by using a stochastic gradient descent method; and stopping training when a loss function of the verification set converges. 7.The deep learning-based isolator remaining useful life prediction method of claim 4, wherein, The operation of the feature compression layer comprises: calculating a variance contribution rate of spliced features; screening feature dimensions with a variance contribution rate greater than a set threshold; mapping the screened features to a low-dimensional space through a linear projection matrix; and outputting a reduced fusion feature vector; The implementation mode of the residual connection path comprises: adding output features of a previous hidden layer to input features of a next hidden layer; adjusting a scale of the added features through a learnable weight matrix; inputting the adjusted features into a gated recurrent unit; and transmitting a final residual output to a next time sequence calculation node.
Citation Information
Patent Citations
Platform scale data analysis method based on Internet of Things technology
CN119670026A
In-orbit satellite remaining service life prediction method based on deep learning
CN119917798A