A method for predicting the remaining life of industrial equipment by integrating pre-trained large language model
By integrating pre-trained large language models to extract and fine-tune the model of industrial equipment timing data, the problems of data complexity and noise in the remaining life prediction of industrial equipment are solved, and higher prediction accuracy and stability are achieved.
Patent Information
- Application Number
- CN202411813374.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-11
AI Technical Summary
The remaining life prediction of industrial equipment faces problems such as unstructured, complex and diverse data, and noise-free data. Traditional machine learning models have limitations on generalization capabilities and prediction effects when processing such data.
Using the method of fusion pre-training large language models, the remaining life prediction of the equipment is generated by preprocessing and input embedding of the timing data of industrial equipment, and feature extraction and model fine-tuning is used for Transformer encoder blocks and additional network blocks.
It improves the prediction accuracy of industrial equipment health status and residual life information, enhances the prediction accuracy and stability of the model, and provides higher quality decision support for equipment health management.
Smart Images

Figure CN119272641B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning, large language models, time series data feature extraction, etc., and specifically relates to a method for predicting the remaining life of industrial equipment by integrating a pre-trained large language model. Background Art
[0002] The remaining life prediction of industrial equipment is of great significance in intelligent operation and maintenance and equipment management. However, the equipment operation data collected in actual industrial scenarios are often unstructured, complex, diverse, and noisy. Especially during long-term operation, the data will be affected by a variety of external factors, resulting in non-stationarity and strong nonlinearity, which brings great challenges to the modeling of equipment degradation process and the prediction of remaining life. In addition, the fault data samples are often small and unevenly distributed, which further limits the generalization ability and prediction effect of traditional machine learning models.
[0003] With the rapid development of pre-trained large language models, these models have shown great potential in processing multi-dimensional data feature extraction, time series modeling, and cross-domain knowledge transfer. Pre-trained large language models have the ability to extract deep features from large-scale data, and can transfer knowledge in specific fields through fine-tuning, thereby capturing key features in the equipment degradation process with limited samples. In addition, the attention mechanism in the pre-trained large language model can effectively model long dependencies in equipment time series data, help identify subtle changes in different states of the equipment, and provide more accurate information support for remaining life prediction. Traditional remaining life prediction methods are usually based on shallow feature extraction or manually constructed features, which have obvious limitations when dealing with the high-dimensional characteristics and heterogeneity of complex industrial data. Summary of the invention
[0004] In order to solve the shortcomings of the prior art and achieve the purpose of improving the accuracy of prediction of the health status and remaining life information of industrial equipment, the present invention adopts the following technical solutions:
[0005] A method for predicting the remaining life of industrial equipment by integrating a pre-trained large language model comprises the following steps:
[0006] Step 1: Collect the time series data of industrial equipment and perform preprocessing and input embedding in sequence; preprocess the time series data into segmented degraded sequence data, and then align the degraded segmented sequence data with the feature space of the pre-trained large language model through input embedding;
[0007] Step 2: The embedded data is input into the pre-trained large language model. The model contains a set of Transformer encoder blocks. Training a large language model from scratch usually leads to performance degradation. In order to maintain the large language model's data-independent representation learning ability, most parameters are frozen, especially the parameters of the multi-head attention layer and the feedforward fully connected layer in the Transformer encoder block, which play a key role in representation learning. The query, key, and value matrices are processed and most of their original parameters are frozen. Only a small number of trainable low-rank matrices are fine-tuned, thereby effectively optimizing parameter efficiency, reducing training burden and improving sensitivity to local features. After the multi-head attention layer, the data is subjected to layer normalization and feedforward fully connected layers. By fine-tuning the affine transformation parameters in the layer normalization, the amount of calculation is further reduced and the generalization ability of the model is enhanced, so that the large language model can process time series data and finally predict segmented time series data.
[0008] Step 3: Input the parameters of the pre-trained large language model into the additional network. The additional network contains a set of parallel and independent additional network blocks. Each additional network block works in parallel with the Transformer encoder block in the pre-trained large language model. After being processed by the additional network, the data enters the output layer to generate the remaining predicted life of the equipment.
[0009] Furthermore, in step 1, the preprocessing process of the time series data includes the following steps:
[0010] Step 1.1.1: Perform pre-fault diagnosis on the collected industrial equipment time series data to obtain the pre-fault time t. Since most industrial equipment failures occur randomly, the failure characteristics are weak and difficult to detect. The signal-to-noise ratio of the time series signal at this stage is low, and it contains less degradation information and more interference noise. Since the root mean square will increase with the severity of component degradation, the root mean square detection method can eliminate interference noise to the greatest extent and retain degradation information;
[0011] Step 1.1.2: Divide the industrial equipment time series data into several time series according to the t value before the fault;
[0012] Step 1.1.3: Pre-extract features of the obtained time series to obtain a time-frequency feature map;
[0013] Step 1.1.4: Connect all time-frequency feature maps to obtain a two-dimensional feature map of the degraded data;
[0014] Step 1.1.5: In order to improve the consistency and reliability of processing different time series data sets, it is necessary to perform instance normalization on the two-dimensional feature map of the degraded data to generate normalized time series data with a mean of zero;
[0015] Step 1.1.6: Since learning local aggregate information is more helpful for large models to understand the long-term trend of industrial equipment degradation, the two-dimensional feature map of the normalized degradation data needs to be segmented to capture higher-level degradation information that cannot be learned at a single time point, thereby enhancing locality.
[0016] Furthermore, in step 1.1.1, the formula for diagnosis before the fault is:
[0017]
[0018] in, represents the input signal, t represents the given time, Indicates input signal The mean value at time t-1 is The range of variance, represents the average value of the root mean square RMS, represents the variance, Represents the index used to traverse all time points before time t;
[0019] If the RMS exceeds The threshold range is exceeded when three consecutive RMS values exceed :
[0020]
[0021] Then time t is considered to correspond to the initial fault diagnosis time value, otherwise continue to traverse time t until the conditions are met and find the pre-fault time in the time series data.
[0022] Furthermore, in step 1.1.3, the feature is pre-extracted by short-time Fourier transform, and the formula of short-time Fourier transform is:
[0023]
[0024] in, represents the signal corresponding to the window function position, i represents the imaginary unit, e represents the base of the natural logarithm, represents the window function, f represents the frequency, represents the position of the window function, Represents the time-frequency characteristics of the input signal, and then uses the energy distribution of the original signal As the input of the subsequent spectrum diagram, the formula for energy distribution is:
[0025]
[0026] After the segmented data is subjected to short-time Fourier transform and energy distribution extraction, each subsequence is converted into a time-frequency feature map of size m×n, where m is the dimension of the frequency axis and n is the dimension of the time axis.
[0027] Furthermore, in step 1.1.6, the two-dimensional feature map of the normalized degradation data is expressed as:
[0028]
[0029] Among them, F represents the two-dimensional feature map of the normalized degraded data, Represents the data from 1 to L in the first row, D represents the number of rows in the graph, and L represents the number of columns in the graph. Represents a matrix containing L time steps and D features in each time step; then each column of the feature map is divided into segmented data of length P and interval step length S. After segmentation, the number of segmented data N in each column is:
[0030]
[0031] For time series whose length cannot be evenly divided into equal lengths, in each original series Use padding at the end of The last value of is repeated S times and added to the end of the original sequence so that the length of each segment data is P, preventing the last segment data from being discarded due to insufficient length, thereby losing key degraded data.
[0032] Furthermore, in step 1, the input embedding process includes the following steps:
[0033] Step 1.2.1: Through the token embedding coding, the convolutional layer is used to map the preprocessed segmented degradation sequence data to the latent space of the hidden layer so that it can be processed by the large model;
[0034] Step 1.2.2: Encode the absolute position information of the segmented degraded sequence data through position embedding encoding; since the multi-head attention mechanism in the large model does not have the ability to actively encode the absolute position relationship of the markers, it needs to be added manually and explicitly by using a simple and effective sine-cosine position encoding to reduce the burden of unnecessary parameter learning;
[0035] Step 1.2.3: Through time embedding coding, the time information of the segmented degradation series data is encoded in relative position to enhance the model's ability to capture relative time information in time series data. The attention mechanism of relative time position is implemented by rotating the encoding matrix to encode relative time information in the embedded features, thereby better extracting time series features in the degradation process and facilitating long-term prediction;
[0036] Step 1.2.4: The result of token embedding encoding is added to the results of position embedding encoding and time embedding encoding to align the degraded segmented sequence data with the large language model feature space.
[0037] Furthermore, the position embedding encoding formula in step 1.2.2 is as follows:
[0038]
[0039] Among them, n is the absolute position of each mark in the segment sequence, j is used to distinguish odd and even positions, and d is the hidden layer dimension of the large model;
[0040] The relative time embedding coding formula in step 1.2.3 is as follows:
[0041]
[0042] in, Indicates The relative time attention vector at time steps, Softmax represents the activation function used to calculate the attention weight, represents the query matrix, defined as , Indicates Segmented time series data of locations, represents the weight of the query matrix, Represents the rotation matrix, which is used to introduce relative time position information. represents the key-value matrix, defined as , represents the weight of the key-value matrix, Represents the hidden layer dimension of the large model, used for scaling factor, represents the value matrix, defined as , Represents the weights of the value matrix.
[0043] Furthermore, in step 2, the input embedding is adjusted to the required embedding dimension The embedding vector , input into the pre-trained large language model, and each containing After processing by the Transformer encoder block of the layer, the embedding vector is obtained ;
[0044] The embedding vector z formula is as follows:
[0045]
[0046] in, represents the processed embedding result, Indicates the length of the time series data, represents the adjusted embedding dimension, TFs(·) represents the processing operation of the pre-trained Transformer block;
[0047] Embedding vector Through the linear output layer , map the embedding back to the segmented time series data to get the predicted target ;
[0048] Target The formula is as follows:
[0049]
[0050] in, Represents the predicted segmented time series data, Represents a linear mapping matrix;
[0051] To ensure that the prediction accurately reconstructs the actual offset segment data target value , the predicted segmented time series data and the target value of the segmented time series data use the mean square error as the loss function for time series alignment:
[0052]
[0053] Among them, LOSS represents the loss function of time series alignment, and MSE represents the mean square error, which is used to measure the difference between the predicted value and the target value.
[0054] Furthermore, the step three comprises the following steps:
[0055] Step 3.1: During fine-tuning, only a small number of trainable parameters in the additional network block are adjusted to reduce the need for labeled data samples and avoid the problem of increasing the depth of the model. For each additional network block, it is fused with the large language model network through addition operations. The specific calculation formula is as follows:
[0056]
[0057] in, Represents the first The output of the additional network block, and Represent the mapping operations of the Transformer block and the additional network block, respectively. and Indicates The input of the layer Transformer block and the additional network block is adjusted by adjusting the parameters , weighted fusion is performed between the output of the Transformer block and the additional network block to achieve the purpose of flexibly controlling the weights of the two;
[0058] Step 3.2: The data after the additional network is input to the output layer. The output layer is a fully connected layer that maps the data to a specific number of output nodes. The size of the output layer depends on the selected prediction target. After passing through the output layer, the remaining life of the industrial equipment is finally output.
[0059] Furthermore, in the step three, by adding upper and lower projection layers, the expressive power of the pre-trained large language model for time series features is maximized; in each additional network block, a lower projection layer and an upper projection layer are included, and the lower projection layer includes a mapping layer, a batch normalization layer and a nonlinear ReLU layer, which projects the input features into a low-dimensional space to reduce the computational complexity; then, the upper projection layer reprojects the reduced-dimensional features back to the original dimension for subsequent fusion with the large language model; the additional network shares the same input as the large language model, that is, receives the pre-trained weight information, and after entering the additional network block, the features are first processed by the lower projection layer, then restored to the original dimension by the upper projection layer, and then output to the next additional network block or Transformer block to form layer-by-layer accumulated features; this structure places the additional network block in parallel with the large language model block, so that the two can achieve functional expansion while maintaining the overall depth of the model, thereby eliminating the need to increase the additional model depth. In addition, due to the parallel structure of the additional network, only a small number of parameters need to be fine-tuned during its training, which will not significantly increase the computational overhead.
[0060] The advantages and beneficial effects of the present invention are:
[0061] The present invention provides a method for predicting the remaining life of industrial equipment that integrates a pre-trained large language model, and aims to utilize technologies such as deep learning and large language models to enhance the prediction accuracy and stability of the model. Through the knowledge transfer and feature extraction capabilities of the natural language processing model, the model's understanding of the degradation process of industrial equipment and its prediction accuracy are further improved, providing higher-quality decision support for equipment health management, thereby more accurately obtaining the health status and remaining life information of industrial equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a flow chart of a method in an embodiment of the present invention.
[0063] Figure 2 It is a flowchart executed by the time series data preprocessing module in an embodiment of the present invention.
[0064] Figure 3 It is a structural diagram of the pre-trained large model parameter fine-tuning module in an embodiment of the present invention.
[0065] Figure 4 It is a structural diagram of the pre-trained large model additional network coding module in an embodiment of the present invention. DETAILED DESCRIPTION
[0066] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0067] like Figure 1 As shown, a method for predicting the remaining life of industrial equipment integrating a pre-trained large language model includes the following steps:
[0068] Step 1: Collect the time series data of industrial equipment and input it into the time series data preprocessing module and input embedding module to convert the original time series data into features suitable for large model input; the execution process of the time series data preprocessing module is as follows: Figure 2 As shown, the following steps are included:
[0069] Step 1.1.1: Perform pre-fault diagnosis on the time series data of industrial equipment. Since most industrial equipment failures occur randomly, the failure characteristics are weak and difficult to detect. The signal-to-noise ratio of the time series signal at this stage is low, and it contains less degradation information and more interference noise. Since the root mean square will increase with the severity of component degradation, the root mean square detection method can eliminate interference noise to the greatest extent and retain degradation information. The formula for pre-fault diagnosis is:
[0070]
[0071] in, represents the average value of the root mean square RMS, represents variance, t represents time, is the index used to traverse all time points before time t, represents the input signal, Indicates input signal The mean value at time t-1 is If the RMS exceeds The threshold range is:
[0072]
[0073] Indicates that three consecutive RMS values have exceeded , then time t is considered to correspond to the initial fault diagnosis time value, otherwise continue to traverse t until the conditions are met and the pre-fault time in the time series data is found.
[0074] Step 1.1.2: According to the pre-fault time t value obtained in step 1.1.1, the industrial equipment time series data is divided into several time series.
[0075] Step 1.1.3: Pre-extract features of the obtained time series through short-time Fourier transform, where the formula of short-time Fourier transform is:
[0076]
[0077] in, Indicates signal, is the window function, f is the frequency, represents the position of the window function, is the time-frequency characteristic of the input signal. Then the energy distribution of the original signal is used As the input of the subsequent spectrum diagram, the formula for energy distribution is:
[0078]
[0079] After the segmented data is subjected to short-time Fourier transform and energy distribution extraction, each subsequence is converted into a time-frequency feature map of size m×n, where m is the dimension of the frequency axis and n is the dimension of the time axis.
[0080] Step 1.1.4: Connect all the time-frequency feature maps to obtain a two-dimensional feature map of the degraded data, where D represents the time dimension of the degradation feature.
[0081] Step 1.1.5: In order to improve the consistency and reliability of processing different time series data sets, it is necessary to perform instance normalization on the two-dimensional feature map of the degraded data to generate normalized time series data with a mean of zero.
[0082] Step 1.1.6: Since learning local aggregate information is more helpful for large models to understand the long-term trend of industrial equipment degradation, it is necessary to segment the two-dimensional feature map of the normalized degradation data to capture higher-level degradation information that cannot be learned at a single time point, thereby enhancing locality. The two-dimensional feature map of the normalized degradation data can be expressed as:
[0083]
[0084] Among them, F represents the two-dimensional feature map of the normalized degraded data, Represents the data from 1 to L in the first row, D represents the number of rows in the graph, and L represents the number of columns in the graph. Represents a matrix containing L time steps and D features in each time step. Then each column of the feature map is divided into segmented data with a length of P and an interval step of S. After segmentation, the number of segmented data N in each column is:
[0085]
[0086] For time series whose length cannot be evenly divided into equal lengths, in each original series Use padding at the end of . The last value of is repeated S times and added to the end of the original sequence, thereby ensuring that the length of each segment data is P, preventing the last segment data from being discarded due to insufficient length, thereby losing key degraded data.
[0087] The input embedding module includes a tag embedding encoder, a position embedding encoder, and a time embedding encoder. The tag embedding encoder fully extracts local features and global degradation features, and then adds the results of the position encoder and the time encoder to align the degraded segmented sequence data with the large language model feature space. The execution process of the input embedding module includes the following steps:
[0088] Step 1.2.1: The token embedding encoder maps the segmented degraded sequence data to a latent space of hidden layer dimension d through the convolutional layer so that it can be processed by a large model;
[0089] Step 1.2.2: The role of the position embedding encoder is to encode the absolute position information. Since the multi-head attention mechanism in the large model does not have the ability to actively encode the absolute position relationship of the tag, it needs to be added manually and explicitly. By using a simple and effective sine-cosine position encoding, the burden of unnecessary parameter learning can be reduced. The position embedding formula is as follows:
[0090]
[0091] Among them, n is the absolute position of each mark in the segmented sequence, j is used to distinguish odd and even positions, and d is the hidden layer dimension of the large model.
[0092] Step 1.2.3: The role of the time embedding encoder is to encode the relative position of the time information to enhance the model's ability to capture the relative time information in the time series data. The attention mechanism of the relative time position is implemented by rotating the encoding matrix to encode the relative time information in the embedded features, so as to better extract the time series features in the degradation process and facilitate long-term prediction. The relative time embedding formula is as follows:
[0093]
[0094] in, Indicates Relative temporal attention vector at time steps, Softmax: activation function used to calculate attention weights. is the query matrix, defined as ,in Indicates Segmented time series data of locations, is the weight of the query matrix, is a rotation matrix used to introduce relative time position information. is a key-value matrix, defined as ,in is the weight of the key-value matrix. Indicates the hidden layer dimension of the large model, used for scaling factor. is a value matrix, defined as in is the weight of the value matrix.
[0095] Step 1.2.4: The result of the token position encoder is added to the results of the position encoder and the temporal encoder to align the degenerate segmented sequence data with the large language model feature space.
[0096] Step 2: The pre-processed data is input into the pre-trained large model parameter fine-tuning module. During the fine-tuning process, the time series data is aligned by freezing some parameters and fine-tuning other parameters, so that the large language model can process the time series data; Figure 3 As shown in the figure, the execution process of the pre-trained large model parameter fine-tuning module is as follows:
[0097] Training a large language model from scratch usually leads to performance degradation. In order to maintain the large language model's ability to learn representations independently of data, most parameters are frozen, especially the parameters of the multi-head attention layer and feed-forward fully connected layer in the Transformer block, which play a key role in representation learning.
[0098] In the pre-trained large language model, in order to adjust only a small number of trainable parameters, a small number of trainable parameters are selectively adjusted or introduced.
[0099] First, the embedding vector processed by the embedding module enters the multi-head attention layer of the Transformer block, where the query, key, and value matrices are processed, most of their original parameters are frozen, and only a small number of trainable low-rank matrices are fine-tuned, thereby effectively optimizing parameter efficiency, reducing the training burden, and improving sensitivity to local features. After the multi-head attention layer, the data is passed through layer normalization and a feed-forward fully connected layer. By fine-tuning the affine transformation parameters in the layer normalization, the amount of computation is further reduced and the generalization ability of the model is enhanced. The specific processing flow is as follows:
[0100] Given an embedding vector adjusted to the required embedding dimension D after three encoding layers , which is fed into a pre-trained large language model. The model consists of a series of pre-trained Transformer blocks, each of which contains Layer, after being processed by the Transformer block, the final embedding vector is obtained , the formula is as follows:
[0101]
[0102] in, represents the processed embedding result, Indicates the length of the time series data, is the adjusted embedding dimension, and TFs(·) represents the processing operation of the pre-trained Transformer block. Embedding vector After pre-training the large language model, a linear output layer is used Map the embedding back to the segmented time series data to get the predicted target The formula is as follows:
[0103]
[0104] in, Represents the predicted segmented time series data, Represents the linear mapping matrix, in order to ensure that the prediction accurately reconstructs the actual offset segment data , using mean square error as the loss function:
[0105]
[0106] Among them, LOSS represents the loss function of time series alignment, and MSE represents the mean square error, which is used to measure the difference between the predicted value and the target value.
[0107] Step 3: Input the weights of the pre-trained large model into the pre-trained large model additional network coding module. This module further improves the model's ability to express time series features by adding upper and lower projection layers. After being processed by the additional coding module, the data enters the output layer to generate the remaining predicted life of the device; Figure 4 As shown in FIG. 1 , the execution process of the pre-trained large model additional network coding module includes the following steps:
[0108] Step 3.1: The additional network is a lightweight parallel module, which is structured as a series of parallel independent modules, each of which is called an additional network block. Each additional network block works in parallel with the Transformer block in the large language model, aiming to maximize the feature representation capabilities of the pre-trained large language model. During fine-tuning, only a small number of trainable parameters in the additional network block are adjusted, thereby reducing the need for labeled data samples and avoiding the problem of increasing the depth of the model.
[0109] For each additional network block, it is fused with the large language model network through an addition operation. The specific calculation formula is as follows:
[0110]
[0111] in, Represents the first The output of the additional network block, and Represent the mapping operations of the Transformer block and the additional network block, respectively. and Indicates The input of the layer Transformer block and the additional network block. By adjusting the parameters , weighted fusion can be performed between the output of the Transformer block and the additional network block to achieve the purpose of flexibly controlling the weights of the two. In each additional network block, there is a down-projection layer and an up-projection layer. The down-projection layer consists of a mapping layer, a batch normalization layer, and a nonlinear ReLU layer, which projects the input features into a low-dimensional space to reduce the computational complexity; then, the up-projection layer reprojects the reduced-dimensional features back to the original dimension for subsequent fusion with the large language model. The additional network shares the same input as the large language model, that is, it receives pre-trained weight information. After entering the additional network block, the features are first processed by the down-projection layer, then restored to the original dimension by the up-projection layer, and then output to the next additional network block or Transformer block to form layer-by-layer accumulated features. This structure places the additional network block in parallel with the large language model block, so that the two can achieve functional expansion while maintaining the overall depth of the model, without increasing the additional model depth. In addition, due to the parallel structure of the additional network, only a small number of parameters need to be fine-tuned during its training, which does not significantly increase the computational overhead.
[0112] Step 3.2: The data after the additional network module is input to the output layer. The output layer is a fully connected layer that maps the data to a specific number of output nodes. The size of the output layer depends on the selected prediction target, such as predicting the RUL of a single sample or the average RUL of a group of samples. After passing through the output layer, the remaining life of the industrial equipment is finally output.
[0113] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting the remaining life of industrial equipment by integrating a pre-trained large language model, characterized in that The steps include: Step 1: Collect the time series data of industrial equipment and perform preprocessing and input embedding in sequence; preprocess the time series data into segmented degraded sequence data, and then align the degraded segmented sequence data with the feature space of the pre-trained large language model through input embedding; Step 2: The embedded data is input into a pre-trained large language model. The model contains a set of encoder blocks. Most of the original parameters of the multi-head attention layer and the feedforward fully connected layer in the encoder block are frozen, and only a small number of trainable low-rank matrices are fine-tuned. After the multi-head attention layer, the data passes through layer normalization and a feedforward fully connected layer, and the affine transformation parameters in the layer normalization are fine-tuned to finally predict the time series data. Step 3: Input the parameters of the pre-trained large language model into the additional network, where the additional network includes a set of parallel and independent additional network blocks, each of which works in parallel with the encoder block in the pre-trained large language model to generate the remaining predicted life of the device; By adding upper and lower projection layers, the expressive power of the pre-trained large language model for temporal features is maximized; in each additional network block, there is a lower projection layer and an upper projection layer, the lower projection layer includes a mapping layer, a batch normalization layer and a nonlinear ReLU layer, which projects the input features into a low-dimensional space; then, the upper projection layer reprojects the reduced-dimensional features back to the original dimensions; the additional network shares the same input as the large language model, that is, it receives the pre-trained weight information. After entering the additional network block, the features are first processed by the lower projection layer, then restored to the original dimension by the upper projection layer, and then output to the next additional network block or encoder block to form layer-by-layer accumulated features.
2. According to claim 1, a method for predicting the remaining life of industrial equipment integrating a pre-trained large language model is characterized by: In step 1, the preprocessing process of time series data includes the following steps: Step 1.1.1: diagnose the pre-fault time of the collected industrial equipment time series data to obtain the pre-fault time; Step 1.1.2: Divide the industrial equipment time series data into several time series according to the time before the fault; Step 1.1.3: Pre-extract features of the obtained time series to obtain a time-frequency feature map; Step 1.1.4: Connect all time-frequency feature maps to obtain a two-dimensional feature map of the degraded data; Step 1.1.5: Perform instance normalization on the two-dimensional feature map of the degraded data to generate normalized time series data with a mean of zero; Step 1.1.6: Segment the normalized 2D feature map of the degraded data to capture higher-level degradation information that cannot be learned at a single time point.
3. The method for predicting the remaining life of industrial equipment integrating a pre-trained large language model according to claim 2 is characterized in that: In step 1.1.1, the formula for diagnosis before the fault is: Where x represents the input signal, t represents the given time, and y t Indicates that the mean of the input signal x at time t-1 is within 3σ t-1 The range of variance, represents the average value of the root mean square RMS, represents the variance, i represents the index used to traverse all time points before time t; If the RMS exceeds the 3σ threshold, that is, three consecutive RMS values exceed y t : {RMS t+1 ,RMS t+2 >y t ∣RMS t >y t } Then time t is considered to correspond to the initial fault diagnosis time value, otherwise continue to traverse time t until the conditions are met and find the pre-fault time in the time series data.
4. The method for predicting the remaining life of industrial equipment integrating a pre-trained large language model according to claim 2 is characterized in that: In step 1.1.3, the feature is pre-extracted by short-time Fourier transform, and the formula of short-time Fourier transform is: Among them, x(τ) represents the signal corresponding to the position of the window function, i represents the imaginary unit, e represents the base of the natural logarithm, w(τ-t) represents the window function, f represents the frequency, τ represents the position of the window function, X(t,f) represents the time-frequency characteristics of the input signal, and then the energy distribution E(t,f) of the original signal is used as the input of the subsequent spectrum diagram, where the formula for energy distribution is: E(t,f)=∣X(t,f)∣ 2 After the segmented data is subjected to short-time Fourier transform and energy distribution extraction, each subsequence is converted into a time-frequency feature map of size m×n, where m is the dimension of the frequency axis and n is the dimension of the time axis.
5. The method for predicting the remaining life of industrial equipment integrating a pre-trained large language model according to claim 2 is characterized in that: In step 1.1.6, the two-dimensional feature map of the normalized degradation data is expressed as: Among them, F represents the two-dimensional feature map of the normalized degraded data, Represents the data from 1 to L in the first row, D represents the number of rows in the graph, and L represents the number of columns in the graph. Represents a matrix containing L time steps and D features in each time step; then each column of the feature map is divided into segmented data of length P and interval step length S. After segmentation, the number of segmented data N in each column is: For time series whose length cannot be evenly divided into equal lengths, in each original series Use padding at the end of The last value of is repeated S times and added to the end of the original sequence so that the length of each segment data is P.
6. The method for predicting the remaining life of industrial equipment integrating a pre-trained large language model according to claim 1 is characterized in that: In step 1, the input embedding process includes the following steps: Step 1.2.1: Map the preprocessed segmented degradation sequence data to the latent space of the hidden layer through token embedding coding; Step 1.2.2: Encode the absolute position information of the segmented degraded sequence data through position embedding coding; Step 1.2.3: Through time embedding coding, the time information of the segmented degraded sequence data is encoded in relative position, and the attention mechanism of the relative time position is implemented by rotating the encoding matrix to encode the relative time information in the embedded features; Step 1.2.4: The result of token embedding encoding is added to the results of position embedding encoding and time embedding encoding to align the degraded segmented sequence data with the large language model feature space.
7. The method for predicting the remaining life of industrial equipment integrating a pre-trained large language model according to claim 6 is characterized in that: The position embedding encoding formula in step 1.2.2 is as follows: Among them, n is the absolute position of each mark in the segment sequence, j is used to distinguish odd and even positions, and d is the hidden layer dimension of the large model; The relative time embedding coding formula in step 1.2.3 is as follows: in, represents the relative temporal attention vector at the i-th time step, Softmax represents the activation function used to calculate the attention weight, and Q represents the query matrix, which is defined as represents the segmented time series data at the i-th position, represents the weight of the query matrix, R represents the rotation matrix, which is used to introduce relative time position information, and K represents the key value matrix, which is defined as represents the weight of the key-value matrix, d represents the hidden layer dimension of the large model, which is used for scaling factors, and V represents the value matrix, which is defined as Represents the weights of the value matrix.
8. The method for predicting the remaining life of industrial equipment integrating a pre-trained large language model according to claim 1 is characterized in that: In the step 2, the embedding vector adjusted to the required embedding dimension by input embedding is input into the pre-trained large language model, and the embedding vector is obtained after being processed by each Transformer encoder block; The embedding vector is passed through a linear output layer to map the embedding back to the segmented time series data to obtain the predicted target; The predicted segmented time series data and the target value of the segmented time series data use the mean square error as the loss function for time series alignment.
9. The method for predicting the remaining life of industrial equipment integrating a pre-trained large language model according to claim 1, characterized in that: The step three comprises the following steps: Step 3.1: During fine-tuning, only a small number of trainable parameters in the additional network block are adjusted to reduce the need for labeled data samples. For each additional network block, it is fused with the large language model network through an addition operation. The specific calculation formula is as follows: Among them, H n represents the output of the nth additional network block after fusion with the Transformer block, Tk n and Pk n Denote the mapping operations of the Transformer block and the additional network block, respectively, and X T,n-1 and H P,n-1 Represents the input of the nth layer Transformer block and the additional network block. By adjusting the parameter α, a weighted fusion is performed between the output of the Transformer block and the additional network block. Step 3.2: The data after the additional network is input to the output layer. The output layer is a fully connected layer that maps the data to a specific number of output nodes. The size of the output layer depends on the selected prediction target. After passing through the output layer, the remaining life of the industrial equipment is finally output.
Citation Information
Patent Citations
Health index curve extraction and life prediction method for mechanical equipment
CN113434970A
Method, system and equipment for predicting residual life of aero-engine and medium
CN115952724A