Straddle type monorail train gearbox vibration prediction method based on large language model
Through the vibration prediction method of the gearbox of a cross-seat monorail train based on a large language model, the complexity and error problems of vibration signal prediction in the prior art are solved, and higher prediction accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510094420.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively capture the complex timing characteristics of the vibration signal of the gearbox of a monorail train, resulting in insufficient signal feature extraction, large volatility in the prediction results, significant trend prediction errors, and fewer fault data samples, which limits the training depth and breadth of the model, and thus affects the model's sensitivity to abnormal signals.
The vibration prediction method of the gearbox of a cross-seat monorail train based on a large language model is adopted. By predicting the trend of the vibration signal and constructing the corresponding prompt word template, the mark embedding vector is generated; the vibration signal data is embedded in text to obtain the text embedding vector; the mark embedding vector and the text embedding vector are spliced, and the vibration signal is predicted in the pre-trained large language model to improve the accuracy and reliability of the prediction.
Through the prediction method of the large language model, the complex timing characteristics of the vibration signal of the monorail gearbox can be captured more accurately, which improves the accuracy and reliability of prediction, reduces trend prediction errors, and enhances the model's sensitivity on abnormal signals.
Smart Images

Figure CN120012747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet big data and vibration signal prediction, and in particular to a method for predicting vibration of a straddle-type monorail train gearbox based on a large language model. Background Art
[0002] As an efficient and environmentally friendly urban rail transit tool, monorail has been widely used in many cities around the world due to its advantages such as small footprint, strong climbing ability and small turning radius.
[0003] The smooth operation of monorail trains is highly dependent on the normal operation of their key components, especially the gearbox, which is responsible for transmitting the torque of the motor to the wheels to provide power for the train. If the gearbox fails, such as gear wear or bearing damage, it will lead to increased vibration and noise, and have a negative impact on the power transmission efficiency of the vehicle. In addition, the vibration problem of the gearbox will also reduce the smoothness of the train operation and affect the comfort of passengers. Especially when running at high speeds, the increase in vibration and noise may aggravate passenger discomfort and even threaten the safety of the train, leading to more serious mechanical failures. Therefore, timely maintenance and detection of the gearbox status are crucial to ensure the normal operation of the monorail train. Vibration signal prediction (i.e. predicting the vibration signal at a certain moment in the future) plays a key role in the intelligent operation and maintenance of mechanical equipment. Through time series data analysis, signal change trends and early signs of failure can be captured in advance, providing data support for abnormal detection and fault diagnosis. The main goal is to identify potential failure risks, reveal failure trends and patterns, and provide important early warnings as a pre-stage of fault diagnosis.
[0004] The methods of vibration signal prediction cover multiple analysis dimensions: time domain analysis, such as root mean square value and variance extraction for monitoring signal changes; frequency domain analysis, such as Fourier transform to identify periodic components; time-frequency domain analysis, such as empirical mode decomposition to extract signal intrinsic mode functions; and machine learning methods, such as LSTM for long time series dependency modeling. Compared with the traditional fixed window method, large-scale language models (LLMs) dynamically adjust the dependencies between time steps through the self-attention mechanism, identify key time points in the data, and can effectively capture long-term and short-term dependencies, thereby better processing non-stationary signals. In addition, models such as LSTM, GRU, Wavenet, and Informe can simultaneously focus on the short-term fluctuations and long-term trends of the signal and capture complex spectral features. In the study of Time-LLM, the advantages of large language models in processing long time series data are demonstrated, especially in providing scalable solutions for prediction tasks with a small number of samples. Despite this, these studies have not been widely used in the field of vibration signal prediction, especially in the prediction of monorail gearbox vibration signals.
[0005] Compared with ordinary gearboxes, monorail gearboxes have special structural requirements. They need to adapt to installation requirements in narrow spaces and withstand more frequent starts and stops as well as load fluctuations under complex working conditions. When operating under complex working conditions, especially during frequent starts and stops, the vibration signal of the monorail gearbox exhibits typical nonlinear and non-stationary characteristics, accompanied by significant dynamic load and speed changes, presenting complex time-frequency characteristics. In order to address the challenges of mixed low-frequency and high-frequency components, unclear features, and signal non-stationarity in the vibration signal of the monorail gearbox, the advantages of LLM in multimodal data processing, automatic feature extraction, and time series data modeling can be utilized to introduce LLM technology into the field of vibration signal prediction.
[0006] The applicant found that the current prediction of monorail gearbox vibration signals faces the following challenges: 1) Existing methods are difficult to effectively capture the complex time series characteristics of monorail gearbox vibration signals, resulting in insufficient signal feature extraction; 2) The prediction results are highly volatile, the trend prediction error is significant, and the predicted value often lags behind the actual value; 3) There are few fault data samples, which limits the training depth and breadth of the model, thereby affecting the model's sensitivity to abnormal signals. Therefore, how to design a straddle-type monorail gearbox vibration prediction method based on a large language model is a technical problem that needs to be solved urgently. Summary of the invention
[0007] In view of the deficiencies of the above-mentioned prior art, the technical problem to be solved by the present invention is: how to provide a method for predicting the vibration of a gearbox of a straddle-type monorail train based on a large language model. First, a tag embedding vector is generated by predicting the trend of the vibration signal and constructing a corresponding prompt word template; then, the vibration signal data is subjected to text embedding mapping to obtain a text embedding vector; then, the tag embedding vector and the text embedding vector are concatenated and input into a pre-trained large language model to predict the vibration signal. This process aims to improve the accuracy and reliability of the prediction of the vibration signal of the gearbox of the straddle-type monorail train.
[0008] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0009] The gearbox vibration prediction method of straddle-type monorail train based on large language model includes:
[0010] S1: Acquire vibration signal data of the gearbox of a straddle-type monorail train;
[0011] S2: Input the vibration signal data into the trained trend prediction model to perform vibration signal trend prediction, and output the trend prediction result of the vibration signal;
[0012] S3: Construct corresponding prompt word templates based on the domain characteristics and prediction requirements of the prediction task;
[0013] S4: Input the trend prediction result of the vibration signal and the prompt word template into the Prompt module for processing to generate the corresponding tag embedding vector;
[0014] S5: Input the vibration signal data into the Input module for text embedding mapping to generate the corresponding text embedding vector;
[0015] S6: concatenate the tag embedding vector and the text embedding vector and input them into the pre-trained large language model, outputting the corresponding vibration signal fusion embedding representation; inputting the vibration signal fusion embedding representation into the prediction head for digitization, and generating the corresponding vibration signal prediction result;
[0016] S7: Outputting the vibration signal prediction result as the vibration signal prediction result of the straddle-type monorail train gearbox.
[0017] Preferably, in step S2, the processing steps of the trend prediction model include:
[0018] S201: Building a trend prediction model including a multi-layer network based on a deep learning algorithm and performing model training;
[0019] S202: applying time position coding to the vibration signal data, embedding the time information into the vibration signal to generate a vibration signal representation containing the time series position information;
[0020] The formula is:
[0021]
[0022] Where: X i,m represents input data; PE(·) represents temporal position encoding; MHSA represents multi-head self-attention; DCC represents dilated causal convolution operation; pos represents the position of the time step, i represents the dimension of the position encoding; d model =32 represents the hidden layer dimension of the trend prediction model;
[0023] S203: Execute the following steps in each network layer of the trend prediction model:
[0024] S2031: Input the vibration signal representation containing the time series position information into the dilated causal convolution layer and the multi-head self-attention layer to alternately extract local and global features, and generate the query and key of the attention head;
[0025] The formula is:
[0026]
[0027] Where: Q m ,K m ,V m,b m represents the query, key, value, and bias of m attention heads; N = 6 represents the number of network layers; W n Represents the weight of the nth layer network; W m represents the attention weight matrix of the mth head; MHSA represents multi-head self-attention; DCC represents the dilated causal convolution operation;
[0028] S2032: Inputting the vibration signal representation containing the time series position information into the gated recurrent unit to perform time series information modeling and generate the value of the attention head;
[0029] S2033: Input the query, key and value of the attention head into the multi-head cross attention layer for feature fusion, and output the fusion features of the network in this layer;
[0030] The formula is:
[0031]
[0032] Where: GRU stands for gated recurrent unit; h t-1 represents the hidden state of the previous time step of GRU; MHCA represents multi-head cross attention; x t represents the input characteristics of the vibration signal at time t;
[0033] S204: Aggregate the fusion features of each layer of the network as the trend prediction result of the vibration signal.
[0034] Preferably, in step S2, vibration signal trend prediction is a classification task, and the classification results include:
[0035] continued to upward: continued to upward;
[0036] continued to downward: continued to downward;
[0037] first upward then downward: first upward then downward;
[0038] first downward then upward: first downward then upward.
[0039] Preferably, in step S3, the prompt word template includes the following content:
[0040] 1) Context information: used to describe the application scenario and background of vibration signal data;
[0041] 2) Historical data description: Provide the observation range, frequency and total number of observation points of the vibration signal data time series;
[0042] 3) Statistical summary: summarize the mean, standard deviation, maximum and minimum values of the vibration signal data;
[0043] 4) Recent trends: Capture the latest changes and change rates of vibration signal data;
[0044] 5) Forecasting requirements: clarify the forecasting step and external influencing factors of the forecasting task;
[0045] 6) Special instructions: used to specify external factors to be considered in the forecasting process.
[0046] Preferably, in step S4, the calculation formula of the Prompt module is expressed as:
[0047]
[0048] Where: P E represents the tag embedding vector; Y DBC Indicates the trend prediction result of vibration signal; LLM E represents the embedder, which is used to generate the tag embedding vector; φ(·) is the nonlinear transformation feature processing operation function; G n Indicates a general prompt for input; W l ″ represents the weight coefficient of the trend prediction part; Δp n Represents the offset or compensation value; l represents the current iteration index; L represents the number of input features involved in feature fusion.
[0049] Preferably, in step S5, the processing steps of the Input module are as follows:
[0050] S501: After normalizing the vibration signal data, the normalized data is decomposed into a number of time segments of fixed lengths through a patching operation;
[0051] S502: Convert each time segment into a corresponding high-dimensional feature vector through a slice embedding module;
[0052] S503: Utilize W through linear mapping layer f ·z+b f Convert the high-dimensional feature vectors of each time segment into corresponding text embedding vectors.
[0053] Preferably, in step S501, the normalized formula is expressed as:
[0054]
[0055] Where: x represents the original vibration signal data; Represents normalized data.
[0056] Preferably, in step S502, the processing steps of the slice embedding module are as follows:
[0057] S5021: Map each time segment to a high-dimensional space through linear transformation to generate a vibration signal time series slice;
[0058] S5022: Map pre-trained word embedding vectors to text prototypes;
[0059] S5023: Generate corresponding embedding as high-dimensional feature vector for each vibration signal time series slice using multi-head cross attention through text prototype;
[0060] The formula is:
[0061]
[0062] Where: Q h , K h 、V h , W h They represent the query, key, value vector and attention weight matrix of the h=8th head of multi-head cross attention respectively; Q h,P , K h,P Provided by the generated text prototype; Provided by time slice;d k Indicates the dimension of the key vector.
[0063] Preferably, in step S6, the calculation formula of the vibration signal prediction result is expressed as:
[0064] FCP signal (T) = W·PH·[Flatten(LLM B (α·I E (T),β·P E (T)))+γ·T];
[0065] Where: FCP signal represents the vibration signal prediction result; PH(·) represents the prediction head; W represents the mapping layer weight; α and β are the proportional coefficients of the Input module and the Prompt module respectively; γ is the weight coefficient of the prediction time step T; LLM B Represents the main part of the pre-trained large language model, which is used to embed the text into the vector I E and the token embedding vector P E Combining and generating high-dimensional features, namely, vibration signal fusion embedding representation; Flatten(·) means expanding high-dimensional features into one-dimensional vectors.
[0066] Preferably, in step S6, the fusion embedding representation of the vibration signal output by the large language model includes: the predicted statistical variable values of the future vibration signal and the relevant description of the trend of the vibration signal.
[0067] Compared with the prior art, the gearbox vibration prediction method for straddle-type monorail train based on the large language model in the present invention has the following beneficial effects:
[0068] The present invention performs trend prediction on vibration signals through the DCBiformerNet trend prediction model, and constructs a prompt word template based on the domain characteristics and requirements of the prediction task. Subsequently, the trend prediction result of the vibration signal and the prompt word template are input into the Prompt module to generate a tag embedding vector. First, trend prediction helps to identify abnormal changes that may exist in vibration signals, thereby providing support for the prediction of vibration signals of straddle-type monorail train gearboxes. Secondly, by constructing prompt word templates closely related to the prediction task, it can help the large language model better understand the background and prediction requirements of the vibration signal data, thereby improving the accuracy of the prediction. The introduction of prompt word templates related to domain characteristics not only provides rich contextual information, but also clarifies task instructions, thereby helping the model to better adapt to vibration signal prediction in different scenarios.
[0069] The present invention obtains a text embedding vector by performing text embedding mapping on vibration signal data, that is, converting a vibration signal sequence into a text prototype representation suitable for language model processing. Through text embedding mapping, high-dimensional vibration signal data can be converted into a low-dimensional embedding vector, thereby simplifying the data processing process and significantly improving the calculation efficiency. At the same time, the text embedding vector can effectively capture the key feature information in the vibration signal and retain the main mode and change trend of the signal. After the vibration signal data is converted into a text embedding vector, it can be more conveniently input into a pre-trained large language model for processing. In this way, the powerful ability of the large language model in processing complex data and making predictions can be fully utilized, and the accuracy and reliability of vibration signal prediction can be further improved.
[0070] The present invention splices the tag embedding vector and the text embedding vector, inputs them into the pre-trained large language model, outputs the fused embedding representation of the vibration signal, and performs numerical processing. First, by splicing the tag embedding vector and the text embedding vector, multi-source information closely related to the prediction task can be fused to fully reflect the characteristics and trends of the vibration signal. Secondly, the fused embedding representation as the model input can make full use of the powerful representation ability of the large language model, thereby significantly improving the accuracy and reliability of the prediction of the gearbox vibration signal of the straddle-type monorail train. Finally, the prediction result of the vibration signal sequence is generated by the prediction head, which is not only convenient for subsequent analysis and processing, but also through the fused embedding representation, the source and basis of the model prediction results can be more intuitively analyzed, thereby enhancing the interpretability and credibility of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0072] Figure 1 Schematic diagram of the workflow of the gearbox vibration prediction method for straddle-type monorail train based on the large language model.
[0073] Figure 2 Model framework for the vibration signal prediction large language model (VSP-LLM).
[0074] Figure 3 Schematic diagram of the application of large language models in sequence data prediction.
[0075] Figure 4 Network architecture of the trend prediction (DCBiformerNet) model.
[0076] Figure 5 This is the algorithm flow chart of the trend prediction (DCBiformerNet) model.
[0077] Figure 6 Schematic diagram of input embedding.
[0078] Figure 7 The error distribution diagram of the five methods.
[0079] Figure 8 The prediction results of Autoformer, DLinear and VSP-LLM* models under the working condition of 78 and T = {48, 96, 192, 384, 768, 1536}.
[0080] Fig. 9 The following is a comparison chart of the convergence results of the five models under different step size conditions. DETAILED DESCRIPTION
[0081] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0082] The following is a further detailed description through specific implementation methods:
[0083] Example:
[0084] This embodiment discloses a method for predicting vibration of a straddle-type monorail train gearbox based on a large language model.
[0085] like Figure 1 As shown, the gearbox vibration prediction method for straddle-type monorail train based on a large language model includes:
[0086] S1: Acquire vibration signal data of the gearbox of a straddle-type monorail train;
[0087] S2: Input the vibration signal data into the trained trend prediction model to perform vibration signal trend prediction, and output the trend prediction result of the vibration signal;
[0088] S3: Construct corresponding prompt word templates based on the domain characteristics and prediction requirements of the prediction task;
[0089] S4: S4: Input the trend prediction result of the vibration signal together with the prompt word template into the Prompt module for processing to generate the corresponding tag embedding vector;
[0090] In the Prompt module: the trend prediction results of the vibration signal and the prompt word template are input into the Promp module, and after being processed by Tokenization and Token Embedder, the corresponding OutputToken Embeddings are generated. These embedding vectors are passed to the subsequent model as prompt information.
[0091] S5: Input the vibration signal data into the Input module for text embedding mapping to generate a corresponding text embedding vector;
[0092] In the Input module: The vibration signal data is input into the Input module, and instance normalization is first performed. Then the vibration signal is segmented into time series slices through Patching, and then the time series embedding is generated through Patch Embedder. Then the embedding vector interacts with Text Prototypes through Multi-Head Cross Attention, and finally the processing is completed through the Linear layer to generate the corresponding text embedding vector.
[0093] S6: Combine Figure 2 As shown, the tag embedding vector and the text embedding vector are concatenated and input into the pre-trained large language model, and the corresponding vibration signal fusion embedding representation is output; the vibration signal fusion embedding representation is input into the prediction head for digitization to generate the corresponding vibration signal prediction result; wherein the vibration signal fusion embedding representation output by the large language model includes: the predicted statistical variable value of the future vibration signal and the relevant description of the vibration signal trend.
[0094] In this embodiment, the output token embedding vector (Output Token Embeddings) output by the prompt module is concatenated with the text embedding vector (Vibration Signal Embeddings) output by the input module and then input into the pre-trained large language model body (Pre-trained LLM Body). After being processed by the large language model, a fused embedding representation (a vector) containing global information is output, and then further calculated through the output prediction layer (Output Prediction Layer), and finally the vibration signal fusion embedding representation is input into the prediction head for numerical and visual representation to generate the corresponding vibration signal prediction result.
[0095] S7: Outputting the vibration signal prediction result as the vibration signal prediction result of the straddle-type monorail train gearbox.
[0096] The present invention performs trend prediction on vibration signals through the DCBiformerNet trend prediction model, and constructs a prompt word template based on the domain characteristics and requirements of the prediction task. Subsequently, the trend prediction result of the vibration signal and the prompt word template are input into the Prompt module to generate a tag embedding vector. First, trend prediction helps to identify abnormal changes that may exist in vibration signals, thereby providing support for the prediction of vibration signals of straddle-type monorail train gearboxes. Secondly, by constructing prompt word templates closely related to the prediction task, it can help the large language model better understand the background and prediction requirements of the vibration signal data, thereby improving the accuracy of the prediction. The introduction of prompt word templates related to domain characteristics not only provides rich contextual information, but also clarifies task instructions, thereby helping the model to better adapt to vibration signal prediction in different scenarios.
[0097] The present invention obtains a text embedding vector by performing text embedding mapping on vibration signal data, that is, converting a vibration signal sequence into a text prototype representation suitable for language model processing. Through text embedding mapping, high-dimensional vibration signal data can be converted into a low-dimensional embedding vector, thereby simplifying the data processing process and significantly improving the calculation efficiency. At the same time, the text embedding vector can effectively capture the key feature information in the vibration signal and retain the main mode and change trend of the signal. After the vibration signal data is converted into a text embedding vector, it can be more conveniently input into a pre-trained large language model for processing. In this way, the powerful ability of the large language model in processing complex data and making predictions can be fully utilized, and the accuracy and reliability of vibration signal prediction can be further improved.
[0098] The present invention splices the tag embedding vector and the text embedding vector, inputs them into the pre-trained large language model, outputs the fused embedding representation of the vibration signal, and performs numerical processing. First, by splicing the tag embedding vector and the text embedding vector, multi-source information closely related to the prediction task can be fused to fully reflect the characteristics and trends of the vibration signal. Secondly, the fused embedding representation as the model input can make full use of the powerful representation ability of the large language model, thereby significantly improving the accuracy and reliability of the prediction of the gearbox vibration signal of the straddle-type monorail train. Finally, the prediction result of the vibration signal sequence is generated by the prediction head, which is not only convenient for subsequent analysis and processing, but also through the fused embedding representation, the source and basis of the model prediction results can be more intuitively analyzed, thereby enhancing the interpretability and credibility of the model.
[0099] Through comprehensive experiments on the DGLC dataset and comparison with models such as Autoformer, Informer, DLinear and TIME-LLM, it is shown that the method of the present invention performs significantly better than other models in the prediction effect of the DGLC dataset, and the generalization ability of the model is verified by application in the public dataset of Case Western Reserve University. The model of the present invention has the ability to be promoted and applied in the field of vibration signal prediction with similar nonlinear and non-stationary characteristics such as high-speed railway transmission systems and wind power generation equipment.
[0100] In order to better introduce the technical solution of the present invention, this embodiment is described through the following parts.
[0101] 1. Research on LLMs in vibration signal prediction
[0102] The gearbox of a monorail train is a power output and necessary deceleration device. Frequent starting and stopping, variable loads and the need to pass through curves make the vibration signal during its operation relatively complex, consisting of a mixture of low-frequency and high-frequency components, fuzzy features, and nonlinear and non-steady-state modes. To accurately predict these signals, a model capable of multimodal data processing, automatic feature extraction and time analysis is necessary. The applicant found that large-scale language models (LLMs) have the advantages of multimodal fusion and can integrate data from different sensors to improve prediction accuracy, making them very suitable for this task. LLM can effectively capture complex nonlinear relationships, which is also proved by the "Universal Approximation Theorem", which states that deep networks can approximate any nonlinear function. This theoretical basis is crucial for vibration signal modeling. In addition, the self-attention mechanism of LLM allows flexible processing of time dependencies and capture of long-range relationships. As an autoregressive model, LLMs can dynamically use historical data for prediction, while deep representation learning can enhance robustness to noise and non-stationarity, thereby ensuring high accuracy in practical applications.
[0103] The main connection between language models and time series models is that the input data for both is sequence data. Figure 3 The figure on the left shows an example of converting a vibration signal sequence containing 96 time stamps into a sequence of length 5, where each step in the sequence is represented by a 4-dimensional feature vector, indicating that the vibration signal can be segmented by a sliding window and discretized to extract statistical values (such as mean, standard deviation, minimum, maximum, etc.) to represent each window. Therefore, the vibration signal can also be decomposed into a series of text symbols. Figure 3The figure on the right shows how a large language model processes and predicts text data, including encoding the input text, processing it through the attention mechanism, and finally outputting the prediction result through a multi-layer network structure. The way this model processes text data is actually very similar to processing vibration signal data.
[0104] First, the vibration signal data can be viewed as a "language" in which each signal segment is equivalent to a "word" in the language. In this view, the prediction of vibration signals is similar to the prediction of words in language models. Therefore, a similar model architecture can be used to encode vibration signals, capture patterns and dependencies in the signal time series, and then use the model to predict future signal states.
[0105] 2. DCBiformerNet model (trend prediction model) architecture
[0106] In order to accurately predict the vibration signal trend of the monorail train gearbox, this embodiment proposes a DCBiformerNet (Deformable Convolutional Bi-directional Gated Recurrent Unit Network based on Informer) model, which is an improved version of the Informer model. Figure 4 As shown. The DCBiformerNet model mainly consists of four parts: Input, Informer Encoder, Informer Decoder and Output. In this embodiment, the task of DCBiformerNet trend prediction is regarded as a classification task, which is processed by embedding causal convolution, gated recurrent unit (GRU), multi-head self-attention (MHSA) and multi-head cross attention (MHCA) modules.
[0107] Combination Figure 5 As shown, the processing steps of the trend prediction model include:
[0108] S201: Building a trend prediction model including a multi-layer network based on a deep learning algorithm and performing model training;
[0109] In this embodiment, the trend prediction model is trained by existing means.
[0110] S202: Applying time position encoding (Positional Encoding) to the vibration signal data, embedding the time information into the vibration signal to generate a vibration signal representation containing the time series position information.
[0111] The characteristics of vibration signals are often significantly affected by the temporal position, so it is crucial to retain and utilize this position information during the modeling process. To this end, a position encoding mechanism is introduced in the input preprocessing of DCBiformerNet. Through this encoding method, the model can maintain sensitivity to the temporal information of the sequence, ensuring that the relationship between time steps is properly handled in subsequent feature extraction.
[0112] The formula is:
[0113]
[0114] Where: X i,m represents input data; PE(·) represents temporal position encoding; MHSA represents multi-head self-attention; DCC represents dilated causal convolution operation; pos represents the position of the time step, i represents the dimension of the position encoding; d model =32 represents the hidden layer dimension of the trend prediction model;
[0115] S203: Execute the following steps in each network layer of the trend prediction model:
[0116] S2031: Input the vibration signal representation containing the time series position information into the dilated causal convolution layer DCC and the multi-head self-attention MHSA layer to alternately extract local and global features, and generate the query (Q) and key (K) of the attention head;
[0117] Traditional sequence modeling methods, such as RNN or CNN, perform poorly when processing long time series. To overcome this limitation, this embodiment introduces dilated causal convolution DCC and multi-head self-attention MHSA mechanism in the encoder of DCBiformerNet. The causal convolution layer can capture important features over a long time range while maintaining causality and avoiding leakage of future information. This combination allows the model to focus on key features in the sequence more effectively and weight important information when necessary, thereby improving the perception of trend changes in vibration signals.
[0118] The formula is:
[0119]
[0120] Where: Q m ,K m ,V m ,b m represents the query, key, value, and bias of m (M=8) attention heads; N=6 represents the number of network layers; W n Represents the weight of the nth layer network; W m represents the attention weight matrix of the mth head; MHSA represents multi-head self-attention; DCC represents the dilated causal convolution operation;
[0121] In the attention mechanism, Q, K, and V represent:
[0122] Q (Query): Query vector, used to represent the information that the current input wants to query, or the expression of "what is needed".
[0123] K (Key): Key vector, which represents the identifier of candidate information and is used to match the query vector (Query) to determine which information is relevant.
[0124] V (Value): Value vector, which represents the actual information corresponding to the key vector (Key). The value vector extracted based on the matching degree between the query and the key is output.
[0125] S2032: Inputting the vibration signal representation containing the time series position information into the gated recurrent unit to perform time series information modeling and generate the value of the attention head;
[0126] To further enhance the model's understanding of temporal dependencies, DCBiformerNet integrates GRU and crisscross attention mechanisms (MHCA). GRU has the advantage of effectively capturing long-term dependencies in a sequence with a small number of parameters without losing important historical information. Combined with the crisscross attention mechanism, the model is able to interactively process the forward and backward dependencies of the time series, further improving the accuracy of the prediction. This bidirectional processing ensures that the model can recognize and capture complex patterns in vibration signals, whether these patterns are gradually developed or sudden.
[0127] S2033: Input the query, key and value of the attention head into the multi-head cross attention layer for feature fusion, and output the fusion features of the network in this layer;
[0128] The formula is:
[0129]
[0130] Where: GRU stands for gated recurrent unit; h t-1 represents the hidden state of the previous time step of GRU, which is used to retain information in the time series; MHCA represents multi-head cross attention; x t Represents the input features of the vibration signal at time t, which is a time point in the time series data. Through GRU and multi-head cross attention mechanism, combined with context information, further features are extracted for trend prediction.
[0131] S204: Aggregate the fusion features of each layer of the network as the trend prediction result of the vibration signal.
[0132] In this embodiment, trend prediction can be regarded as a classification task, and the classification results include four categories: continued to upward, continued to downward, first upward then downward, and first downward then upward.
[0133] Through this comprehensive feature processing and prediction mechanism, the trend prediction (DCBiformerNet) model of the present invention can adapt to vibration signals under various working conditions and provide high-precision and high-reliability prediction results.
[0134] 3. VSP-LLM (Vibration Signal Prediction Large Language Model) Architecture
[0135] In order to extract more time series features, the Prompt and Input modules are designed, and the vibration signal features and trends of the monorail train gearbox are embedded into the Prompt. To this end, this embodiment proposes a VSP-LLM architecture, such as Figure 2 As shown in the figure, the pre-trained large language model is frozen, and the powerful semantic understanding and time series modeling capabilities of the large language model (LLM) are used to solve the problem that it is difficult to capture the complex time series characteristics of non-stationary data such as vibration signals, and thus it is difficult to effectively extract signal features. The high-precision prediction of the monorail gearbox vibration signal is achieved by combining the trend of the monorail gearbox vibration signal, feature embedding and carefully designed prompts.
[0136] 1. Prompt module
[0137] Prompt design is a key step to ensure that VSP-LLM can effectively understand the vibration signal prediction task. In this embodiment, a Prompt template specifically used for vibration signal prediction is designed in the VSP-LLM architecture. This Prompt template contains the key elements in Table 1 below:
[0138] Table 1 Data type examples
[0139]
[0140] The VSP-LLM model embeds the content in Table 1 into the Prompt module to form Figure 1 The prompt template shown in (step 2) includes several key parts in Table 1:
[0141] 1) Context information: used to describe the application scenario and background of vibration signal data (such as signal type, acquisition conditions, etc.);
[0142] 2) Historical data description: Provide the observation range, frequency and total number of observation points of the vibration signal data time series;
[0143] 3) Statistical summary: summarize the main statistical characteristics of the vibration signal data, such as the mean, standard deviation, maximum and minimum values;
[0144] 4) Recent trends: Capture the latest changes and change rates of vibration signal data;
[0145] 5) Forecasting requirements: clarify the forecasting step and external influencing factors of the forecasting task;
[0146] 6) Special instructions: used to specify external factors to be considered in the forecasting process.
[0147] Through these modularly designed prompt contents, the prompt module plays the following roles in the architecture: 1) Enhance model understanding: Through carefully designed prompts, the model can more accurately interpret the characteristics of the input data, especially when processing multimodal or non-stationary signals, which can significantly improve the extraction of input features; 2) Optimize prediction accuracy: Clearly writing the task objectives into the prompt template (such as the number of prediction steps, external variables, etc.) can guide the model to focus on the key points of the task, thereby reducing noise interference and improving prediction accuracy; 3) Unify the multi-task processing framework: Through the modular prompt template design, it can adapt to prediction tasks in different scenarios (such as trend prediction, fault diagnosis, etc.) to achieve universal and refined prediction functions. These prompt elements are processed by natural language and combined with the DCBiformerNet network prediction result Y DBC , generate the embedding vector P E , forming a unified input representation.
[0148] The calculation formula of the Prompt module is expressed as:
[0149]
[0150] Where: P E represents the tag embedding vector; Y DBC Indicates the trend prediction result of vibration signal; LLM E represents the embedder of the pre-trained LLM, which is used to generate Output Token Embeddings (output tag embedding vector); φ(·) nonlinear transformation feature processing operation function; G n Refers to General Prompt, which indicates the general prompt for input; W l ″ represents the weight coefficient of the trend prediction part obtained by the DCBiformerNet model, which is used for weighted summation; Δp nRepresents an offset or compensation value, which is used to adjust the merged embedding vector (the vector obtained by concatenating the output token embedding vector output by the prompt module and the text embedding vector output by the input module); l represents the current iteration index, which is used to gradually process features or calculate weights from the input; L represents the number of input features involved in feature fusion. This combination not only retains more temporal feature information, but also incorporates the semantic information in the natural language prompt into the model, thus laying the foundation for subsequent processing of LLM.
[0151] 2. Input module
[0152] The design of the Input module is a key step to ensure that LLM can effectively process vibration signal data and extract vibration signal features. In the VSP-LLM architecture, a module dedicated to processing vibration signal data and extracting vibration signal features is designed, namely the Input module. This module goes through the following series of detailed operation steps to ensure that the input vibration signal can be fully understood and utilized by the model.
[0153] Specifically, the processing steps of the Input module are as follows:
[0154] S501: After normalizing the vibration signal data, the normalized data is decomposed into a number of time segments of fixed lengths through a patching operation;
[0155] Vibration signals are often affected by noise and amplitude differences, which may cause the model's prediction accuracy to decrease. To address this challenge, the input vibration signal needs to be instance normalized. This process can eliminate the amplitude differences between signals and ensure that the statistical characteristics of the signals are consistent. The normalized signal This can provide more stable input characteristics in the subsequent processing. Then, the normalized signal is patched into multiple time segments of fixed length, and the signal is broken down into easy-to-process segments through the patching operation. These segments retain the local characteristics of the time series and can effectively capture the detailed information in the signal.
[0156] The normalized formula is expressed as:
[0157]
[0158] Where: x represents the original vibration signal data; represents the normalized data;
[0159] S502: Convert each time segment into a corresponding high-dimensional feature vector through a patch embedding module.
[0160] For each time segment, the slice embedding module is responsible for converting it into a high-dimensional feature vector. This module maps the segment into a high-dimensional space through linear transformation to generate vibration signal time series patches (Time Series Patches). These embedding vectors capture rich information in the time segment. Next, the pre-trained word embedding (Pre-trained Word Embeddings) vector is mapped to the prototype (Text Prototype): that is, typical and representative sentences or paragraphs are extracted from a large amount of text data. And the input embedding vector for the time series patch is generated through the multi-head cross attention mechanism (MHCA). This process is as follows Figure 6 As shown in the figure, it illustrates the process of mapping pre-trained word embeddings to prototypes, which are then used to generate input embeddings of time series patches through multi-head criss-cross attention (MHCA).
[0161] The processing steps of the slice embedding module are as follows:
[0162] S5021: Map each time segment to a high-dimensional space through linear transformation to generate vibration signal time series slices (Patches);
[0163] S5022: Map pre-trained word embedding vectors to text prototypes;
[0164] In this embodiment, mapping the pre-trained word embedding vector to the text prototype means that the semantic information of each specific word (such as "late", "early", "up", "down", etc.) is mapped to a higher-level text prototype (Prototypes) representation by using pre-trained word embeddings. This text prototype is an abstract representation of the semantics of words, which is used to refine and summarize the relationships and characteristics between words. The role is to be used for multimodal interaction. These text prototypes can be combined with time series patches (Time Series Patches) and information fusion is performed through the multi-head cross attention mechanism (MHCA), which ultimately helps the model to better complete the prediction task.
[0165] Pre-trained Word Embeddings are a commonly used method for representing words in natural language processing. They convert words into low-dimensional dense vectors, capturing both their semantic and contextual information. These vectors are typically trained by deep learning models on large-scale corpora and can represent semantic relationships and similarities between words.
[0166] Pre-training: Trained on a large-scale corpus using models such as Word2Vec, GloVe, BERT, etc., and can be directly used for downstream tasks.
[0167] Example:
[0168] The embedding vectors of "king" and "queen" are close because they are similar in the semantic space.
[0169] The embedding vectors of "up" and "down" capture opposite direction information.
[0170] Text Prototypes are a higher-level abstract representation that summarizes task-related features in pre-trained word embeddings. Text Prototypes do not directly correspond to specific words but summarize general attributes in certain categories, features, or semantics, such as "up", "down", "steady", etc., which describe time trends.
[0171] Semantic aggregation: Maps semantically similar words to the same prototype representation. For example, both "up" and "rise" can be mapped to the prototype representing "ascending".
[0172] Examples:
[0173] Both "up" and "rise" are mapped to the text prototype of "ascending".
[0174] Both "steady" and "stable" are mapped to the text prototype of "stable".
[0175] Specific implementation of the mapping operation:
[0176] MHCA (Multi-Head Cross Attention) is used to implement the mapping between Text Prototypes and Time Series Patches.
[0177] 1) Input data:
[0178] On the left are pre-trained word embeddings, which are mapped to generate text prototypes.
[0179] On the right are multiple time series fragments (Patches) formed after the time series data is sliced.
[0180] 2) Mapping mechanism:
[0181] Multi-head criss-cross attention (MHCA) is the core module that connects text prototypes with time series slices.
[0182] Cross-attention means:
[0183] A query is a time series slice.
[0184] Key and Value are text primitives.
[0185] By calculating the similarity between Query and Key (dot product attention), features related to Query are extracted from Value.
[0186] 3) Function: The function of the MHCA module is to align time series information with text semantic information and generate a fused representation containing time information and semantic features.
[0187] 4) Output results: The final mapping results are semantically enhanced representations of time series segments, which will be further used in the downstream prediction tasks of the model.
[0188] S5023: Generate corresponding embedding as high-dimensional feature vector for each vibration signal time series slice using multi-head cross attention through text prototype;
[0189] The formula is:
[0190]
[0191] Where: Q h , K h 、V h , W h They represent the query, key, value vector and attention weight matrix of the h=8th head of multi-head cross attention respectively; Q h,P , K h,P Provided by the generated text prototype; Provided by time slice;d kRepresents the dimension of the key vector; through MHCA processing, the correlation between the time series segments of the vibration signal and the text prototype is comprehensively modeled. The model can identify the key features in the time series data and use the information of the text prompts to enhance the prediction ability.
[0192] S503: Utilize W through linear mapping layer f ·z+b f The high-dimensional feature vectors of each time segment are converted into corresponding text embedding vectors. After being processed by the multi-head cross attention mechanism, the generated high-dimensional feature embedding contains rich temporal information and semantic information. In order to make these embedding vectors further applied to prediction tasks, it is necessary to connect a linear mapping layer behind the generated high-dimensional feature vectors. f ·z+b f Convert high-dimensional features into the final model input form.
[0193] 3. Pre-training large language models and prediction
[0194] After the vibration signal data is processed by the previous Prompt module and the Input module, it is sent to the pre-trained large language model (LLM) for processing. The embedded vector after the multi-layer processing of the LLM is passed to the prediction generation module. At this stage, the task of the model is to convert the high-dimensional embedded vector into an interpretable prediction result. To this end, the output first passes through a linear projection layer (Projection Layer) to project the high-dimensional vector into the output space.
[0195] The mathematical expression of this process is as follows:
[0196] FCP signal (T) = W·PH·[Flatten(LLM B (α·I E (T),β·P E (T)))+γ·T];
[0197] Where: FCP signal represents the prediction result of the vibration signal; PH(·) represents the prediction head; W represents the mapping layer weight; α and β are the scale coefficients of the Input module and the Prompt module respectively; γ is the weight coefficient of the prediction time step T; Flatten(·) represents the expansion of high-dimensional features into a one-dimensional vector for subsequent linear operations or projections; LLM B Represents the body of the pre-trained LLM, which is used to embed the input into I E and P E Combined for processing and generate high-dimensional features; I E and P EDenote the text embedding vector and the tag embedding vector respectively. This formula introduces the time step T into the prediction generation process, ensuring that the model can consider the impact of future time when generating predictions.
[0198] 4. Experimental Description
[0199] 1. Experimental details and evaluation indicators
[0200] In order to verify the effectiveness of the VSP-LLM architecture, experiments were conducted on the straddle-type monorail train gearbox vibration dataset (DGLC) collected on the test bench. The dataset is divided into three categories: inner race fault, assembly error, and normal gearbox vibration signal. The sampling frequency of the measured data is 10240Hz, and the dataset is divided into training set, validation set, and test set in a ratio of 7:2:1.
[0201] In the experiment, 5 sets of data from the DGLC data set and one set of data from the CWRU data set (drive-end fault data with a sampling frequency of 12kHz, numbered: Drive_end_299.mat) were selected for testing. Among them, the normal working condition data set numbered 7 was collected under the conditions of theoretical speed 2000rpm, load 252Nm, and actual speed 1946rpm; the normal working condition data set numbered 13 was collected under the conditions of theoretical speed 1500rpm, load 334Nm, and actual speed 1460rpm; the inner ring fault working condition data set numbered 211 was collected under the conditions of theoretical speed 2000rpm, load 375Nm, and actual speed 1945rpm; the assembly error working condition data sets numbered 78 and 87 were collected under the conditions of theoretical speed 2000rpm, load 375Nm, actual speed 1945rpm and theoretical speed 2000rpm, load 252Nm, and actual speed 1945rpm respectively. Ablation experiments and comparative experiments were conducted on 78 working condition datasets. The prediction effects of the five DGLC datasets and the CWRU dataset were verified on the VSP-LLM model. The prediction effects of the TIME-LLM, DLinear, Autoformer, and Informer models were compared on the CWRU dataset to analyze the generalization ability of the models.
[0202] In terms of experimental details, in order to ensure that the model achieves optimal performance, this experiment carefully adjusted and recorded all key hyperparameters. Including learning rate, batch size, number of layers, and regularization parameters. The specific hyperparameter settings are detailed in Table 2. Since the vibration signal data has complex time series characteristics and the model complexity is high, especially when the memory is limited, choosing a batch size of 2 helps to capture the subtle features in the data more carefully. The learning rate gradually decays from 0.01 to 0.001 to prevent gradient explosion or model overfitting, so as to maintain a balance between early rapid convergence and later fine-tuning. The hidden layer dimension d_model is set to 32, and the feedforward network dimension d_ff is set to 128 to meet the requirements of model complexity and balanced performance. GELU is selected as the activation function because it has better stability than ReLU when processing complex nonlinear features and can maintain good gradient propagation. In terms of loss function, the experiment uses a combination of MSE, RMSE, and MAE to comprehensively measure the prediction performance of the model from multiple perspectives.
[0203] Table 2 Experimental hyperparameter settings of DGLC dataset
[0204]
[0205] The performance of the prediction model is evaluated based on three indicators: mean absolute error (MAE), root mean square error (RMSE) and mean square error (MSE). The specific calculation formula is as follows:
[0206]
[0207] Where n is the prediction length, Indicates the actual value, y i is the predicted value. The lower the value of the three evaluation indicators, the better the model prediction performance.
[0208] 2. Analysis of ablation experiment results
[0209] Table 3 Ablation experiment model categories
[0210]
[0211] In order to evaluate the impact of different components in the VSP-LLM architecture on prediction performance, an ablation experiment model was set up to evaluate the model, including the statistical variable settings of different prompt templates in Table 3, and experimental evaluation of whether the DCBiformerNet module is embedded.
[0212] Table 4 Model prediction step length = 96, ablation experiment results
[0213]
[0214] Table 4 shows the ablation experiment results of different prediction models when the model prediction step size is 96. By comparing the performance indicators of different models, we can clearly see the performance of each model in the three key indicators of MSE (mean square error), MAE (mean absolute error) and RMSE (root mean square error).
[0215] As can be seen from Table 4, the F5 model performs well in all three indicators, especially in MSE, MAE and RMSE, reaching 0.1744, 0.3228 and 0.4176 respectively, which is significantly better than other models. The excellent performance of the F5 model fully demonstrates the effectiveness of the VSP-LLM architecture in the monorail train gearbox vibration signal prediction task. Compared with other models, the F5 model better captures the complex time series characteristics in the vibration signal and significantly improves the prediction accuracy. The experimental results show that the proposed VSP-LLM architecture can provide accurate prediction results in vibration signal prediction, which not only optimizes the learning ability of the model, but also reduces the prediction error and improves the reliability of the overall prediction.
[0216] 3. Comparative experiment
[0217] In order to comprehensively evaluate the effectiveness of the model proposed in this paper, a comparative experiment was conducted using the DGLC dataset collected in the laboratory. Table 5 systematically presents the performance of different models at various prediction steps, including three evaluation indicators: mean absolute error (MAE), mean square error (MSE), and root mean square error (RMSE). In order to fully understand the performance of the model in different step-size prediction tasks, we focus on comparing Informer (from JIN M, WANG S, MA L, et al. Time-LLM: Timeseries forecasting by reprogramming large language models[J].arXiv preprintarXiv:2310.01728,2023.), Autoformer (from WU H, XU J, WANG J, et al. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting[C] / / Advances in Neural Information Processing Systems, 2021, 34:22419-22430.), and DLinear (from ZENG A, CHEN M, ZHANG L, et al. Are transformers effective for time series forecasting[C] / / Proceedings of the AAAI Conference on Artificial Intelligence). Intelligence.2023,37(9):11121-11128.), VSP-LLM and its fine-tuned version VSP-LLM* model.
[0218] Table 5 Prediction performance of different models at different steps (VSP-LLM* model is a fine-tuned version of VSP-LLM, and the training parameters are adjusted to improve performance)
[0219]
[0220]
[0221] As can be seen from Table 5, in the short time step prediction task (T = 48 and T = 96), the VSP-LLM* model performs well, with MAE and MSE of 0.2517 and 0.1014 respectively when the step length is 48, which is significantly better than Autoformer (MAE: 0.7198, MSE: 0.8301), showing its high-precision prediction ability for short-step vibration signals. In the medium time step prediction task (T = 192 and T = 384), the VSP-LLM* series model still maintains its advantage. When the step length is 192, the MAE of VSP-LLM* is 0.2793 and the MSE is 0.1235, which is much lower than DLinear (MAE: 0.4530, MSE: 0.3461), showing its stability and robustness in multi-step prediction tasks. In the ultra-long step prediction tasks (T=768 and T=1536), VSP-LLM* continued to perform well, with a MAE of 0.2882 and an MSE of 0.1317 at a step length of 768, significantly ahead of Informer and Autoformer, indicating that it can better capture complex features and achieve the best prediction effect when processing long time series signals.
[0222] Under different experimental settings, the change in the performance of the VSP-LLM model is mainly due to the effect of the step length on the prediction complexity. Under shorter step lengths (such as 48 and 96), the model only needs to capture the time series features in a smaller range, and the prediction difficulty is lower, so the error is smaller. However, as the step length increases (such as 192 and 384), the model needs to model more complex long-term dependencies and trend changes, resulting in a significant increase in the difficulty of prediction. Especially when the step length is 192 and 384, it is difficult for the model to balance short-term features and long-term trends, and the mean square error increases to more than 0.2. It is worth noting that under longer step lengths (such as 768 and 1536), the overall trend of the vibration signal is smoother, making it easier for the model to capture macroscopic distribution characteristics; at the same time, the trend information incorporated into the prompt word has a stronger constraint on the model, further improving the accuracy of the overall trend prediction, so the model performance is improved. To further improve the performance, the VSP-LLM model is fine-tuned (marked as VSP-LLM*). By adjusting the model structure and learning strategy (such as resetting d model =128, optimized the number of encoder and decoder layers to 3 and 2 respectively, and redesigned the prompt template based on the knowledge of vibration signal domain), the VSP-LLM* model showed comprehensive advantages in multiple tasks. Especially in long-step prediction, the VSP-LLM* model achieved significant improvements in capturing abnormal trends and enhancing expression capabilities, but there is still room for optimization. These factors jointly determine the performance of the model under different experimental settings and provide a clear direction for future improvements.
[0223] Figure 7The figure shows the prediction error distribution curves of the five models used in the comparative experiment at different step sizes, which can intuitively understand the performance differences of different models in the vibration signal prediction task.
[0224] The prediction results of Autoformer, DLinear, and VSP-LLM* models are shown in Figure 2. Figure 8 As shown in the figure. The blue curve represents the actual value, and the orange curve represents the predicted value of the model. In the short time step prediction task, the VSP-LLM model performs well and can accurately capture the rapid changes of the vibration signal, showing its excellent prediction ability under short time steps. In contrast, the predicted value of the Autoformer model has obvious lag, cannot accurately track the signal changes, and performs relatively weakly. Although the DLinear model has improved in trend capture, it is still insufficient in areas of drastic changes. In the medium time step prediction task, the VSP-LLM model still maintains high accuracy and stability and can cope with more complex signal changes. However, the Autoformer lags more significantly under these step sizes and has difficulty capturing the rapid fluctuation characteristics in the signal. DLinear performs well in some stable areas, but the prediction error is still large when dealing with large fluctuations. In the ultra-long time step prediction task, the VSP-LLM model shows extremely strong prediction ability, can accurately track the complex characteristics of the vibration signal, and maintain a low error in the entire ultra-long time series. In contrast, the prediction curve of Autoformer is too smooth to capture the drastic fluctuations of the signal, while the DLinear model performs well in some areas, but its cumulative error increases significantly with the increase of time steps.
[0225] Figure 8 The analysis results clearly show the excellent performance of the VSP-LLM model in the vibration signal prediction task with different time steps. Whether it is a short step (T=48, T=96), a medium step (T=192, T=384), or an ultra-long step (T=768, T=1536), VSP-LLM always maintains a high prediction accuracy and a keen ability to capture complex signal features. Under a short time step, VSP-LLM quickly tracks the drastic changes in the signal to avoid lagging prediction values; under medium and ultra-long time steps, VSP-LLM exhibits a strong trend capture capability, not only accurately predicting the overall trend of the vibration signal, but also responding to high-frequency fluctuations in a timely manner, significantly outperforming the Autoformer and DLinear models.
[0226] In order to comprehensively evaluate the effectiveness of the VSP-LLM architecture, we conducted an in-depth analysis of its prediction performance at different time steps and discussed in detail the impact of system complexity on model performance. The complexity of the architecture mainly comes from the multi-head attention calculation in the MHCA module. Assume that the number of input time series segments is N q , the feature dimension is d q , then the computational complexity of MHCA is Despite the high complexity, by optimizing the number of sampling points and weight distribution strategy in the attention mechanism, the system achieves high computational efficiency while maintaining high prediction accuracy. This enables the VSP-LLM architecture to meet the real-time requirements of industrial environments while providing sufficient prediction accuracy in practical applications. Models such as Autoformer, Informer, DLinear, VSP-LLM and VSP-LLM* were selected for multi-step prediction sensitivity and experimental results analysis. When analyzing the training loss of the model at different time steps, the changing trends of each model at different time steps T can be observed. Overall, VSP-LLM and VSP-LLM* exhibited superior performance at most time steps, as evidenced by a faster decrease in training loss and stabilization at lower loss values, such as Fig. 9 shown.
[0227] In the short time step prediction task, the VSP-LLM and VSP-LLM* models show significant advantages. In the early iterations (Epoch0-2), these two models quickly reduce the training loss and eventually stabilize it at a low level (about 0.2-0.3). In contrast, other models such as Informer and Autoformer perform relatively weakly. Although the Dlinear model also shows good convergence, its final loss is slightly higher than that of VSP-LLM*. This shows that the VSP-LLM and VSP-LLM* models have higher prediction accuracy under short time steps.
[0228] As the time step increases, VSP-LLM and VSP-LLM* still maintain good performance, with smooth and steadily decreasing training loss curves. The results reflect the prediction stability and robustness of these two models when dealing with longer time steps. In contrast, Autoformer's loss decreases more slowly at these time steps and has higher loss levels at higher time steps, reflecting its limitations in long time step prediction tasks. Although Informer's performance is close to that of the VSP-LLM series of models, its final converged loss value is still higher than that of VSP-LLM and VSP-LLM*.
[0229] At very long time steps, the VSP-LLM and VSP-LLM* models still effectively reduce training loss and converge at a lower loss level. Especially at time step T = 768, the performance of VSP-LLM* is significantly better than other models, further verifying its excellent performance in long time step prediction tasks. In contrast, the Autoformer model performs the worst at these time steps, and the loss does not decrease significantly, exposing its limitations in long time step prediction tasks. Informer and Dlinear also perform worse than the VSP-LLM series models at these very long time steps, especially at T = 1536, the final training loss of VSP-LLM and VSP-LLM* is significantly lower than that of other models.
[0230] Table 6 summarizes the prediction performance of the VSP-LLM* model on the DGLC dataset, covering different fault conditions (7, 13, 78, 87, 221) and prediction steps (48 to 1536), including MSE, RMSE and MAE. From the experimental results, it can be seen that the VSP-LLM* model performs well under different sampling points and step sizes. In the DGLC dataset 7, VSP-LLM* achieves the best performance in MSE, RMSE and MAE indicators for all step sizes. Overall, the proposed VSP-LLM model shows good vibration signal prediction capabilities under different working conditions and different prediction steps.
[0231] Table 6 shows the vibration signal prediction results of the VSP-LLM* model under different working conditions.
[0232]
[0233] Table 7 Prediction performance of each model on the CWRU dataset
[0234]
[0235]
[0236] 4. Verification of generalization ability
[0237] In order to verify the generalization ability of the model, the public dataset of Case Western Reserve University was selected for experimental verification. The experimental results are shown in Table 7. On the CWRU dataset, the VSP-LLM* model showed low mean square error (MSE) and mean absolute error (MAE) at all step sizes. In particular, at 48 and 768 step sizes, the VSP-LLM* model reached the lowest value in the table, and also performed well at 1536 step sizes. In terms of average indicators (Avg), the VSP-LLM model has an MSE of 0.2585 and a MAE of 0.4045, which are better than other comparison models, such as DLinear's MSE of 0.2740 and TIME-LLM's MAE of 0.4023. In the best performance statistics,
[0238] VSP-LLM* achieved the best performance twice in the test with 6 step lengths. Although DLinear achieved the best result in 3 step lengths, the stability advantage of VSP-LLM* over multiple step lengths is more obvious, and other models (such as Autoformer and Informer) did not perform well at any step length.
[0239] In summary, the MSE and MAE indicators of the VSP-LLM framework proposed in the present invention fluctuate little at different step sizes, showing stability and robustness to different time steps and fault conditions. It can accurately capture the complex timing characteristics of vibration signals, has strong generalization capabilities in the field of rotating machinery vibration signal prediction, and can adapt to a variety of fault signals and time series tasks.
[0240] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the technical solution. Those skilled in the art should understand that those modifications or equivalent substitutions of the technical solution of the present invention that do not depart from the purpose and scope of the technical solution should be included in the scope of the claims of the present invention.
Claims
1. A gearbox vibration prediction method for straddle-type monorail train based on a large language model, characterized in that: include: S1: Acquire vibration signal data of the gearbox of a straddle-type monorail train; S2: Input the vibration signal data into the trained trend prediction model to perform vibration signal trend prediction, and output the trend prediction result of the vibration signal; S3: Construct corresponding prompt word templates based on the domain characteristics and prediction requirements of the prediction task; S4: Input the trend prediction result of the vibration signal and the prompt word template into the Prompt module for processing to generate the corresponding tag embedding vector; S5: Input the vibration signal data into the Input module for text embedding mapping to generate the corresponding text embedding vector; S6: concatenate the tag embedding vector and the text embedding vector and input them into the pre-trained large language model, output the corresponding vibration signal fusion embedding representation; input the vibration signal fusion embedding representation into the prediction head for digitization, and generate the corresponding vibration signal prediction result; S7: Outputting the vibration signal prediction result as the vibration signal prediction result of the straddle-type monorail train gearbox.
2. The method for predicting gearbox vibration of straddle-type monorail train based on large language model according to claim 1, characterized in that: In step S2, the processing steps of the trend prediction model include: S201: Building a trend prediction model including a multi-layer network based on a deep learning algorithm and performing model training; S202: applying time position coding to the vibration signal data, embedding the time information into the vibration signal to generate a vibration signal representation containing the time series position information; The formula is: Where: X i,m represents input data; PE(·) represents temporal position encoding; MHSA represents multi-head self-attention; DCC represents dilated causal convolution operation; pos represents the position of the time step, i represents the dimension of the position encoding; d model =32 represents the hidden layer dimension of the trend prediction model; S203: Execute the following steps in each network layer of the trend prediction model: S2031: Input the vibration signal representation containing the time series position information into the dilated causal convolution layer and the multi-head self-attention layer to alternately extract local and global features, and generate the query and key of the attention head; The formula is: Where: Q m ,K m ,V m ,b m represents the query, key, value, and bias of m attention heads; N = 6 represents the number of network layers; W n Represents the weight of the nth layer network; W m represents the attention weight matrix of the mth head; MHSA represents multi-head self-attention; DCC represents the dilated causal convolution operation; S2032: Inputting the vibration signal representation containing the time series position information into the gated recurrent unit to perform time series information modeling and generate the value of the attention head; S2033: Input the query, key and value of the attention head into the multi-head cross attention layer for feature fusion, and output the fusion features of the network in this layer; The formula is: Where: GRU stands for gated recurrent unit; h t-1 represents the hidden state of the previous time step of GRU; MHCA represents multi-head cross attention; x t represents the input characteristics of the vibration signal at time t; S204: Aggregate the fusion features of each layer of the network as the trend prediction result of the vibration signal.
3. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 2, characterized in that: In step S2, vibration signal trend prediction is a classification task, and the classification results include: continued to upward: continued to upward; continued to downward: continued to downward; first upward then downward: first upward then downward; first downward then upward: first downward then upward.
4. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 1, characterized in that: In step S3, the prompt word template includes the following content: 1) Context information: used to describe the application scenario and background of vibration signal data; 2) Historical data description: Provide the observation range, frequency and total number of observation points of the vibration signal data time series; 3) Statistical summary: summarize the mean, standard deviation, maximum and minimum values of the vibration signal data; 4) Recent trends: Capture the latest changes and change rates of vibration signal data; 5) Forecasting requirements: clarify the forecasting step and external influencing factors of the forecasting task; 6) Special instructions: used to specify external factors to be considered in the forecasting process.
5. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 1, characterized in that: In step S4, the calculation formula of the Prompt module is expressed as: Where: P E represents the tag embedding vector; Y DBC Indicates the trend prediction result of vibration signal; LLM E represents the embedder, which is used to generate the tag embedding vector; φ(·) is the nonlinear transformation feature processing operation function; G n Indicates a general prompt for input; W l ″ represents the weight coefficient of the trend prediction part; Δp n Represents the offset or compensation value; l represents the current iteration index; L represents the number of input features involved in feature fusion.
6. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 1, characterized in that: In step S5, the processing steps of the Input module are as follows: S501: After normalizing the vibration signal data, the normalized data is decomposed into a number of time segments of fixed lengths through a patching operation; S502: Convert each time segment into a corresponding high-dimensional feature vector through a slice embedding module; S503: Utilize W through linear mapping layer f ·z+b f Convert the high-dimensional feature vectors of each time segment into corresponding text embedding vectors.
7. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 6, characterized in that: In step S501, the normalized formula is expressed as: Where: x represents the original vibration signal data; Represents normalized data.
8. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 6, characterized in that: In step S502, the processing steps of the slice embedding module are as follows: S5021: Map each time segment to a high-dimensional space through linear transformation to generate a vibration signal time series slice; S5022: Map pre-trained word embedding vectors to text prototypes; S5023: Generate corresponding embedding as high-dimensional feature vector for each vibration signal time series slice using multi-head cross attention through text prototype; The formula is: Where: Q h , K h 、V h , W h They represent the query, key, value vector and attention weight matrix of the h=8th head of multi-head cross attention respectively; Q h,P , K h,P Provided by the generated text prototype; Provided by time slice;d k Indicates the dimension of the key vector.
9. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 1, characterized in that: In step S6, the calculation formula of the vibration signal prediction result is expressed as: FCP signal (T)=W·PH·[Flatten(LLM B (α·I E (T),β·P E (T)))+γ·T]; Where: FCP signal represents the vibration signal prediction result; PH(·) represents the prediction head; W represents the mapping layer weight; α and β are the proportional coefficients of the Input module and the Prompt module respectively; γ is the weight coefficient of the prediction time step T; LLM B Represents the main part of the pre-trained large language model, which is used to embed the text into the vector I E and the token embedding vector P E Combining and generating high-dimensional features, namely, vibration signal fusion embedding representation; Flatten(·) means expanding the high-dimensional features into a one-dimensional vector.
10. The method for predicting gearbox vibration of straddle-type monorail train based on large language model as claimed in claim 1, characterized in that: In step S6, the fusion embedding representation of the vibration signal output by the large language model includes: the predicted statistical variable values of the future vibration signal and the relevant description of the trend of the vibration signal.
Citation Information
Cited By
Power consumption prediction method based on Time-LLM
CN120448714A