Track irregularity prediction method based on TCN-BiLSTMs-Attention model

By using the TCN-BiLSTMs-Attention model in orbital uneven prediction, the problem of low prediction accuracy in the existing methods is solved, and more accurate orbital uneven trend prediction is achieved, providing an early warning effect.

CN120067561APending Publication Date: 2025-05-30XIAN UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411981541.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing track uneven sequence data, the existing track uneven prediction methods fail to fully consider the different impacts of various track uneven indexes on track unevenness, resulting in low prediction accuracy.

Method used

The orbital uneven prediction method based on the TCN-BiLSTMs-Attention model is used to extract the preprocessed orbital uneven multivariate input data through a time convolutional neural network (TCN). Combining the multi-layer bidirectional long and short-term memory network (BiLSTMs) and attention mechanism, multi-scale features and nonlinear changes in the data are captured.

Benefits of technology

It improves the accuracy of track uneven trend prediction, can more effectively capture the impact of various indicators on track unevenness, provide more accurate prediction results, and achieve early warning effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067561A_ABST
    Figure CN120067561A_ABST
Patent Text Reader

Abstract

The invention discloses a track irregularity prediction method based on a TCN-BiLSTMs-Attention model, and the method comprises the specific steps: obtaining TQI data, and carrying out the preprocessing of the data; inputting the data into a TCN layer to carry out expansion causal convolution operation so as to extract data features; inputting the training set into a BiLSTM model for training, adding an attention mechanism, and carrying out weighted average on all obtained prediction subsequences to obtain a final irregularity trend prediction result; and carrying out evaluation on the trained TCN-BiLSTMs-Attention model by adopting the test set. The method provided by the invention has more accurate track irregularity trend prediction capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rail transit, and particularly relates to a method for predicting track irregularity based on the TCN-BiLSTMs-Attention model. Background Art

[0002] The smoothness of the track is one of the key factors to ensure the safe operation of trains. An uneven track will accelerate the wear of wheels and tracks, thereby affecting the driving performance of trains, reducing the riding comfort, and increasing the maintenance cost. In extreme cases, it may also pose safety risks. Therefore, accurate prediction of the track irregularity trend is crucial for track maintenance, train safety, and improving passenger comfort.

[0003] Remarkable progress has been made in the field of track irregularity research at home and abroad, mainly focusing on the exploration of new methods, the construction of theoretical foundations, the development of evaluation tools, and the application of deep learning technologies. Researchers have provided theoretical support for the prediction of track irregularity by analyzing the geometric characteristics of wheel-rail contact and dynamic simulation, and used the Track Quality Index (TQI) as an evaluation index. At the same time, deep learning technologies, especially models combining Convolutional Neural Networks (CNNs) and Long Short-Term Memory Networks (LSTMs), have shown great potential in processing spatio-temporal feature data and have become a hot topic in current research.

[0004] However, the data sequences generated by track irregularities exhibit complex non-linear and non-stationary characteristics, and the information and laws contained in these data are not easily identifiable. Traditional prediction techniques usually have difficulty in accurately capturing the inherent properties of the data. In addition, predictions based on only a single index may ignore the influence of other relevant factors, which may lead to a decrease in the accuracy of the prediction results. Summary of the Invention

[0005] The object of the present invention is to provide a method for predicting track irregularity based on the TCN-BiLSTMs-Attention model, which solves the problems that the existing prediction methods do not consider the different influences of various track irregularity indicators on track irregularity and have low prediction accuracy when processing track irregularity sequence data.

[0006] For the above purpose, the technical solution adopted by the present invention is: a method for predicting track irregularity based on the TCN-BiLSTMs-Attention model, which is specifically implemented according to the following steps:

[0007] Step 1, obtain TQI data and perform preprocessing. After preprocessing, it is divided into a training set and a test set;

[0008] Step 2: Input the data obtained in Step 1 into the TCN layer for dilated causal convolution operation to extract data features, thereby capturing the influence of various indicators on the TQI value, providing richer inputs for the subsequent model, and constructing the feature vectors output by the TCN into a sequence form as the input of the BiLSTMs network;

[0009] Step 3: Input the data processed in Step 2 into a multi-layer BiLSTMs model for training;

[0010] Step 4: Apply the attention mechanism module to the results obtained through training in Step 3 for information extraction;

[0011] Step 5: Use the test set to evaluate the model processed in Step 4.

[0012] As a preferred technical solution of the present invention, in Step 1, the preprocessing includes data normalization and outlier removal.

[0013] As a preferred technical solution of the present invention, Step 1 is specifically as follows:

[0014] Step 1.1: Set an L-meter track section as a unit section, and calculate the sum of the standard deviations of various track geometric irregularity indicators on the unit section, that is, obtain the TQI data;

[0015] Step 1.2: Perform normalization processing on the TQI data obtained in Step 1.1;

[0016] Step 1.3: Remove outliers from the data normalized in Step 1.2.

[0017] As a preferred technical solution of the present invention, in Step 2:

[0018] To avoid applying information from the next moment in the convolution operation, causal convolution performs zero-padding at the beginning of the sequence to ensure that at any time point t, the output depends only on the input at time point t and before, and does not depend on the input after t;

[0019] Dilated convolution enlarges the receptive field of the convolutional layer by filling "holes" in the convolutional kernel, so that there are intervals between the convolutional kernel elements, thereby covering a wider range, enabling the dilated convolution to enlarge the receptive field without increasing model parameters or computational costs.

[0020] As a preferred technical solution of the present invention, in Step 2:

[0021] For the feature extraction of the TCN, through the above dilated causal convolution operation, the TCN can extract multi-scale features in the data at different levels, providing richer inputs for the subsequent model.

[0022] As a preferred technical solution of the present invention, in the step 3, training is performed using a BiLSTM model. The bidirectional long short-term memory network (BiLSTM) can capture more comprehensive context relationships by considering both forward and backward information in the sequence.

[0023] As a preferred technical solution of the present invention, the attention mechanism is applied to the BiLSTMs network to promote information extraction in a probability-weighted manner.

[0024] As a preferred technical solution of the present invention, in the step 5, the evaluation metrics used are: root mean square error (RMSE), mean square error (MSE), mean absolute percentage error (MAPE), mean absolute error (MAE), and R-squared (R 2 )

[0025] The beneficial effects of the present invention are as follows: (1) The high-speed railway track irregularity prediction method based on the TCN-BiLSTMs-Attention model of the present invention uses a temporal convolutional neural network (TCN) to extract features from the preprocessed multi-variable input data of track irregularities, which can extract multi-scale features in the data at different levels, thereby capturing the influence of various indicators on the TQI value and providing richer inputs for the subsequent model;

[0026] (2) The high-speed railway track irregularity prediction method based on the TCN-BiLSTMs-Attention model of the present invention uses a multi-layer bidirectional long short-term memory network (BiLSTMs) model. The LSTM network can learn the non-linear changes and long-term dependence relationships of the TQI sequence data, and the bidirectional LSTM can process past and future information simultaneously. In addition, the model depth is increased, further improving the learning ability;

[0027] (3) In the high-speed railway track irregularity prediction method based on the TCN-BiLSTMs-Attention model of the present invention, the attention mechanism is applied to the BiLSTMs network to promote information extraction in a probability-weighted manner. The core of this probability-weighted mechanism lies in evaluating the importance of different information units and assigning different weights accordingly.

[0028] (4) The high-speed railway track irregularity prediction method based on the TCN-BiLSTMs-Attention model of the present invention has a more accurate track irregularity trend prediction ability, which can provide a reference for the railway department and achieve the effect of early warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0030] Figure 1 is the overall framework diagram of the high-speed railway track irregularity prediction method based on the TCN-BiLSTMs-Attention model of the present invention;

[0031] Figure 2 is the example diagram of TCN dilated causal convolution in the method of the present invention;

[0032] Figure 3 is the structure diagram of the LSTM model in the method of the present invention;

[0033] Figure 4 is the attention mechanism diagram in the method of the present invention;

[0034] Figure 5 is the TQI data graph adopted in Embodiment 6 of the present invention;

[0035] Figure 6 is the distribution graph of TQI detection outliers in Embodiment 6 of the present invention;

[0036] Figure 7 is the TQI data graph after outlier removal in Embodiment 6 of the present invention;

[0037] Figure 8 is the prediction result graph of the training set of the TCN-BiLSTMs-Attention model in Embodiment 6 of the present invention;

[0038] Figure 9 is the prediction result graph of the test set of the TCN-BiLSTMs-Attention model in Embodiment 6 of the present invention. Detailed implementation manners

[0039] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0040] Embodiment 1

[0041] A method for predicting track irregularity based on the TCN-BiLSTMs-Attention model of the present invention, as Figure 1 shown, is specifically implemented according to the following steps:

[0042] Step 1, obtain TQI (Track Quality Index) data, preprocess the data, and divide the preprocessed TQI data into a training set and a test set, where the training set accounts for 80%;

[0043] Step 2: Subsequently, the data is input into the TCN layer for dilated causal convolution operation to extract data features, thereby capturing the influence of various indicators on the TQI value, and the feature vectors output by the TCN are constructed into a sequence form as the input of the BiLSTMs network;

[0044] Step 3: The data processed in Step 2 is input into the BiLSTM (Bidirectional Long Short-Term Memory Network) model for training, and the model outputs a predicted subsequence;

[0045] Step 4: The attention mechanism is applied to the BiLSTMs network to promote information extraction through probability weighting. The core of this probability weighting mechanism lies in evaluating the importance of different information units and assigning different weights accordingly. All the predicted subsequences obtained in Step 3 are weighted and averaged to obtain the final prediction result of the roughness trend;

[0046] Step 5: The model processed in Step 4 is evaluated using the test set.

[0047] The present invention uses a Temporal Convolutional Network (TCN) to extract features from the preprocessed multi-variable input data of track roughness, which can extract multi-scale features in the data at different levels, thereby capturing the influence of various indicators on the TQI value and providing richer inputs for subsequent models.

[0048] The present invention uses a multi-layer Bidirectional Long Short-Term Memory Network (BiLSTMs) model, which utilizes the characteristics of the LSTM network that can learn the non-linear changes and long-term dependence relationships of TQI sequence data, as well as the ability of the bidirectional LSTM to process past and future information simultaneously, and increases the depth of the model, further improving the learning ability.

[0049] The present invention applies the attention mechanism to the BiLSTMs network to promote information extraction through probability weighting. The core of this probability weighting mechanism lies in evaluating the importance of different information units and assigning different weights accordingly.

[0050] The TCN-BiLSTMs-Attention model of the present invention has a more accurate track roughness trend prediction ability, which can provide a reference for the railway department and achieve the effect of early warning.

[0051] Embodiment 2

[0052] Different from Embodiment 1, in the track roughness prediction method based on the TCN-BiLSTMs-Attention model of the present invention in Embodiment 2, Step 1 is specifically as follows:

[0053] Step 1.1: Set an L-meter track section as a unit section, calculate the sum of the standard deviations of various track geometric irregularity indexes on the unit section, and thus obtain the TQI data;

[0054] The track geometric irregularity indexes include level, left elevation, right elevation, left alignment, right alignment, cross level, and gauge;

[0055] Step 1.2: Normalize the TQI data obtained in Step 1.1;

[0056] Specifically, the maximum-minimum normalization method is used to scale the TQI data and scale the data to the interval [0, 1]. Normalizing the TQI data can unify the dimension of the data and improve the accuracy and convergence speed of the model to a certain extent;

[0057] Step 1.3: Remove outliers from the normalized data in Step 1.2;

[0058] Specifically, the Isolation Forest algorithm is used to detect the outliers existing in the data to reduce the interference of the outliers on the normal data, thereby affecting the prediction accuracy.

[0059] Example 3

[0060] Different from Example 2, in Example 3 of the track irregularity prediction method based on the TCN-BiLSTMs-Attention model of the present invention, as Figure 2 shown, the specific process of Step 2 is as follows:

[0061] Step 2.1: Application of causal convolution: To avoid applying information from the next moment in the convolution operation, causal convolution performs zero-padding at the beginning of the sequence. This padding method ensures that at any time point t, the output only depends on the input at time point t and before it, rather than on the input after t.

[0062] Step 2.2: Implementation of dilated convolution: Dilated convolution enlarges the receptive field of the convolutional layer by filling "holes" in the convolutional kernel. This method creates intervals between the convolutional kernel elements, thereby covering a wider range. Dilated convolution realizes enlarging the receptive field without increasing the model parameters or computational cost.

[0063] Step 2.3: Feature extraction of TCN: Through the above dilated causal convolution operation, the Temporal Convolutional Network (TCN) can extract multi-scale features in the data at different levels. This multi-scale feature extraction provides richer inputs for the subsequent model.

[0064] Example 4

[0065] Different from Embodiment 3, in Embodiment 4 of the track irregularity prediction method based on the TCN-BiLSTMs-Attention model of the present invention, the specific process of Step 3 is as follows:

[0066] The structure of the LSTM (Long Short-Term Memory Network) model is as Figure 3 shown;

[0067] The calculation formulas of each module of the LSTM model are as follows:

[0068] f t = σ(W f · [h t-1 , X t + b f ) (1)

[0069] i t = σ(W i · [h t-1 , X t + b i ) (2)

[0070] o t = σ(W o · [h t-1 , X t + b o ) (3)

[0071]

[0072] In the formula, f t , i t , o t respectively represent the forget gate, input gate, and output gate, W f , W i , W o respectively represent the weight coefficient matrices of the forget gate, input gate, and output gate, h t-1 and X t respectively represent the hidden layer state at the previous moment and the input at the current moment, b f , b i , b o respectively represent the bias terms of the forget gate, input gate, and output gate, σ represents the activation function, and C t represents the value of the current cell state;

[0073] Through the functions of the forget gate, input gate, and output gate, the LSTM model can achieve good prediction results. Specifically, the forget gate can determine which information should be forgotten at the current time step, the input gate can determine which new information should be remembered at the current time step, and the output gate can determine which information should be output at the current time step. Under the combined action of these three gating mechanisms, the LSTM network can better capture the long-term dependencies in the sequence data, selectively remember and forget information, and can transmit the effective information to the next time step, thereby improving the model's ability to model time series data and prediction accuracy.

[0074] The bidirectional long short-term memory network (BiLSTM) can capture more comprehensive context relationships by considering both the forward and backward information in the sequence simultaneously. This structure enables the BiLSTM to not only understand the dependencies before the current information when processing sequence data but also predict the impact of future information on the current state, thereby providing richer context information than the unidirectional LSTM. In addition, by stacking multiple BiLSTM layers, the depth of the model can be increased, enabling the model to learn more complex feature representations and improving its learning ability and generalization ability for data. This deep structure helps the model to better capture long-distance dependencies when facing complex sequence prediction problems, enhancing the prediction accuracy and robustness of the model.

[0075] Example 5

[0076] Different from Example 4, in the method for predicting track irregularity based on the TCN-BiLSTMs-Attention model in Example 2 of the present invention, in step 5, the evaluation uses root mean square error (RMSE), mean square error (MSE), mean absolute percentage error (MAPE), mean absolute error (MAE), and R-squared (R 2 ) evaluation indicators to reflect the prediction accuracy of the experiment, and the formulas are as follows:

[0077]

[0078] In the formula, y i represents the predicted value output by the model, represents the true value in the test set, represents the mean of the true values in the dataset, and n refers to the number of data in the predicted value and the true value.

[0079] The smaller the values of root mean square error (RMSE), mean square error (MSE), mean absolute percentage error (MAPE), and mean absolute error (MAE), the better the training effect of the model. The closer the value of R 2 is to 1, the better the fitting effect between the predicted value and the true value.

[0080] Tune the model according to the evaluation results, such as adjusting the model structure, reselecting hyperparameters, etc., to improve the performance and stability of the model.

[0081] Example 6

[0082] Different from Example 5, in the method for predicting track irregularity based on the TCN-BiLSTMs-Attention model of the present invention, the TQI data used is as Figure 5 shown, and the preprocessed TQI data obtained after being processed by Step 1 is as Figure 7 shown.

[0083] Figure 5 is the original TQI data graph. The TQI dataset in the experiment comes from the actual track, the interval between data points is 0.25 m, and the collected data is taken with a 200 m track section as the unit section, and the sum of the standard deviations of various track irregularity indexes on the unit section is calculated to obtain the TQI data.

[0084] Figure 6 is the distribution graph of outliers detected after normalizing the original TQI data. After normalization processing, the TQI data is scaled to the interval [0, 1]. Normalizing the TQI data can unify the dimension of the data and improve the accuracy and convergence speed of the model to a certain extent; after normalization, outlier detection and elimination are performed on the data, Figure 7 is the TQI data graph after outlier elimination, to reduce the interference of outliers on normal data, thereby affecting the prediction accuracy. Among them, the darkly marked ones are outliers, and these outliers will be eliminated for subsequent experiments.

[0085] The track irregularity dataset on July 24, 2021 is divided into a training set and a test set, where the training set accounts for 80% and the test set accounts for 20%. The original data is preprocessed and then input into the TCN-BiLSTMs-Attention model for training. After multiple experiments, the selected model hyperparameters are as follows: BiLSTM Layers = 4; lookback_window = 8, epochs = 200, batch_size = 64, LSTM_units = 64, 32, 32, 32, and the learning rate uses the Adam optimization algorithm, which is an adaptive learning rate optimization algorithm. By adaptively adjusting the learning rate, the parameter update becomes more stable and efficient. In addition, a Dropout layer and an Early Stopping strategy are set to prevent the model from overfitting, save computing resources, and improve the generalization ability of the model.

[0086] In the present invention, the XGBoost model, BP model, LSTM model, LSTM-Attention model, and TCN-BiLSTMs-Attention model (the method of the present invention) prediction models are respectively used to predict the TQI data. The results are shown in Table 1 and Figure 8 and Figure 9 as shown

[0087] Table 1 Prediction Results of Different Models

[0088]

[0089]

[0090] In addition, to verify the robustness of the present model, we conducted experiments on another four track irregularity data sets on August 7, 2021, August 28, 2021, September 17, 2021, and September 24, 2021. The sizes of the data sets are 1048576 rows * 10 columns, 1044580 rows * 10 columns, 1048576 rows * 10 columns, and 1044728 rows * 10 columns respectively. Similarly, the seven track irregularity indicators are used as multivariate input features, and the TQI data is used as the target to be predicted. 80% of the data is used as the training set, and 20% is used as the test set for the experiment. The experimental results are shown in Table 2. It can be seen that the present model has achieved good prediction effects on different data sets, proving that the present algorithm has strong robustness.

[0091] Table 2 Experimental Results of the TCN-BiLSTMs-Attention Model on Different Data Sets

[0092]

[0093] The above description shows and describes several preferred embodiments of the invention. However, as mentioned above, it should be understood that the invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the inventive concept described herein through the above teachings or the technology or knowledge in the relevant field. And any changes and variations made by those skilled in the art without departing from the spirit and scope of the invention shall fall within the protection scope of the appended claims of the invention.

Claims

1. Track irregularity prediction method based on TCN-BiLSTMs-Attention model, characterized by: Follow the steps below to implement it: Step 1, obtain TQI data, preprocess it, and divide it into training set and test set after preprocessing; Step 2: Input the data obtained in step 1 into the TCN layer for dilated causal convolution to extract data features, thereby capturing the impact of various indicators on the TQI value, providing richer input for the model, and constructing the feature vector output by TCN into a sequence form as the input of the BiLSTMs network; Step 3, input the data processed in step 2 into the multi-layer BiLSTMs model for training; Step 4: Use the attention mechanism module to extract information from the results obtained through step 3 training; Step 5: Use the test set to evaluate the model processed in step 4.

2. The track irregularity prediction method based on the TCN-BiLSTMs-Attention model according to claim 1 is characterized in that: In step 1, the preprocessing includes data normalization and outlier removal.

3. The track irregularity prediction method based on the TCN-BiLSTMs-Attention model according to claim 2 is characterized in that: The step 1 is specifically: Step 1.1, set the L-meter track section as the unit section, calculate the sum of the standard deviations of various track geometric irregularity indicators on the unit section, and obtain the TQI data; Step 1.2, normalizing the TQI data obtained in step 1.1; Step 1.3, remove outliers from the normalized data in step 1.

2.

4. The track irregularity prediction method based on the TCN-BiLSTMs-Attention model according to claim 3 is characterized in that: In step 2: In order to avoid applying information to the next moment in the convolution operation, causal convolution performs 0 padding at the beginning of the sequence to ensure that at any time point t, the output only depends on the input at time point t and before, and does not depend on the input after t; Dilated convolution enlarges the receptive field of the convolution layer by filling the convolution kernel with "holes", so that there are gaps between the convolution kernel elements, which can cover a wider range and enable dilated convolution to enlarge the receptive field without increasing model parameters or computational costs.

5. The track irregularity prediction method based on the TCN-BiLSTMs-Attention model according to claim 4 is characterized in that: In step 2: Feature extraction of TCN,Through the above-mentioned dilated causal convolution operation, TCN can extract multi-scale features in the data at different levels,,providing richer input for the subsequent model.

6. The track irregularity prediction method based on the TCN-BiLSTMs-Attention model according to claim 5, characterized in that: In step 3, BiLSTM model training is adopted, and the bidirectional long short-term memory network (BiLSTM) can capture more comprehensive contextual relationships by simultaneously considering the forward and backward information in the sequence.

7. The track irregularity prediction method based on the TCN-BiLSTMs-Attention model according to claim 6, characterized in that: In step 4, the attention mechanism is applied to the BiLSTMs network to promote information extraction in a probability weighted manner.

8. The track irregularity prediction method based on the TCN-BiLSTMs-Attention model according to claim 7, characterized in that: In step 5, the evaluation indicators used are: root mean square error (RMSE), mean square error (MSE), mean absolute percentage error (MAPE), mean absolute error (MAE) and R square (R 2 ).

Citation Information

Cited By

  • Large-scale MIMO channel estimation method fusing double attention mechanism and TCN-BiLSTM network

    CN121000559A

  • Large-scale MIMO channel estimation method fusing dual attention mechanism and TCN-BiLSTM network

    CN121000559B