Dynamic and static track irregularity mapping method
By combining CNN and Transformer networks, the problems of low frequency and insufficient accuracy in dynamic and static track irregularity detection were solved, achieving efficient and accurate track irregularity detection and ensuring the safe operation of high-speed railways.
Patent Information
- Application Number
- CN202510792897.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-26
AI Technical Summary
Existing dynamic and static track irregularity detection technologies have problems such as low detection frequency, low efficiency, and difficulty in accurately mapping dynamic and static data, especially in long time series and multi-scale feature extraction.
A convolutional neural network (CNN) is used to extract local features of track irregularities, and combined with a Transformer network to capture time series relationships. The model is optimized through data preprocessing and a multi-step mapping mechanism, a dynamic and static data mapping model is constructed, and a mask mechanism is used to process sequences of unequal lengths.
It improves the frequency and accuracy of track irregularity detection, reduces labor costs, ensures the safe operation of high-speed railways, and achieves higher prediction accuracy.
Smart Images

Figure CN120705469A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of railway tracks, and in particular to a method for mapping dynamic and static track irregularities. Background Art
[0002] Dynamic and static track irregularity data represent the actual track geometry. Due to differences in measurement principles, the two measurement results differ. When a train passes, the track bears the vehicle load. If there are defects in the track substructure, the dynamic irregularity measurement results may differ significantly from the static irregularity measurement. This is because the wheel-rail contact force generates additional vibration, and defects cause additional deformation of the track. Static irregularity is often measured using manual stringing or track inspection trolleys. Within a short range, the rails and sleepers do not bend due to deformation and defects in the substructure. Therefore, conditions such as empty sleepers cannot be accurately reflected in static irregularity. They can only reflect the uneven residual deformation of the track accumulated over a long period of time. In addition, the frequency of static data collection is less than that of dynamic data.
[0003] Multilayer Perceptron (MLP) is a classic neural network structure, which is widely used in tasks such as classification and regression, and is often used as a basic model for deep learning. MLP consists of different layers, which can be divided into three types according to their functions: input layer, hidden layer, and output layer. Each layer consists of several neurons, and the relationship between different layers is connected by weight matrices. The input layer accepts data, and each neuron corresponds to a dimension of the feature, such as the left high and low data in track unevenness; the hidden layer uses an activation function to perform a nonlinear transformation on the received data, and increases the dimension of the data to extract features. Specifically, it may be the mean difference, peak value, cross rate, spectrum and other features of track unevenness. Of course, these features are not all explicit features. As the number of hidden layers increases and the number of neurons increases, the extracted features will gradually change from concrete to abstract. The output layer outputs the corresponding predicted unevenness data, such as Figure 1 shown.
[0004] The input data is passed layer by layer through forward propagation, and each layer is transformed using the following formula:
[0005] u i =f(w i x i +b i )
[0006] where u i represents the output of layer i, w i represents the weight matrix of the i-th layer, b irepresents the bias vector, and f is the activation function (such as ReLu, Sigmoid, etc.), which transforms the equation from linear transformation to nonlinear transformation. The data is forward propagated and the predicted value is output, but the w of the MLP network is i Initially, the matrix is a random unknown matrix, so the predicted values differ significantly from the actual values. Therefore, training is performed using the backpropagation algorithm. The difference between the predicted and actual values is typically represented by a loss function. The gradient of each layer is calculated using the chain rule, and the model parameters are updated using an optimization algorithm (such as SGD or Adam).
[0007] The MLP model is gradually being replaced by other models in complex tasks. For example, compared with the CNN model, it cannot capture the local features and spatial structure of the input data. Compared with the Transformer model, it lacks the ability to model sequence data. However, as the cornerstone of neural networks, it has laid a theoretical foundation for subsequent more complex neural network architectures through its simple structure and powerful expression capabilities.
[0008] Recurrent Neural Network (RNN) is a neural network that specializes in processing sequence data. Unlike traditional feedforward neural networks, RNN has a "memory" function that can capture the temporal dependencies in the sequence. RNN contains an input layer, a hidden layer, and an output layer. However, unlike MLP, RNN is special in that a hidden state is set in the network, which is transferred and updated between different time steps to capture the dependencies in the sequence. The hidden layer not only accepts input at the current moment, but also accepts hidden states from more previous moments as input. The hidden state is transferred between time steps to form a "memory." In other words, when performing prediction tasks, not only the state at the most recent moment is considered, but also the state at previous times. Therefore, it has a natural advantage when processing time series, language and text, etc.
[0009] The basic structure of RNN is as follows Figure 2 As shown, W X Represents the weight matrix parameter of the current input, W H Represents the shared weight matrix of the current moment in the recurrent layer (determined by the time steps before the current moment), W Y Represents the output weight matrix parameters corresponding to the output layer. The RNN network accepts the current input x t , combined with the hidden layer state h at the previous moment t-1 To calculate the current moment y t , and update the hidden layer state h according to the calculation results t , followed by the hidden layer h t will be the input x of the next time step t+1 Take them together as input, and calculate y at the next moment t+1and the hidden layer state h t+1 Therefore, the latest input will be affected by all previous inputs, but as the input sequence increases, the influence of the initial data will gradually decrease. The specific formula is as follows:
[0010] h t =f(W H h t-1 +W X x t +b h )
[0011] y t =g(W y h t +b y )
[0012] Where f and g represent the corresponding activation functions, b h and b y Corresponding to the bias term between hidden state and output.
[0013] Although RNNs are good at processing sequential data, they still have some problems when faced with very long sequences (such as data on uneven tracks), mainly including the two cases of gradient vanishing and gradient exploding. Gradient vanishing means that when the gradient is passed layer by layer between time steps, the gradient value gradually shrinks to zero with back propagation, making it impossible for the model to learn long-term dependencies; while gradient exploding means that the gradient value gradually increases to a very large value with back propagation, causing the model parameter update to fluctuate greatly. Both problems are due to the fact that the gradient is obtained by multiplying the previous gradient and the Jacobian matrix of the current time step. Therefore, if the Jacobian matrix is too large or too small, it will show exponential explosion or decay. At the same time, the activation function will also have a significant impact on the gradient problem. In addition, the hidden state formula of RNN cannot explicitly control which information needs to be retained and which information can be discarded. As the time step increases, the early information will gradually be forgotten, making it difficult for the model to remember distant information.
[0014] In order to solve the limitations of RNN networks, domestic and foreign scholars have formed a new recurrent neural network variant, the Long Short-Term Memory (LSTM) model, by adding gating mechanisms and cell states. The LSTM model improves the traditional RNN network, such as Figure 3 As shown. C t-1 and C tThis horizontal line represents the cell state, which is the "transmission belt" of long-term memory. Since there are no updates to the weight coefficients and only a small amount of linear interaction, it is easy for information to flow and remain unchanged. LSTM introduces a structure called a "gate" to control the increase and decrease of information in the cell state. The gate unit contains a sigmoid network layer and a pointwise multiplication operation. Specifically, there are three units: "forget gate", "input gate", and "output gate". The "forget gate" determines how much old information is retained in the cell state, the "input gate" determines what new information is added to the cell state, and the "output gate" generates a hidden state based on the current cell state and combines long-term and short-term memory. The specific formula is as follows:
[0015] f t =σ(W f ·[h t-1 , x t ]+b f )
[0016] i t =σ(W i ·[h t-1 , x t ]+b i )
[0017]
[0018] Where W f 、W i is the weight matrix corresponding to the gate unit, f t It is a vector between 0 and 1, indicating the proportion of information retained in the cell state, and cleaning up unimportant information in the cell state. Represents the candidate value, which is determined by the input data x t Transformed, the input gate passes through i t Determine which parts of the candidate values need to be written into the cell state while ignoring irrelevant noise. As shown in the following formula, O t is the output of the output gate, the hidden state h t It is the final output of LSTM to the outside world and is used to generate the final prediction result:
[0019] O t =σ(W o ·[h t-1 , x t ]+b o )
[0020] h t =O t tanh(C t )
[0021] With the support of cell states, the LSTM network's gradient will not explode or disappear, and it can remember information from longer sequences. It is suitable for most long time sequences, but it still has defects when facing uneven tracks.
[0022] (1) Although it is superior to traditional RNN in terms of long-term sequence dependency, when dealing with very long sequence problems (such as track unevenness data of several thousand meters), it becomes difficult to capture the dependencies between sequences and it is difficult to capture the complex patterns within the range. At the same time, it will cause the number of LSTM parameters to be large and the computational complexity to increase linearly.
[0023] (2) LSTM is only applicable to sequence features at a single scale and is difficult to capture multi-scale information at the same time. It mainly focuses on the overall pattern of the sequence and has a weak ability to extract local features. When considering track irregularities, both time and mileage information should be considered simultaneously.
[0024] (3) The internal state of LSTM (cell state and hidden state) is a black box and difficult to explain intuitively. When hyperparameters need to be adjusted, physical features cannot be used to improve them. Summary of the Invention
[0025] In order to solve the problems existing in the prior art, the purpose of the present invention is to provide a dynamic and static track irregularity mapping method. The present invention can not only optimize the existing detection technology and increase the detection frequency, but also further improve the detection efficiency and ensure the safe operation of high-speed railways.
[0026] To achieve the above object, the present invention adopts a technical solution: a dynamic and static track irregularity mapping method, comprising the following steps:
[0027] Step 1: Collect dynamic and static data and preprocess them;
[0028] Step 2: The preprocessed data is used as input to the dynamic-static data mapping model. The dynamic data is sorted according to the mileage and time dimensions, and a CNN network is used to extract local features based on mileage.
[0029] Step 3: Rely on the transformer network to capture time series relationships, thereby achieving higher prediction accuracy.
[0030] As a further improvement of the present invention, in step 1, the preprocessing includes filtering, data reconstruction, mileage alignment, standardization, logarithmic transformation and normalization.
[0031] As a further improvement of the present invention, the step 2 is specifically as follows:
[0032] Assume there is a two-dimensional image input X, the convolution kernel is K, and a two-dimensional convolution operation is performed on the input to output the feature map Z, as shown in the following formula:
[0033]
[0034] Where Z[i, j] is the (i, j)th position on the output feature map, X[i+m, j+n] represents the local area of the input map where the operation is performed, K[m, n] represents the weight of the convolution kernel, and b is the bias term. The entire process is repeated on the input image to eventually generate a feature map that represents the convolution result of the convolution kernel on the input map. The convolution process also includes stride and padding. Stride refers to the step size of the convolution kernel movement, and padding refers to the zero padding at the edge of the input. The two together control the size of the output map, thereby obtaining appropriate features.
[0035] The convolution result is processed by a nonlinear activation function to enhance the expressiveness of the model. The feature map is then downsampled to reduce the spatial resolution and overfitting risk while retaining important feature information, as shown in the following formula:
[0036] P[i, j] = max m,n∈window A[i·s+m,j·s+n]
[0037] They are maximum pooling and average pooling respectively. i·s and j·s determine the starting position of the pooling window in the feature map, where s is the step size; m and n represent the offsets of traversal within the pooling window;
[0038] Finally, the pooled data is flattened and the extracted features are converted into the final result through the fully connected layer.
[0039] As a further improvement of the present invention, the dynamic and static data mapping model includes a CNN network, a position encoding, a multi-head attention mechanism, an FFN network and a fully connected layer; the CNN network, the position encoding, and the multi-head attention mechanism are used to perform feature extraction and encoding during the mapping process, and the encoder layer is composed of a multi-head attention and a feedforward network layer, as follows:
[0040] LayerNorm(X+MultiHeadAttention(X))
[0041] LayerNorm(X+FeedForward(X))
[0042] The output of each layer is residually connected and normalized, and six stacked encoders are connected to a fully connected layer.
[0043] As a further improvement of the present invention, when constructing the mapping sample, the data surrounding the target dynamic data is included in the mapping relationship. To ensure the correspondence of the data, the adjacent areas of the target dynamic data are covered by a mask mechanism, and only the central area is retained to achieve sequence prediction of unequal lengths, and a multi-step matching mechanism is constructed to include the dynamic data corresponding to the target static data into the mapping sample.
[0044] The beneficial effects of the present invention are:
[0045] In actual engineering maintenance, static data peaks are still one of the inspection indicators. Therefore, by establishing a mapping relationship between dynamic and static track irregularities based on measured data, the reliance on static irregularity measurements can be reduced, reducing labor costs and thus the cost of data collection. At the same time, this method can not only optimize existing detection technology and increase detection frequency, but also further improve detection efficiency and ensure the safe operation of high-speed railways. The present invention proposes a time-mileage multi-dimensional model to predict static irregularities based on measured dynamic irregularity data. Based on precisely aligned dynamic and static data, a CNN network is used to extract local features at mileage, while a transformer network is relied upon to capture time series relationships, thereby achieving higher prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of a simple MLP network structure;
[0047] Figure 2 A schematic diagram of a simple RNN infrastructure;
[0048] Figure 3 Schematic diagram of LSTM model;
[0049] Figure 4 Schematic diagram of the Transformer model structure using multi-head attention;
[0050] Figure 5 Schematic diagram of the self-attention mechanism;
[0051] Figure 6 Schematic diagram of the encoder-decoder attention mechanism;
[0052] Figure 7 Schematic diagram of padding mask (left) and causal mask (right);
[0053] Figure 8 Schematic diagram of FFN structure;
[0054] Figure 9 This is a comparison diagram of the dynamic and static left height irregularity data at the beam end (top) and mid-span (bottom) in an embodiment of the present invention;
[0055] Figure 10 The comparison diagram of the dynamic and static left height irregularity data after preprocessing in the embodiment of the present invention is the beam end (top) and the mid-span (bottom);
[0056] Figure 11 Schematic diagram of multi-dimensional dynamic and static data of the same shape in an embodiment of the present invention;
[0057] Figure 12 Schematic diagram of the mapping relationship between dynamic and static data in an embodiment of the present invention;
[0058] Figure 13 Schematic diagram of the structure of the mapping model in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0060] Example
[0061] like Figure 1 As shown in Figure 1, a dynamic and static track irregularity mapping method is presented. First, the convolutional neural network (CNN) and Transformer model are explained:
[0062] 1. Convolutional Neural Network (CNN):
[0063] Convolutional Neural Networks (CNNs) are a type of model that specializes in processing multi-layered structures and have achieved remarkable results in fields such as computational vision, images, sound, and natural language. By mimicking the receptive field of the human visual system, they can automatically and effectively learn spatial hierarchical features from data. A convolutional layer consists of several core components: an input layer, a convolutional layer, an activation function, a pooling layer, a fully connected layer, and an output layer. The convolutional layer is the most significant feature that distinguishes it from other networks. The convolution operation involves element-wise multiplication of a window with different weights for different regions with the image and then adding the result. The window can be viewed as a specific filter or convolution kernel. This operation is called "convolution" and is primarily used to extract features.
[0064] Assume there is a two-dimensional image input X, the convolution kernel is K, and a two-dimensional convolution operation is performed on the input to output the feature map Z, as shown in the following formula:
[0065]
[0066] Where Z[i, j] is the (i, j)th position on the output feature map, X[i+m, j+n] represents the local region of the input image where the operation is performed, K[m, n] represents the weight of the convolution kernel, and b is the bias term. This entire process is repeated on the input image, ultimately generating a feature map representing the convolution result of the convolution kernel on the input image. The convolution process also involves stride and padding. Stride refers to the step size of the convolution kernel, while padding refers to the addition of zeros to the edges of the input. Together, these two factors control the size of the output image, thereby obtaining appropriate features.
[0067] The convolution result is usually processed by a nonlinear activation function to enhance the expressiveness of the model. The feature map is then downsampled to reduce the spatial resolution and overfitting risk while retaining important feature information, as shown in the following formula:
[0068] P[i, j] = max m,n∈window A[i·s+m,j·s+n]
[0069] They are maximum pooling and average pooling respectively. i·s and j·s determine the starting position of the pooling window in the feature map, where s is the step size; m and n represent the offsets of traversal within the pooling window.
[0070] Finally, the pooled data is flattened and the extracted features are converted into the final result through the fully connected layer.
[0071] Convolutional neural networks (CNNs) have significant advantages in processing track irregularity data due to their ability to extract local features. The laser sensors on track inspection vehicles are fixed to specific locations on the vehicle body, and there is a spatiotemporal coupling effect during the actual inspection process: when the train load acts on the rails, causing instantaneous deformation, the laser sensors are unable to capture this deformation in real time due to differences in spatial position. By the time the sensors move to the detectable area, the recorded state of the rail deformation has already been delayed for a certain period of time. Due to the continuity of the track structure, this deformation propagates along the track in the form of waves, forming a feature distribution with spatiotemporal correlation. This characteristic is highly consistent with the CNN mechanism of extracting spatial features through local receptive fields, enabling it to effectively identify the waveform propagation characteristics of track irregularities. This is why this model was chosen to extract local features from track data.
[0072] 2. Transformer model:
[0073] The Transformer model is a neural network architecture based on the self-attention mechanism. It was first proposed by a Google team for sequence-to-sequence (Seq2Seq) tasks, such as machine translation and question-answering systems. It has been a huge success since its introduction, and popular industry projects such as ChatGPT and DeepSeek are based on this model. The core innovations of the Transformer network lie in its encoder-decoder structure, as well as innovations such as the attention mechanism and mask padding. These features significantly differentiate it from traditional neural networks (RNNs and CNNs) and are key to modeling sequence mapping relationships and improving computational performance.
[0074] The Transformer model mainly consists of input embedding, position encoding, encoder, decoder, and output layer. The input of Transformer is a word sequence. First, each word is mapped to a vector of fixed dimension through the embedding layer, so that each word is a "unique" vector; the Transformer model has no sequential processing capabilities, and its core mechanism (such as the self-attention mechanism) is calculated in parallel. It cannot directly perceive the sequential information of the input sequence. When facing a time series, it is necessary to position encode the data to obtain the specific position information of each data. The common practice is to generate position encoding based on sine and cosine functions. By calculating the mapping values of sine and cosine functions of different frequencies at each position, a position-related vector is generated for each input element. After obtaining the result, the position information is integrated into the target data by vector addition. The formula is as follows:
[0075]
[0076] Where t represents the position of the current value in the sequence, is the i-th element in the position vector, and d_model is the dimension of the embedding vector. After position encoding, a multi-head attention mechanism is used to process the data, enhancing the model's understanding of the relationships between different vectors.
[0077] The attention mechanism is the most essential feature that distinguishes transformer neural networks from other network models. Attention is similar to how the human eye observes the world: when observing something, it tends to focus more on the key points and only partially pay attention to less important things. The single-head attention mechanism divides the input data into Q, K, and V, as shown below:
[0078] Q=Linear(data pos )=data pos W Q
[0079] K=Linear(datapos )=data pos W K
[0080] V=Linear(data pos )=data pos W V
[0081]
[0082] Where Q, K, and V are query matrix (Query), key matrix (Key), and value matrix (Value) respectively; data pos is the vector after position coding, W Q 、W K 、W V is the linear transformation matrix. Softmax is the normalization function, n head is the number of attention heads.
[0083] The above is a single-head attention mechanism, but the general Transformer model uses multi-head attention to calculate data weights, that is, single-head attention is used on multi-dimensional data. In other words, there are multiple dimensions Q, K, and V that calculate weights independently in parallel. Different dimensions focus on different aspects, so W Q 、W K 、W V Different features are extracted, and the extracted features are also different. Finally, the different results are spliced and output through the linear layer. The structure is as follows Figure 4 As shown in Figure 2, this structure is an extension of the attention mechanism and can be viewed as multiple single-head attention modules operating in parallel, with each attention head focusing on a different part of the input data to extract diverse feature information. This design also effectively balances the potential bias introduced by a single attention mechanism, ensuring the preservation of data features and thus improving the overall performance of the model.
[0084] The self-attention mechanism establishes global dependencies at different positions in the sequence, thereby capturing the correlation between vectors and using this as the core to establish the connection between input and output, such as Figure 5 As shown in the figure, a i Represents each element in the input sequence, and the arrows represent the calculation of the correlation between the current element and other elements, generating a set of attention weights. These weights determine the importance of other elements to this element. Finally, the self-attention mechanism will sum these weights to generate a new representation that integrates the information of all elements in the sequence.
[0085] The self-attention mechanism is used to capture the dependencies between input elements and learn the transformation rules of internal data. For different target sequences, it is necessary to learn the mapping relationship between them and the input sequence. Therefore, the encoder-decoder attention mechanism exists in the transformer structure, which is used to establish a relationship between the input sequence and the target sequence. Figure 6 As shown, b i is an element in the target sequence. Similar to the self-attention mechanism, an element a i Will calculate b1-b n The correlation between them is used to generate the corresponding weight matrix.
[0086] Each input of the network model is of fixed length, but in reality the input is not a sequence of equal length. The model cannot read data of different lengths. The introduction of the masking mechanism allows the model to effectively deal with sequences of unequal lengths. The principle is to use negative infinite weight coefficients to fill the sequence to make the input data of equal length, such as Figure 7 When nonlinear mapping is introduced, since the padding data in the sequence is negative infinity, the padded sequence will not be assigned a weight and will not affect the attention matrix of the model.
[0087] In addition, in autoregressive tasks, the model gradually outputs the target sequence. Therefore, each moment can only rely on previously generated words and cannot rely on future words. (This is also true for track irregularities. The future waveform is the track irregularities in the distance and has nothing to do with the irregularities in the area.) Masking ensures that the output data is logical by shielding information from future time steps.
[0088] After the input sequence passes through the attention mechanism, it is re-encoded into a new sequence, where the vector at each position incorporates the contextual information of the entire sequence. To ensure that local information is not lost and alleviate the problem of vanishing gradients, residual connections and normalization (Add & Norm) are introduced. Residual connections directly add the input to the output, forming a skip connection. This combines the original input data with the transformed output, allowing the model to capture both low-level and high-level information.
[0089] Despite the attention mechanism and residual connection processing, the data is still in a linear transformation. Therefore, in order to make the model have stronger learning ability, it is necessary to add another layer of feedforward neural network (FFN) to the data. Similar to the traditional neural network model, the Transformer model also uses FFN to introduce nonlinear transformation. FNN consists of a simple two-layer fully connected network. It first increases the dimension of the data to improve the expression ability, and then reduces the dimension to restore the original dimension, which is compatible with the subsequent decoder. The specific process is as follows Figure 8 shown.
[0090] The Transformer model generally uses a decoder-encoder approach for analysis. However, when mapping track irregularities between static and dynamic, the input and output sequences are of the same length, and the sequence length is relatively long. To reduce computational overhead, the model was modified so that the decoder uses a linear layer for output instead of a decoder.
[0091] 3. Model construction:
[0092] 3.1 Data preprocessing:
[0093] Both dynamic and static data contain spurious wavelengths, which affect the representation of the actual track geometry. These wavelengths include high-frequency noise and long waves, which affect the actual mapping relationship between dynamic and static data, resulting in reduced model accuracy and deviations in the predicted static waveform data. Directly inputting the original data into the model will interfere with the prediction results. Therefore, to improve the accuracy of the prediction model and reduce the complexity of model training, this embodiment filters, reconstructs, mileage-aligns, standardizes, and logarithmically transforms the dynamic and static data. The resulting data is used as model input. The unprocessed original images are shown in Figure 9, showing the beam end data and mid-span data, respectively.
[0094] Normalization is an important step in data preprocessing. Its purpose is to unify different input data into the same range. If the numerical differences between inputs are too large, the model will attach greater attention to larger features, which is contrary to the goal. We hope that the model can treat each element more equally. Z-Score normalization scales the data by mean and standard deviation. Similar to the logic of track irregularity data detection, Z-SCORE normalization is used, as shown in the following formula, where u is the mean, σ is the standard deviation, and Z is the scaled element:
[0095]
[0096] The data is corrected for mileage to ensure that the dynamic and static data can correspond one to one, and then normalized to process the data, such as Figure 10 shown.
[0097] 3.2 Multi-period matching and multi-step mapping mechanism:
[0098] Dynamic data is sampled by a track inspection vehicle, and static data is inferred by determining the specific coordinates of the track through GPS measurement of CPIII points. When using macro and micro mileage corrections, the static data is high-pass filtered and up-sampled so that the frequency of dynamic and static data is measured 4 times per meter, ensuring that the number of data in the two sequences is equal. Although the mileage correction has greatly matched the waveforms of dynamic and static data, due to the different collection methods and the limitations of the model itself, some peaks are still not completely aligned, which means that there is a "mileage lag" phenomenon in a short area, which can be seen from Figure 9 Therefore, in order to eliminate the influence of this difference on the dynamic and static mapping, the surrounding data of the current data should be taken into account during the mapping, so the data in the short area should be fused to ensure the integrity of the features.
[0099] Previous studies of dynamic and static mapping have not considered time as a key factor, instead analyzing only mileage. However, on long-span bridges, time has a significant impact on track irregularity. This is because the thermal expansion and contraction of large spans can cause waveform changes, and the bridge is also affected by seasonal wind loads. Therefore, time cannot be excluded as a factor.
[0100] The dynamic data is sorted by mileage and time and presented in a two-dimensional matrix. A two-dimensional convolution is performed on the data. Since the impact of the convolution kernel size on the model effect is not clear, a 3×3 convolution kernel is used for feature extraction. Since CNN will increase the dimension of the data feature channel, the static data is increased in dimension using a linear layer to ensure that the shape of the dynamic and static data is the same, such as Figure 11 shown.
[0101] Because the raw data incorporates the time dimension compared to previous studies, the input data has an (N×T) shape. After processing the data through convolutional and linear layers, the output data has an (N×T×D) shape, where N represents the number of times, T is the sequence length, and D is the number of channels in the feature dimension. The Transformer is a sequence-based modeling architecture, and its input is typically a two-dimensional tensor of (T×D) shape. Therefore, feature-extracted data cannot be directly imported into the model structure; dimensionality reduction is required. Pooling and flattening the data yields a two-dimensional tensor.
[0102] Both dynamic and static data increase their feature expression capabilities through dimensionality increase, reduce computational complexity through dimensionality reduction, and extract high-level features, but the mapping relationship between the two remains to be discussed. When constructing the mapping sample, the data surrounding the target dynamic data is considered to be included in the mapping relationship. Therefore, in order to ensure the corresponding relationship of the data, the adjacent areas of the target dynamic data are covered through a masking mechanism, and only the central area is retained to achieve sequence prediction of unequal lengths. Since the data has been fused in the initial stage, the target static data is considered to be the 5×5 data around the corresponding position when mapping the corresponding original dynamic data, such as Figure 12 shown.
[0103] The improved model for dynamic and static data mapping in this embodiment is mainly composed of CNN network, position encoding, multi-head attention mechanism, FFN network and fully connected layer. Among them, CNN network, position encoding and multi-head attention mechanism play the role of feature extraction and encoding in the mapping process. The encoder layer is composed of multi-head attention and feedforward network layer. The output of each layer will be residual connected and normalized, such as Figure 13 As shown on the right, some data in this structure is directly used in residual and normalization calculations without passing through the multi-head attention mechanism and feedforward neural network layers. This is to mitigate the impact of vanishing and exploding gradients, making deep Transformer models easier to train while retaining some of the original data's features and enhancing learning capabilities. The two-layer network in the encoder is shown in the equation.
[0104] LayerNorm(X+MultiHeadAttention(X))
[0105] LayerNorm(X+FeedForward(X))
[0106] The Transformer model generally adopts a decoder-encoder structure. However, in actual testing, it was found that if the decoder is selected as the decoder, the error will continue to accumulate along the mileage and time series, the loss will remain high, and the prediction results will be unsatisfactory. Therefore, this embodiment uses six stacked encoders and then uses a fully connected layer instead of a decoder to output the uneven mapping sequence, avoiding the linear accumulation of errors. The overall flow chart of the model is shown below. Figure 13 shown.
[0107] The improved Transformer model in this embodiment takes dynamic unevenness data as input and static unevenness data as output. After filtering, upsampling, data reconstruction, and mileage correction algorithms, the data achieves waveform alignment. To address the "mileage lag" issue, a CNN is used to fuse the dynamic data, ensuring that no point in the data is isolated. Masking is also introduced, creating a multi-step matching mechanism that incorporates the dynamic data surrounding the target static data into the mapping samples, thereby constructing a 1×1 matrix of static data to 3×3 matrix of dynamic data. However, the Transformer model structure dictates that it can only process data of the same shape. Therefore, negative infinity values are inserted around the static data labels, equal to the shape of the dynamic data, resulting in a 3×3 matrix of the same shape. However, since the mask is filled with negative infinity values, the model, through the softmax function, only focuses on the target static data and excludes the surrounding data. Because the data processed by the self-attention mechanism is unordered, a position vector must be added to the mapped data. Generally, the Transformer model adds position vectors to one-dimensional sequence data in time order, which cannot be directly applied to the two-dimensional vector data of this embodiment. Therefore, position vectors are added to the mileage and time dimensions respectively, and a single data contains two-dimensional position information. The mapped data will be passed through W Q 、W K 、W V A linear transformation is performed to obtain Q, K, and V, and the output is calculated through the self-attention mechanism. At the same time, part of the original data is retained and residually connected with the obtained output. Then, the same steps are used to pass through the feedforward network layer. The data passes through six stacked encoder layers and finally undergoes feature integration and compression in the fully connected layer. The final prediction sequence is obtained by the dimensionality change operation.
[0108] 4. Model parameter optimization:
[0109] 4.1 Model Evaluation Metrics
[0110] This embodiment uses four indicators to evaluate the quality of the model, namely, Mean Absolute Percentage Error (MAPE), Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Coefficient of determination (R 2 ), the formula is as shown in the formula, where n is the total number of data, y i is the i-th predicted value, y i is the actual true value of the ith value.
[0111]
[0112]
[0113] 4.2 Hyperparameter Analysis
[0114] During model construction, in addition to some unchangeable model parameters (such as positional encoding), many hyperparameters are also set. These hyperparameters are generally adjusted manually, and therefore can be subjective and affect the model's performance. To determine the optimal parameter configuration, we use a controlled variable analysis method to compare the model performance under different parameter settings. The main parameters involved include convolution kernel size, feature dimension, number of multi-head attention mechanism heads, and batch size.
[0115] The size of different convolution kernels directly affects the convolutional network's ability to extract features. It determines the "receptive field" of each convolution operation, that is, the size of the input area that the network can see. Small convolution kernels tend to extract local features, while large convolution kernels tend to extract the overall features of the data. The purpose of using convolution on the data in this embodiment is to reduce the waveform error caused by "mileage lag", so smaller convolution kernels meet the requirements of the article. For the extraction of overall features, the Transformer structure is more often used. If the convolution kernel is too large, it will lead to the loss of detailed information and the training speed will be slower. The statistical results are shown in the table. It can be seen that the best effect is achieved when the convolution kernel is 3×3, with an RMSE of 0.217. As the size of the convolution kernel increases, the error also increases. It can be seen that the various indicators in the table gradually deteriorate as the convolution kernel becomes larger. When it reaches 9×9, the data drops sharply, and the model training effect is the worst.
[0116]
[0117] In addition to the convolution kernel, feature dimension also influences the model's learning performance. Feature dimension is the dimension used to transform the input into an embedding vector. Similar to the convolution kernel, the model should choose an appropriate feature dimension. Too small a dimension will result in insufficient model expressiveness and excessive compression of the original information, leading to loss of original information and inaccurate predictions. Too large a dimension will make it difficult for the model to converge and also increase the risk of overfitting. To this end, we analyzed different numbers of dimensions, and the conclusions are shown in the table below. It can be seen that a feature dimension of 128 best meets the prediction requirements of this paper, while a dimension of 512 results in almost zero predictive power.
[0118]
[0119] Feature dimension and the number of attention heads are two closely related hyperparameters in the transformer model, which together determine the dimension size d of each head. h, as shown in the formula, d is the size of the feature dimension and h is the number of attention heads:
[0120]
[0121] Each attention head can independently learn the characteristics of the target data, and through multi-head parallel processing, long-range dependencies can be better understood. Too many attention heads will reduce the model's generalization ability and increase computational costs; too few will lead to insufficient expressiveness. Therefore, according to the feature dimensions obtained, different numbers of attention heads are analyzed. In practice, it is found that the number of attention heads does not significantly affect the performance of the model. This may be due to redundancy between attention heads, as shown in the table below. However, increasing the number of attention heads will significantly increase computational costs and reduce model training speed. Therefore, the number of attention heads is selected as 4 in this embodiment.
[0122]
[0123] In neural network models, batch size is extremely important, directly impacting the model's training efficiency, convergence speed, and ultimate performance. When training neural networks, due to the large dataset size, it's impossible to calculate the loss for the entire dataset with each parameter update. Therefore, the data is divided into small batches, and the loss gradient is calculated using this small batch for each parameter update. Batch size refers to the number of data samples processed simultaneously by the model during a training session. Therefore, choosing an appropriate batch size can reduce computational overhead while balancing stability and generalization. The calculation results for different batch size metrics are shown in the table below.
[0124]
[0125] This example investigates a dynamic and static track irregularity mapping model for long-span high-speed railway bridges. Starting with the fundamental neural network architecture (MLP), this paper gradually explains why traditional neural network models such as MLP, RNN, and LSTM are not suitable for this data. Because the raw data is discrete, random, and has long sequences, it undergoes preprocessing through filtering, reconstruction, mileage correction, and normalization to meet the model's input requirements. Furthermore, a CNN is employed to extract local features, while a Transformer model is employed to analyze the dynamic and static sequence relationships, and corresponding improvements are made to the model.
[0126] To address the "mileage lag" problem, this example utilizes the mask mechanism in the Transformer model to complete multi-step mapping and achieve complete model construction. Furthermore, the following key parameter configurations are determined through hyperparameter comparison to optimize model performance: CNN uses a 3×3 convolution kernel, the feature vector size is set to 128, the number of attention heads is set to 4, and the batch size is set to 4. These parameters optimize the model and minimize the model error, with the RMSE, MAPE, and R being the lowest at 0.065, 0.347, and 0.045, respectively. 2 It is 0.989, which shows that the improved transformer model in this paper can effectively capture the corresponding dependency between dynamic and static data and perform high-precision modeling based on the time dimension.
[0127] The above-described embodiments merely represent specific implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, and all such variations and improvements fall within the scope of protection of the present invention.
Claims
1. A method for mapping dynamic and static track irregularities, characterized in that: The following steps are involved: Step 1: Collect dynamic and static data and preprocess them; Step 2: The preprocessed data is used as input to the dynamic-static data mapping model. The dynamic data is sorted according to the mileage and time dimensions, and a CNN network is used to extract local features based on mileage. Step 3: Rely on the transformer network to capture time series relationships, thereby achieving higher prediction accuracy.
2. The method for mapping dynamic and static track irregularities according to claim 1, characterized in that: In step 1, the preprocessing includes filtering, data reconstruction, mileage alignment, standardization, logarithmic transformation and normalization.
3. The method for mapping dynamic and static track irregularities according to claim 1, characterized in that: The step 2 is specifically as follows: Assume there is a two-dimensional image input X, the convolution kernel is K, and a two-dimensional convolution operation is performed on the input to output the feature map Z, as shown in the following formula: Where Z[i, j] is the (i, j)th position on the output feature map, X[i+m, j+n] represents the local area of the input map where the operation is performed, K[m, n] represents the weight of the convolution kernel, and b is the bias term. The entire process is repeated on the input image to eventually generate a feature map that represents the convolution result of the convolution kernel on the input map. The convolution process also includes stride and padding. Stride refers to the step size of the convolution kernel movement, and padding refers to the zero padding at the edge of the input. The two together control the size of the output map, thereby obtaining appropriate features. The convolution result is processed by a nonlinear activation function to enhance the expressiveness of the model. The feature map is then downsampled to reduce the spatial resolution and overfitting risk while retaining important feature information, as shown in the following formula: P[i,j]=max m,n∈wind o w A[i·s+m,j·s+n] They are maximum pooling and average pooling respectively. i·s and j·s determine the starting position of the pooling window in the feature map, where s is the step size; m and n represent the offsets of traversal within the pooling window; Finally, the pooled data is flattened and the extracted features are converted into the final result through the fully connected layer.
4. The method for mapping dynamic and static track irregularities according to claim 3, characterized in that: The dynamic and static data mapping model includes a CNN network, position encoding, a multi-head attention mechanism, an FFN network, and a fully connected layer; the CNN network, position encoding, and multi-head attention mechanism are used to extract and encode features during the mapping process. The encoder layer is composed of a multi-head attention and a feedforward network layer, as follows: LayerNorm(X+MultiHeadAttention(X)) LayerNorm(X+FeedForward(X)) The output of each layer is residually connected and normalized, and six stacked encoders are connected to a fully connected layer.
5. The method for mapping dynamic and static track irregularities according to claim 4, characterized in that: When constructing the mapping sample, the data surrounding the target dynamic data is included in the mapping relationship. To ensure the correspondence of the data, the adjacent areas of the target dynamic data are covered through the mask mechanism, and only the central area is retained to achieve sequence prediction of unequal lengths. A multi-step matching mechanism is constructed to include the dynamic data corresponding to the target static data into the mapping sample.
Citation Information
Cited By
Heavy-load track irregularity mean value prediction method and device
CN120911320A