Dynamic response prediction method of high-speed train based on physical driving
The graph structured and multi-scale analysis of the axle coupling system is performed through the graph neural network model, which solves the problems of low computational efficiency and insufficient prediction accuracy of the axle coupling system, and realizes efficient, accurate prediction and real-time monitoring of the dynamic response of high-speed trains.
Patent Information
- Application Number
- CN202510477056.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art has low computational efficiency and slow processing speed in axle coupling system, and traditional models are difficult to accurately capture complex nonlinear behavior, resulting in reduced prediction accuracy and reliability.
The graph neural network model is used to graphically structure the axle coupling system. Through the multi-head attention mechanism and graph attention network, combined with multi-scale analysis technology, the dynamic response of high-speed trains is predicted and the nonlinear and time-varying characteristics of the system are captured.
It improves computing efficiency and prediction accuracy, can monitor the dynamic response of the axle coupling system in real time, supports bridge health monitoring and fault diagnosis, and optimizes train operation scheduling and bridge maintenance.
Smart Images

Figure CN120409215A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle-bridge coupling systems, and particularly to a method for predicting the dynamic response of high-speed trains based on physical driving. Background Art
[0002] The vehicle-bridge coupling system often faces problems of low computational efficiency and slow processing speed, mainly due to its highly nonlinear and complex multi-degree-of-freedom characteristics. The vehicle-bridge coupling system not only involves the interaction between the train and the bridge, but also includes multiple dynamic factors, such as the train's traveling speed, the natural frequency of the bridge, road surface unevenness, load distribution, etc. The change of each factor may cause a significant change in the dynamic response of the system, thus increasing the complexity of the analysis.
[0003] In addition, the modeling of the vehicle-bridge coupling system usually requires detailed geometric descriptions and complex physical properties, such as the contact mechanics between the train wheels and the track, the vibration modes of the bridge structure, etc. Traditional numerical methods, such as finite element analysis and multi-body dynamics simulation, although able to provide highly accurate solutions, often require extremely large computational resources and time. This makes real-time analysis and prediction very difficult, especially in application scenarios that require a large number of simulation calculations, such as design optimization, bridge health monitoring, traffic prediction, etc.
[0004] In addition, with the aging of the bridge structure and the increase in train loads, the long-term dynamic behavior of the vehicle-bridge coupling system has become increasingly difficult to predict. Traditional models usually rely on linear assumptions or simplified physical models, and these assumptions often cannot accurately capture the complex nonlinear behavior of the system, resulting in reduced prediction accuracy and reliability. Against this background, how to improve computational efficiency, speed up processing speed and at the same time maintain high prediction accuracy has become a key challenge in the research of vehicle-bridge coupling systems.
[0005] To address this problem, recent research has started to adopt advanced technologies such as deep learning and graph neural networks (GNNs). These methods model the nonlinear dynamic behavior of the vehicle-bridge coupling system in a data-driven manner and can effectively reduce the computational cost and improve the prediction efficiency. However, although these methods have shown good results in some scenarios, they still face problems such as poor generalization ability and insufficient computational efficiency when dealing with long-time series prediction and complex scenarios. Especially in practical applications, the accuracy and robustness of long-term prediction are often limited by the quality of training data, the complexity of the model, and the limitation of computational resources. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method for predicting the dynamic response of a high-speed train based on physical driving. Using a graph neural network model to predict the dynamic response of a high-speed train can not only capture the complex interactions between various components in the vehicle-bridge coupling system, but also significantly improve the calculation efficiency, meeting the dual requirements of real-time and accuracy in engineering applications.
[0007] The technical solution of the present invention is as follows:
[0008] A method for predicting the dynamic response of a high-speed train based on physical driving specifically includes the following steps:
[0009] (1) Graph-structurize the vehicle-bridge coupling system to generate a graph structure G = (V, E), where V represents the nodes in the graph structure G, and E represents the edges connecting the nodes in the graph structure G. The nodes in the graph structure G include train nodes, bridge nodes, and pier nodes in the vehicle-bridge coupling system. The data corresponding to the train nodes is driving data, the data corresponding to the bridge nodes is bridge span information, and the data corresponding to the pier nodes is seismic wave data;
[0010] (2) Project the vectors of all nodes at each moment t into the same dimension, and then splice them along the time dimension to synthesize the feature matrix of the entire graph structure;
[0011] (3) Construct a graph neural network model for predicting the dynamic response of a high-speed train. Specifically, first embed the feature vectors of all nodes in the graph structure into a unified dimensional space and downsample the embedded time series; then input the feature vectors corresponding to the original time series and the downsampled time series into the encoder respectively, extract the feature vectors through the multi-head attention mechanism, and then input the feature vectors corresponding to the encoded time series into the graph attention network, and splice them along the time series dimension to obtain the graph embedding result. After repeating the encoder and the graph attention network alternately N times, the encoding process is completed; finally, combine the output of the encoder with the input of the decoder for decoding, and then map the decoded output to the dimension of the observed time series through a linear layer to obtain the prediction result;
[0012] (4) Train the graph neural network model to obtain a trained graph neural network model, and then input the driving data, bridge span information, and seismic wave data in the vehicle-bridge coupling system into the trained graph neural network model for prediction to obtain the dynamic response of the high-speed train.
[0013] The driving data includes the acceleration value and displacement value of the train, and the feature vector of the train node is where, represents the acceleration value in the lateral direction of the train, represents the acceleration value in the longitudinal direction of the train, Represents the displacement value in the lateral direction of the train, Represents the displacement value in the longitudinal direction of the train; the bridge span information includes the acceleration value and displacement value of each span area of the bridge, and the feature vector of the bridge node is
[0014] Represents the acceleration value in the vertical direction of the bridge, Represents the acceleration value in the lateral direction of the bridge, Represents the acceleration value in the longitudinal direction of the bridge, Represents the displacement value in the vertical direction of the bridge, Represents the displacement value in the lateral direction of the bridge, Represents the displacement value in the longitudinal direction of the bridge, and the feature vector of the bridge node is initialized to zero; the feature vector of the pier node is Represents the acceleration value of the seismic wave in the vertical direction, Represents the acceleration value of the seismic wave in the lateral direction, Represents the acceleration value of the seismic wave in the longitudinal direction.
[0015] The vectors of all nodes at each moment t are projected into the same dimension through the linear layer Linear, and then concatenated in the time dimension to synthesize the feature matrix of the entire graph structure. The specific processing process is shown in the following formula (1), and formula (2) represents the projection and concatenation of the feature vectors of all moments in the temporal length T of the node input;
[0016]
[0017] In formula (1), Represents the feature vector of the train node at moment t, Represents the feature vector of the bridge node at moment t, Represents the feature vector of the pier node at moment t, Linear represents the linear layer, and Concat represents concatenation;
[0018]
[0019] In formula (2), EmbeddingLayer represents projection for all moments and then concatenation in the time dimension, x i Train Represents the feature vector of any train node, x i Bridge Represents the feature vector of any bridge node, x i Pier Represents the feature vector of any pier node, X graph Is the feature matrix of the entire graph structure.
[0020] Embed the feature vectors of all nodes in the graph structure into a unified dimensional space, and downsample the embedded time series. The specific processing process is shown in the following formula (3):
[0021] X down_sample = X graph [::2] (3);
[0022] Formula (3) represents downsampling the feature matrix X corresponding to the original time series with a slice time length of 2 steps graph to obtain the feature vector X corresponding to the downsampled time series down_sample .
[0023] Respectively input the feature vectors corresponding to the original time series and the downsampled time series into the encoder, and extract the feature vectors through the multi-head attention mechanism. The specific processing process is shown in the following formula (4) and formula (5):
[0024]
[0025] In formula (4) and formula (5), Encoder represents the encoder, mask represents the mask, and are both the outputs of the encoder;
[0026] Respectively input the feature vectors corresponding to the encoded time series into the graph attention network, and then splice them along the time series dimension to obtain the graph embedding result. The specific processing process is shown in the following formula (6) and formula (7);
[0027]
[0028] X dense = Concat(X' graph , X' down_sample ) (7);
[0029] In formula (6), A ij represents the adjacency matrix, e ij represents the edge connecting node v i and node v j in the graph structure G. GATs represents the graph attention network; in formula (7), Concat represents splicing.
[0030] The processing process of the encoder is as follows: first, the input time series are projected to the hidden layer dimension through three linear transformation matrices Wq, Wk, and Wv to obtain Q, K, and V; then Q and K are dot-producted and projected to an interval with a sum of one through a softmax activation function to obtain the attention weight, and the obtained attention weight is weighted and summed with V to obtain the attention output, as shown in the following formula (8):
[0031]
[0032] In formula (8), d k Represents the feature dimensions of Q and K, head i is the output of an attention head;
[0033] The multi-head attention mechanism is to use the attention mechanism of formula (8) simultaneously and splice the final results together;
[0034] Then the output of the multi-head attention mechanism X attention After passing through the first residual and layer normalization, the linear layer, and the second residual and layer normalization, it is input into the distillation layer, thus completing one layer of encoding, as shown in the following formula (9):
[0035]
[0036] In formula (9), X attention represents the output of the multi-head attention mechanism, X res1 Represents the result of a residual and layer normalization, LayerNorm represents layer normalization, X res2 Represents the result after quadratic residual and layer normalization, Linear represents the linear layer, X ditllblock Represents the output result of the distillation layer, that is, the output of one layer of encoding, Elu represents the ELU activation function, Conv1d represents one-dimensional convolution, MaxPool represents maximum pooling, and stride=2 represents the stride of maximum pooling is 2.
[0037] The processing process of the graph attention network is as follows: first, each node is initialized as a feature vector, and then the graph attention layer in the graph attention network aggregates the features of its neighboring nodes according to the attention coefficient of each node, as shown in the following formula (10), where the attention coefficient is The calculation process of is shown in the following formula (11):
[0038]
[0039] In formula (10) and formula (11), represents the output of the graph attention layer, σ is the sigmoid function, W represents the weight matrix, The feature vector of node i at time t, The feature vector of node j adjacent to node i at time t, The feature vector of node k adjacent to node i, The set of all nodes adjacent to node i, | represents the concatenation operation, a T The weight vector, and LeakyReLU represents the non-linear activation function.
[0040] The input of the decoder is constructed by an auxiliary learning method, and the calculation process is shown in Equation (12) below:
[0041]
[0042] In Equation (12), X auxiliary represents the auxiliary feature vector, X graph is the feature matrix corresponding to the original time series, L auxiliary_length is the length of the auxiliary learning, -1 represents indexing from the end of the vector forward, and the length of the indexing field is L auxiliary , the step size is 1, X0 represents the initial vector that has not been decoded, that is, the part where the vector is 0, L y represents the length of the initial vector that has not been decoded, N represents the number of nodes, D represents the feature dimension of each node, X dec is the input of the decoder.
[0043] Advantages of the present invention:
[0044] (1). By introducing a physics-driven model, the present invention collects real-time data of train nodes, bridge nodes, and pier nodes in the vehicle-bridge coupling system, accurately simulates the complex interaction between high-speed trains and bridges, not only considers the dynamic characteristics of the trains, but also comprehensively describes the dynamic factors such as the vibration characteristics of the bridges, road surface unevenness, and vehicle driving speed, ensuring that the non-linear behavior of the vehicle-bridge coupling system can be accurately captured.
[0045] (2). The present invention uses a graph neural network model to predict the dynamic response of high-speed trains. The various components (trains, bridges, and piers) of the vehicle-bridge coupling system are represented in a graph structure, and the influence between each part is dynamically adjusted through a graph attention mechanism, enabling the model to capture the non-linear and time-varying characteristics of the vehicle-bridge coupling system globally and locally. The graph attention mechanism can automatically weight the interaction effects between various components according to different input conditions, improving the adaptability of the vehicle-bridge coupling system to complex and changeable working conditions.
[0046] (3) The present invention adopts a multi-scale analysis technique to perform hierarchical analysis on the eigenvectors corresponding to the original time series and the downsampled time series, which can effectively capture the dynamic changes at different time scales, enabling the graph neural network model to learn and analyze data from multiple time perspectives, thereby enhancing its ability to understand time series data. A short time window can capture rapidly changing trends, while a longer time window helps to identify more stable or periodic patterns. By splicing the features from different time windows, the response laws of the vehicle-bridge coupling system in the short term and the long term can be better understood, and optimized prediction can be carried out by combining multi-scale information, which can not only accurately capture the rapid response in the short term but also effectively predict the steady-state behavior of the system in the long term, thus improving the accuracy and robustness of the prediction.
[0047] (4) The present invention not only realizes the prediction of the dynamic response of high-speed trains but also can analyze the response data of the vehicle-bridge coupling system for bridge health monitoring and fault diagnosis. By real-time monitoring the dynamic response of the vehicle-bridge coupling system, abnormal conditions in the structure can be detected in a timely manner and early warnings can be provided, thereby optimizing the train operation scheduling and bridge maintenance strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] See Figure 1 , a method for predicting the dynamic response of a high-speed train based on physical driving, specifically including the following steps:
[0051] (1) Perform graph structuring on the vehicle-bridge coupling system to generate a graph structure G=(V, E), where V represents the nodes in the graph structure G, and E represents the edges connecting the nodes in the graph structure G. The nodes in the graph structure G include train nodes, bridge nodes, and pier nodes in the vehicle-bridge coupling system. The data corresponding to the train nodes is train operation data, the data corresponding to the bridge nodes is bridge span information, and the data corresponding to the pier nodes is seismic wave data;
[0052] The train operation data includes the acceleration value and displacement value of the train, and the eigenvector of the train node is Among them, represents the acceleration value in the lateral direction of the train, Represents the acceleration value in the longitudinal direction of the train, Represents the displacement value in the lateral direction of the train, Represents the displacement value in the longitudinal direction of the train; The bridge span information includes the acceleration value and displacement value in each span area of the bridge, and the feature vector of the bridge node is
[0053] Represents the acceleration value in the vertical direction of the bridge, Represents the acceleration value in the lateral direction of the bridge, Represents the acceleration value in the longitudinal direction of the bridge, Represents the displacement value in the vertical direction of the bridge, Represents the displacement value in the lateral direction of the bridge, Represents the displacement value in the longitudinal direction of the bridge, and the feature vector of the bridge node is initialized to zero; The feature vector of the pier node is Represents the acceleration value of the seismic wave in the vertical direction, Represents the acceleration value of the seismic wave in the transverse direction, Represents the acceleration value of the seismic wave in the longitudinal direction;
[0054] (2) Project the vectors of all nodes at each moment t onto the same dimension through the linear layer Linear, and then concatenate them in the time dimension to synthesize the feature matrix of the entire graph structure. The specific processing process is shown in the following formula (1), and formula (2) represents the projection and concatenation of the feature vectors of all moments in the time sequence length T of the node input;
[0055]
[0056] In formula (1), Represents the feature vector of the train node at moment t, Represents the feature vector of the bridge node at moment t, Represents the feature vector of the pier node at moment t, Linear represents the linear layer, and Concat represents concatenation;
[0057]
[0058] In formula (2), EmbeddingLayer represents projection for all moments and then concatenation in the time dimension, x i Train Represents the feature vector of any train node, x i Bridge Represents the feature vector of any bridge node, x i Pier Represents the feature vector of any pier node, X graph Is the feature matrix of the entire graph structure;
[0059] (3) Construct a graph neural network model to predict the dynamic response of high-speed trains. The processing process of the graph neural network model is as follows:
[0060] S31. Embed the feature vectors of all nodes in the graph structure into a unified dimensional space, and downsample the embedded time series. The specific processing process is shown in the following formula (3):
[0061] X down_sample =X graph [::2] (3);
[0062] Formula (3) represents downsampling the feature matrix X corresponding to the original time series with a slice time length of step 2 graph to obtain the feature vector X corresponding to the downsampled time series down_sample ;
[0063] S32. Input the feature vectors corresponding to the original time series and the downsampled time series into the encoder respectively, and extract the feature vectors through the multi-head attention mechanism. The specific processing process is shown in the following formula (rac{4}) and formula (5):
[0064]
[0065] In formula (4) and formula (5), Encoder represents the encoder, mask represents the mask, and are both the outputs of the encoder;
[0066] The processing process of the encoder Encoder is as follows: First, project the input time series sequence through three linear transformation matrices Wq, Wk, and Wv to the hidden layer dimension to obtain Q, K, and V; then perform a dot product operation on Q and K, and project it to an interval with a sum of one through the softmax activation function to obtain the attention weight, and perform a weighted sum of the obtained attention weight and V to obtain the attention output. The specific process is shown in the following formula (8):
[0067]
[0068] k In formula (8), d i represents the feature dimension of Q and K, and head attention
[0069] is the output of one attention head;
[0070] The multi-head attention mechanism synchronously uses the attention mechanism of formula (8) and splices the final results together; attentionAfter passing through a first residual and layer normalization, a linear layer (Linear), and a second residual and layer normalization in sequence, it is input into the distillation layer, and thus one layer of encoding is completed, as shown in the following formula (9):
[0071]
[0072] In formula (9), X attention represents the output of the multi-head attention mechanism, X res1 represents the result after the first residual and layer normalization, LayerNorm represents layer normalization, X res2 represents the result after the second residual and layer normalization, Linear represents the linear layer, X ditllblock represents the output result of the distillation layer, that is, the output of one layer of encoding, Elu represents the ELU activation function, Conv1d represents one-dimensional convolution, MaxPool represents max pooling, and stride = 2 represents the stride of max pooling is 2;
[0073] Input the feature vectors corresponding to the encoded time series into the graph attention network, and then concatenate them along the time series dimension to obtain the graph embedding result. The processing process is specifically shown in the following formulas (6) and (7);
[0074]
[0075] X dense = Concat(X′ graph , X′ down_sample ) (7);
[0076] In formula (6), A ij represents the adjacency matrix, e ij represents the edge connecting node v i and node v j in the graph structure G, and GATs represents the graph attention network; in formula (7), Concat represents concatenation;
[0077] The processing process of the graph attention network (GATS) is as follows: First, initialize each node as a feature vector, and then the graph attention layer in the graph attention network aggregates the features of its neighbor nodes according to the attention coefficients of each node, as shown in the following formula (10), and the calculation process of the attention coefficient is specifically shown in the following formula (11):
[0078]
[0079] In formulas (10) and (11), represents the output of the graph attention layer, σ is the sigmod function, W represents the weight matrix, The feature vector of node i at time t, The feature vector of node j adjacent to node i at time t, The feature vector of node k adjacent to node i, The set of all nodes adjacent to node i, | represents the concatenation operation, a T The weight vector, LeakyReLU represents the non-linear activation function;
[0080] After the last encoder and the graph attention network are alternately repeated N times according to the above steps, the encoding process is completed;
[0081] S33. Combine the output X of the encoder dense with the input X of the decoder dec , perform decoding, and then map the decoding output to the dimension of the observed time series through the linear layer Linear to obtain the prediction result. The specific calculation process is shown in the following formula (13):
[0082]
[0083] In formula (13), mask represents the mask, and Linear means that each node is connected to a linear layer;
[0084] The input of the decoder is constructed by the auxiliary learning method. The calculation process is shown in the following formula (12):
[0085]
[0086] In formula (12), X auxiliary represents the auxiliary feature vector, X graph is the feature matrix corresponding to the original time series, L auxiliary_length is the length of the auxiliary learning, -1 represents indexing from the end of the vector forward, and the length of the indexing field is L auxiliary , the step size is 1, X0 represents the initial vector that has not been decoded, that is, the part where the vector is 0, L y represents the length of the initial vector that has not been decoded, N represents the number of nodes, D represents the feature dimension of each node, X dec is the input of the decoder;
[0087] (4) Train the graph neural network model to obtain a trained graph neural network model, and then input the driving data, bridge span information, and seismic wave data in the vehicle-bridge coupling system into the trained graph neural network model for prediction to obtain the dynamic response of the high-speed train.
[0088] Model performance analysis:
[0089] 1. To evaluate the prediction performance of the graph neural network model (GAttention) on the vehicle-bridge coupling system dataset and compare it with other models (LSTM, GNN, GATS, GATS+LSTM, GNBlock), all models are developed using PyTorch;
[0090] When training the graph neural network model, the dataset is divided into 60% for training and 40% for testing. The training conditions are as follows: the batch size is 32, the dropout rate is 0.2, the hidden layer dimension is 256, and the time lengths of the encoder and decoder are both 36. The neural network architecture of the graph neural network model consists of two layers, which are consistent with the encoder and decoder. The auxiliary learning length is set to 18, the 8-head multi-head attention mechanism is adopted, and the learning rate is 3e-5; all models are trained for 1000 epochs to ensure the stability of performance. The mean absolute error (MAE) is selected as the standard to evaluate the prediction accuracy and generalization ability of the model. The prediction performance (MAE) of all models in Table 1 will be compared under the same vehicle-bridge coupling system (TBCs) conditions. 3, 4, 5, 6, 7, 9, 10, and 11 in Table 1 represent the spans of the bridges, that is, the prediction is carried out on the TBCs datasets with different spans.
[0091] Table 1
[0092] Model 3 4 5 6 7 9 10 11 LSTM 0.194 0.214 0.237 0.232 0.193 0.214 0.232 0.239 GNN 0.302 0.326 0.380 0.376 0.302 0.336 0.367 0.381 GATS 0.186 0.195 0.224 0.221 0.178 0.199 0.204 0.211 GATS + LSTM 0.182 0.193 0.222 0.220 0.180 0.194 0.221 0.198 GNBlock 0.215 0.230 0.318 0.284 0.230 0.252 0.235 0.281 GAttention 0.095 0.097 0.115 0.122 0.107 0.108 0.117 0.118
[0093] As can be seen from Table 1, in the vehicle-bridge coupling system dataset with the same span, the graph attention network (GATs) shows better prediction performance than the traditional graph neural network (GNNs). Through the unique mechanism of dynamically selecting and updating node features, GATs can more flexibly and accurately adapt to the changes in the data, thus improving the prediction accuracy and efficiency. In addition, the GATtention model performs excellently in the long-term sequence prediction scenario. The GATtention model integrates an improved attention mechanism, which can effectively capture time-dependent relationships, especially showing strong capabilities when dealing with data with long time periods. By accurately understanding and analyzing time relationships, the GATtention model is particularly important in predicting the structural behavior of bridge spans, especially in applications with high requirements for reliability and precision.
[0094] II. To evaluate the generalization ability of the GAttention model, it will be trained on the TBCs datasets with 3 spans and 11 spans, and tested on the TBCs datasets containing 4, 5, 6, 7, 9, and 10 spans to demonstrate the model's generalization ability. MAE will continue to be used as the evaluation metric, as shown in Tables 2 and 3. Table 2 shows the MAE obtained by testing the model trained on the TBCs dataset with 3 spans, and Table 3 shows the MAE obtained by testing the model trained on the TBCs dataset with 11 spans.
[0095] Table 2
[0096] Model 4 5 6 7 9 10 LSTM 0.858 0.778 0.907 0.811 0.916 0.992 GNN 0.405 0.445 0.473 0.437 0.523 0.559 GATS 0.381 0.395 0.409 0.431 0.449 0.430 GATS + LSTM 0.449 0.412 0.450 0.462 0.497 0.405 GNBlock 0.391 0.384 0.387 0.388 0.330 0.362 GAttention 0.256 0.339 0.324 0.264 0.302 0.316
[0097] Table 3
[0098] Model 4 5 6 7 9 10 LSTM 0.839 0.810 0.779 0.610 0.575 0.383 GNN 0.414 0.451 0.455 0.407 0.447 0.421 GATS 0.344 0.334 0.312 0.306 0.350 0.257 GATS + LSTM 0.439 0.0.416 0.411 0.412 0.437 0.288 GNBlock 0.444 0.384 0.312 0.386 0.450 0.283 GAttention 0.347 0.365 0.341 0.294 0.302 0.216
[0099] As can be seen from Tables 2 and 3, the Graph Attention Networks (GATs) perform exceptionally well in generalization tasks. This performance is attributed to its powerful graph attention mechanism. It is worth noting that different attention heads focus on different aspects: some prioritize node attributes, while others emphasize neighboring nodes. This difference in the degree of aggregation of neighboring nodes reflects a close connection with the concept of stiffness. Such a mechanism enables GATs to effectively process complex graph-structured data and extract key features and patterns.
[0100] III. Due to the unique encoding characteristics of the graph structure, the hybrid training method allows for the simultaneous training of different vehicle-bridge coupling system structures, enabling the GAttention model to have scalable prediction capabilities. This allows the model to be extended to adapt to different vehicle-bridge coupling systems. The hybrid training method not only enhances the stability and accuracy of the model when processing long-term data but also effectively prevents overfitting that may be caused by over-focusing on a specific dataset. In addition, the hybrid training method promotes the mutual learning and fusion between different data features, enabling the GATtention model to exhibit superior performance in multi-task and multi-environment settings. The prediction performance (MAE) of various models under the hybrid training method is shown in Table 4 below:
[0101] Table 4
[0102] Model 3 4 5 6 7 9 10 11 LSTM 0.235 0247 0.282 0.279 0.264 0.290 0.296 0.308 GNN 0.341 0.370 0.396 0.398 0.352 0.386 0.393 0.388 GATS 0.209 0.241 0.246 0.249 0.242 0.276 0.263 0.266 GATS + LSTM 0.184 0.207 0.212 0.207 0.226 0.256 0.240 0.248 GNBlock 0.284 0.257 0.262 0.297 0326 0.246 0.280 0.258 GAttention 0.132 0.149 0.151 0.164 0.107 0.126 0.157 0140
[0103] As can be seen from Table 4, the GATtention model using the hybrid training method achieves better prediction accuracy and generalization ability on multiple TBCs datasets. This indicates that the hybrid training method effectively overcomes the limitations of traditional single-dataset training methods and provides a more robust and reliable solution for prediction tasks in complex data environments.
[0104] Through the prediction performance analysis of the above three items, the GATtention model of the present invention was evaluated. The results show that, supported by the hybrid training method, the GAttention model exhibits powerful long-term prediction performance and excellent generalization ability, effectively overcoming the challenges of accuracy and generalization in long-term sequence prediction.
[0105] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for predicting the dynamic response of a high-speed train based on physical driving, characterized in that: Specifically, it includes the following steps: (1) Structure the vehicle-bridge coupling system into a graph to generate a graph structure G = (V, E), where V represents the nodes in the graph structure G, and E represents the edges connecting the nodes in the graph structure G. The nodes in the graph structure G include train nodes, bridge nodes, and pier nodes in the vehicle-bridge coupling system. The data corresponding to the train nodes is train operation data, the data corresponding to the bridge nodes is bridge span information, and the data corresponding to the pier nodes is seismic wave data; (2) Project the vectors of all nodes at each moment t onto the same dimension, and then concatenate them along the time dimension to synthesize the feature matrix of the entire graph structure; (3) Construct a graph neural network model for predicting the dynamic response of high-speed trains. Specifically, first embed the feature vectors of all nodes in the graph structure into a unified dimensional space and downsample the embedded time series; then input the feature vectors corresponding to the original time series and the downsampled time series into the encoder respectively, extract the feature vectors through the multi-head attention mechanism, and then input the feature vectors corresponding to the encoded time series into the graph attention network and concatenate them along the time series dimension to obtain the graph embedding result. After repeating the encoder and the graph attention network alternately N times, the encoding process is completed; finally, combine the output of the encoder with the input of the decoder for decoding, and then map the decoded output to the dimension of the observed time series through a linear layer to obtain the prediction result; (4) Train the graph neural network model to obtain a trained graph neural network model, and then input the train operation data, bridge span information, and seismic wave data in the vehicle-bridge coupling system into the trained graph neural network model for prediction to obtain the dynamic response of the high-speed train.
2. The dynamic response prediction method of a high-speed train based on physical driving according to claim 1, wherein: The driving data described above includes the acceleration value and displacement value of the train, and the feature vector of the train node is wherein, represents the acceleration value of the train in the lateral direction, represents the acceleration value of the train in the longitudinal direction, represents the displacement value of the train in the lateral direction, represents the displacement value of the train in the longitudinal direction; the bridge span information described above includes the acceleration value and displacement value of each span area of the bridge, and the feature vector of the bridge node is represents the acceleration value of the bridge in the vertical direction, represents the acceleration value of the bridge in the lateral direction, represents the acceleration value of the bridge in the longitudinal direction, represents the displacement value of the bridge in the vertical direction, represents the displacement value of the bridge in the lateral direction, represents the displacement value of the bridge in the longitudinal direction, and the feature vector of the bridge node is initialized to zero; the feature vector of the pier node is represents the acceleration value of the seismic wave in the vertical direction, represents the acceleration value of the seismic wave in the lateral direction, represents the acceleration value of the seismic wave in the longitudinal direction.
3. A method for predicting the dynamic response of a high-speed train based on physical driving according to claim 1, characterized in that: The process of projecting the vectors of all nodes at each moment t onto the same dimension through a linear layer Linear and then concatenating them along the time dimension to synthesize the feature matrix of the entire graph structure is specifically shown in the following formula (1), and formula (2) represents projecting and concatenating the feature vectors at all moments in the time sequence length T of the node input; In formula (1), represents the feature vector of the train node at time t, represents the feature vector of the bridge node at time t, represents the feature vector of the pier node at time t, Linear represents the linear layer, and Concat represents concatenation; In Equation (2), EmbeddingLayer represents projecting at all times and then concatenating in the time dimension, x i Train represents the feature vector of any train node, x i Bridge represents the feature vector of any bridge node, x i Pier represents the feature vector of any pier node, X graph is the feature matrix of the entire graph structure.
4. A method for predicting the dynamic response of a high-speed train based on physical driving according to claim 3, characterized in that: The process of embedding the feature vectors of all nodes in the graph structure into a unified dimensional space and downsampling the embedded time series is specifically shown in the following formula (3): X down_sample = X graph [::2] (3); Equation (3) represents downsampling the feature matrix X corresponding to the original time series with a slice time length of step 2 graph to obtain the feature vector X corresponding to the time series after downsampling down_sample .
5. A dynamic response prediction method for a high-speed train based on physical driving according to claim 4, characterized in that: The process of inputting the feature vectors corresponding to the original time series and the downsampled time series into the encoder respectively and extracting the feature vectors through the multi-head attention mechanism is specifically shown in the following formulas (4) and (5): In Equation (4) and Equation (5), Encoder represents the encoder, and mask represents the mask, and are both the outputs of the encoder; The process of inputting the feature vectors corresponding to the encoded time series into the graph attention network and concatenating them along the time series dimension to obtain the graph embedding result is specifically shown in the following formulas (6) and (7); X dense = Concat(X′ graph , X′ down_sample ) (7); In formula (6), A ij represents the adjacency matrix, e ij represents the edge connecting node v i and node v j in the graph structure G, and GATs represents the graph attention network; in formula (7), Concat represents concatenation.
6. The dynamic response prediction method of a high-speed train based on physical driving according to claim 5, wherein: The processing process of the described encoder Encoder is specifically as follows: First, the input time series are projected into the hidden layer dimension through three linear transformation matrices Wq, Wk, and Wv to obtain Q, K, and V; then, a dot product operation is performed on Q and K, and it is projected into an interval with a sum of one through the softmax activation function to obtain the attention weights. The obtained attention weights and V are weighted and summed to obtain the attention output, as shown in the following formula (8): In Equation (8), d k represents the feature dimension of Q and K, and head i is the output of one attention head; The multi-head attention mechanism synchronously uses the attention mechanism of formula (8) and concatenates the final results together; Then the output X of the multi-head attention mechanism attention After passing through a residual and layer normalization, a linear layer (Linear), and a second residual and layer normalization in sequence, it is input into the distillation layer, thus completing one layer of encoding. See the following formula (9) for details: In Equation (9), X attention represents the output of the multi-head attention mechanism, X res1 represents the result after one residual and layer normalization, LayerNorm represents layer normalization, X res2 represents the result after a second residual and layer normalization, Linear represents the linear layer, X ditllblock represents the output result of the distillation layer, that is, the output of one layer of encoding, Elu represents the ELU activation function, Conv1d represents one-dimensional convolution, MaxPool represents max pooling, and stride = 2 represents the stride of max pooling is 2.
7. A method for predicting the dynamic response of a high-speed train based on physical driving according to claim 5, characterized in that: The processing process of the graph attention network is as follows: First, each node is initialized as a feature vector, and then the graph attention layer in the graph attention network aggregates the features of its neighbor nodes according to the attention coefficient of each node, as shown in the following formula (10) specifically. The calculation process of the attention coefficient is shown in the following formula (11) specifically: In formulas (10) and (11), represents the output of the graph attention layer, σ is the sigmoid function, and W represents the weight matrix, represents the feature vector of node i at time t, represents the feature vector of node j adjacent to node i at time t, represents the feature vector of node k adjacent to node i, represents the set of all nodes adjacent to node i, | represents the concatenation operation, a T represents the weight vector, and LeakyReLU represents the non-linear activation function.
8. A method for predicting the dynamic response of a high-speed train based on physical driving according to claim 5, characterized in that: The input of the described decoder is constructed through an auxiliary learning method, and the calculation process is shown in the following formula (12): In formula (12), X auxiliary represents the auxiliary feature vector, and X graph is the feature matrix corresponding to the original time series. L auxiliary_length is the auxiliary learning length. -1 represents indexing from the end of the vector forward, with the length of the indexing field being L auxiliary and the step size being 1. X0 represents the initial vector that has not been decoded, that is, the part where the vector is 0. L y represents the length of the initial vector that has not been decoded. N represents the number of nodes, D represents the feature dimension of each node, and X dec is the input to the decoder.