Novel motor train unit traction motor temperature intelligent prediction method and system

By adopting variable position coding, causal convolution multi-head self-attention mechanism and non-autoregressive output strategy in the temperature prediction of EMU traction motors, the problem of local context information being ignored and accumulated errors in the prior art is solved, and higher temperature prediction accuracy is achieved.

CN120067569APending Publication Date: 2025-05-30DALIAN JIAOTONG UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510067919.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When predicting the temperature of the traction motor of the EMU, the prior art ignores the cumulative error caused by local context information and autoregression strategies, resulting in low temperature prediction accuracy.

Method used

Variable position coding is used instead of absolute position coding, a causal convolution multi-head self-attention mechanism is designed to capture global and local features, and a non-autoregressive output strategy is adopted, combining a multi-layer perceptron output layer to predict temperature sequences at one time.

Benefits of technology

It effectively improves the accuracy of traction motor temperature prediction, eliminates cumulative errors, and improves the model's ability to extract local context information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067569A_ABST
    Figure CN120067569A_ABST
Patent Text Reader

Abstract

The invention discloses a novel motor train unit traction motor temperature intelligent prediction method and system, and relates to the technical field of deep learning. Comprising the following steps: in position coding, changing an absolute position code into a variable position code so as to better adapt to temperature sequences with different lengths; in an encoder and a decoder, designing causal convolution multi-head self-attention, applying the causal convolution multi-head self-attention to a multi-head self-attention layer, and extracting global features and local features of traction motor temperature data at the same time; based on a non-autoregressive output strategy, taking the input data as the input of a decoder, and eliminating an accumulated error generated by taking the output of the encoder as the input; the encoder extracts a traction motor temperature memory information matrix, and the decoder captures the traction motor temperature change at the approaching moment, and performs attention calculation on the traction motor temperature change and the traction motor temperature memory information matrix to predict the temperature trend at the future moment; the problem of low temperature prediction precision caused by neglecting local context information and accumulative errors is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly relates to a novel intelligent prediction method and system for the temperature of a traction motor of a bullet train. Background Art

[0002] As a core component of the traction drive system of a bullet train, the operating state of the traction motor is crucial for the running safety of the bullet train. To ensure the normal operation of the traction motor, accurately predicting its temperature using real-time sensor data has become a key measure for preventing failures and maintaining train safety. Currently, the prediction methods for the temperature of the traction motor are mainly divided into three categories: mechanism model-based methods, machine learning-based methods, and deep learning-based methods. However, due to the complexity of the traction motor structure and the diversity of the working environment, it is difficult to accurately model using physical model-based methods, resulting in large deviations in temperature prediction; machine learning methods are limited by their model structures and fail to fully capture the deep time series relationships in temperature changes. In contrast, deep learning methods, with the hierarchical structure of their neural networks and the advantages of non-linear functions, can more effectively extract the time series information in the temperature of the traction motor.

[0003] Among many deep learning algorithms, Transformer is still widely used in natural scenarios. Therefore, the Transformer model is selected to perform intelligent prediction of the temperature of the traction motor of the bullet train. The Transformer model consists of four parts: positional encoding, encoder, decoder, and output layer, as Figure 1 shown.

[0004] The said positional encoding: By adopting trigonometric positional encoding to help the model understand the order relationship of elements at different positions in the input sequence. This encoding method is based on the periodic characteristics of sine and cosine functions, and generates a positional encoding matrix through a predefined function, assigning a fixed and unique vector to each position of the input sequence. These vectors not only mark the sequence order in the time dimension but also enable the model to capture the relative position relationships of elements in the sequence, thereby more accurately predicting the temperature of the traction motor. For each time series position pos and each dimension i, the trigonometric positional encoding matrix PE is calculated using the following formula:

[0005]

[0006] where sin(·) and cos(·) are the sine and cosine functions respectively; 2i and 2i + 1 represent the even and odd dimensions respectively; d is the width of the Transformer model; and 10000 represents the empirical base of the trigonometric positional encoding.

[0007] The encoder and decoder: They are stacked by multiple encoder layers or decoder layers with the same structure. The encoder layer is used to encode the input matrix into a series of hidden representations, and the decoder layer is used to decode the autoregressive output sequence and calculate the attention scores between it and the encoder output. Each encoder layer is mainly composed of a multi-head self-attention layer and a feed-forward neural network layer. The multi-head self-attention layer calculates the similarity between feature vectors, represents the importance and correlation of each position in the sequence, and forms a multi-head self-attention score matrix. The feed-forward neural network layer further maps this matrix to a higher-dimensional feature space, introducing non-linear transformations to enhance the model's expressive power and non-linear characteristics. The decoder layer additionally includes a masked multi-head self-attention layer and an encoder-decoder attention layer. The former adds a mask to block information at future time steps, and the latter uses the output of the encoder to help the decoder better align and understand the input sequence, thereby generating a more accurate target sequence.

[0008] The multi-head attention mechanism: It allows the model to establish associations between different positions in the input sequence and allocate different attention weights between different positions, enabling the model to better understand the dependencies between different elements in the input sequence. Its structure is as Figure 2 shown. In the self-attention mechanism, the original input sequence is multiplied by three different weight matrices to obtain three new sets of representations: the query matrix, the key matrix, and the value matrix. The similarity between each query and key is calculated through the dot product of the matrices. After the dot product calculation, it is divided by the square root of the input vector dimension to ensure that the calculation of attention weights is not affected by the input dimension, making the distribution of attention weights more stable and controllable. Then, it is normalized by the Softmax function to obtain the degree of association between each query and the keys at other positions. Finally, the corresponding Values are weighted and summed using the attention weights to aggregate and integrate the information at different positions in the input sequence, obtaining the final self-attention score. When applying the self-attention mechanism in the decoder part, a mask needs to be added after scaling to limit the model to only focus on the current position and previous information and not access future information. The Transformer model usually adopts the multi-head attention mechanism to improve the model's expressive power. Its structure is as Figure 3 shown. The multi-head attention mechanism calculates different attention weights in different representation spaces by simultaneously learning multiple different query matrices, key matrices, and value matrices. Finally, the attention outputs of each head are concatenated and passed through a linear transformation to obtain the multi-head self-attention output, enhancing the model's ability to capture temporal feature changes.

[0009] The output layer: Usually consists of a linear layer and a Softmax activation function. The linear layer converts the output of the decoder into a dimension that matches the target vector, while the Softmax layer converts it into probability values, representing the probability distribution of the predicted values at the next moment. The output layer uses an autoregressive strategy to sequentially generate prediction results, and at the same time uses these prediction results as the input of the decoder to jointly predict the output of the decoder at the next moment with the output of the encoder.

[0010] Although the Transformer model has been widely used in the field of prediction, the traction motor temperature data has significant global and local correlation characteristics, and the traditional Transformer model has deficiencies in capturing these local feature relationships. In addition, the autoregressive strategy of the Transformer model may generate large cumulative errors when generating the output sequence, thus affecting the accuracy of temperature prediction. Summary of the Invention

[0011] The object of the present invention is to propose a new intelligent prediction method and system for the temperature of the traction motor of a bullet train to solve the problem of low temperature prediction accuracy caused by ignoring local context information and cumulative errors.

[0012] According to the first aspect of the embodiments of the present disclosure, a new intelligent prediction method for the temperature of the traction motor of a bullet train is provided, including the following steps:

[0013] In the position encoding, change the absolute position encoding to variable position encoding to better adapt to temperature sequences of different lengths;

[0014] In the encoder and decoder, design a Causal Convolutional Multi-Head Self-Attention (CA), apply it to the multi-head self-attention layer, and simultaneously extract the global features and local features of the traction motor temperature data;

[0015] Based on the non-autoregressive output strategy (NAR), use the input data as the input of the decoder to eliminate the cumulative error generated by using the encoder output as the input;

[0016] The encoder extracts the traction motor temperature memory information matrix, the decoder captures the temperature change of the traction motor at adjacent moments, and performs an attention calculation with the traction motor temperature memory information matrix to predict the temperature trend at future moments;

[0017] In one embodiment, in the output layer, the linear layer and the Softmax layer are changed to a multi-layer perceptron to output the predicted temperature sequence at one time, increasing the non-linear expression ability of the model and mapping high-dimensional features to the model output to improve the prediction accuracy of the model.

[0018] In one embodiment, the variable position encoding: To improve the prediction accuracy of the traction motor temperature prediction model, two weight matrices and biases with learnable parameters are designed to map the input data into a variable position encoding matrix, enabling the model to better capture the long-distance dependence relationship of the traction motor temperature sequence during the encoding and decoding stages. The calculation process of the variable position encoding is as follows:

[0019]

[0020] In the formula, W even and W odd represent the weight matrices of the even and odd columns respectively.

[0021] In one embodiment, the causal convolutional multi-head self-attention: Although the self-attention mechanism of the Transformer model can calculate the similarity between the query matrix Q and the key matrix K point by point, it fails to fully utilize local information, such as the shape characteristics of the traction motor temperature curve. Therefore, to enhance the model's ability to extract local context information, a causal convolutional multi-head self-attention structure as shown in Figure 5 is proposed. It includes two parts: causal convolution and self-attention mechanism. The causal convolution is used to capture local context information, and the self-attention mechanism helps to obtain global information. In the self-attention mechanism, performing causal convolution on the query matrix Q and the key matrix K can better help the model learn the local temporal dependence relationship in the traction motor input sequence. Performing causal convolution on the value V may lead to information confusion. Therefore, to maintain the effectiveness and correctness of the model, the designed causal convolutional self-attention module uses the causal convolution of the query matrix Q and the key matrix K and their residual connections to input the self-attention mechanism to better calculate the attention weights. The causal convolution of Q and K and their residual connections can assign a greater weight to the current moment, which is beneficial to the fitting of the model.

[0022] In one embodiment, through the causal convolutional multi-head self-attention mechanism, different features and context relationships are captured in parallel to improve the feature expression ability of the model. The calculation process of the causal convolutional multi-head self-attention is as follows:

[0023]

[0024] In the formula, W Q and W K and W K are weight matrices, and d kis the dimension of K, the softmax(·) function converts the dot product result into a probability distribution, C(*) represents the causal convolution operation, o(*) represents the residual operation, Q i , K i , V i respectively represent the query, key, and value matrices of the i-th attention head, and CAhead i represents the output of the i-th attention head, and the Concat(*) function is used to concatenate the outputs of each head, and W 0 is the combined weight matrix, and m represents the number of heads of the multi-head attention.

[0025] In one embodiment, the non-autoregressive output strategy: When the traditional Transformer model constructs the target sequence, it has an autoregressive characteristic. The model needs to wait for the prediction value of the previous moment to be generated before it can continue to generate the subsequent prediction values. This autoregressive strategy will lead to a large cumulative error and cause the prediction quality of the target sequence to decline. Therefore, the present invention proposes a non-autoregressive strategy as Figure 6 shown. The sequence of the first s minutes of the source sequence is used as the input of the encoder, and the sequence of the subsequent l minutes is used as the input of the decoder; by reorganizing the inputs of the encoder and decoder through the non-regressive strategy, the CAformer model has the same input structure in the training and prediction stages, improving the temperature prediction accuracy of the model.

[0026] In one embodiment, the output layer: The train traction motor temperature prediction belongs to a multi-step prediction task. The output layer of the traditional Transformer uses the softmax activation function to convert the output of the decoder into the probability distribution of the prediction result. The accuracy of the output temperature result is not high, and it can only predict one time step. To improve the one-time prediction of the temperature sequence of the traction motor and increase the non-linear expression ability of the model. CAformer uses a multi-layer perceptron as the output layer. The number of neurons in the hidden layer is set to the number of output features of the traction motor. Each neuron shares the output of the decoder in the same layer. The output dimension of the linear layer is set to the prediction length. The temperature sequence Y of the traction motor is predicted at one time p , reducing the cumulative error of the traction motor temperature prediction model and improving the rapidity of the prediction model. The calculation process is as follows:

[0027]

[0028] In the formula, j is the number of neurons in the hidden layer, which is also the number of output features of the traction motor; Y D represents the output matrix of the decoder, Z is the output matrix of the hidden layer; W j , b j are the weight matrix and bias vector of the neuron; W p , b prespectively represent the weight matrix and bias vector of the linear layer; Y p is the predicted traction motor temperature vector.

[0029] According to the second aspect of the embodiments of the present disclosure, a novel intelligent prediction system for the temperature of traction motors of multiple units is provided, including:

[0030] A position encoding module that changes the absolute position encoding to variable position encoding;

[0031] A causal convolution module that designs a causal convolution multi-head self-attention in the encoder and decoder, applies it to the multi-head self-attention layer, and simultaneously extracts the global features and local features of the traction motor temperature data;

[0032] A non-autoregressive module that, based on the output strategy of non-autoregressive, uses the input data as the input of the decoder to eliminate the cumulative error generated by using the encoder output as the input;

[0033] A prediction module that the encoder extracts the traction motor temperature memory information matrix, and the decoder captures the traction motor temperature changes at adjacent moments, and performs an attention calculation with the traction motor temperature memory information matrix to predict the temperature trend at future moments.

[0034] According to the third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program running on the memory, and when the processor executes the program, it implements the novel intelligent prediction method for the temperature of traction motors of multiple units.

[0035] According to the fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the novel intelligent prediction method for the temperature of traction motors of multiple units.

[0036] The above technical solutions adopted by the present invention, compared with the prior art, have the following advantages: The present invention uses a deep learning model to deeply mine the time information in the operation state data of high-speed train traction motors. For this purpose, a novel intelligent prediction method for the temperature of traction motors of multiple units is designed; specifically, this method first uses variable position encoding to replace the traditional absolute position encoding to improve the adaptability of the Transformer model to input data of different lengths. Secondly, a causal convolution multi-head self-attention is designed, which can simultaneously capture the global features and local features of the temperature data, thereby enhancing the model's ability to extract local context information. Finally, we adopt a non-autoregressive strategy. By reconstructing the input method of the decoder, that is, using the input sequence and combining with a multi-layer perceptron, the predicted traction motor temperature sequence is output at one time. This approach effectively eliminates the cumulative error and further improves the prediction accuracy. Brief Description of the Drawings

[0037] The attached drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application.

[0038] Figure 1 is the structure diagram of the Transformer;

[0039] Figure 2 is the structure diagram of the self-attention mechanism;

[0040] Figure 3 is the structure diagram of the multi-head self-attention;

[0041] Figure 4 is the structure diagram of the intelligent prediction method for the temperature of the traction motor of the new multiple unit;

[0042] Figure 5 is the structure diagram of the CA;

[0043] Figure 6 is the structure diagram of the non-autoregressive strategy;

[0044] Figure 7 is the heat map of the influence of hyperparameters m and n on the model prediction accuracy;

[0045] Figure 8 is the iterative loss graph of the CAformer model with different output window lengths;

[0046] Figure 9 is the heat map of the self-attention score matrix and the CA score matrix;

[0047] Figure 10 is the iterative loss graph of each ablation model under different prediction scenarios;

[0048] Figure 11 is the line graph of the MAE and MSE values of different models under different output windows;

[0049] Figure 12 is the temperature prediction effect diagram of each model when the traction motor is working normally;

[0050] Figure 13 is the temperature prediction effect diagram of each model when the traction motor fails. Detailed Embodiments

[0051] The present disclosure will be further described below in conjunction with the drawings and embodiments.

[0052] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0053] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0054] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. Similarly, it should be noted that each block in the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented using a dedicated hardware-based system for performing the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.

[0055] This embodiment provides a novel intelligent prediction method for the temperature of the traction motor of a multiple unit train. The specific content is as follows:

[0056] Each carriage of the multiple unit train contains 4 traction motors to provide power for the train, and the train combination is "4 motors + 4 trailers". Therefore, the data set selected in this embodiment is the stator temperature, drive-end shaft temperature, non-drive-end shaft temperature, running speed of the train, and ambient temperature of the No. 1-4 traction motors recorded in carriages 2, 4, 5, and 7 of a certain type of multiple unit train. In addition, through a preliminary observation of the data set, it is found that during the train operation, there is a connection between the traction motor temperature and the train running speed and ambient temperature. Therefore, a total of 50-dimensional data is also used as the model input. Since the data collected by the sensors is of poor quality, the data is first preprocessed. The process includes: linear interpolation and maximum-minimum normalization processing. Linear interpolation can interpolate the missing temperature data, and the Z-score normalization method can reduce the influence of the differences in different dimensions in the data. The formula is:

[0057]

[0058] In the formula, x and x norm are the original value and the normalized value of the sample respectively, and mean(x) and std(x) are the standard deviation and the average value of all sample data respectively.

[0059] The time span of the collected data is the temperature of the traction motor bearing, the train running speed, and the ambient temperature from October 2017 to December 2021, and the total sampling length is 915261.

[0060] The Mean Squared Error (MSE) is the average of the squares of the differences between the predicted values and the true values. If there are points with abnormally high increases in the traction motor temperature, the square of the prediction error will amplify the influence of these outliers, making the MSE more sensitive to abnormal data distributions and better reflecting the overall error size of the predicted traction motor temperature, which is used to evaluate the stability of the temperature prediction model. The calculation formula is as follows:

[0061]

[0062] In the formula, n represents the total number of data, and y i represent the predicted value and the true value of the data respectively.

[0063] The Mean Absolute Error (MAE) is the average of the absolute values of the errors between the predicted values and the true values. Using the absolute value to represent the size of the prediction error for each sample, it is not sensitive to outliers and is more robust, which is used to evaluate the accuracy of the temperature prediction model. The calculation formula is as follows:

[0064]

[0065] 60% of the data in the dataset is divided into the training set, 20% into the validation set, and the remaining 20% into the test set. The training uses the Z-score normalization method to normalize the input data, adopts the L2 loss function, the optimizer is Adam, the initial learning rate is 0.001, the number of iteration rounds is 100, the training batch size is 128, the number of layers of both the encoder and the decoder is 6, and the number of heads of the multi-head attention is 8. The time step of the window size is 1 min, the weight matrix is randomly initialized, and the early stopping model training technique is used to avoid overfitting and improve the generalization ability of the model. The model training is stopped when the loss does not decrease for 5 consecutive epochs.

[0066] 1. Hyperparameter Analysis

[0067] (1) Causal Convolution Kernel Size Analysis

[0068] To explore the influence of the causal convolution kernel size on feature extraction, the unnormalized MAE evaluation metric was used to analyze the relationship between the two hyperparameters, the causal convolution kernel size \(m\) of the encoder and the causal convolution kernel size \(n\) of the decoder, and the prediction accuracy of the CAforemr model. By traversing different parameter combinations, the optimal model hyperparameter combination was selected. When the input window lengths of both the encoder and the decoder were 30 minutes, the influence of the hyperparameters \(m\) and \(n\) on the model prediction accuracy is as Figure 7 shown.

[0069] As Figure 7 can be seen, when the values of the causal convolution kernel sizes of both the encoder and the decoder are 3, the color in the corresponding heatmap is the lightest, the unnormalized MAE value is the smallest, which is 0.033, and the CAforemr model has better temperature prediction performance.

[0070] (2) Comparison of Encoder and Decoder Inputs

[0071] To explore the influence of the input lengths of the encoder and the decoder on the prediction accuracy under the non-autoregressive strategy. The causal convolutions of both the encoder and the decoder were set to \(1\times3\) in the experiment. Three different input window lengths were designed, which were 30 minutes, 60 minutes, and 90 minutes respectively, and the output window length was fixed at 30 minutes. The influence of the hyperparameters \(X\) s and \(X\) l on the model prediction accuracy is shown in Table 1.

[0072] Table 1 Influence of Hyperparameters \(X\) s and \(X\) l on the Model Prediction Accuracy

[0073]

[0074] As can be seen from Table 1, the longer the input window length, the lower the MAE and MSE values of the CAformer model, indicating that as the input window increases, the CAformer model can obtain more feature information and improve the prediction accuracy. When the input window length remains unchanged, when the lengths of the encoder input \(X\) s and the decoder input \(X\) l are the same, the MAE and MSE values of the CAformer model are the lowest, and the predicted values of the CAforemr model can better fit the actual values, with higher prediction accuracy. This indicates that when \(X\) s is equal to \(X\) l , the CAformer model can better align the position information and capture the feature information.

[0075] 2. Comparison of Different Output Window Sizes

[0076] The prediction step, i.e., the output window size, is an important parameter for evaluating the prediction performance of the model. On the premise of ensuring the temperature prediction accuracy, the larger the prediction step, the earlier the faults of the traction motor can be identified, leaving more time for fault maintenance. Therefore, experiments are conducted to analyze the influence of different output window sizes on the temperature prediction accuracy. The input window length is set to 60 min in the experiment, and the output window lengths are 15 min, 30 min, 45 min, and 60 min respectively. In the validation set, the iterative loss of the CAformer model is as Figure 8 shown.

[0077] It can be Figure 8 seen that as the number of iterative steps increases, the loss values all show a downward trend and finally tend to be stable, indicating that the CAforemr model has good generalization ability for unknown data. The smaller the output window length, the lower the loss of the CAforemr model, and vice versa. This shows that increasing the output window length will reduce the generalization ability and prediction accuracy of the CAformer model.

[0078] To further observe the prediction effect of the CAformer model under different output window lengths, MSE and MAE are used as evaluation indicators, and the prediction results for different output window lengths are shown in Table 2.

[0079] Table 2 Prediction effects for different output window lengths

[0080]

[0081] It can be seen from Table 2 that the smaller the output window length, the smaller the MAE and MSE values of the CAformer model. Specifically, when the output window length is 15 min, the values of MAE and MSE are 1.097 and 2.869 respectively. When the output time window length increases to 60 min, the values of MAE and MSE rise to 1.959 and 6.976 respectively. The experimental results show that increasing the output window length will reduce the prediction accuracy of the CAformer model. In practical applications, it is crucial to select an appropriate output window length, and the relationship between prediction accuracy and output window length should be balanced according to the requirements of specific tasks.

[0082] 3. Ablation experiment

[0083] (1) Verification of the attention mechanism

[0084] To analyze the ability of the causal convolutional self-attention mechanism in the CAformer model to extract the temperature characteristics of the traction motor, it is compared with the self-attention mechanism in the Transformer model. The heatmap of the attention score matrix is as Figure 9As shown in the figure. Among them, (a) is the heat map of the self-attention mechanism, and (b) is the heat map of the causal convolutional self-attention mechanism.

[0085] From Figure 9 It can be seen that the color change of the heat map of the self-attention mechanism is small, indicating that when the self-attention mechanism calculates the weights, it only focuses on the global information, resulting in similar weights in the self-attention score matrix. The color change of the heat map of the causal convolutional attention mechanism is large because when the causal convolutional attention mechanism calculates the weights, it considers both local information and global information, making the attention weight distribution sparser, which can better focus on the key parts of the traction motor temperature change trend, thus better predicting the traction motor temperature.

[0086] When the traction motor is working normally, its temperature will not change violently. When a fault is about to occur, the temperature of a certain key point will rise abnormally. When facing this situation, the self-attention mechanism of the Transformer model is difficult to quickly capture the local trend of the time point with abnormal temperature rise. The designed causal convolutional self-attention mechanism of the present invention can accurately capture the key information of the traction motor temperature rise according to the temperature change trend, improving the prediction accuracy of abnormal temperature.

[0087] (2) Ablation experiment and analysis

[0088] To clarify the necessity of each improved module in the CAformer model, ablation experiments are carried out using the Transformer model, the CA-Transformer model, the NAR-Transformer model and the CAformer model. Among them, the CA-Transformer model is the Transformer model using causal convolution, and the NAR-Transformer model is the Transformer model using the non-autoregressive strategy. The parameter settings of each model are the same. The experiment designs 4 prediction scenarios, the input window length is fixed at 60 min, and the output window lengths are 15 min, 30 min, 45 min and 60 min respectively. The iterative losses of each model in the validation set are as Figure 10 shown.

[0089] From Figure 10It can be seen that in the 4 scenarios, the convergence loss value of the Transformer model is the largest. The convergence loss values of the CA-Transformer model and the NAR-Transformer model are both lower than that of the Transformer model. The loss value of the CAformer model is the smallest, and as the output window length increases, the decrease in the loss value of the CAformer model is more obvious. This indicates that the proposed CAforemr model can effectively reduce the loss value of the Transformer model by adding causal convolutional self-attention and non-autoregressive strategies. The CAforemr model has a stronger fitting ability.

[0090] To further analyze the performance of each improved module, the MSE and MAE evaluation metrics were used for comparison, and the results are shown in Table 3. To more intuitively observe the comparison of the prediction effects, line charts of the MAE and MSE values of different models under four output windows were drawn, as Figure 11 shown.

[0091] Table 3 Results of ablation experiments

[0092]

[0093] As can be seen from Table 3 and Figure 11 it can be seen that when the input window length remains unchanged, as the output window length increases, the MAE and MSE of the model gradually increase, and the prediction accuracy decreases. Compared with the Transformer model, the CA-Transformer model and the NAR-Transformer model, the CAformer model has the highest prediction accuracy. When the output window lengths are 15 min, 30 min, 45 min, and 60 min respectively, the CAformer model proposed in the present invention reduces by at least 10.38% and 12.47%, 17.19% and 30.07%, 13.67% and 32.69%, 11.03% and 24.18% respectively in terms of the MAE and MSE metrics compared with other models. The causal convolutional self-attention mechanism of the CAformer model enhances the ability to capture global and local features, and the non-autoregressive strategy eliminates the cumulative error, thereby improving the accuracy of traction motor temperature prediction.

[0094] 3. Comparative experiments

[0095] To verify the accuracy and stability of the CAformer model in predicting the temperature of traction motors, the ARIMR, SVM, LSTM, GRU, and Transformer models were used for comparative analysis with the model of the present invention. When the input window length was fixed at 60 min and the output window lengths were 15 min, 30 min, 45 min, and 60 min respectively, the performance of each model was analyzed using the MAE and MSE evaluation metrics, and the results are shown in Table 4.

[0096] Table 4 MAE and MSE values of different models

[0097]

[0098] As can be seen from Table 4, at different prediction steps, compared with other models, the CAformer model has lower MAE and MSE values. Specifically, when the output window lengths are 15 min, 30 min, 45 min, and 60 min respectively, the MAE and MSE values of the CAformer model are reduced by at least 21.92%, 25.21%, 33.03%, and 32.49% respectively.

[0099] There are significant differences in the temperature data of traction motors under normal and faulty conditions. When the traction motor is operating normally, the temperature fluctuates slowly within a very small range. When a fault occurs in the traction motor, a phenomenon of temperature rise will occur. The difference in temperature can, to a certain extent, reflect the state of the traction motor. To analyze the effect of each model in predicting temperature, taking the stator end of the No. 3 traction motor in Carriage 5 of the CR300BF EMU as the research object, data of the traction motor under normal and faulty conditions were selected for experiments, and the cumulative prediction length was 300 min. From 6:00 to 11:00 on November 24th, the traction motor was in a normal operating state, and its temperature prediction results are as Figure 12 shown. Starting from 6:00 to start running, the temperature of the motor rises from the outdoor temperature to the temperature in the normal operating state during the starting period.

[0100] From Figure 12 it can be seen that when the traction motor is operating normally, the CAformer model can learn the change trend in a short time, accurately track the temperature, and completely predict the change trend of the temperature fluctuation of the traction motor. The fitting degree between the predicted value and the true value is relatively high, and the prediction ability is stronger; the prediction effects of other models are poor, and there is an obvious lag.

[0101] The temperature prediction results of the traction motor from 5:00 to 10:00 on October 1st are as Figure 13As shown. At 5:45, the train made a temporary stop at the station for 15 minutes. During this period, the train's running speed was 0 km / h, and the temperature at the stator end decreased slightly. Subsequently, the train continued to run, and the temperature at the stator end rose to the normal operating temperature. At 8:25, the real temperature at the stator end had a rapid temperature rise, exceeding the normal operating temperature.

[0102] It can be seen from Figure 13 that when a traction motor fails, the real temperature of the traction motor has a rapid temperature rise. It is difficult for ARIMA, SVR, LSTM, GRU, and Transformer models to track the sudden change in temperature, and the fitting degree between the predicted value and the real value is poor, with obvious prediction deviations. The CAformer model can more accurately capture the rapid upward trend of temperature and has a high prediction accuracy.

[0103] In summary, through hyperparameter analysis and comparison of different output window sizes, the rationality of the selected parameters is illustrated; the effectiveness of each improved part is verified through ablation experiments; design comparison experiments verify the prediction accuracy and stability of the model under different prediction scenarios; the proposed intelligent temperature prediction method for the traction motor of the new EMU is significantly superior to the existing classic temperature prediction methods in terms of prediction robustness, accuracy, and efficiency, and is more suitable for the temperature prediction of traction motors.

[0104] Those skilled in the art should understand that the above-mentioned various modules or steps of the present disclosure can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by the computing device, so that they can be stored in the storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. The present disclosure is not limited to any specific combination of hardware and software.

[0105] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0106] Although the specific implementation manners of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that based on the technical solutions of the present disclosure, various modifications or deformations that can be made without creative labor by those skilled in the art are still within the protection scope of the present disclosure.

Claims

1. A new intelligent prediction method for the temperature of traction motor of EMU, characterized in that: The following steps are involved: In position coding, change the absolute position coding to variable position coding; In the encoder and decoder, a causal convolutional multi-head self-attention is designed and applied to the multi-head self-attention layer to extract the global and local features of the traction motor temperature data at the same time; Based on the non-autoregressive output strategy, the input data is used as the input of the decoder to eliminate the accumulated error generated by the encoder output as the input; The encoder extracts the traction motor temperature memory information matrix, and the decoder captures the traction motor temperature changes at nearby moments, and performs attention calculation with the traction motor temperature memory information matrix to predict the temperature trend at future moments.

2. According to the novel intelligent prediction method for the temperature of the traction motor of a train set as claimed in claim 1, it is characterized in that: In the output layer, the linear layer and the Softmax layer are changed to a multi-layer perceptron, which outputs the predicted temperature sequence at one time and maps the high-dimensional features as the model output.

3. According to the novel intelligent prediction method for the temperature of the traction motor of a train set as claimed in claim 1, it is characterized in that: The variable position encoding designs two learnable parameter weight matrices and biases to map the input data into a variable position encoding matrix. The process is as follows: Where W even , W odd Represent the weight matrices of even and odd columns respectively.

4. According to claim 1, a new type of intelligent prediction method for traction motor temperature of EMU is characterized in that: The causal convolution multi-head self-attention includes two parts: causal convolution and self-attention mechanism. The causal convolution is used to capture local context information, and the self-attention mechanism helps to obtain global information. The self-attention mechanism takes the causal convolution of the query matrix Q and the key matrix K and their residual connection as input, giving a greater weight to the current moment.

5. A novel intelligent prediction method for the temperature of traction motor of a train set according to claim 1 or 4, characterized in that: Through the causal convolution multi-head self-attention mechanism, different features and contextual relationships are captured in parallel. The process is as follows: Where W Q , W K , W K is the weight matrix, d k is the dimension of K, the softmax(·) function converts the dot product result into a probability distribution, C(*) represents the causal convolution operation, o(*) represents the residual operation, Q i , K i 、V i denote the query, key, and value matrices of the ith attention head, CAhead i Indicates i The Concat(*) function is used to concatenate the outputs of each attention head. 0 is the combined weight matrix, and m represents the number of multi-head attention heads.

6. According to claim 1, a new type of intelligent prediction method for traction motor temperature of EMU is characterized by: The non-autoregressive output strategy is to use the first s minutes of the source sequence as the input of the encoder and the last 1 minute of the source sequence as the input of the decoder. The inputs of the encoder and decoder are reorganized through the non-regressive strategy so that the CAformer model has the same input structure in the training and prediction stages.

7. According to claim 1, a new type of intelligent prediction method for the temperature of the traction motor of the EMU is characterized by: The output layer adopts a multi-layer perceptron, the number of neurons in the hidden layer is set to the output feature number of the traction motor, each neuron in the same layer shares the output of the decoder, and the output dimension of the linear layer is set to the prediction length; the temperature sequence Y of the traction motor is predicted at one time p , reduce the cumulative error of the traction motor temperature prediction model; the output layer calculation process is: In the formula, j is the number of neurons in the hidden layer, which is also the output characteristic number of the traction motor; Y D represents the output matrix of the decoder, Z is the output matrix of the hidden layer; W j 、b j is the weight matrix and bias vector of the neuron; W p 、b p Represent the weight matrix and bias vector of the linear layer respectively; Y p is the predicted traction motor temperature vector.

8. A new type of intelligent prediction system for traction motor temperature of EMU, characterized by: include: Position coding module, which changes absolute position coding to variable position coding; Causal convolution module, in the encoder and decoder, a causal convolution multi-head self-attention is designed and applied to the multi-head self-attention layer to extract the global and local features of the traction motor temperature data at the same time; The non-autoregressive module uses the input data as the input of the decoder based on the non-autoregressive output strategy, eliminating the accumulated error generated by the encoder output as the input; In the prediction module, the encoder extracts the traction motor temperature memory information matrix, and the decoder captures the traction motor temperature changes at nearby moments, and performs attention calculation with the traction motor temperature memory information matrix to predict the temperature trend at future moments.

9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the novel method for intelligently predicting the temperature of the traction motor of an EMU is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, a novel intelligent temperature prediction method for traction motors of EMUs is implemented.

Citation Information

Cited By

  • Power terminal equipment credibility evaluation method and system based on multi-source data fusion

    CN120449019A

  • Physical guidance multi-time-sequence model integrated bridge dynamic reliability evaluation method

    CN122414260A

  • A bridge dynamic reliability evaluation method for physically guided multi-sequential model integration

    CN122414260B