Urban rail train pantograph abrasion loss lightweight prediction method based on twin data

By building a lightweight convolutional neural network combined with twin data fine-tuning, the problems of low accuracy and high cost of prediction of pantograph wear amount are solved, and fast and accurate prediction of wear amount is achieved, reducing operation and maintenance costs are reduced, and the safety and reliability of the rail transit system are improved.

CN120257784APending Publication Date: 2025-07-04BEIJING UNIV OF CHEM TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510278776.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the prediction of the wear amount of pantograph depends on the experience of professionals, the prediction accuracy is low and the cost is high, the deep learning model parameters are large, the calculation complexity is high, and the pantograph wear data is difficult and the cycle is long, which affects the safety and reliability of train operations.

Method used

Using a lightweight prediction method based on twin data, a lightweight convolutional neural network is built and combined with twin data fine-tuning, the rapid and accurate prediction of pantograph wear is achieved. The method includes data acquisition and preprocessing, lightweight model construction and pre-training, twin data fine-tuning and prediction result output, and the model tuning is generated using carbon skateboard material parameters for model tuning.

Benefits of technology

It realizes rapid and accurate prediction of pantograph wear, reduces operation and maintenance costs, improves the practicality and promotion of predictions, and ensures the safety and reliability of the rail transit system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120257784A_ABST
    Figure CN120257784A_ABST
Patent Text Reader

Abstract

An urban rail train pantograph abrasion loss lightweight prediction method based on twin data is used for urban rail train pantograph abrasion loss prediction. The invention provides a lightweight model which is high in calculation efficiency and simple in structure aiming at the problem of wear loss prediction of a pantograph carbon contact strip. According to the model, the prediction robustness is ensured, and meanwhile, the wear loss prediction of the pantograph carbon contact strip can be quickly and accurately realized. In addition, twin data of pantograph abrasion are calculated through cyclic simulation by means of material parameters of the carbon slide plate, the weight and offset of the model are adjusted and optimized, and the prediction precision is improved. According to the method, firstly, collected wear data are preprocessed and then input into a lightweight network for pre-training; and then, performing parameter tuning on the pre-trained lightweight model by using twin data, and finally storing the adjusted parameters to complete prediction of the wear life. The method achieves the precise prediction of the wear loss of the pantograph carbon slide plate, and has the advantages of being small in parameter quantity and short in operation time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of life prediction of mechanical equipment components, and relates to a method combining a lightweight model and twin data fine-tuning for predicting the wear amount of pantographs of urban rail trains. Background Art

[0002] As an electrical device for obtaining electrical energy from the catenary by a rail train, a pantograph is generally installed on its roof. The entire pantograph system relies on the carbon slider and the contact wire for continuous and stable sliding contact to provide powerful power for the high-speed operation of an electrified train. The catenary wire and the pantograph carbon slider are respectively used as static and dynamic friction pairs, and are prone to wear under the sliding electrical contact mode, which directly affects the running condition and service life of the train and even causes safety accidents. Therefore, accurately predicting the wear amount of the pantograph and evaluating its wear life have become important problems to be solved urgently in the field of rail transit. Traditional life prediction methods usually rely on the experience and knowledge of professionals, with low prediction accuracy and high cost. With the acceleration of the industrial modernization process and the rapid development of technologies such as machine learning, deep learning, as the most popular branch of machine learning, is widely used in the field of life prediction due to its powerful data processing and feature learning capabilities. With the increasing complexity of deep learning model structures, the accuracy of life prediction has been significantly improved. However, this improvement is also accompanied by a significant increase in the number of model parameters, computational complexity, and training time. In addition, obtaining pantograph wear data faces many challenges, including difficulties in data collection, long cycles, and high costs. Therefore, developing a lightweight prediction model and combining it with twin data of pantograph wear for efficient prediction is of great significance for improving the practicality and popularization of the prediction method. Summary of the Invention

[0003] The present invention proposes a lightweight prediction method based on twin data for predicting the wear amount of pantographs of urban rail trains. Aiming at the problems of large number of parameters, high computational complexity, and long training time of existing deep learning models proposed in the background art, a lightweight model with high computational efficiency and a concise structure is proposed. While ensuring the prediction robustness, this model can quickly and accurately predict the wear amount of the pantograph carbon slider. In addition, aiming at the problems of difficult data collection, long cycle, and high cost of pantograph wear data, twin data of pantograph wear are generated through simulation calculation using the material parameters of the carbon slider, and the weights and biases of the model are finely tuned, significantly improving the prediction accuracy. The proposed method not only provides important technical support for the safety and reliability of the rail transit system, but also effectively reduces the operation and maintenance costs and risks, and has important engineering application value.

[0004] The lightweight prediction method based on twin data proposed by the present invention is jointly composed of a data acquisition and preprocessing module, a lightweight model construction and pre-training module, a twin data fine-tuning module, and a prediction result output module. The wear data collected by the sensor is preprocessed and then input into the constructed lightweight convolutional neural network to complete pre-training, and the parameter values of the pre-trained model are saved. Then, the twin data of the pantograph wear is input into the pre-trained lightweight model to fine-tune the model parameters. Finally, the parameters are saved, and the wear life is predicted.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] S1 Data acquisition and preprocessing module

[0007] Obtain the remaining thickness data of the pantograph carbon slide through a monitoring sensor; the wear of the carbon slide is caused by scraping with the contact wire on the left and right. By the mileage and speed of the urban rail train in one day, calculate the number of wear times of the carbon slide in one day. Then, according to the unit mass, elastic modulus, Poisson's ratio of the contact wire, and the density, elastic modulus, Poisson's ratio of the carbon slide, combined with the contact pressure, current, and operating speed parameters, use ANSYS Workbench for cyclic simulation to calculate the wear amount per wear, so as to obtain the wear thickness in one day. Subtract the wear thickness from the original thickness as the starting amount for the next day and then perform simulation, and so on, until the thickness of the carbon slide reaches its use threshold and stops, so as to obtain the twin data of the pantograph wear. Standardize and preprocess the collected monitoring data and twin data to eliminate the dimensional difference between the data, and normalize all feature values to the interval [0, 1]. The maximum-minimum normalization method is used for processing, and its calculation formula is:

[0008]

[0009] Among them, X represents the original data, min(X) represents the minimum value in the data set, max(X) is the maximum value in the data set, and Normalized(X) is the data after maximum-minimum normalization.

[0010] Make the processed data into a sample set by the method of sliding window sampling, and set a sliding window with a length of W and a step size of N . Sequentially sample each complete time series data. Assume that the total length of the data is L, and finally divide the entire time series data into samples, and the length of each sample is WThe data is divided into a dataset for pre-training the network and a dataset for fine-tuning the network. The pre-training network dataset is the pantograph wear data obtained by monitoring sensors, which is further divided into a training set and a validation set. The dataset for fine-tuning the network comes from the twin data obtained based on simulation and the pantograph wear data obtained by monitoring sensors. The twin data of pantograph wear is used as the training set, and the real data of pantograph wear is used as the validation set and the test set. Labels are set for the above-obtained pantograph wear data and twin data. Since the remaining thickness wear amount of the pantograph continues to wear, only the label Y for the degradation stage is set, and the formula is as follows:

[0011]

[0012] where t c represents the current degradation moment, t w represents the operating time of the pantograph from the start of wear to complete failure, and t s represents the time point of the start of wear.

[0013] S2 Lightweight Model Construction Module

[0014] Construct a lightweight convolutional neural network: The lightweight convolutional neural network includes a feature learning module, a feature enhancement module, and a prediction module. The feature learning module consists of a depthwise separable convolution improved based on the standard convolution. The feature enhancement module consists of two inverted residual structures. The prediction module consists of a fully connected layer and a linear regression layer.

[0015] Furthermore, the construction of the lightweight convolutional neural network specifically includes: The feature learning module first performs an initial convolution operation on the input data using a depth convolution layer with a convolution kernel size of 3*3, a padding of 1, and a stride of 1 to extract low-level feature information. After the depth convolution operation is completed, the ReLU activation function is applied to the output data to set all negative values to zero and keep all positive values unchanged. Then, a point convolution layer with a 1*1 convolution kernel, no padding, and a stride of 1 is used to further extract the features. The calculation process is shown in formula (3). This feature learning module significantly reduces the number of parameters and the computational cost while effectively extracting features.

[0016]

[0017] where x represents a sample of the dataset for pre-training the network or a sample of the dataset for fine-tuning the network and serves as the input data; represents the 3*3 depth convolution operation. The formula for the ReLU activation function is max{x, 0}, represents the 1*1 point convolution operation, and y1 represents the output feature of the feature learning module.

[0018] The feature enhancement module is to learn richer feature representations while maintaining computational efficiency and further extract feature information. The output features of the feature learning module first pass through the first inverted residual structure. In this structure, the features first pass through a lightweight dilated convolutional layer with a kernel size of 1*1, no padding, and a stride of 1 to increase the receptive field without adding extra parameters. Then, they pass through a depthwise convolutional layer with a kernel size of 3*3, padding of 0, a stride of 1, and a group number of 8, which convolves each input channel independently. Finally, the features pass through a pointwise convolutional layer with a kernel size of 1*1, no padding, and a stride of 1 to mix channel information and reduce the number of channels. The calculation process is shown in Equation (4):

[0019]

[0020] Among them, y1 represents the input features, represents the 1*1 lightweight dilated convolution operation, and y2 represents the output features of the first inverted residual structure of the feature enhancement module. Similarly, the second inverted residual structure is composed of a lightweight dilated convolutional layer with a kernel size of 1*1, no padding, and a stride of 1, a depthwise convolutional layer with a kernel size of 3*3, no padding, a stride of 2, and a group number of 4, and a pointwise convolutional layer with a kernel size of 1*1, no padding, and a stride of 1 stacked in sequence. While learning deep features, it further optimizes computational efficiency. The calculation process is shown in Equation (5):

[0021]

[0022] Among them, y2 represents the input features, represents the 3*3 depthwise convolution operation, and y3 represents the output features of the second inverted residual structure of the feature enhancement module.

[0023] The prediction module consists of a fully connected layer and a linear regression layer with 16 input features and 1 output feature. The output features passing through the two inverted residual structures are output as the RUL prediction results through the fully connected layer and the linear regression layer.

[0024] S3 Model Pre-training Module

[0025] Update the model parameters through the model pre-training module and perform model pre-training: Input the dataset samples of the pre-trained network made by S1 into the lightweight network model constructed by S2 for pre-training. Set the model iteration times, training batches, and learning rate. After each round of training, perform backpropagation according to the mean squared error (MSE) loss value obtained by Equation (6), and continuously adjust the training parameters using the Adam optimizer.

[0026]

[0027] where N represents the number of training samples, and RUL predicted,i represents the predicted value of the i-th sample, and RUL actual,i represents the true value of the i-th sample. The MSE function plays an important role in minimizing the prediction error, thus ensuring that the model can predict the RUL more accurately.

[0028] After iterating to the set number of iterations, it is verified using the validation set. When the root mean square error (RMSE) of the verification result ≤ 0.10, it is considered that the model has extracted the features related to the degradation process at this time, that is, it meets the requirements, and the training of the model is completed. If it meets the requirements, the model parameters of the pre-trained network module are fixed and saved, and are ready to be transmitted to the next module. The calculation process of the root mean square error (RMSE) is shown in formula (7):

[0029]

[0030] where n represents the number of samples, and y i represents the predicted value of the network model, represents the true label value.

[0031] S4 Twin Data Fine-tuning Module

[0032] The parameters of the pre-trained network model are optimized through the twin data fine-tuning module, and the twin data generated by simulation calculation is used to finely adjust the model parameters, aiming to improve the prediction robustness of the model. Specifically, first, the parameters of the pre-trained network model saved in S3 are loaded into the initialized lightweight network model; subsequently, the parameters of the feature learning module and the feature enhancement module are frozen, and only the parameters of the linear regression layer are fine-tuned, and the loss value of the wear life prediction is calculated through the mean square error (MSE) to guide the parameter optimization process.

[0033] Taking the MSE loss function as the training objective, the Adam optimizer is used to update the network parameters of the linear regression layer; the number of iterations, training batches, and learning rate of the model are set. After each round of training, backpropagation is performed according to the obtained MSE loss value, and the training parameters are continuously adjusted. After iterating to the set number of iterations, it is verified using the validation set. When the root mean square error (RMSE) of the verification result ≤ 0.03, the weights of the depthwise separable convolution, the two inverted residual structures, as well as the weights and biases of the linear regression layer are saved.

[0034] S5 Prediction Result Output Module

[0035] The test set in the dataset of the twin data fine-tuning module is input into the network model saved in S4, and the current wear life prediction value is output under the condition that the parameters are fixed, that is, the weights w and biases b are not updated. The cumulative wear life prediction value is calculated and visualized, and a wear life prediction graph is drawn to complete the wear life prediction. Brief Description of the Drawings

[0036] To more clearly illustrate the embodiments of the present invention and their design solutions, the accompanying drawings required for this embodiment will be briefly introduced below.

[0037] Figure 1 Schematic diagram of the algorithm flow for predicting the wear amount of the pantograph provided by the present invention;

[0038] Figure 2 Structural diagram of the lightweight model for predicting the wear amount of the pantograph provided by the present invention;

[0039] Figure 3 Graph of the predicted result of the wear life of the pantograph in the example; Detailed Description of the Invention

[0040] To more intuitively and elaborately elaborate on the technical core and significant advantages of the present invention, we took the wear degradation of the pantograph of Chongqing urban rail trains as a practical case and deeply carried out the research work on predicting its wear life.

[0041] Figure 1 As shown in the schematic diagram of the algorithm flow of the present invention, the specific implementation manner of the present invention will be further described below in conjunction with the accompanying drawings and examples. As Figure 1 shown, the present invention includes four modules: (1) data collection and preprocessing module, (2) lightweight model construction and pre-training module, (3) twin data fine-tuning module, and (4) prediction result output module.

[0042] In step (1), the present invention collects and records the wear conditions of the pantograph carbon slides of urban rail transit trains under actual operation conditions to obtain a pantograph carbon slide wear data set, which includes trains 01037, 01038, 01039, 01040, and 01041. Each train has four sets of wear data, namely the remaining thickness of the front slide of pantograph 1, the remaining thickness of the rear slide of pantograph 1, the remaining thickness of the front slide of pantograph 2, and the remaining thickness of the rear slide of pantograph 2. The wear of the carbon slide is caused by scraping against the contact wire left and right. By the daily mileage and speed of the urban rail train, the daily wear times of the carbon slide are estimated. According to the unit mass, elastic modulus, Poisson's ratio of the contact wire, and the density, elastic modulus, Poisson's ratio of the carbon slide, combined with the contact pressure, current, and operating speed parameters, cyclic simulation is carried out using ANSYS Workbench to calculate the wear amount per wear, so as to obtain the total wear thickness per day. Then, the original thickness is subtracted from the total wear thickness to obtain the starting thickness of the next day, and then simulation calculation is carried out, and so on, until the thickness of the carbon slide reaches the use threshold, and finally the twin data of pantograph wear is generated. According to the actual operation conditions of the pantograph carbon slide, when the thickness of the carbon slide is less than 24 mm, it is a failure. Therefore, the use threshold of the twin data is 24 mm. After that, preprocessing is carried out on the original data, including standardization, sliding window sampling, label annotation, sample division and production, etc. According to formula (8), the collected monitoring data and twin data are preprocessed by standardization to eliminate the dimensional difference between the data, and all eigenvalue are normalized to the interval [0,1]. The formula for max-min standardization is:

[0043]

[0044] where X represents the original data, min(X) represents the minimum value in the data set, max(X) is the maximum value in the data set, and Normalized(X) is the data after max-min standardization.

[0045] The sliding window sampling method is used to obtain samples. In this example, the size of the sliding window is 8 and the step size is 1. Then the data is divided into a data set for pre-training the network and a data set for fine-tuning the network. The data set for pre-training the network includes the wear data of the pantographs of 5 trains. The wear data of the pantographs of the first four trains are used as the training set, and the wear data of the pantograph of the last train is used as the validation set. The data set for fine-tuning the network comes from the wear data of the pantograph of the last train and the obtained twin data. The twin data of pantograph wear is used as the training set, and the wear data of the pantograph of the last train is used as the validation set and the test set. Labels are set for the above-obtained pantograph wear data and twin data. Since the wear amount of the remaining thickness of the pantograph continues to wear, only the label Y of the degradation stage is set. The formula is as follows:

[0046]

[0047] Among them, t c represents the current degradation moment, t w represents the operating time of the pantograph from the start of wear to complete failure, and t s represents the time point when wear starts.

[0048] Step (2) is to construct a lightweight network model, pre-train the network model, extract key features related to the device health status from the original sensor data, and save the network parameters after the training is completed.

[0049] The lightweight convolutional neural network in this example includes a feature learning module, a feature enhancement module, and a prediction module. The feature learning module consists of a depthwise separable convolution, the feature enhancement module consists of two inverted residual structures, and the prediction module consists of a fully connected layer and a linear regression layer. The structure of this lightweight network model is as Figure 2 shown. Specifically, the pre-training of the network includes the following steps:

[0050] 1) Use the four columns of wear data of the first four trains in the samples made in step (1) as the training set, and the four columns of wear data of the remaining one train as the test set, and then input them into the network for pre-training.

[0051] 2) First, perform an initial convolution operation on the input data using a depth convolution layer with a convolution kernel size of 3*3, padding of 1, and stride of 1 to extract low-level feature information; after the depth convolution operation is completed, apply the ReLU activation function to the output data, set all negative values to zero, and keep all positive values unchanged. Then, further extract the features through a point convolution layer with a 1*1 convolution kernel, no padding, and stride of 1. The calculation process is shown in formula (10). This feature learning module significantly reduces the number of parameters and the computational cost while effectively extracting features.

[0052]

[0053] Among them, x represents the dataset sample of the pre-trained network or the dataset sample of the fine-tuned network, and is used as the input data. represents the 3*3 depth convolution operation, and the ReLU activation function formula is max{x, 0}. represents the 1*1 point convolution operation, and y1 represents the output feature of the feature learning module.

[0054] 3) To learn richer feature representations while maintaining computational efficiency, the feature enhancement module further extracts feature information. The output features of the feature learning module first pass through the first inverted residual structure. In this structure, the features first pass through a lightweight dilated convolutional layer with a kernel size of 1*1, no padding, and a stride of 1 to increase the receptive field without adding extra parameters; then through a depthwise convolutional layer with a kernel size of 3*3, padding of 0, a stride of 1, and a group number of 8 to perform convolution independently on each input channel; finally, the features pass through a pointwise convolutional layer with a 1*1 kernel, no padding, and a stride of 1 to mix channel information and reduce the number of channels. The calculation process is shown in Equation (11):

[0055]

[0056] Among them, y1 represents the input features, represents the 1*1 lightweight dilated convolution operation, and y2 represents the output features of the first inverted residual structure of the feature enhancement module. Similarly, the second inverted residual structure is composed of a lightweight dilated convolutional layer with a kernel size of 1*1, no padding, and a stride of 1, a depthwise convolutional layer with a kernel size of 3*3, no padding, a stride of 2, and a group number of 4, and a pointwise convolutional layer with a kernel size of 1*1, no padding, and a stride of 1 stacked in sequence. While learning deep features, it further optimizes computational efficiency. The calculation process is shown in Equation (12):

[0057]

[0058] Among them, y2 represents the input features, represents the 3*3 depthwise convolution operation, and y3 represents the output features of the second inverted residual structure of the feature enhancement module.

[0059] The prediction module consists of a fully connected layer and a linear regression layer. The output features passing through the two inverted residual structures pass through the fully connected layer and a linear regression layer with 16 input feature numbers and 1 output feature number to output the RUL prediction result.

[0060] 4) In this example, the number of iterations of the model is set to 50, the training batch is set to 16, and the learning rate is set to 0.001. After each round of training, backpropagation is performed according to the obtained mean squared error (MSE) loss value, and the Adam optimizer is used to update the parameters of the network model. The definition of the mean squared error (MSE) loss function is shown in Equation (13):

[0061]

[0062] Among them, N represents the number of training samples, RUL predicted,i represents the predicted value of the model, RUL actual,iRepresents the actual label value.

[0063] After iterating to the set number of iterations, the validation set is used for validation. When the root mean square error (RMSE) of the validation result ≤ 0.10, it is considered that the model has extracted the features related to the fault degradation process at this time, that is, it meets the requirements and the training of the model is completed. If the requirements are met, the model parameters of the pre-trained network module are fixed and saved, and prepared to be transmitted to the next module. The calculation process of the root mean square error (RMSE) is shown in formula (14):

[0064]

[0065] Among them, n represents the number of samples, and y i represents the predicted value of the network model, represents the true label value.

[0066] In step (3), the twin data fine-tuning module is used to optimize the parameters of the pre-trained network model, and the twin data generated by simulation calculation is used to finely adjust the model parameters, aiming to improve the prediction robustness of the model. Specifically, first, the model parameters of the pre-trained network saved in step (2) are loaded into the initialized lightweight network model, and then the parameters of the feature learning module and the feature enhancement module are frozen. The frozen network is trained with twin data, and only the parameters of the linear regression layer are fine-tuned. The mean square error (MSE) loss function is used to calculate the loss value of the wear life prediction to guide the parameter optimization process.

[0067] Taking the MSE loss function as the training target, the Adam optimizer is used to update the network parameters of the linear regression layer. The iteration number of the model is set to 200, the training batch is 16, and the learning rate is 0.001. After each round of training, backpropagation is performed according to the obtained MSE loss value, and the training parameters are continuously adjusted. After iterating to the set number of iterations, it is verified with the validation set. When the root mean square error of the verification result ≤ 0.03, the weights of the depthwise separable convolution, the two inverted residual structures, and the weights and biases of the linear regression layer are saved.

[0068] After that, the test set in the fine-tuned network dataset is input into the network for testing, and the root mean square error (RMSE) is used to evaluate the prediction results. In order to illustrate the advantages of the lightweight convolutional neural network proposed by the present invention, it is compared with other classic life prediction models, namely CNN-LSTM and GRU. The comparison results are shown in Table 1, and the comparison is made from three aspects: mean square error, running time, and number of parameters. It can be seen from Table 1 that the lightweight method proposed by the present invention has lower number of parameters, shorter running time, lower root mean square error, and higher prediction accuracy compared with the classic prediction models.

[0069] Table 1 Comparison of the lightweight degree and prediction accuracy of different models

[0070]

[0071] In step (4), the wear life prediction result is visually output. The dataset to be predicted is sequentially input into the saved model in the way of extracting time series sample sets with a sliding window. Without updating the fixed parameters, that is, the weight w and the bias b, the current wear life prediction value is output. Then, the cumulative wear life prediction values are processed visually, a life prediction graph is drawn, and the wear life prediction is completed. In this example, the wear data of four trains, namely Train 01037, Train 01038, Train 01039, and Train 01040, are selected as the training set, and the wear data of four trains of Train 01041 are selected as the test set, and the samples of this test set are visually displayed. The life prediction graph is as Figure 3 shown.

[0072] In the example of the present invention, through the signal acquisition and preprocessing module, the lightweight model construction and pre-training module, the twin data fine-tuning module, and the wear life prediction result output module, the train wear data collected by the sensor is preprocessed and then input into the constructed lightweight convolutional neural network to complete the pre-training of the model and use the twin data of the pantograph wear to tune the parameters of the model. Finally, the life prediction value is output. The results show that the example of the present invention can reduce the number of parameters and the amount of calculation, reduce the running time, and ensure the robustness of the prediction, demonstrating the innovation and practical engineering application value of the invention.

Claims

1. A lightweight prediction method for the wear amount of pantographs of urban rail trains based on twin data, characterized in that, The following steps are involved: The data acquisition and preprocessing module is used to collect the full life cycle data of the pantograph carbon slide wear degradation failure and its pantograph wear twin data, and preprocess the original data, including standardization, sliding window sampling, labeling, and sample preparation and division steps; The lightweight model construction and pre-training module uses the processed full life cycle data of pantograph carbon slide wear degradation failure to pre-train the lightweight network model; the process includes initializing network parameters, calculating prediction loss, performing iterative training and validation set verification, saving the pre-trained model parameters, and passing them to the next module, the twin data fine-tuning module; the lightweight network model is composed of a feature learning module, a feature enhancement module and a prediction module; wherein the feature learning module adopts a deep separable convolution, which is composed of a deep convolution layer with a convolution kernel size of 3*3, a padding of 1, and a step size of 1, a ReLU activation function and a point convolution layer with a convolution kernel size of 1*1, no padding, and a step size of 1; the feature enhancement module consists of two The first inverted residual structure consists of a lightweight dilated convolution layer with a convolution kernel size of 1*1, no padding, and a stride of 1, a deep convolution layer with a convolution kernel size of 3*3, a padding of 0, a stride of 1, and a grouping number of 8, and a point convolution layer with a convolution kernel size of 1*1, no padding, and a stride of 1; the second inverted residual structure consists of a lightweight dilated convolution layer with a convolution kernel size of 1*1, no padding, and a stride of 1, a deep convolution layer with a convolution kernel size of 3*3, a padding of 0, a stride of 2, and a grouping number of 4, and a point convolution layer with a convolution kernel size of 1*1, no padding, and a stride of 1; the prediction module consists of a fully connected layer and a linear regression layer with an input feature number of 16 and an output feature number of 1; The twin data fine-tuning module optimizes the parameters of the pre-trained lightweight network model using the twin data of pantograph wear. Specifically, first, the network model parameters saved in the pre-training module are loaded into the initialized lightweight network model. Then, the parameters of the feature learning module and the feature enhancement module are frozen, and only the parameters of the linear regression layer are fine-tuned. By calculating the prediction loss, iterative training, and validation set verification, the parameter optimization process is guided. Finally, the fine-tuned model parameters are saved and transmitted to the next module, i.e., the prediction result output module. The twin data of pantograph wear is obtained in the following way: The wear of the carbon slide plate is caused by the left and right scraping of the contact wire. By the daily driving mileage and speed of the urban rail train, the daily wear times of the carbon slide plate are calculated. According to the unit mass, elastic modulus, Poisson's ratio of the contact wire, and the density, elastic modulus, Poisson's ratio of the carbon slide plate, combined with the contact pressure, current, and operating speed parameters, cyclic simulation is carried out using ANSYS Workbench to calculate the wear amount of each wear, so as to obtain the total wear thickness per day. Then, the original thickness is subtracted from the total wear thickness to obtain the starting thickness of the next day, and then simulation calculation is carried out, and so on, until the thickness of the carbon slide plate reaches the use threshold, and finally the twin data of pantograph wear is generated. According to the actual operation of the pantograph carbon slide plate, when the thickness of the carbon slide plate is less than 24 mm, it is a failure. The prediction result output module inputs the test set in the dataset of the twin data fine-tuning module into the network model after parameter fine-tuning. Without updating the fixed parameters, i.e., the weight w and the bias b, the current wear life prediction value is output, and the prediction results are accumulated. Finally, the wear life prediction value is visualized to draw a wear life prediction graph, thus completing the prediction of the wear life.

2. The pantograph wear prediction method for urban rail trains as described in claim 1, characterized in that The specific steps are as follows: S1: Data collection and preprocessing: Obtain the remaining thickness data of the pantograph carbon slide plate through the monitoring sensor. The wear of the carbon slide plate is caused by the left and right scraping of the contact wire. By the daily driving mileage and speed of the urban rail train, the daily wear times of the carbon slide plate are calculated. According to the unit mass, elastic modulus, Poisson's ratio of the contact wire, and the density, elastic modulus, Poisson's ratio of the carbon slide plate, combined with the contact pressure, current, and operating speed parameters, cyclic simulation is carried out using ANSYS Workbench to calculate the wear amount of each wear, so as to obtain the total wear thickness per day. Then, the original thickness is subtracted from the total wear thickness to obtain the starting thickness of the next day, and then simulation calculation is carried out, and so on, until the thickness of the carbon slide plate reaches the use threshold, and finally the twin data of pantograph wear is generated. According to the actual operation of the pantograph carbon slide plate, when the thickness of the carbon slide plate is less than 24 mm, it is a failure. The collected monitoring data and twin data are preprocessed by standardization to eliminate the dimensional difference between the data, and all feature values are normalized to the interval [0,1]. The maximum-minimum standardization method is used for processing, and its calculation formula is: Among them, X represents the original data, min(X) represents the minimum value in the dataset, max(X) is the maximum value in the dataset, and Normalized(X) is the data after max-min normalization; The processed data above is made into a sample set by the method of sliding window sampling. A sliding window with a length of W and a step size of N is set, and each complete time series data is sampled in sequence. Assuming the total length of the data is L, the entire time series data is finally segmented into samples, and the length of each sample is W; the data is divided into a dataset for pre-training the network and a dataset for fine-tuning the network. The pre-training network dataset is the pantograph wear data obtained by monitoring sensors, and is further divided into a training set and a validation set; the dataset for fine-tuning the network comes from the twin data obtained based on simulation and the pantograph wear data obtained by monitoring sensors. The twin data of pantograph wear is used as the training set, and the real data of pantograph wear is used as the validation set and the test set; labels are set for the above-obtained pantograph wear data and twin data. Since the remaining thickness wear amount of the pantograph continues to wear, only the label Y of the degradation stage is set, and the formula is as follows: where t c represents the current degradation moment, t w represents the operating time of the pantograph from the start of wear to complete failure, t s represents the time point of the start of wear; S2: Construct a lightweight convolutional neural network: including a feature learning module, a feature enhancement module, and a prediction module; the feature learning module is composed of a depthwise separable convolution improved based on a standard convolution, the feature enhancement module is composed of two inverted residual structures, and the prediction module is composed of a fully connected layer and a linear regression layer; Among them, the feature learning module first uses a depth convolution layer with a convolution kernel size of 3*3, padding of 1, and stride of 1 to perform an initial convolution operation on the input data to extract low-level feature information; after the depth convolution operation is completed, the ReLU activation function is applied to the output data to set all negative values to zero and keep all positive values unchanged. Then, a point convolution layer with a 1*1 convolution kernel, no padding, and stride of 1 is used to further extract the features. The calculation process is shown in formula (3). This feature learning module greatly reduces the number of parameters and computational cost while effectively extracting features; where x represents the dataset samples of the pre-trained network or the dataset samples of the fine-tuned network, and serves as the input data. represents a 3*3 depth convolution operation, and the ReLU activation function formula is max{x, 0}. represents a 1*1 point convolution operation, and y1 represents the output features of the feature learning module. The feature enhancement module is to learn richer feature representations and further extract feature information while maintaining computational efficiency; the output features of the feature learning module first pass through the first inverted residual structure of the feature enhancement module. In this structure, the features first pass through a lightweight dilated convolution layer with a convolution kernel size of 1*1, no padding, and stride of 1 to increase the receptive field without adding additional parameters; then, a depth convolution layer with a convolution kernel size of 3*3, padding of 0, stride of 1, and number of groups of 8 is used to perform convolution independently on each input channel; finally, the features pass through a point convolution layer with a 1*1 convolution kernel, no padding, and stride of 1; the calculation process is shown in formula (4): Among them, y1 represents the input feature, represents a 1×1 lightweight dilated convolution operation, and y2 represents the output feature of the first inverted residual structure of the feature enhancement module; similarly, the second inverted residual structure is composed of a 1×1 convolutional kernel, a lightweight dilated convolutional layer without padding and with a stride of 1, a depthwise convolutional layer with a convolutional kernel size of 3×3, without padding, with a stride of 2, and a group number of 4, and a point convolutional layer with a convolutional kernel size of 1×1, without padding, and with a stride of 1, which are stacked in sequence; the calculation process is shown in formula (5): Among them, y2 represents the input feature, represents a 3*3 depth convolution operation, and y3 represents the output feature of the second inverted residual structure of the feature enhancement module; The prediction module is composed of a fully connected layer and a linear regression layer with 16 input features and 1 output feature. The output features after passing through two inverted residual structures are output as the RUL prediction result through the fully connected layer and the linear regression layer; S3: Update the model parameters through the model pre-training module and perform model pre-training: input the dataset samples of the pre-trained network made into the lightweight convolutional neural network model constructed in S2 for training. Set the model iteration times, training batches, and learning rate. After each round of training, perform backpropagation according to the mean square error MSE loss value obtained by formula (6), and use the Adam optimizer to continuously adjust the training parameters; where N represents the number of training samples, and RUL predicted,i represents the predicted value of the i-th sample, and RUL actual,i represents the true value of the i-th sample; When the iteration reaches the set iteration times, verify with the validation set; when the root mean square error RMSE of the verification result ≤ 0.10, it is considered that the module has extracted the features related to the degradation process at this time, that is, it meets the requirements, and the training of the model is completed; if it meets the requirements, fix and save the model parameters of the pre-trained network module and prepare to transmit them to the next module; the calculation process of the root mean square error is shown in formula (7): Among them, n represents the number of samples, and y i represents the predicted value of the network model, and represents the true label value; S4: Optimize the parameters of the pre-trained network model saved in S3 through the twin data fine-tuning module. The specific optimization process is as follows: Use the twin data of pantograph wear obtained in S1 to tune the weights and biases of the model. Specifically, first load the model parameters of the pre-trained network saved in S3 into the lightweight network model constructed in S2. Subsequently, freeze the parameters of the feature learning module and the feature enhancement module, and only fine-tune the parameters of the linear regression layer in the prediction module. Use the twin data of pantograph wear as the training set, calculate the loss value of wear life prediction through the mean square error (MSE), and use the Adam optimizer to update the network parameters of the linear regression layer. Set the number of iterations, training batches, and learning rate of the model. After each round of training, perform backpropagation according to the obtained MSE loss value and continuously adjust the training parameters. When the iteration reaches more than the set number of iterations, i.e., 200, verify with the validation set. When the root mean square error of the verification result ≤ 0.03, save the weights of the depthwise separable convolution, the two inverted residual structures, as well as the weights and biases of the linear regression layer. S5: Input the test set in the twin data fine-tuning module dataset into the network model saved in S4, output the current wear life prediction value without updating the fixed parameters, i.e., the weights w and biases b, accumulate the wear life prediction values and perform visualization processing, draw the wear life prediction graph, and complete the wear life prediction.

Citation Information

Cited By

  • Pantograph carbon slide plate friction wear characteristic prediction method and system

    CN120724712A

  • A method and system for predicting the friction and wear characteristics of a pantograph carbon strip

    CN120724712B

  • Pantograph fault prediction method, system, medium, equipment and program product

    CN120874611A

  • Simulation running time prediction method and system, storage medium and program product

    CN121638083A