Train braking distance prediction method considering meteorological and line conditions
By expanding the data volume and filtering the samples in the original sample set, calculating the elimination coefficient using local coverage and boundary proximity, constructing a training set, and training a deep learning model, the problems of insufficient generalization ability and abnormal sample influence of the train braking distance prediction model under complex working conditions are solved, and more accurate braking distance prediction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU RAILWAY VOCATIONAL & TECH COLLEGE
- Filing Date
- 2026-03-18
- Publication Date
- 2026-06-16
AI Technical Summary
Existing train braking distance prediction models lack generalization ability under complex operating conditions, and the augmented samples of generative adversarial networks (GANs) contain mode collapse and anomalous samples, which affect prediction accuracy.
By expanding the data volume and screening the samples in the original sample set, a generative adversarial network is used to expand the samples outside the clusters. The first elimination coefficient is calculated by combining the local coverage and boundary proximity, and the second elimination coefficient is obtained by using a nonlinear regression model. A comprehensive elimination coefficient is constructed to screen the regenerated samples, forming a training set, and a deep learning model is trained for prediction.
This improves the prediction accuracy and generalization ability of the train braking distance prediction model under complex operating conditions, reduces the impact of mode collapse and anomalous samples in generative adversarial networks (GANs), and achieves more accurate braking distance prediction.
Smart Images

Figure CN122220877A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of train braking distance prediction technology, specifically to a train braking distance prediction method that takes into account weather and track conditions. Background Technology
[0002] With the rapid development of rail transit, accurately predicting train braking distance has become crucial for ensuring safe train operation and precise stopping, especially under complex conditions such as rain, snow, or steep gradients. The development of deep learning technology has provided a new research direction for train braking distance prediction. By training a deep learning model on multi-source data collected by an intelligent sensing system, a train braking distance prediction model adaptable to different operating conditions can be established. This model, combined with real-time multi-source data collected by the intelligent sensing system, enables real-time prediction of train braking distance. Compared to traditional methods based on physical laws, this approach can achieve accurate prediction of train braking distance under complex operating conditions.
[0003] However, data from trains operating under complex conditions such as rain, snow, or steep gradients are typically scarcer and more important than data from trains operating under normal conditions. Therefore, to improve the generalization ability of train braking distance prediction models, existing methods often augment small sample data from multi-source train data to enhance data diversity, thereby enabling the trained train braking distance prediction model to maintain high prediction accuracy even under complex conditions. For example, CN114852129A, "Train Emergency Stopping Method Based on Small Sample Data Augmented Ensemble Learning," uses a Generative Adversarial Network (GAN) to augment a small sample dataset formed from historical train braking sample data. Then, it uses the augmented training dataset to train a stacking ensemble learning model combining multiple statistical and deep learning models, resulting in a braking distance prediction ensemble learning model that predicts the maximum common braking distance of trains. However, this method ignores the mode collapse phenomenon present in traditional GANs and the abnormal samples generated due to the use of random variables, which may not conform to the physical laws of train braking distance, thus affecting the accuracy and generalization ability of the final trained train braking distance prediction model. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a train braking distance prediction method that considers weather and track conditions, thereby resolving the existing issues.
[0005] The train braking distance prediction method considering weather and track conditions proposed in this application adopts the following technical solution: One embodiment of this application provides a train braking distance prediction method that takes into account weather and track conditions, including the following steps: Obtain the original sample set for training the train braking distance prediction model. Each sample consists of train data, track data, and meteorological data for each braking moment during the historical operation of the target train. The data volume of the samples in the original sample set is expanded to obtain an enhanced sample set, which is divided into initial samples and regenerated samples. The spatial distribution characteristics between each regenerated sample and the original data are used to obtain the first elimination coefficient of each regenerated sample. The second elimination coefficient of each regenerated sample is obtained through the nonlinear relationship between the braking distance in the train data of each regenerated sample and other types of data. Then, the first elimination coefficient is combined to obtain the comprehensive elimination coefficient, which is used to screen and eliminate regenerated samples to obtain the training set. A neural network model is trained using the training set and used as a train braking distance prediction model to predict the braking distance of the target train.
[0006] Preferably, the step of augmenting the data volume of samples in the original sample set to obtain an enhanced sample set includes: The original sample set is clustered, and the cluster with the largest number of samples is used as the base sample set. Generative adversarial networks are used to expand the sample data in each cluster other than the base sample set to obtain each expanded sample set. All expanded sample sets are combined into an enhanced sample set, and the samples in the enhanced sample set that do not belong to the original sample set are used as regenerated samples.
[0007] Preferably, the spatial distribution characteristics between each regenerated sample and the original data include: For each regenerated sample, the regenerated sample and the original sample set are clustered as a whole, and the local density of the regenerated sample is used as the local coverage of the regenerated sample. The proximity of the boundaries of the regenerated samples is constructed by the distance relationship between the center points of each cluster after the regenerated samples and the original sample sets are clustered.
[0008] Preferably, the process of obtaining the boundary proximity of the regenerated sample specifically includes: The Euclidean distance between the center point of the cluster v after the regenerated sample and the original sample set is recorded as the first distance. From all the boundary points of the cluster v, the boundary point with the smallest Euclidean distance to the regenerated sample is selected, and the Euclidean distance between the selected boundary point and the center point of the cluster v is recorded as the second distance. The boundary proximity of the regenerated sample is calculated by combining the first distance and the second distance.
[0009] Preferably, the ratio of the first distance to the second distance is calculated, and the mean of the corresponding ratios between the regenerated sample and all clusters is used as the boundary proximity of the regenerated sample.
[0010] Preferably, the first rejection coefficient of each regenerated sample is obtained by using the local coverage degree and the boundary proximity degree, wherein the first rejection coefficient is positively correlated with the local coverage degree and negatively correlated with the boundary proximity degree.
[0011] Preferably, the process of obtaining the second elimination coefficient for each regenerated sample includes: A nonlinear regression model is obtained between the train braking distance of the regenerated sample and any type of data other than braking distance. The residual of the nonlinear regression model is used as the statistical deviation factor of the regenerated sample under any type of data. The second elimination coefficient of the regenerated sample is obtained by combining the statistical deviation factors of the regenerated sample under all other types of data except braking distance.
[0012] Preferably, the mean of the statistical deviation factors of the regenerated sample under all data types other than braking distance is used as the second elimination coefficient of the regenerated sample.
[0013] Preferably, the comprehensive elimination coefficient is positively correlated with both the first elimination coefficient and the second elimination coefficient.
[0014] Preferably, the minimum value of the comprehensive elimination coefficient of the regenerated samples in the expanded sample set is used as the comprehensive elimination coefficient of the original samples in the expanded sample set. Samples in the expanded sample set are eliminated in descending order of comprehensive elimination coefficient until the number of samples in the expanded sample set after elimination is the same as that in the basic sample set. All the expanded sample sets after elimination and the basic sample set are combined to form the training set.
[0015] This application has at least the following beneficial effects: This application divides all sample data in the original sample set acquired by the intelligent sensing system of the target train, and expands the sample data under the train industrial control conditions with a small amount of sample data in the original sample set. Then, it analyzes the regenerated samples in the enhanced sample set to construct a comprehensive elimination coefficient, and eliminates all the expanded sample sets. Finally, it uses the train braking distance prediction model trained on the training set obtained after elimination and the real-time data collected by the intelligent sensing system of the target train to realize the braking distance prediction of the target train. Compared to traditional methods that directly use the augmented training dataset obtained by Generative Adversarial Networks (GANs) to train train braking distance prediction models, this application can effectively reduce the impact of pattern collapse in GANs and abnormal samples that do not conform to the physical laws of train braking distance due to the use of random variables in sample generation on the accuracy and generalization ability of the finally trained train braking distance prediction model, thereby improving the accuracy of the final train braking distance prediction results. Attached Figure Description
[0016] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating the steps of a train braking distance prediction method that takes into account weather and track conditions, provided in this application. Detailed Implementation
[0018] To further illustrate the technical means and effects adopted by this application to achieve the intended purpose of the invention, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a train braking distance prediction method considering weather and track conditions proposed in this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0019] Unless otherwise defined, terms such as “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a circuit structure, article, or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such an article or device. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the article or device that includes said element. Furthermore, the term “and / or” as used herein includes any and all combinations of one or more of the associated listed items. All technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0020] The following description, in conjunction with the accompanying drawings, details a specific scheme for a train braking distance prediction method that takes into account weather and track conditions, as provided in this application.
[0021] One embodiment of this application provides a train braking distance prediction method that considers weather and track conditions. For details, please refer to [link to specific implementation details]. Figure 1 This includes the following steps: The train braking distance predicted in this application is used for the target distance indication of smooth stopping in the station by the train driver assistance system (DAS) or for offline operation energy consumption simulation analysis, and is not used to replace the emergency braking safety calculation of the automatic train protection system (ATP).
[0022] Step 1: Obtain the historical braking sample dataset for subsequent training of the train braking distance prediction model, and preprocess the sample data in the obtained historical braking sample dataset.
[0023] First, a historical braking sample dataset of the target train is compiled. Each sample in this dataset consists of train data, track data, and meteorological data from each braking instance during the target train's historical operation. In this embodiment, the train data includes braking distance, pressure, and initial braking speed; the track data includes track gradient (positive for uphill slopes and negative for downhill slopes) and track curvature; and the meteorological data includes precipitation and temperature. The data acquisition sensors installed on the train mainly include GPS positioning sensors, pressure sensors, speed sensors, rainfall sensors, and temperature sensors. The implementer can select and set the data types for the target train according to the actual application scenario; this embodiment does not impose any special restrictions on this.
[0024] The GPS positioning sensor is used to acquire track data such as braking distance, track gradient, and track curvature. The pressure, speed, rainfall, and temperature sensors are used to acquire pressure, initial braking speed, precipitation, and temperature data, respectively. The braking distance can be obtained based on the initial and stopping positions of the train during braking collected by the GPS positioning sensor. The track gradient and track curvature can be obtained by matching the initial braking position of the train collected by the GPS positioning sensor with the track database (which stores the track gradient and track curvature at various positions on the track where the train is located).
[0025] Further, all sample data in the acquired historical braking sample dataset are sequentially subjected to data cleaning and data normalization. The resulting dataset is denoted as the original sample set A. The data cleaning process includes missing value imputation and outlier handling. Missing value imputation uses the mean imputation method, and outlier handling uses the Z-score method. The data normalization process uses the Min-Max normalization method, where the upper and lower bounds of the value range for each type of train data, line data, and meteorological data are used as the maximum and minimum values for normalization when the Min-Max normalization method is applied to each type of train data, line data, and meteorological data in the sample data. The upper and lower bounds of the value range for train data can be determined according to the "Railway Technical Management Regulations," and the upper and lower bounds of the value range for line data can be determined according to the "Railway Line Design Specification" (TB). The upper and lower limits of the range of meteorological data (10098-2017) can be determined based on the historical precipitation and temperature data of the area covered by the target train line. Data cleaning and normalization are well-known technologies, and the specific process will not be described in detail.
[0026] Step 2: Expand the data volume of the samples in the original sample set to obtain an enhanced sample set, and divide it into initial samples and regenerated samples. Utilize the spatial distribution characteristics between each regenerated sample and the original data to obtain the first elimination coefficient of each regenerated sample. Obtain the second elimination coefficient of each regenerated sample by using the nonlinear relationship between the braking distance in the train data of each regenerated sample and other types of data. Then combine the first elimination coefficient to obtain the comprehensive elimination coefficient, which is used to screen and eliminate regenerated samples to obtain the training set.
[0027] Since mode collapse in Generative Adversarial Networks (GANs) refers to the phenomenon of generating a large number of similar samples due to the mismatch between the minimax optimization objective domain and the iterative optimization method of the GAN, directly using a GAN to augment all the sample data in the original sample set A may easily lead to excessive redundancy of sample data under certain train operating conditions in the augmented original sample set A. If this augmented sample set is directly used to train the train braking distance prediction model, the model will overfit those redundant sample data under train operating conditions and will not be able to fully learn the data distribution characteristics and distinguishing features under different train operating conditions, thereby reducing the model's generalization ability.
[0028] Therefore, this embodiment divides all sample data in the acquired original sample set A to specifically expand the sample data under train control conditions with a small amount of sample data in the original sample set A, and analyzes and removes all sample data under each expanded train control condition so that the model can better learn the real data distribution characteristics and distinguishing features under different train control conditions.
[0029] S1 augments the data volume of the samples in the original sample set to obtain an enhanced sample set.
[0030] S1.1 Specifically, each type of data in the original sample set A is treated as a dimension. A multidimensional data space V1 is constructed using all the sample data in the original sample set A. This space is used to distinguish the sample data in the original sample set A that belong to different train control conditions of the target train. The sample data in the original sample set A corresponds one-to-one with the data points in the multidimensional data space V1.
[0031] S1.2 The K-Means clustering algorithm is used to cluster all data points in the multidimensional data space V1 to obtain multiple clusters in the multidimensional data space V1. These clusters are used to represent the data point clusters composed of all sample data belonging to different train control conditions of the target train in the original sample set A. They are used to obtain sample data in the original sample set A that need to be expanded. The number of clusters in the K-Means clustering algorithm is determined by the silhouette coefficient method. The K-Means clustering algorithm and the determination of the number of clusters are well-known techniques, and the specific process will not be described in detail.
[0032] S1.3 Select the cluster with the largest number of samples as the basic sample set a (if there are multiple clusters, choose one). Use a generative adversarial network (GAN) to expand the sample data in each cluster other than the basic sample set a, so that the number of sample data in each cluster after expansion is a preset multiple of the number of sample data in the basic sample set a, to obtain each expanded sample set. Combine all the expanded sample sets into an enhanced sample set B, and complete the expansion processing of the sample data under train control conditions with a small amount of sample data in the original sample set A. In this embodiment, the preset multiple is 5, which can be set by the implementer. The generative adversarial network is a well-known technology, and the specific process will not be described in detail.
[0033] S2 divides the augmented sample set into initial samples and regenerated samples, constructs the first and second elimination coefficients for each regenerated sample, and then obtains the comprehensive elimination coefficient, which is used to screen and eliminate regenerated samples to obtain the training set.
[0034] S2.1 Specifically, based on whether the sample data in the enhanced sample set B belongs to the original sample set A, the sample data in the enhanced sample set B is divided into initial samples and regenerated samples. This is used to evaluate which regenerated samples need to be removed from the enhanced sample set B, so as to ensure that the train braking distance prediction model learns the true data distribution pattern when it is trained. Among them, the samples in the enhanced sample set that belong to the original sample set are called initial samples, and otherwise they are called regenerated samples.
[0035] S2.2 All sample data in the enhanced sample set B are also mapped to the multidimensional data space V1 to obtain the mapped multidimensional data space V2, which is used for subsequent analysis of the data distribution characteristics of the regenerated samples in the enhanced sample set B and the sample data in the original sample set A in the multidimensional data space V2. For the convenience of subsequent processing, in this embodiment, the data points corresponding to the sample data in the original sample set A and the regenerated samples in the enhanced sample set B are respectively denoted as the original sample point and the regenerated sample point.
[0036] S2.3 Generally, Generative Adversarial Networks (GANs) are prone to generating a large number of repetitive samples clustered at the distribution center (i.e., pattern collapse) when training is insufficient. When augmenting the sample data under train control conditions in the original sample set A, which has a relatively small amount of sample data, the augmented sample data usually needs to better fill the blank areas or sparse areas covered by the sample data in the data space of the sample data under train control conditions. Furthermore, the augmented sample data needs to better fill the boundary areas between sample data under different train control conditions in the original sample set A, so that the subsequent training of the train braking distance prediction model can better learn the data distribution characteristics and distinguishing features under different train control conditions. Therefore, in this embodiment, excessively clustered redundant samples in the center of the augmented sample set B are removed, while effective sample data with higher information content distributed at the edges is retained. This ensures that the sample data in the resulting sample dataset can better cover the blank, sparse, and boundary areas of the data space of the sample data in the original sample set A.
[0037] S2.3.1 Specifically, taking any regenerated sample b in the enhanced sample set B as an example, the local coverage degree of the regenerated sample b is calculated based on the spatial distribution of the regenerated sample point V2(b) and all the original sample points in the multidimensional data space V2. This is used to evaluate the degree to which the local data space where the regenerated sample b is located is covered by the sample data in the original sample set A. The more original sample points exist in the local space where the regenerated sample point V2(b) is located in the multidimensional data space V2, the greater the local coverage degree. In this embodiment, the process of obtaining the local coverage degree is as follows: Using the coordinates of all original sample points in the multidimensional data space V2 and the regenerated sample point V2(b) as input, the density peak clustering (DPC) algorithm is used to calculate the local density of the regenerated sample point V2(b), and the local density is used as the local coverage of the regenerated sample point V2(b). The larger the local density, the more original sample points exist in the local space of the regenerated sample point V2(b) in the multidimensional data space V2, and the greater the local coverage.
[0038] It should be noted that, during density peak clustering, the cutoff distance parameter used to calculate local density is obtained by calculating the distance between all pairs of original sample points in the multidimensional data space, sorting the distances in descending order, and taking the distance value at the 2nd percentile as the cutoff distance for clustering. In practical applications, implementers can set their own existing methods for calculating the cutoff distance; this embodiment does not impose any special limitations on this.
[0039] S2.3.2 For each cluster in the multidimensional data space V1, that is, each cluster after the original sample set is divided, in this embodiment, taking any cluster v as an example, the Euclidean distance between the coordinates of the cluster center point of cluster v and the coordinates of the regenerated sample point V2(b) is denoted as the first distance between the regenerated sample point V2(b) and cluster v, which is used to characterize the distance between the regenerated sample point V2(b) and the cluster center point of cluster v. At the same time, the Euclidean distance from all original sample points in cluster v to the cluster center point is calculated, and the distances are arranged in descending order. The top 5% of the original sample points with the largest distances are taken to form the boundary point set of cluster v. From all the boundary points of cluster v, the boundary point with the smallest Euclidean distance to the regenerated sample point V2(b) is selected (if there are multiple, one is selected). In this embodiment, the selected boundary point is further... The Euclidean distance between the coordinates of the boundary point and the coordinates of the cluster center point is denoted as the second distance between the regenerated sample point V2(b) and the cluster v. This distance characterizes the distance between the boundary point of cluster v and the cluster center point of cluster v along the ray direction with the cluster center point of cluster v and the regenerated sample point V2(b) as the starting and ending points, respectively. The ratio of the first distance to the second distance is denoted as the degree of distance between the regenerated sample point V2(b) and the cluster v from the regional center in this embodiment for ease of understanding. This is used to evaluate the degree to which the regenerated sample point V2(b) is far from the center of the spatial region where cluster v is located. The greater the degree of distance from the regional center, the more the regenerated sample b is located in the boundary region of the data space where the sample data under the train control conditions corresponding to cluster v is located.
[0040] It should be noted that in this embodiment, in order to avoid the denominator being zero when calculating the ratio, a preset minimum positive number (such as 0.001) is added to the denominator. In actual application scenarios, the implementer can set the specific value of the preset minimum positive number as they see fit. This embodiment does not impose any special restrictions on this.
[0041] The mean value of the distance between the regional centers of the regenerated sample point V2(b) and all clusters in the multidimensional data space V1 is denoted as the boundary proximity of the regenerated sample point V2(b). This is used to evaluate the degree to which the location of the regenerated sample b in the data space is close to the boundary region of the sample data in the original sample set A under various train control conditions. The larger the mean value, the greater the boundary proximity and the greater the degree of proximity to the boundary region.
[0042] S2.3.3 The local coverage and boundary proximity of all regenerated samples in the enhanced sample set B are normalized using the Min-Max normalization method to map the values of the local coverage and boundary proximity to [0,1] for subsequent data processing. Taking regenerated sample b as an example, the first elimination coefficient of regenerated sample b is calculated based on the Min-Max normalization results of the local coverage and boundary proximity of regenerated sample b. The first elimination coefficient is positively correlated with the local coverage and negatively correlated with the boundary proximity. The first elimination coefficient is used to evaluate whether to remove regenerated sample b from the enhanced sample set B. The larger the Min-Max normalization result of the local coverage and the smaller the Min-Max normalization result of the boundary proximity, the larger the first elimination coefficient. This indicates that removing regenerated sample b has less impact on the learning of the real data distribution characteristics and distinguishing features under different train control conditions during the subsequent training of the train braking distance prediction model, and therefore, it is more necessary to remove regenerated sample b from the enhanced sample set B.
[0043] It should be noted that a negative correlation means that the two variables have opposite trends, while a positive correlation means that the two variables have the same trend. In this embodiment, the first elimination coefficient is calculated as the ratio between the normalized result of the local coverage degree and the normalized result of the boundary proximity degree. It should be noted that during the ratio calculation, a very small positive number (0.001 in this embodiment) is added to the denominator to avoid the denominator being 0. The Min-Max normalization method is a well-known technique, and the specific process will not be described in detail.
[0044] S2.4 Because Generative Adversarial Networks (GANs) use random variables to generate samples, the augmented sample set B obtained by the GAN may contain anomalous sample data that does not conform to the physical laws of train braking distance. Therefore, this application analyzes the physical laws of train braking distance in the sample data of the augmented sample set B to remove anomalous sample data that does not conform to the physical laws of train braking distance. This allows the subsequent train braking distance prediction model to learn the real data distribution characteristics and distinguishing features under different train control conditions during training.
[0045] S2.4.1 Specifically, under normal circumstances, the braking distance of a train usually exhibits obvious physical laws with other train data, track data, and meteorological data. For example, the faster the initial velocity of the train during braking, the longer the braking distance; the greater the pressure during braking, the greater the braking force; the faster the deceleration, the shorter the braking distance; the greater the track gradient during braking (the gradient in the uphill direction is positive, and the gradient in the downhill direction is negative); the greater the component of the train's gravity along the slope, the shorter the braking distance; the greater the precipitation during braking; and the smaller the friction between the train and the track, the longer the braking distance. Therefore, in order to effectively remove abnormal sample data in the enhanced sample set B that does not conform to the physical laws of train braking distance generated by the Generative Adversarial Network (GAN), the following processing is performed.
[0046] Taking the regenerated sample b as an example, the cluster v(b) with the smallest Euclidean distance between the cluster center point and the regenerated sample point V2(b) is selected from all clusters in the multidimensional data space V1. This cluster is used to characterize the clusters corresponding to all sample data of a certain train control condition in the original sample set A in the multidimensional data space V1.
[0047] S2.4.2 Considering that the braking distance of a train usually exhibits obvious physical patterns compared to other train data, track data, and meteorological data, for the braking distance of a train, taking any one of the other train data (excluding braking distance), track data, and meteorological data, C, as an example, the two dimensions corresponding to the braking distance and data C in the multidimensional data space V1 are obtained. Based on the magnitude of the coordinate components of all original sample points in the cluster v(b) in the above two dimensions, a nonlinear regression model between the two variables corresponding to the two dimensions is constructed. The nonlinear regression model uses a second-order polynomial for regression fitting. The residual of the regenerated sample point V2(b) in the nonlinear regression model is denoted as the statistical deviation factor of the regenerated sample b under data C, which is used to evaluate whether the regenerated sample b conforms to the physical patterns of braking distance and data C under the train control conditions to which the regenerated sample b belongs. The larger the residual, the larger the statistical deviation factor, and the less the regenerated sample b conforms to the physical patterns. The construction of the nonlinear regression model is a well-known technique, and the specific process will not be elaborated further.
[0048] Based on all statistical deviation factors of the regenerated sample b under the train data (excluding braking distance), track data, and meteorological data, a second elimination coefficient for the regenerated sample b is calculated. This coefficient is used to assess whether the regenerated sample b should be removed from the augmented sample set B. The larger the statistical deviation factors, the more likely the regenerated sample b is to be an abnormal sample data generated by the generative adversarial network (GAN) in the augmented sample set B that does not conform to the physical laws of train braking distance. Therefore, the regenerated sample b needs to be removed more, i.e., the second elimination coefficient is larger. Thus, there is a positive correlation between the second elimination coefficient and the statistical deviation factors. The second elimination coefficient is calculated as the mean or weighted sum of all statistical deviation factors. In this embodiment, the mean is preferred.
[0049] S2.5 Further, the first and second elimination coefficients of all regenerated samples in the enhanced sample set B are normalized using the Min-Max normalization method to map the values of the first and second elimination coefficients to [0,1] for subsequent data processing. Taking regenerated sample b as an example, the comprehensive elimination coefficient of regenerated sample b is calculated based on the Min-Max normalization results of the first and second elimination coefficients of regenerated sample b. This coefficient is used to evaluate whether regenerated sample b should be eliminated from the enhanced sample set B. The comprehensive elimination coefficient is positively correlated with the first and second elimination coefficients. That is, the larger the second and first elimination coefficients are, the larger the comprehensive elimination coefficient is. Otherwise, the smaller the second and first elimination coefficients are, the smaller the comprehensive elimination coefficient is. The Min-Max normalization method is a well-known technique, and the specific process will not be described in detail.
[0050] It should be noted that the calculation method of the comprehensive elimination coefficient can be the mean, product or weighted sum of the second elimination coefficient and the first elimination coefficient. Preferably, the mean is selected in this embodiment.
[0051] The minimum comprehensive elimination coefficient of the regenerated samples in the expanded sample set is taken as the comprehensive elimination coefficient of the original samples in the expanded sample set. For each expanded sample set, samples in the expanded sample set are eliminated in descending order of comprehensive elimination coefficient until the number of samples in the expanded sample set after elimination is the same as the number of samples in the basic sample set. The dataset formed by all the expanded sample sets after elimination and the basic sample set a is recorded as the training set, which is used to train the train braking distance prediction model.
[0052] Step 3: Use the training set to train a deep learning model and use it as a train braking distance prediction model to predict the braking distance of the target train.
[0053] Furthermore, the braking distance of the sample data in the training set is used as the output of the deep learning model. The historical sequence train data, line data, and meteorological data of the target train within a set time period are uniformly constructed into a fixed-dimensional feature vector matrix, which is used as the input of the train braking distance prediction model to train the train braking distance prediction model. The deep learning model is a convolutional neural network or a gated recurrent neural network. In this embodiment, a convolutional neural network is selected. The performance evaluation index of the deep learning model is accuracy or precision. In this embodiment, accuracy is selected. The training of the train braking distance prediction model based on the deep learning model is a well-known technique, and the specific process will not be described in detail.
[0054] Furthermore, when the target train begins to brake, all types of train data (excluding braking distance), track data, and meteorological data of the target train are converted into the data format used as input to the deep learning model for training the train braking distance prediction model. This data serves as input to the trained train braking distance prediction model, and the braking distance prediction result of the target train for this braking action is output, thus enabling the prediction of train braking distance.
[0055] It is understood that references to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include the specific features, structures, or characteristics described in connection with that embodiment. Therefore, the appearance of phrases such as "in one embodiment," "in some embodiments," "in other embodiments," or "in still other embodiments" in different parts of this specification does not necessarily refer to the same embodiment, but rather means "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0056] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous. Moreover, the sequence numbers of the steps in the embodiments do not imply a specific order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments in this specification.
[0057] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for predicting train braking distance considering meteorological and track conditions, characterized in that, Includes the following steps: Obtain the original sample set for training the train braking distance prediction model. Each sample consists of train data, track data, and meteorological data for each braking moment during the historical operation of the target train. The data volume of the samples in the original sample set is expanded to obtain an enhanced sample set, which is divided into initial samples and regenerated samples. The spatial distribution characteristics between each regenerated sample and the original data are used to obtain the first elimination coefficient of each regenerated sample. The second elimination coefficient of each regenerated sample is obtained through the nonlinear relationship between the braking distance in the train data of each regenerated sample and other types of data. Then, the first elimination coefficient is combined to obtain the comprehensive elimination coefficient, which is used to screen and eliminate regenerated samples to obtain the training set. A neural network model is trained using the training set and used as a train braking distance prediction model to predict the braking distance of the target train.
2. The train braking distance prediction method considering weather and track conditions as described in claim 1, characterized in that, The process of augmenting the original sample set to obtain an enhanced sample set includes: The original sample set is clustered, and the cluster with the largest number of samples is used as the base sample set. Generative adversarial networks are used to expand the sample data in each cluster other than the base sample set to obtain each expanded sample set. All expanded sample sets are combined into an enhanced sample set, and the samples in the enhanced sample set that do not belong to the original sample set are used as regenerated samples.
3. The train braking distance prediction method considering meteorological and track conditions as described in claim 2, characterized in that, The spatial distribution characteristics between each regenerated sample and the original data include: For each regenerated sample, the regenerated sample and the original sample set are clustered as a whole, and the local density of the regenerated sample is used as the local coverage of the regenerated sample. The proximity of the boundaries of the regenerated samples is constructed by the distance relationship between the center points of each cluster after the regenerated samples and the original sample sets are clustered.
4. The train braking distance prediction method considering meteorological and track conditions as described in claim 3, characterized in that, The process of obtaining the boundary proximity of the regenerated sample specifically includes: The Euclidean distance between the center point of the cluster v after the regenerated sample and the original sample set is recorded as the first distance. From all the boundary points of the cluster v, the boundary point with the smallest Euclidean distance to the regenerated sample is selected, and the Euclidean distance between the selected boundary point and the center point of the cluster v is recorded as the second distance. The boundary proximity of the regenerated sample is calculated by combining the first distance and the second distance.
5. The train braking distance prediction method considering meteorological and track conditions as described in claim 4, characterized in that, The ratio of the first distance to the second distance is calculated, and the mean of the corresponding ratios between the regenerated sample and all clusters is used as the boundary proximity of the regenerated sample.
6. The train braking distance prediction method considering meteorological and track conditions as described in claim 5, characterized in that, The first rejection coefficient of each regenerated sample is obtained by using the local coverage degree and the boundary proximity degree, wherein the first rejection coefficient is positively correlated with the local coverage degree and negatively correlated with the boundary proximity degree.
7. The train braking distance prediction method considering meteorological and track conditions as described in claim 1, characterized in that, The process of obtaining the second elimination coefficient for each regenerated sample includes: A nonlinear regression model is obtained between the train braking distance of the regenerated sample and any type of data other than braking distance. The residual of the nonlinear regression model is used as the statistical deviation factor of the regenerated sample under any type of data. The second elimination coefficient of the regenerated sample is obtained by combining the statistical deviation factors of the regenerated sample under all other types of data except braking distance.
8. The train braking distance prediction method considering weather and track conditions as described in claim 7, characterized in that, The mean of the statistical deviation factors of the regenerated sample under all data types except braking distance is used as the second elimination coefficient of the regenerated sample.
9. The train braking distance prediction method considering meteorological and track conditions as described in claim 1, characterized in that, The overall elimination coefficient is positively correlated with both the first elimination coefficient and the second elimination coefficient.
10. The train braking distance prediction method considering meteorological and track conditions as described in claim 2, characterized in that, The minimum value of the comprehensive elimination coefficient of the regenerated samples in the expanded sample set is taken as the comprehensive elimination coefficient of the original samples in the expanded sample set. Samples in the expanded sample set are eliminated in descending order of comprehensive elimination coefficient until the number of samples in the expanded sample set after elimination is the same as that in the basic sample set. All the expanded sample sets after elimination and the basic sample set are combined to form the training set.