Rail transit train driving automatic control method and system and storage medium
Through the deep reinforcement learning algorithm DRL-NNC and fully connected feedforward neural network FNN combined with ε-greedy strategy, the problems of local optimal solution, high computational complexity and poor adaptability in the automatic control method of rail transit train driving are solved, and more stable, efficient and energy-saving automatic control effects are achieved.
Patent Information
- Application Number
- CN202510177642.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
AI Technical Summary
The existing automatic control method for driving rail transit trains has problems such as local optimal solutions, high computational complexity and poor adaptability to complex environments. The prediction algorithm is insufficient when dealing with uncertainty and lacks real-time feedback correction.
The deep reinforcement learning algorithm DRL-NNC is used to combine the fully connected feedforward neural network FNN and ε-greedy strategies to collect data in real time to perform balanced calculations of the recommended speed difference p and the limit speed difference q, generate control instructions, and feedback and corrections in real time to adapt to the operating routes of different locations and lengths.
It improves the stability and efficiency of automatic train control, has good energy-saving effects, can adapt to complex environments more simply and quickly, and handle emergencies in real time.
Smart Images

Figure CN120122431A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rail transit, and particularly to an automatic train driving control method, system and storage medium for rail transit. Background Art
[0002] Currently, the following problems mainly exist in the automatic train driving control method for rail transit: 1. In terms of optimization algorithms, being trapped in local optimal solutions: When many automatic control algorithms solve the optimization problem of train operation control, they are prone to being trapped in local optimal solutions. For example, in algorithms based on traditional gradient descent, a seemingly optimal solution may be found in a certain local area, but in fact, it is not the global optimum. This may lead to a train operation plan that is not the most efficient or energy-consuming overall.
[0003] High computational complexity: In order to accurately describe the complex dynamic process of train operation, some algorithms need to consider numerous constraint conditions and variables, which greatly increases the computational complexity of the algorithms. For example, when a large-scale mixed integer programming algorithm is used for train operation scheduling, as the problem scale increases, the amount of calculation increases exponentially, which may lead to too long algorithm solving time and difficulty in meeting the requirements of real-time control.
[0004] Poor adaptability to complex environments: In reality, the train operation environment is complex and changeable, including different line conditions, weather conditions, passenger flow changes, etc. Some existing optimization algorithms are often designed based on ideal conditions or simplified models. When the actual environment does not match the preset conditions, the optimization effect of the algorithm will be greatly reduced. For example, when encountering sudden line failures or extreme weather, the original optimization algorithm may not be able to adjust the strategy in time, resulting in a decrease in train operation efficiency or potential safety hazards.
[0005] 2. In terms of prediction algorithms, insufficient handling of uncertainties: There are many uncertain factors during train operation, such as signal transmission delays, wear and failures of vehicle components, etc. Existing prediction algorithms often have deficiencies in handling these uncertainties. For example, prediction algorithms based on deterministic models may not be able to accurately predict the train operation state when facing random factors such as signal transmission delays, thus affecting the accuracy of control decisions.
[0006] Lack of real-time feedback correction: When some prediction algorithms make predictions, they do not fully utilize the real-time obtained train operation data for feedback correction. During train operation, the actual situation may differ from the prediction result. However, if the algorithm cannot adjust the prediction model in time according to new data, the prediction error will continue to increase. For example, when the actual running speed of the train is different from the predicted speed, if the subsequent prediction cannot be adjusted in time, it may cause the control strategy to deviate from the actual demand. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an automatic control method, system and storage medium for the driving of rail transit trains.
[0008] The purpose of the present invention is achieved through the following technical solutions: In the first aspect of the present invention, there is provided: an automatic control method for the driving of rail transit trains, characterized by comprising the following steps: S1: Real-time collect the train running position, running speed and running time data through sensors; S2: Input the train running position, running speed and running time data into the train controller of the fully connected feedforward neural network FNN containing the deep reinforcement learning algorithm DRL-NNC for the equilibrium calculation of the recommended speed difference p and the restricted speed difference q ; S3: Generate control instructions with the ε-greedy strategy, and select the best control instructions to guide the Sigmoid activation function to adapt to the operating routes of different sections and lengths; S4: Calculate the probability values of the recommended speed difference p and the restricted speed difference q respectively, and select the recommended speed difference with the maximum speed, position reward value of the train operation and the lowest energy consumption p and the restricted speed difference q for automatic train operation control.
[0009] Preferably, the neural network of the train controller includes an input layer, a hidden layer and an output layer, wherein the input layer contains three neurons, namely the speed, slope and time of the train; the activation function of the train controller is the Sigmoid activation function; the input calculation formula of the train controller is as follows: , where p and q represent the recommended speed difference and the restricted speed difference respectively; is the output value of the jth neuron in the lth layer; and represent the number of neurons in the lth layer and the number of layers of the FNN network respectively; is the weight of the jth neuron in the lth layer; at this time k the output of is as follows: , where σ represents the Sigmoid activation function, , where e represents the running time.
[0010] Preferably, step S3 further includes the following steps: Randomly pair through ε p andq Combine them, and calculate the optimal output value of the train controller and the corresponding action at this time with a probability of 1 - ε. The calculation formula for this process is as follows: , Where and respectively represent the train operation control instruction at time t and the train operation control instruction output by the deep reinforcement learning algorithm DRL-NNC; s t represents the train displacement at time t; A represents the p output by the deep reinforcement learning algorithm DRL-NNC q and the number of balanced instruction combinations.
[0011] Preferably, step S4 further includes the following steps: When the line and the stop are greater than the preset value, the deep reinforcement learning algorithm DRL-NNC calculates the recommended speed difference p and the restricted speed difference q by continuously updating the controller parameters, and selects the recommended speed difference with the maximum speed, position reward value and the lowest energy consumption for the train operation p and the restricted speed difference q for automatic train operation control; thereby accumulating the reward value to reflect the control effect of the previous train controller; the train control strategy trajectory of the entire line is as follows: , where y t and r t respectively represent the output instruction and the reward value of the deep reinforcement learning algorithm DRL-NNC at time t; the reward calculation for each stop at this time is as follows: , where λ represents the reward value coefficient; E t represents the energy consumption at time t; T and T' respectively represent the total running time of the train on the line and the total stop time at the stations.
[0012] The second aspect of the present invention provides: An automatic train driving control system for rail transit, used to implement any of the above automatic train driving control methods for rail transit, including: An acquisition module, used to collect train running position, running speed and running time data in real time through sensors; A speed difference calculation module, which is used to input the train running position, running speed and running time data into a train controller of a fully connected feedforward neural network FNN containing a deep reinforcement learning algorithm DRL-NNC for recommending speed differences p and restricted speed differences q for equilibrium calculation; A control instruction generation module, which is used to generate control instructions with an ε-greedy strategy, select the best control instructions to guide the Sigmoid activation function to adapt to the running routes of different sections of the environment and length; An automatic control module, which is used to calculate the recommended speed differences p and restricted speed differences q for probability values, select the recommended speed differences with the maximum speed, position reward values of the train operation and the lowest energy consumption p and restricted speed differences q for automatic train operation control.
[0013] The third aspect of the present invention provides: a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, any of the above-mentioned rail transit train driving automatic control methods is realized.
[0014] The beneficial effects of the present invention are: 1) Based on DRL, control optimization is carried out through FNN and ε-greedy strategy, speed difference budget is realized, the optimal solution of controller instructions is determined, the stability and efficiency of train automatic control are improved, and good energy-saving effects are achieved.
[0015] 2) The constructed model is simpler, the calculation and reaction speed is faster, and it can adapt to more complex environments.
[0016] 3) Through FNN for control and introducing ε-greedy strategy to optimize instruction generation, real-time feedback correction can be carried out to handle more emergencies. Description of the Drawings
[0017] Figure 1 is a flow chart of the rail transit train driving automatic control method; Figure 2 is a schematic diagram of the neural network model structure of the DRL-NNC controller; Figure 3 is a schematic diagram of test results of different parameters ε and λ; Figure 4 is a schematic diagram of test comparison of reward values of different algorithms; Figure 5 is a schematic diagram of test results of train speed and traction / braking of different control methods. Specific Embodiments
[0018] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0019] Refer to Figures 1-5 , the first aspect of the present invention provides: an automatic control method for driving a rail transit train, characterized in that it includes the following steps: S1: Real-time collect the train running position, running speed and running time data through sensors; S2: Input the train running position, running speed and running time data into the train controller of the fully connected feedforward neural network FNN containing the deep reinforcement learning algorithm DRL-NNC for balanced calculation of the recommended speed difference p and the restricted speed difference q ; S3: Generate control instructions with an ε-greedy strategy, and select the best control instructions to guide the Sigmoid activation function to adapt to the operating routes of different sections and lengths; S4: Calculate the probability values of the recommended speed difference p and the restricted speed difference q respectively, and select the recommended speed difference with the largest speed, position reward value and the lowest energy consumption for the train operation p and the restricted speed difference q for automatic train operation control.
[0020] In this embodiment, the activation function of the train controller is the Sigmoid activation function. Compared with other activation functions, the Sigmoid activation function has the characteristics of being continuous, differentiable, and having an output range between (0, 1), which makes it perform well in model convergence and gradient calculation. To avoid the Sigmoid activation function falling into local optima, an ε-greedy strategy is introduced to generate control instructions, and the best instructions are selected to guide the Sigmoid activation function to adapt to the operating lines of different sections and lengths. It is simpler, more stable, and more efficient than traditional methods, addressing multiple challenges faced by the automatic control of rail transit trains in complex environments, such as signal interference, emergencies, local optima, etc. The research optimizes control through the FNN and ε-greedy strategies based on DRL, achieving speed difference budgeting and determining the optimal solution of the controller instructions. Experimental results show that when the parameter ε of the ε-greedy strategy takes a value of 0.75 and the reward coefficient λ takes a value of 1.5, the fluctuation range of the sensitivity value of the DRL-NNC model is the smallest, with a minimum range of [-1, 1.5]. Quantitative data shows that the average reward value of the DRL-NNC algorithm is the highest, approaching 6. This indicates that the proposed method in the research achieves the best energy-saving effect while ensuring the safety and comfort of the train. Taking the section between Wansheng and Guanghua Park Subway Stations on Chengdu Rail Transit Line 4 as an example, simulation tests show that the maximum running speed of the train under DRL-NNC control is 76 km / h, and the minimum number of speed control inflection points is 16. In the section test, the running time of DRL-NNC is 161.37 seconds, the late arrival time is 0.38 seconds, the parking error is 0.33 meters, and the traction energy consumption is 17.52 kW·h, all significantly superior to other methods. In summary, the DRL-NNC algorithm improves the stability and efficiency of train automatic control and has a good energy-saving effect.
[0021] The following are the performance tests for the present invention: Build a suitable experimental environment, set the CPU to Intel Xeon E5-2680, the GPU to NVIDIA Tesla V100, the memory to 64GB, and the operating system to Ubuntu 20.04. Set the number of iterations to 1000 times, the number of training steps per round to 300, to 30, and to 16. Use the High-Speed Automated Rail Transit System Dataset (HARTS) and the City Rail Transit Management Dataset (CRTM) as the test data sources. First, test the parameters ε and λ that have the greatest impact on DRL-NNC to determine the optimal hyperparameter values for subsequent model comparison. The test results are asFigure 3 as shown
[0022] Figure 3 Figure (a) shows the test results for different parameters ε, Figure 3 and figure (b) shows the test results for different parameters λ. It can be seen that when ε is selected as 0.75, the fluctuation range of the model sensitivity is the smallest, indicating that the DRL-NNC model has a more stable output state at this time. It also shows that both smaller or larger ε values will affect the accuracy of the DRL-NNC model for Figure 3 and p and q for combined probability prediction. In addition, when the change of the reward coefficient λ is similar to that of ε, there are unstable outputs of the ε-greedy policy instructions of the model under both larger and smaller λ values. Only when λ is 1.5, the effectiveness of the model is more obvious and the instruction output is more stable. It can be seen that when ε is 0.75 and λ is 1.5, the difference prediction and instruction output of the DRL-NNC model are more effective, and subsequent tests will be based on these parameters. The research introduced more advanced control algorithms of the same type for comparison, such as the Deep Deterministic Policy Gradient algorithm, the Proximal Policy Optimization algorithm, and the Double Deep Q-Network. Taking the reward value as the index for testing, multiple tests were conducted and the average value was taken. The test results are as Figure 4 shown
[0023] Figure 4 Figure (a) shows the test results of the four algorithms in the HARTS dataset, Figure 4 and figure (b) shows the test results of the four algorithms in the CRTM dataset. In the HARTS dataset, the DRL-NNC algorithm showed a rapid increase in the reward value at the initial stage of training, exceeding other algorithms, and maintained a high reward value level in the later stage of training. The Deep Deterministic Policy Gradient algorithm and the Proximal Policy Optimization algorithm also had good performances in the middle stage, but finally failed to reach the level of DRL-NNC. The Double Deep Q-Network showed relatively stable performance throughout the process, but the reward value was slightly lower. Quantitative data found that the average reward value of the DRL-NNC algorithm was the highest, approaching 6; the average reward values of the Deep Deterministic Policy Gradient algorithm and the Proximal Policy Optimization algorithm were the highest at 3 and 4 respectively; the average reward value of the Double Deep Q-Network was the highest at 5. This shows that the proposed DRL-NNC algorithm has obvious performance advantages in comparison with other algorithms in the same field, and its model output effect is more stable.
[0024] The following is a simulation test for the present invention: Taking the subway Line 4 in Chengdu as an example, the test section is randomly selected from Wansheng Station to Guanghua Park Station, with a total length of 5 kilometers. The section includes Nanxun Avenue Station, Fengxi River Station and Yangliu River Station. The train type is set as Type B, the train mass is 288 tons, the maximum traction force is 360 kN, the maximum braking force is 340 kN, and the maximum operating speed is 80 km / h. Taking the operating speed and braking distance as indicators, continue to test the train operation data under the control of the above four algorithms. The test results are as Figure 5 shown.
[0025] Figure 5 (a) shows the results of the train operating speed changes under the four control methods, Figure 5 (b) shows the results of the train traction / braking changes under the four control methods. As Figure 5 can be seen, the speed of DRL-NNC always remains the highest, with the highest driving speed of 76 km / h, and the number of speed control inflection points is the least, only 16, especially showing the best performance in the Fengxi River - Nanxun Avenue section. In the traction / braking test, the number of inflection points of the traction force and braking force of DRL-NNC is significantly the least, only 14. The number of inflection points of the Deep Deterministic Policy Gradient algorithm, Proximal Policy Optimization algorithm and Double Deep Q-Network are 20, 17 and 16 respectively. It shows that the DRL-NNC algorithm can better balance the relationship between train traction and braking, effectively saving control resources. The research conducted a comparative test with indicators such as running time, delay time, parking error, and traction energy consumption. The results are shown in Table 1.
[0026] Table 1 Multi-index test results of each method
[0027] As can be seen from Table 1, compared with the other three types of control methods, the DRL-NNC algorithm performs optimally in both sections, showing the shortest running time, the smallest delay time and parking error, and the lowest traction energy consumption. For example, in the Yangliu River - Fengxi River section, the running time of DRL-NNC is 161.37 seconds, the delay time is 0.38 seconds, the parking error is 0.33 meters, and the traction energy consumption is 17.52 kW·h. Other algorithms such as the Deep Deterministic Policy Gradient algorithm, Proximal Policy Optimization algorithm and Double Deep Q-Network perform relatively close in each index, but none of them can comprehensively exceed DRL-NNC. Especially in terms of energy consumption and accuracy, DRL-NNC is significantly better than other algorithms, demonstrating its high efficiency and accuracy in train control.
[0028] In some embodiments, the neural network of the train controller includes an input layer, a hidden layer, and an output layer. The input layer contains three neurons, namely the speed, gradient, and time of the train. The activation function of the train controller is the Sigmoid activation function. The input calculation formula of the train controller is as follows: , where p and q respectively represent the recommended speed difference and the restricted speed difference; is the output value of the j-th neuron in the l-th layer; and respectively represent the number of neurons in the l-th layer and the number of layers of the FNN network; is the weight of the j-th neuron in the l-th layer; at this time k the output of is as follows: , where σ represents the Sigmoid activation function, e represents the running time.
[0029] In some embodiments, S3 further includes the following steps: Randomly combine p and q through ε, and at the same time calculate the optimal output value of the train controller and the corresponding action at this time with a probability of 1 - ε. The calculation formula of this process is as follows: , where and respectively represent the train operation control instruction at time t and the train operation control instruction output by the deep reinforcement learning algorithm DRL-NNC; s t represents the train displacement at time t; A represents the p output by the deep reinforcement learning algorithm DRL-NNC and q equilibrium instruction combination number.
[0030] In some embodiments, S4 further includes the following steps: When the line and the stop point are greater than the preset value, the deep reinforcement learning algorithm DRL-NNC continuously updates the controller parameters, and calculates the probability values of the recommended speed difference p and the restricted speed difference q respectively, and selects the recommended speed difference with the maximum speed, position reward value of the train operation and the lowest energy consumption, p and the restricted speed difference q for automatic train operation control; thus, the cumulative reward value is used to reflect the control effect of the previous train controller; the train control strategy trajectory of the entire line in series is as follows: , where y t and r t respectively represent the output instruction and reward value of the deep reinforcement learning algorithm DRL-NNC at time t; the reward calculation for each stop is as follows: , where λ represents the reward value coefficient; E t represents the energy consumption at time t; T and T' respectively represent the total running time of the train on the line and the total stop time at stations.
[0031] The second aspect of the present invention provides: an automatic train driving control system for rail transit, used to implement any one of the above automatic train driving control methods for rail transit, including: An acquisition module, used to collect train running position, running speed and running time data in real time through sensors; A speed difference calculation module, used to input the train running position, running speed and running time data into the train controller of the fully connected feedforward neural network FNN containing the deep reinforcement learning algorithm DRL-NNC to calculate the recommended speed difference p and limit the speed difference q for balanced calculation; A control instruction generation module, used to generate control instructions with an ε-greedy strategy, and select the best control instruction to guide the Sigmoid activation function to adapt to the operating routes of different sections and lengths; An automatic control module, used to calculate the probability values of the recommended speed difference p and the limited speed difference q respectively, select the recommended speed difference with the maximum speed, position reward value of the train operation and the lowest energy consumption p and the limited speed difference q for automatic train operation control.
[0032] The third aspect of the present invention provides: a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are loaded and executed by a processor, any one of the above automatic train driving control methods for rail transit is implemented.
[0033] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in the relevant field. As long as the changes and variations made by those skilled in the art do not depart from the spirit and scope of the present invention, they should all be within the protection scope of the appended claims of the present invention.
Claims
1. A rail transit train driving automatic control method, characterized in that: The following steps are involved: S1: collects train location, speed and travel time data in real time through sensors; S2: Input the train position, speed and travel time data into the train controller of the fully connected feedforward neural network FNN containing the deep reinforcement learning algorithm DRL-NNC to recommend the speed difference p Speed limit difference q Balance calculation; S3: Generate control instructions using the ε-greedy strategy, select the best control instructions to guide the Sigmoid activation function to adapt to different environments and lengths of running routes; S4: Calculate the recommended speed difference respectively p Speed limit difference q The probability value is used to select the recommended speed difference with the maximum train running speed, position reward value, and minimum energy consumption. p Speed limit difference q Carry out automatic control of train operation.
2. The rail transit train driving automatic control method according to claim 1, characterized in that: The neural network of the train controller includes an input layer, a hidden layer and an output layer, wherein the input layer contains three neurons, which are the speed, slope and time of the train; the activation function of the train controller is a Sigmoid activation function; the input calculation formula of the train controller is as follows: ,in p and q Respectively represent the recommended speed difference and the limited speed difference; is the output value of the 𝑗th neuron in the lth layer; and Respectively represent the number of neurons in the lth layer and the number of FNN network layers; For the lth layer k The weight of the neuron; at this time The output is as follows: , where σ represents the Sigmoid activation function, e Indicates the running time.
3. The rail transit train driving automatic control method according to claim 1, characterized in that: The S3 further comprises the following steps: Through ε random pair p and q Combine them and calculate the optimal output value of the train controller with a probability of 1-ε, as well as the corresponding action at this time. The calculation formula for this process is as follows: , in and They represent the train operation control instructions at time t and the train operation control instructions output by the deep reinforcement learning algorithm DRL-NNC respectively; s t represents the train displacement at time t; A represents the output of the deep reinforcement learning algorithm DRL-NNC p and q Balance the number of instruction combinations.
4. The rail transit train driving automatic control method according to claim 1, characterized in that: The S4 further comprises the following steps: When the route and the stop are greater than the preset value, the deep reinforcement learning algorithm DRL-NNC calculates the recommended speed difference by continuously updating the controller parameters. p Speed limit difference q The probability value is used to select the recommended speed difference with the maximum train running speed, position reward value, and minimum energy consumption. p Speed limit difference q Carry out automatic control of train operation; thereby accumulating reward values to reflect the control effect of the last train controller; the train control strategy trajectory of the entire line in series is as follows: ,in y t and r t They represent the output instructions and reward values of the deep reinforcement learning algorithm DRL-NNC at time t respectively; the reward calculation for each stop is as follows: , where λ represents the reward value coefficient; E t represents the energy consumption at time t; T and T' They represent the total running time of the train line and the total stopping time of the train station respectively.
5. A rail transit train driving automatic control system, characterized in that: The method for realizing the automatic control of rail transit train driving according to any one of claims 1 to 4 comprises: The acquisition module is used to collect the train's running position, running speed and running time data in real time through sensors; The speed difference calculation module is used to input the train position, speed and travel time data into the train controller of the fully connected feedforward neural network FNN containing the deep reinforcement learning algorithm DRL-NNC to recommend the speed difference p Speed limit difference q Balance calculation; The control instruction generation module is used to generate control instructions using the ε-greedy strategy and select the best control instructions to guide the Sigmoid activation function to adapt to different environments and lengths of running routes; Automatic control module for calculating the recommended speed difference respectively p Speed limit difference q The probability value is used to select the recommended speed difference with the maximum train running speed, position reward value, and minimum energy consumption. p Speed limit difference q Carry out automatic control of train operation.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are loaded and executed by the processor, the rail transit train driving automatic control method as described in any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Method and system for planning and controlling train travelling speed
CN102514602A
Method for controlling train operation on basis of train operation grades
CN106828540A
Train automatic driving method and device, electronic equipment and storage medium
CN117022406A
Train energy-saving driving curve calculation method based on batch constraint depth Q learning
CN117922654A
Train autonomous driving calculation method based on deep reinforcement learning
CN118151679A