DQN design method and system combined with high-speed train operation condition decomposition

By decomposing the operating conditions of high-speed trains and designing a DQN system, the problem of high-speed trains with high-speed trains under complex road conditions is solved, and energy saving optimization is achieved.

CN120337398APending Publication Date: 2025-07-18LANZHOU JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510386316.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Under complex road conditions, high-speed trains have a problem of high-speed energy consumption during multi-zone operation.

Method used

By combining high-speed train operating conditions decomposition, a DQN system is designed, including defining a deep reinforcement learning environment, designing Markov decision-making processes, establishing state and action space models, designing reward functions and value functions, and using agent Agent to learn optimal strategies to optimize train operation.

Benefits of technology

On the premise of ensuring that the train is on time and on time, a large-scale comprehensive energy saving has been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337398A_ABST
    Figure CN120337398A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer technology and machine learning, in particular to a DQN design method and system combined with high-speed train operation condition decomposition, and the system comprises a deep reinforcement learning environment unit which defines a deep reinforcement learning DQN environment of high-speed train operation according to the operation environment of a high-speed train; wherein the operation environment of the high-speed train comprises a static environment and a dynamic environment, and the static environment comprises the speed limit, the gradient, the curve radius and the tunnel length of a high-speed railway line; the dynamic environment comprises the running speed, acceleration, running time and displacement of the high-speed train; according to the method, DNQ design of high-speed train operation is carried out by decomposing the four operation conditions of the high-speed train, and by solving the interval energy-saving operation conditions and adjusting the multi-interval operation timetable of the high-speed train, large-range comprehensive energy saving can be achieved under the condition that it is ensured that the starting point and the ending point of the high-speed train are punctual.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer technology and machine learning technology, and in particular, to a DQN design method and system combined with the decomposition of high-speed train operation conditions. Background Art

[0002] The operation process of a high-speed train is to control the operation process of the high-speed train according to four operation conditions of traction - cruise - coasting - braking under the conditions of meeting safety, comfort, high efficiency, punctuality, and on-time. The operation goals of high-speed trains are different, and the combination processes of the four operation conditions of traction - cruise - coasting - braking are also different. If the operation goal is the shortest time and the highest efficiency, then at this time, three operation conditions of maximum traction - cruise - maximum braking can be adopted, but this process is also the process with the largest energy consumption; if the operation goal is energy conservation, then at this time, three operation conditions of traction - coasting - braking can be adopted, but this process has low efficiency and long time. In addition, when a high-speed train operates under different road conditions, different speed limits, and different external disturbances, the four operation conditions of traction - cruise - coasting - braking will have different combinations. Therefore, finding the optimal combination of operation conditions to achieve the optimal goal has always been the purpose of our research.

[0003] Deep Reinforcement Learning (DRL) combines the perception ability of deep learning and the decision-making ability of reinforcement learning, can directly perform control according to the input, and is an artificial intelligence method closer to the human way of thinking. Deep reinforcement learning can be used to handle complex problems with high-dimensional state spaces and action spaces. It uses a deep neural network to approximate the value function or policy function, thereby realizing learning and decision-making for the environment.

[0004] As a typical representative of deep reinforcement learning, the Deep Q Network (DQN) has been widely used in many fields. The DQN can obtain the maximum reward according to the value function of reinforcement learning, enable the intelligent agent to select the optimal action, and achieve the optimal goal. Therefore, the DQN algorithm designed by combining the decomposition of high-speed train operation conditions can provide technical support for fields such as energy-saving operation optimization of high-speed trains. Summary of the Invention

[0005] The present invention provides a DQN design method and system combined with the decomposition of high-speed train operation conditions, which overcomes the above-mentioned deficiencies of the prior art and can effectively solve the problem of large energy consumption existing in the multi-section operation process of high-speed trains under complex road conditions.

[0006] To solve the above problems, one of the technical solutions of the present invention is achieved by the following: A DQN design method combined with the decomposition of high-speed train operation conditions, comprising the following steps:

[0007] Define the deep reinforcement learning DQN environment for the operation of high-speed trains according to the operation environment of high-speed trains. Among them, the operation environment of high-speed trains includes a static environment and a dynamic environment. The static environment includes the speed limit, gradient, curve radius and tunnel length of the high-speed railway line. The dynamic environment includes the speed, acceleration, operation time and displacement of the high-speed train during operation.

[0008] According to the subdivision rules of the operation section of high-speed trains, interact the intelligent agent Agent of deep reinforcement learning DQN with the operation environment of high-speed trains, and design a Markov decision process.

[0009] Combine the dynamic environment of high-speed train operation to establish a state space model of deep reinforcement learning DQN.

[0010] According to the four operation conditions of traction, cruise, coasting and braking of high-speed trains, decompose the four operation conditions into 13 actions, and establish an action space model of deep reinforcement learning DQN.

[0011] According to the process of the intelligent agent Agent interacting with the operation environment of high-speed trains, maximize the cumulative reward to guide the intelligent agent Agent to learn the optimal strategy, and design 6 reward functions. Among them, the sum of the 6 reward functions is the reward obtained in the Markov decision process.

[0012] Design the value function of deep reinforcement learning DQN according to the intelligent agent Agent obtaining the maximum reward corresponding to the reward function by selecting the optimal action.

[0013] Design the Q network according to the state space model and the action space model.

[0014] The result of establishing the state space model of deep reinforcement learning DQN is as follows:

[0015]

[0016] In the formula, S i is the state space; a i , v i , s i , t i are respectively the current running acceleration, speed, distance and time of the high-speed train; v max is the maximum running speed allowed by the line; s, t g , a max are respectively the section length, the given running time of the section and the maximum acceleration of the high-speed train, a max = 1m / s 2 ; Set the initial state as S1 = [0.46, 0, 0, 1], and the termination state as S t = [a i, v i / v max , s i / s, (t g -t i ) / t g 。

[0017] The action space model of the above-established deep reinforcement learning DQN is as follows:

[0018] A s = [n t , 0, n b

[0019] In the formula, n t and n b are the traction level and braking level of the high-speed train in the action space respectively.

[0020] The above six reward functions are the speed limit reward J v , the energy consumption reward J e , the time reward J t , the accurate stop reward J s , the comfort reward J c and other rewards J o ; specifically as follows:

[0021] Speed limit reward J v :

[0022]

[0023] Among them, v is the line running speed; v min is the minimum allowable running speed of the line; v max is the maximum allowable running speed of the line;

[0024] Energy consumption reward J e :

[0025]

[0026] In the formula, F t (v) is the traction force of the high-speed train; M is the mass of the high-speed train; △t is the running time interval of the high-speed train, △t = t(i) - t(i - 1), t(i) is the time of the high-speed train at the current i-th moment, and t(i - 1) is the time of the previous moment of the high-speed train at the current i-th moment; μ1 is the coefficient of the energy consumption reward;

[0027] Time reward J t :

[0028]

[0029] ​Wherein, t is the actual running time of the high-speed train in the interval; Δs is the current moving step of the high-speed train, Δs = s(i) - s(i - 1), s(i) is the displacement of the high-speed train at the current i-th moment, and s(i - 1) is the displacement of the high-speed train at the previous moment of the current i-th moment; μ2 is the coefficient of on-time reward;

[0030] Precise stop reward J s :

[0031]

[0032] Wherein, s e is the difference between the stop point of the high-speed train and the designated point of the platform, s e = s i - s; μ3 is the coefficient of precise stop reward;

[0033] Comfort reward J c :

[0034]

[0035] Wherein, c h is the expected comfort value; a(i) is the action taken by the agent at the current moment, and a(i - 1) is the action taken by the agent at the previous moment of the current moment; μ4 is the coefficient of comfort reward;

[0036] Other rewards J o :

[0037]

[0038] Wherein, J o1 is the single-step update reward; J o2 is the end completion reward; μ5 and μ6 are the single-step update reward coefficient and the end completion reward coefficient respectively.

[0039] The value function of the above-designed deep reinforcement learning DQN is as follows:

[0040]

[0041] Wherein, γ is the discount rate of the reward; the behavior policy π represents the probability value of selecting action A in state S, π = P(A|S); T is the total number of time steps, and R t+k is the reward value.

[0042] The above-designed Q network includes a target Q network and an actual Q network;

[0043] The actual Q-network receives the experience data during the interaction between the agent and the high-speed train operation environment, updates the parameters, selects actions according to the current policy, evaluates the values of different actions by calculating the actual Q-values, and selects the optimal actions based on the actual Q-values to optimize the decision-making process;

[0044] The target Q-network periodically copies the parameters from the actual Q-network, calculates the target Q-values, and provides a stable target for training.

[0045] The second technical solution of the present invention is implemented in the following way: A DQN design system combined with the decomposition of high-speed train operation conditions includes,

[0046] The deep reinforcement learning environment unit defines the deep reinforcement learning DQN environment for the operation of the high-speed train according to the operation environment of the high-speed train; among them, the operation environment of the high-speed train includes a static environment and a dynamic environment, the static environment includes the speed limit, gradient, curve radius and tunnel length of the high-speed railway line; the dynamic environment includes the speed, acceleration, operation time and displacement of the high-speed train operation;

[0047] The Markov unit is designed to interact with the operation environment of the high-speed train based on the agent of the deep reinforcement learning DQN according to the subdivision rules of the high-speed train operation section, and design the Markov decision-making process;

[0048] The state space model unit combines the dynamic environment of the high-speed train operation to establish the state space model of the deep reinforcement learning DQN;

[0049] The action space model unit decomposes the four operation conditions of traction, cruise, inertia and braking of the high-speed train into 13 actions, and establishes the action space model of the deep reinforcement learning DQN;

[0050] The reward function unit is designed to guide the agent to learn the optimal policy by maximizing the cumulative reward during the interaction between the agent and the high-speed train operation environment, and design 6 reward functions; among them, the sum of the 6 reward functions is the reward obtained in the Markov decision-making process;

[0051] The value function unit is designed to design the value function of the deep reinforcement learning DQN according to the agent obtaining the maximum reward corresponding to the reward function by selecting the optimal action;

[0052] The Q-network unit is designed to design the Q-network according to the state space model and the action space model.

[0053] The present invention conducts DNQ design for the operation of high-speed trains by decomposing four operation conditions of high-speed trains. Through solving the energy-saving operation conditions in sections and adjusting the operation timetables of high-speed trains in multiple sections, it can achieve comprehensive energy saving in a large range while ensuring the punctuality of the starting point and the end point of high-speed trains. Description of the Drawings

[0054] The following further elaborates in detail the specific implementation manners of the present invention with reference to the drawings.

[0055] Figure 1 It is the flowchart of the method in Embodiment 1 of the present invention.

[0056] Figure 2 It is the DQN structure diagram designed in Embodiment 1 of the present invention.

[0057] Figure 3 It is the schematic diagram of the Markov decision process in Embodiment 1 of the present invention.

[0058] Figure 4 It is the system block diagram of Embodiment 2 of the present invention. Specific Implementation Manners

[0059] The present invention is not limited by the following embodiments, and the specific implementation manners can be determined according to the technical solutions of the present invention and the actual situation.

[0060] Embodiment 1, as Figure 1 shown, the embodiment of the present invention discloses a DQN design method combined with the decomposition of high-speed train operation conditions, including the following steps:

[0061] S101. Define the deep reinforcement learning DQN environment for the operation of high-speed trains according to the operation environment of high-speed trains. Among them, the operation environment of high-speed trains includes a static environment and a dynamic environment. The static environment includes the speed limit, gradient, curve radius, and tunnel length of the high-speed railway line. The dynamic environment includes the speed, acceleration, operation time, and displacement of high-speed train operation.

[0062] S102. Design a Markov decision process according to the subdivision rules of the high-speed train operation section and the interaction between the intelligent agent Agent of the deep reinforcement learning DQN and the operation environment of the high-speed train.

[0063] S103. Establish a state space model of the deep reinforcement learning DQN in combination with the dynamic environment of high-speed train operation.

[0064] S104. Decompose the four operation conditions of traction, cruise, inertia, and braking of high-speed trains into 13 actions, and establish an action space model of the deep reinforcement learning DQN.

[0065] S105. During the interaction process between the agent and the high-speed train operation environment, six reward functions are designed by maximizing the cumulative reward to guide the agent to learn the optimal strategy. Among them, the sum of the six reward functions is the reward obtained in the Markov decision process.

[0066] S106. According to the agent obtaining the maximum reward corresponding to the reward function by selecting the optimal action, the value function of the deep reinforcement learning DQN is designed.

[0067] S107. According to the state space model and the action space model, a Q-network is designed.

[0068] Among them, in step S101, the static environment includes parameters such as the speed limit, slope, curve radius, and tunnel length of the high-speed railway line, which are used to initialize the line information of the high-speed railway in the deep reinforcement learning. The dynamic environment includes parameters such as the speed, acceleration, running time, and displacement of the high-speed train operation, which are used to update the operation state of the high-speed train in the deep reinforcement learning. It can be seen that the parameters of the static environment and the dynamic environment need to be obtained. For example: Select a section of the high-speed railway line, and its parameters such as the speed limit, slope, curve radius, and tunnel length can be directly obtained from the line data; the parameters such as the speed, acceleration, running time, and displacement of the high-speed train operation are obtained by real-time detection of on-vehicle sensors or by reading the data of the microcomputer monitoring system.

[0069] Among them, in step S102, the subdivision rule of the high-speed train operation interval means that the high-speed train operation interval is subdivided into several small sections (where the small section refers to: the agent sends 1 action per cycle, and the running distance generated by the high-speed train executing this action). Then, the high-speed train operation process can be represented by a Markov decision process, as Figure 3 shown. Its essence is that the high-speed train operation starts from state S1, the deep reinforcement learning agent selects action A1, and interacts with the high-speed train operation environment to obtain reward R1, and then transfers to the next state S2 according to the state transition probability. The deep reinforcement learning agent selects action A2 and interacts with the high-speed train operation environment to obtain reward R2. Repeat this step until the high-speed train completes the entire operation process, which is the Markov decision process. Thus, the high-speed train operation process is executed according to the designed Markov decision process, which can enable the high-speed train to achieve a large range of comprehensive energy saving during multi-interval operation.

[0070] In the above step S103, the state space model of the reinforcement learning is established, and the result is as follows:

[0071]

[0072] In the formula, S i is the state space; ai , v i , s i , t i are respectively the current running acceleration, speed, distance and time of the high-speed train; v max is the maximum running speed allowed by the line; s, t g , a max are respectively the length of the high-speed train section, the given running time of the section and the maximum acceleration, a max = 1 m / s 2 ; Set the initial state as S1 = [0.46, 0, 0, 1], and the termination state as S t = [a i , v i / v max , s i / s, (t g -t i ) / t g .

[0073] Among them, it should be noted that in this step, the maximum acceleration a of the high-speed train is taken max = 1 m / s 2 , set the initial state as S1 = [0.46, 0, 0, 1], and then take the speed, acceleration, displacement and running time data of the high-speed train running state to update the state space until the high-speed train completes the entire running process.

[0074] In the above step S104, an action space model of deep reinforcement learning DQN is established, and the result is as follows:

[0075] A s = [n t , 0, n b

[0076] In the formula, n t and n b are respectively the traction level and braking level of the high-speed train in the action space.

[0077] Among them, in step S104, the four operating conditions are decomposed into 13 actions, that is, the traction level n t ∈ [1 / 6, 2 / 6, 3 / 6, 4 / 6, 5 / 6, 1] 6 actions; the braking level n b ∈ [-1, -5 / 6, -4 / 6, -3 / 6, -2 / 6, -1 / 6] 6 actions; A s = 0 is 1 action; a total of 13 actions. When A s = 0, the high-speed train is in the coasting condition, and the traction condition and the braking condition cannot exist simultaneously.

[0078] n t ​∈[1 / 6, 2 / 6, 3 / 6, 4 / 6, 5 / 6, 1] represents dividing the traction of the high-speed train into 6 levels, where n t = 1 represents the maximum traction level; n t = 5 / 6 represents 5 / 6 times the maximum traction level; n t = 4 / 6 represents 4 / 6 times the maximum traction level; n t = 3 / 6 represents 3 / 6 times the maximum traction level; n t = 2 / 6 represents 2 / 6 times the maximum traction level; n t = 1 / 6 represents 1 / 6 times the maximum traction level.

[0079] n b ∈[-1, -5 / 6, -4 / 6, -3 / 6, -2 / 6, -1 / 6] represents dividing the braking of the high-speed train into 6 levels, where n b = -1 represents the maximum braking level; n b = -5 / 6 represents 5 / 6 times the maximum braking level; n b = -4 / 6 represents 4 / 6 times the maximum braking level; n b = -3 / 6 represents 3 / 6 times the maximum braking level; n b = -2 / 6 represents 2 / 6 times the maximum braking level; n b = -1 / 6 represents 1 / 6 times the maximum braking level.

[0080] Therefore, the cruise condition does not need to be considered here. Only the traction, coasting, and braking operating conditions are decomposed into 13 actions to establish an action space model for deep reinforcement learning. It should be noted that the maximum traction and maximum braking of the high-speed train need to be determined according to the actual vehicle type and its traction and braking characteristics. For example, for the CRH3C type high-speed train, the maximum traction power P = 8800 kW and the maximum braking power P = 8000 kW.

[0081] In the above step S105, the 6 reward functions are respectively the speed limit reward J v , the energy consumption reward J e , the time reward J t , the precise stop reward J s , the comfort reward J c and other rewards J o ; specifically as follows:

[0082] The speed limit reward J v :

[0083]

[0084] Among them, v is the line running speed; v min is the minimum allowable running speed of the line; v maxis the maximum operating speed allowable for the line;

[0085] Energy consumption reward J e :

[0086]

[0087] In the formula, F t (v) is the traction force of the high - speed train; M is the mass of the high - speed train; △t is the running time interval of the high - speed train, △t = t(i) - t(i - 1), t(i) is the time of the high - speed train at the current i - th moment, and t(i - 1) is the time of the previous moment of the high - speed train at the current i - th moment; μ1 is the coefficient of the energy consumption reward;

[0088] Time reward J t :

[0089]

[0090] In the formula, t is the actual running time of the high - speed train in the section; s is the position on the platform or a specified displacement; △s is the current moving step of the high - speed train, △s = s(i) - s(i - 1), s(i) is the displacement of the high - speed train at the current i - th moment, and s(i - 1) is the displacement of the high - speed train at the previous moment of the current i - th moment; μ2 is the coefficient of the punctuality reward;

[0091] Precise stop reward J s :

[0092]

[0093] In the formula, s e is the difference between the stop point of the high - speed train and the specified point on the platform, s e = s(i) - s; μ3 is the coefficient of the precise stop reward;

[0094] Comfort reward J c :

[0095]

[0096] In the formula, c h is the expected comfort value; a(i) is the action taken by the agent at the current moment, and a(i - 1) is the action taken by the agent at the previous moment of the current moment; μ4 is the coefficient of the comfort reward;

[0097] Other rewards J o :

[0098]

[0099] In the formula, J o1 is the single - step update reward; Jo2 is the end - point completion reward; μ5 and μ6 are the single - step update reward coefficient and the end - point completion reward coefficient respectively.

[0100] Among them, in the process of the agent of the deep reinforcement learning DQN interacting with the environment, by maximizing the cumulative reward to guide the agent to learn the optimal policy. Therefore, at a certain moment here, the reward of the agent is the sum of 6 reward values, that is, J = J v + J e + J t + J s + J c + J o ; Figure 3 In the Markov decision process shown, the rewards obtained in each cycle, namely R1, R 2、 ···, R t are respectively the sums of the 6 reward functions corresponding to each cycle.

[0101] In the above step S106, the value function of the deep reinforcement learning DQN is designed, and the result is as follows:

[0102]

[0103] In the formula, γ is the discount rate of the reward; the behavioral policy π represents the probability value of selecting action A in state S, π = P(A|S); T is the total number of time steps, and R t+k is the reward value.

[0104] In the above step S107, as Figure 2 shown, the Q - network is designed, including the target Q - network and the actual Q - network;

[0105] The actual Q - network receives the experience data in the process of the agent interacting with the high - speed train operation environment, updates the parameters, selects actions according to the current policy, evaluates the value of different actions by calculating the actual Q - value, and selects the optimal action according to the actual Q - value to optimize the decision - making process;

[0106] The target Q - network periodically copies the parameters from the actual Q - network, calculates the target Q - value, and provides a stable target for training.

[0107] When designing the Q - network above, the data collected in the process of the agent interacting with the high - speed train operation environment will be stored in the experience replay buffer, and these data will be stored in the form of a quadruple of state, action, reward, and next state; among them, each time the agent interacts with the high - speed train operation environment, new experience data will be added to the experience replay buffer, and the data in the experience replay buffer is:

[0108] P t+i ={S t+i,A t+i ,R t+i ,S t+i+1}

[0109] When the experience replay buffer reaches its capacity limit, the earliest stored data is discarded to make room for new experience data.

[0110] When the actual Q-network updates its parameters, a batch of experience tuples is randomly retrieved from the experience replay buffer for training during each update. Based on the current policy action selection, the actual Q-value is calculated to evaluate the value of different actions, and the optimal action is selected according to the actual Q-value to optimize the decision-making process; by periodically copying the parameters from the actual Q-network, the target Q-value is calculated and provided as a stable target for training.

[0111] When designing the Q-network, it is necessary to make the actual Q-value output by the actual Q-network and the target Q-network close to the target Q-value. That is, by quantifying the error between the actual Q-value and the target Q-value, if the error meets the requirements, the Q-value is output; if the error does not meet the requirements, the weights of the actual Q-value are adjusted by error backpropagation so that the error between the actual Q-value output by the network and the target Q-value meets the requirements; among them, according to the current state S and the current action A, the actual Q-value is calculated through the actual Q-network; according to the next state S′ and the next optimal action A′, the target Q-value is calculated through the target Q-network.

[0112] It should be noted that the Q-network designed in the present invention based on 4 inputs in the state space, 13 outputs in the action space, and 3 hidden layers with 64 nodes in each hidden layer is only a design method.

[0113] In summary, the present invention decomposes the four operating conditions of high-speed trains for the DNQ design of high-speed train operation. This design can achieve comprehensive energy conservation in a large range while ensuring the punctuality of the starting and ending points of high-speed trains by solving the interval energy-saving operation conditions and adjusting the multi-interval operation schedule of high-speed trains.

[0114] Example 2, as Figure 4 shown, the embodiment of the present invention discloses a DQN design system combined with the decomposition of high-speed train operating conditions, including

[0115] A deep reinforcement learning environment unit defines the deep reinforcement learning DQN environment for high-speed train operation according to the operating environment of high-speed trains; among them, the operating environment of high-speed trains includes a static environment and a dynamic environment. The static environment includes the speed limit, gradient, curve radius, and tunnel length of the high-speed railway line; the dynamic environment includes the speed, acceleration, running time, and displacement of high-speed train operation.

[0116] Design a Markov unit. According to the subdivision rules of the operating range of high-speed trains, an intelligent agent of the deep reinforcement learning DQN interacts with the operating environment of high-speed trains to design a Markov decision process;

[0117] A state space model unit, combined with the dynamic environment of high-speed train operation, to establish a state space model of the deep reinforcement learning DQN;

[0118] An action space model unit. According to the four operating conditions of traction, cruising, inertia, and braking of high-speed trains, the four operating conditions are decomposed into 13 actions to establish an action space model of the deep reinforcement learning DQN;

[0119] Design a reward function unit. According to the intelligent agent interacting with the operating environment of high-speed trains, by maximizing the cumulative reward to guide the intelligent agent to learn the optimal strategy, design 6 reward functions; among them, the sum of the 6 reward functions is the reward obtained in the Markov decision process;

[0120] Design a value function unit. According to the intelligent agent obtaining the maximum reward corresponding to the reward function by selecting the optimal action, design the value function of the deep reinforcement learning DQN; design a Q-network unit. According to the state space model and the action space model, design a Q-network.

Claims

1. A DQN design method combined with the decomposition of the operating conditions of high-speed trains, characterized in that, It includes the following steps: Define the deep reinforcement learning DQN environment for the operation of high-speed trains according to the operation environment of high-speed trains. Among them, the operation environment of high-speed trains includes a static environment and a dynamic environment. The static environment includes the speed limit, gradient, curve radius and tunnel length of the high-speed railway line. The dynamic environment includes the speed, acceleration, operation time and displacement of high-speed train operation; According to the subdivision rules of the operation section of high-speed trains, interact the agent of the deep reinforcement learning DQN with the operation environment of high-speed trains, and design a Markov decision process; Combined with the dynamic environment of high-speed train operation, establish a state space model of deep reinforcement learning DQN; According to the four operation conditions of traction, cruise, inertia and braking of high-speed trains, decompose the four operation conditions into 13 actions, and establish an action space model of deep reinforcement learning DQN; According to the process of the agent interacting with the operation environment of high-speed trains, by maximizing the cumulative reward, guide the agent to learn the optimal strategy, and design 6 reward functions. Among them, the sum of the 6 reward functions is the reward obtained in the Markov decision process; Design the value function of deep reinforcement learning DQN according to the agent obtaining the maximum reward corresponding to the reward function by selecting the optimal action; Design a Q network according to the state space model and the action space model; 2. The DQN design method based on the decomposition of the operating conditions of high-speed trains according to claim 1, characterized in that, The result of establishing the state space model of deep reinforcement learning DQN is as follows: Where S i is the state space; a i , v i , s i , t i are respectively the current running acceleration, speed, distance and time of the high-speed train; v max is the maximum allowable running speed of the line; s, t g , a max are respectively the length of the high-speed train section, the given running time of the section and the maximum acceleration, a max = 1 m / s 2 ; Set the initial state as S1 = [0.46, 0, 0, 1], and the termination state as S t = [a i , v i / v max , s i / s, (t g -t i ) / t g .

3. A DQN design method combined with the decomposition of high-speed train operation conditions according to claim 1, characterized in that, The result of establishing the action space model of deep reinforcement learning DQN is as follows: A s = [n t , 0, n b ​ where n t and n b are the traction level and braking level of the high-speed train in the action space respectively.

4. A DQN design method combined with the decomposition of the operating conditions of high-speed trains according to claim 1, characterized in that The six reward functions are respectively the speed limit reward J v , the energy consumption reward J e , the time reward J t , the precise parking reward J s , the comfort reward J c , and other rewards J o ; specifically as follows: Speed limit reward J v : wherein, v is the line operation speed; v min is the minimum allowable operation speed of the line; v max is the maximum allowable operation speed of the line; Energy consumption reward J e : where F t (v) is the traction force of the high-speed train; M is the mass of the high-speed train; Δt is the running time interval of the high-speed train, Δt = t(i) - t(i - 1), t(i) is the time at the current i-th moment of the high-speed train, and t(i - 1) is the time at the previous moment of the current i-th moment of the high-speed train; μ1 is the coefficient of energy consumption reward; Time Reward J t : In the formula, t is the actual operation time of the high-speed train section; △s is the current moving step of the high-speed train, △s = s(i) - s(i - 1), s(i) is the displacement of the high-speed train at the current i-th moment, and s(i - 1) is the displacement of the high-speed train at the previous moment of the current i-th moment; μ2 is the coefficient of the on-time reward; Precise Parking Reward J s : where s e is the difference between the high-speed train stop point and the designated point on the platform, s e = s i - s; μ3 is the coefficient of the precise stop reward; Comfort Reward J c : where c h is the desired comfort value; a(i) is the action taken by the agent at the current moment, and a(i - 1) is the action taken by the agent at the previous moment of the current moment; μ4 is the coefficient of the comfort reward; Other rewards J o : where J o1 is the single-step update reward; J o2 is the end completion reward; μ5 and μ6 are the single-step update reward coefficient and the end completion reward coefficient, respectively.

5. A DQN design method combined with the decomposition of high-speed train operation conditions according to claim 1, characterized in that The result of designing the value function of deep reinforcement learning DQN is as follows: where γ is the discount rate of the reward; the behavioral policy π represents the probability value of selecting action A in state S, π = P(A|S); T is the total number of time steps, and R t+k is the reward value.

6. The DQN design method combined with the decomposition of high-speed train operation conditions according to claim 1, characterized in that The designed Q network includes a target Q network and an actual Q network; The actual Q network receives the experience data in the process of the agent interacting with the operation environment of high-speed trains, updates the parameters, selects actions according to the current strategy, evaluates the values of different actions by calculating the actual Q value, and selects the optimal action according to the actual Q value to optimize the decision-making process; The target Q network periodically copies the parameters from the actual Q network, calculates the target Q value, and provides a stable target for training; 7. A DQN design system combined with the decomposition of the operating conditions of high-speed trains, characterized in that, It includes The deep reinforcement learning environment unit defines the deep reinforcement learning DQN environment for the operation of high-speed trains according to the operation environment of high-speed trains. Among them, the operation environment of high-speed trains includes a static environment and a dynamic environment. The static environment includes the speed limit, gradient, curve radius and tunnel length of the high-speed railway line. The dynamic environment includes the speed, acceleration, operation time and displacement of high-speed train operation; The design Markov unit interacts the agent of the deep reinforcement learning DQN with the operation environment of high-speed trains according to the subdivision rules of the operation section of high-speed trains, and designs a Markov decision process; State space model unit, which combines the dynamic environment of high-speed train operation to establish the state space model of deep reinforcement learning DQN; Action space model unit, which decomposes the four operating conditions of traction, cruise, coasting and braking of high-speed trains into 13 actions according to the four operating conditions, and establishes the action space model of deep reinforcement learning DQN; Reward function design unit, which, in the process of the agent interacting with the high-speed train operation environment, guides the agent to learn the optimal strategy by maximizing the cumulative reward, and designs 6 reward functions; among them, the sum of the 6 reward functions is the reward obtained in the Markov decision process; Value function design unit, which designs the value function of deep reinforcement learning DQN according to the agent obtaining the maximum reward corresponding to the reward function by choosing the optimal action; Q-network unit, which designs the Q-network according to the state space model and the action space model.