An Optimization Method for Multi-objective Control of Freight Trains Combining Pneumatic and Electric Systems
Through reinforcement learning and multi-parameter brake cylinder model, the matching between air braking and electric braking is optimized, and the precise control problem of long downhill sections is solved, achieving improved train operation performance and reduced energy consumption.
Patent Information
- Application Number
- CN202411886730.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In the prior art, traditional empirical models have insufficient precise control in long downhill sections, poor matching of air brakes and electric brakes, and high computational complexity of fluid dynamics models, making it difficult to apply in real-time braking optimization or online control, and lack highly targeted solutions.
The Q-learning algorithm in reinforcement learning is adopted, combined with the detailed model of multi-parameter brake cylinders, and the matching strategy between air braking and electric braking is optimized. By decomposing the action vector and calculating mechanics information, the train status is updated in real time, and an optimization method of air-electric joint multi-objective control is designed.
It significantly improves the train operating performance, improves the average speed by 10%, reduces the energy loss of brake operation by 70%, and improves operating stability and safety.
Smart Images

Figure CN119807769B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rail transit operation control and energy-saving optimization, and particularly relates to a multi-objective operation optimization method for air-electric combined control of freight trains. Background Art
[0002] In the prior art, the force characteristics during the operation of a single train, the specific effects of air braking force and running resistance are studied. A mathematical model that can accurately describe the dynamic behavior of the train is constructed through longitudinal dynamics modeling. A fine equivalent model of the brake cylinder is constructed by simplifying fluid mechanics, and complex non-linear characteristics during the train operation such as historical release factors, altitude, and end-of-train exhaust are considered. Combining reinforcement learning theory, especially Markov decision model (MPD) and Q-learning algorithm, a method for air-electric combined control of long and steep grade freight trains is designed.
[0003] Disadvantages of the prior art:
[0004] 1) The description accuracy of traditional empirical models is insufficient. The simplified processing of empirical models leads to a large deviation from the actual braking performance in practical applications. Especially in the case of long downhill sections with higher requirements for precise control, the model shows obvious limitations.
[0005] 2) Although the air braking model based on fluid dynamics can theoretically describe the braking process more accurately, its computational complexity is high, the solution speed is slow, and it is difficult to be applied to braking optimization or online control in real time.
[0006] 3) The air-electric braking matching strategy lacks optimization. Traditional braking strategies are mainly based on fixed logic rules or simple optimization algorithms, and the synergistic effect between air braking and electric braking is not fully considered.
[0007] Existing research mainly focuses on the optimal design of speed curves, mainly considering some safety factors, but it is relatively simple in practical application scenarios. Especially in the precise control of long and steep grades and complex working conditions, there is a lack of targeted solutions. Summary of the Invention
[0008] In view of the above deficiencies in the prior art, the present invention provides a multi-objective operation optimization method for air-electric combined control of freight trains, which solves the problem of poor matching between air braking and electric braking during the cyclic braking process of heavy-haul trains on long downhill sections, and significantly improves the train operation performance.
[0009] To achieve the above object, the technical solution adopted by the present invention is: A multi-objective operation optimization method for air-electric combined control of freight trains, comprising the following steps:
[0010] S1. Set training parameters, vehicle information, and freight train line information, and initialize the reinforcement learning environment and variables;
[0011] S2. Calculate and select an action using the policy function;
[0012] S3. Based on the selected action, decompose the action vector to obtain the current step traction and braking force and the air brake state, and obtain a new state by calculating the mechanical information;
[0013] S4. Calculate the state value of the new state;
[0014] S5. Update the Q-value related to the current state in the Q-table;
[0015] S6. Determine whether the current step number n exceeds the maximum number of rounds N. If so, go to step S7; otherwise, return to step S3;
[0016] S7. Use the total reward to determine whether the updated Q-table converges. If so, go to step S9; otherwise, go to step S8;
[0017] S8. Determine whether the current number of iterations reaches the maximum number of iterations. If so, go to step S9; otherwise, reset the train variables and return to step S3;
[0018] S9. The agent executes actions according to the Q-table, calculates a new state, obtains the final sequence of working conditions, and completes the multi-objective operation optimization of the air-electric combined control of the freight train.
[0019] The beneficial effects of the present invention are as follows: The present invention uses the Q-learning algorithm in reinforcement learning to optimize the matching strategy of the air brake timing and the electric braking force. Compared with the traditional method, the optimized strategy more precisely balances the roles of air braking and electric braking, significantly reducing the energy loss in the braking operation. Experimental results show that compared with the actual operation of the driver, the present invention can increase the average running speed of the train by about 10%, while reducing the cumulative change of the control force by about 70%, significantly improving the speed following performance and running stability. In addition, the optimized strategy of the present invention is simple to operate in actual operation, can adapt to different braking performance conditions, and significantly reduces the impact of braking on the train stability, thereby improving the safety and reliability of the train operation. Generally speaking, the present invention shows significant technical effects in terms of performance improvement, energy consumption reduction, and running stability.
[0020] Further, the step S2 includes the following steps:
[0021] S201. Randomly generate a random number p ∈ [0, 1], and calculate the probability that the current agent follows the policy function π(A s |S d )
[0022] S202. Compare the random number p with the action selected by the agent according to the policy function π(A s |S d ). If the random number is less than or equal to it, then select the action A with the maximum state value in the current Q-table d * . If is greater than it, then randomly select an action that can be executed in the current state S d .
[0023] The beneficial effect of the above further solution is that through the above random selection steps, the mechanism has the exploration ability for different operations under the same working conditions, ensuring the extensiveness of the learning strategy.
[0024] Furthermore, the expression of the current agent according to the policy function π(A s |S d ) is as follows:
[0025]
[0026] where A d |S d represents taking action A in state S d , A d represents the optimal action under the current conditions, rand(A(S d * )) represents randomly selecting an action that the agent in state S d can execute, δ represents the exploration rate decay rate, ep represents the number of learning iterations, ep d represents the decay period, and ε0 represents the initial value of the random exploration rate. T
[0027] The beneficial effect of the above further solution is that by introducing the comprehensive design of random selection and the policy function π, it can increase the exploration ability of the algorithm while ensuring the execution of the optimal action, avoiding falling into local optimal solutions. In addition, this solution effectively balances the stability of policy execution and learning efficiency by introducing a decay mechanism, further improving the reliability and adaptability of the system.
[0028] Furthermore, step S3 includes the following steps:
[0029] S301. Decompose the action vector based on the selected action to obtain the current-step traction braking force and the air braking state;
[0030] S302. Calculate the mechanical information through the dynamics of the freight train according to the current-step traction braking force and the air braking state;
[0031] S303. Calculate the relevant state variables of the freight train based on the mechanical information to obtain the new state S d+1 .
[0032] The beneficial effect of the above further solution is that by decomposing the operation vector, the states of the traction force and the air braking force can be accurately obtained, thereby providing more accurate data support for the dynamic modeling of the train and strengthening the system's perception ability of the train operation state.
[0033] Furthermore, the expression for decomposing the action vector is as follows:
[0034] A d = [F d , Ab d
[0035] where A d represents the action at the d-th step, F d represents the traction and braking force at the d-th step, and Ab d represents the air braking state at the d-th step.
[0036] The beneficial effect of the above further solution is that it can accurately distinguish the traction force and the air braking force, ensure the refined expression of the mechanical state, and provide more reliable basic data for the dynamic modeling.
[0037] Furthermore, the step S302 includes the following steps:
[0038] S3021. Construct a detailed multi-parameter brake cylinder model based on exponential function fitting and correction function;
[0039] S3022. Calculate the corresponding air braking force according to the detailed multi-parameter brake cylinder model;
[0040] S3023. Calculate the basic running resistance of the freight train and the additional resistance of the whole vehicle based on the train kinematic equation, and combine with the air braking force to obtain the mechanical information; the expression of the train kinematic equation is as follows:
[0041]
[0042] F b = ∑F b,i (t B , t R , v)
[0043] where v represents the train running speed, t represents the system time, M represents the total vehicle weight, F t represents the train traction force, F eb represents the total vehicle air braking force, W0 represents the basic running resistance of the whole vehicle, W j represents the additional resistance of the whole vehicle, Fb Denotes the air braking force of the whole vehicle of the rigid multi-particle model, F b,i Denotes the air braking force of the i-th carriage, t B Denotes the braking duration of air braking, t R Denotes the release duration of air braking, and x denotes the train displacement.
[0044] The beneficial effects of the above further solution are as follows: By establishing a multi-parameter brake cylinder equivalent model considering historical release time and altitude change, the present invention can quickly and accurately solve the brake cylinder pressure and air braking force, thereby more precisely describing the change of air braking performance during the cyclic braking process, and effectively overcoming the deficiency in accuracy of the traditional two-stage empirical model.
[0045] Furthermore, the expression of the multi-parameter brake cylinder detailed model is as follows:
[0046]
[0047] Among them, P i (t) represents the function of the brake cylinder pressure changing with time, P m,i Denotes the initial pressure of the brake cylinder, P s,i Denotes the steady-state pressure of the brake cylinder, β i Denotes the function of the brake cylinder pressure increasing, t represents the system time, t delay,i Denotes the time when the brake cylinder pressure starts to increase;
[0048] The expression of the air braking force is as follows:
[0049]
[0050] Among them, F b,i (t) represents the function of the air braking force pressure changing with time, S h Denotes the piston area, α b Denotes the mechanical structure transmission ratio, eff denotes the braking efficiency, Denotes the brake pad friction coefficient;
[0051] The expressions of the basic running resistance and the additional resistance of the whole vehicle of the freight train are as follows:
[0052] W0 = ∑m i ·g·w0·10 -3
[0053] w0 = a + bv + cv 2
[0054] W j = ∑m i ·g·(w g,i + w r,i+w s,i )·10 -3
[0055] wherein, m i represents the mass of the i-th carriage, g represents the acceleration due to gravity, w0 represents the basic resistance per unit of the carriage, a, b, and c represent empirical coefficients, v represents the running speed of the train, and w g,i represents the ramp resistance of the i-th carriage, and w r,i represents the curve resistance of the i-th carriage when turning, and w s,i represents the air resistance of the i-th carriage.
[0056] The beneficial effects of the above further solution are as follows: Through resistance modeling and mechanical parameter calculation, the accuracy of train dynamics analysis is enhanced, providing important data support for energy-saving operation and control optimization of the train.
[0057] Furthermore, the expression of the new state S d+1 is as follows:
[0058]
[0059] wherein, S d+1 represents the new state of the system at the (d + 1)-th step, represents the brake pad friction coefficient, S d represents the state of the train at the d-th step, g represents the acceleration due to gravity, A d represents the operation performed by the train, acc d () represents the calculation function of the train acceleration, F t represents the train traction force, V d represents the train speed at the d-th step, W0 represents the basic running resistance of the whole train, F eb represents the air braking force of the whole train, W j represents the additional resistance of the whole train, X d represents the position of the train in the current state, γ represents the train rotary mass coefficient, V d+1 represents the train speed at the (d + 1)-th step, ΔX represents the discrete step length, Δt represents the time consumed to travel the distance of ΔX, pos d represents the position of the control handle at the d-th step, pos d-1 represents the position of the control handle at the (d - 1)-th step, X d+1 represents the position of the train at the (d + 1)-th step, t rmax represents the maximum air braking release time, t d represents the system time at the d-th step, t d+1 represents the system time at the (d + 1)-th step, Ab d represents the air braking state, t bmax represents the maximum air braking time.
[0060] The beneficial effects of the above further solution are as follows: By introducing the system state equation and the dynamic calculation of speed and position, the state information of train operation can be updated in real time, ensuring the accuracy and continuity of data. And by judging and adjusting the air braking and relief time, as well as comprehensively considering the system friction coefficient and additional resistance, the solution can adapt to complex operating environments and improve the stability and safety of train operation.
[0061] Furthermore, the expression of the state value in step S4 is as follows:
[0062]
[0063] Among them, U d represents the state value of the current state S d , that is, the expected total reward that can be obtained starting from the current state S d . R d represents the immediate return obtained after executing the action in the current state S d . γ represents the discount factor, and Q(S d+1 , A d+1 ) represents the expected total return when the state S d+1 executes the action A d+1 .
[0064] The beneficial effects of the above further solution are as follows: By introducing the state value function, the expected return brought by executing the action in the current state can be quantified. The solution takes into account both immediate return and long-term benefits, providing an accurate evaluation basis for train operation decision-making.
[0065] Furthermore, step S5 includes the following steps:
[0066] S501. For the continuous state S d , decompose it into the surrounding discrete states S0, S1,..., S n , and calculate the temporal difference error TD of all adjacent discrete states:
[0067] δ i = U d - Q(S i , A d )
[0068] Among them, δ i represents the temporal difference error TD, U d represents the state value of the d-th step, A d represents the action of the d-th step, and S i represents the discrete state after decomposing S d ;
[0069] S502. Update all discrete state Q-tables using the following formula based on the temporal difference error TD of all neighboring discrete states:
[0070] Q π (s i ,a D )←Q π (s i ,a D )+α·ω i ·δ i
[0071] Among them, Q π (s i ,a D ) represents the state-action value, a D represents the action, α represents the regenerative braking efficiency, and ω i represents the weight coefficient corresponding to S i .
[0072] The beneficial effect of the above further solution is that by calculating the TD (temporal difference) error, the system can quickly evaluate the difference between the current state and the target state, provide feedback to accelerate the convergence process of the policy, and the solution can balance the update speed of new and old knowledge, improving the learning efficiency and robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0074] The following describes the specific embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0075] Embodiment
[0076] Before describing the complete technology of the present invention, the theoretical results used later are summarized and described.
[0077] (1) The specific form of the multi-parameter brake cylinder detailed model based on exponential function fitting and correction function is as follows:
[0078]
[0079] Among them, P i (t) represents the function of the pressure of the brake cylinder changing with time, and P m,i represents the initial pressure of the brake cylinder, and Ps,i represents the steady-state pressure of the brake cylinder, β i represents the characteristic function of the pressure increase in the brake cylinder, t represents the system time, t delay,i represents the starting time of the pressure increase in the brake cylinder.
[0080] The corresponding calculation formula for the air braking force is as follows:
[0081]
[0082] where, F b,i (t) represents the function of the air braking force pressure changing with time, P i (t) represents the function of the pressure in the brake cylinder changing with time, S h represents the piston area, α b represents the mechanical structure transmission ratio, eff represents the braking efficiency, represents the brake pad friction coefficient.
[0083] (2) Train kinematic equation:
[0084]
[0085] where, v represents the train running speed, t represents the system time, x represents the train displacement. M represents the total vehicle weight, F t represents the train traction force, F eb represents the total vehicle air braking force, W0 represents the basic running resistance of the whole vehicle, W j represents the additional resistance of the whole vehicle.
[0086] Basic running resistance of the whole vehicle:
[0087] W0 = ∑m i ·g·w0·10 -3 (4)
[0088] where, m i represents the mass of the i-th carriage, g represents the acceleration due to gravity, w0 represents the unit basic resistance of the carriage.
[0089] Unit weight basic resistance:
[0090] w0 = a + bv + cv 2 (5)
[0091] where, a, b, c represent empirical coefficients, selected according to the vehicle type, v represents the train running speed.
[0092] Additional resistance of the whole vehicle:
[0093] W j = ∑m i ·g·(w g,i + wr,i +w s,i )·10 -3 (6)
[0094] Among them, m i represents the mass of the i-th carbody, g represents the acceleration due to gravity, and w g,i represents the ramp resistance of the i-th carbody, and w r,i represents the curve resistance of the i-th carbody when cornering, and w s,i represents the air resistance of the i-th carbody.
[0095] Calculation method of air braking force for rigid multi-particle model:
[0096] F b = ∑F b,i (t B , t R , v) (7)
[0097] Among them, F b represents the air braking force of the whole vehicle for the rigid multi-particle model, and F b,i represents the air braking force of the i-th carbody, t B represents the braking duration of air braking, and t R represents the release duration of air braking, and v represents the running speed of the train.
[0098] Based on the above theoretical analysis, in order to achieve the invention purpose of the present invention, the technical solutions adopted are as follows:
[0099] As Figure 1 shown, the present invention provides a method for optimizing the combined air and electric multi-objective operation of a freight train, and its implementation method is as follows:
[0100] S1. Set training parameters, vehicle information, and freight train line information, and initialize the reinforcement learning environment and variables;
[0101] In this embodiment, set the basic training parameters of the program (maximum number of iterations M, maximum number of steps per episode N, etc.), vehicle information (train weight, train traction / braking characteristics, etc.), and line information (speed limit, gradient, curve, etc.).
[0102] In this embodiment, initialize the reinforcement learning environment (system time t, number of steps per episode n, etc.), variables such as the Q-table.
[0103] S2. Calculate and select an action using the policy function, and its implementation method is as follows:
[0104] S201. Randomly generate a random number p ∈ [0, 1], and calculate the current agent according to the policy function π(A s |S d );
[0105] S202. Compare the random number p with the action selected by the agent according to the policy function π(A s |S d ). If the random number , select the action A with the maximum state value in the current Q-table d * . If , randomly select an action that can be executed in the current state S d .
[0106] In this embodiment, the action A is selected by calculating through the policy function of formula (17). d .
[0107] S3. Based on the selected action, decompose the action vector to obtain the current step traction braking force and the air braking state, and obtain the new state by calculating the mechanical information. The implementation method is as follows:
[0108] S301. Based on the selected action, decompose the action vector to obtain the current step traction braking force and the air braking state;
[0109] In this embodiment, the action vector is decomposed according to formula (15) to obtain the current step traction braking force and the air braking state.
[0110] S302. According to the current step traction braking force and the air braking state, calculate the mechanical information by calculating the dynamics of the freight train. The implementation method is as follows:
[0111] S3021. Construct a detailed multi-parameter brake cylinder model based on exponential function fitting and correction function;
[0112] S3022. Calculate the corresponding air braking force according to the detailed multi-parameter brake cylinder model;
[0113] S3023. Based on the train kinematic equation, calculate the basic running resistance of the freight train and the additional resistance of the whole vehicle, and combine the air braking force to obtain the mechanical information;
[0114] In this embodiment, calculate the train dynamics according to formulas (1)-(7) to obtain the mechanical information.
[0115] S303. According to the mechanical information, calculate the relevant state quantities of the freight train to obtain the new state;
[0116] In this embodiment, calculate the train dynamics according to formulas (1)-(7) to obtain the mechanical information S d+1 .
[0117] S4. Calculate the state value of the new state;
[0118] In this embodiment, the new state S is calculated according to Equation (18). d for the state value of.
[0119] S5. Update the Q value related to the current state in the Q table, and the implementation method is as follows:
[0120] S501. For the continuous state S d decompose it into the surrounding discrete states S0, S1,..., S n , calculate the temporal difference error TD of all neighboring discrete states: δ i = U d - Q(S i , A d );
[0121] where, δ i represents the temporal difference error TD, U d represents the state value at the d-th step, A d represents the action at the d-th step, and S i represents the discrete state after the decomposition of S d ;
[0122] S502. Based on the temporal difference error TD of all neighboring discrete states, use the following formula to update the Q table of all discrete states: Q π (s i , a D ) ← Q π (s i , a D ) + α·ω i ·δ i ;
[0123] where, Q π (s i , a D ) represents the state-action value, a D represents the action, α represents the regenerative braking efficiency, and ω i represents the weight coefficient corresponding to S i ;
[0124] S6. Determine whether the current step number n exceeds the maximum number of rounds N. If so, go to step S7; otherwise, return to step S3;
[0125] S7. Use the total reward to determine whether the updated Q table converges. If so, go to step S9; otherwise, go to step S8;
[0126] S8. Determine whether the current iteration number reaches the maximum iteration number. If so, go to step S9; otherwise, reset the train variables and return to step S3;
[0127] In this embodiment, resetting the train variables includes the position, environment, etc.
[0128] S9. The agent executes actions according to the Q-table, calculates the new state, obtains the final sequence of operating conditions, and completes the multi-objective operation optimization of the air-electric combined control of the freight train.
[0129] In this embodiment, the agent's decision-making is guided by estimating the value function. The goal of reinforcement learning is to maximize the reward during the interaction between the agent and the environment, so as to determine the optimal action and obtain a better air-electric combined control method.
[0130] Since the above calculation model is difficult to linearize and transform into the form of a standard planning algorithm, reinforcement learning, as a widely used model-free optimization method, can achieve multi-objective optimization under complex models through the interaction and learning between the agent and the environment. The target reward function of reinforcement learning is the basis for the agent's learning guidance in the present invention. It gives different rewards according to the different operations made by the agent in different states. The reward function R of the present invention d The specific form is:
[0131] R d (S d ,A d )=w s R s +w v R v +w f (R f +R b )+w e R e (8)
[0132] Wherein, S d represents the current state of the train, A d represents the operation executed by the train, w s represents the safety index weight coefficient, R s represents the safety index, w v represents the speed-up index weight coefficient, R v represents the speed-up index, w f represents the comprehensive index weight coefficient of the control force and the air brake, R f represents the train control force index, R b represents the air brake index.
[0133] Safety index:
[0134]
[0135] Wherein, r v represents the safety index penalty, V d+1 represents the speed of the next state, v lim (X d+1 ) represents the position X of the next stated+1 The line speed limit at indicates the cumulative braking release time of the train in the next state.
[0136] Energy-saving index R e :
[0137]
[0138] where X d represents the position of the train in the current state, and X d+1 represents the position of the train in the next state, F d represents the train speed in the current state, η t , η b represents the traction and braking energy transfer efficiency, and α represents the regenerative braking efficiency.
[0139] Speed-reaching index R v :
[0140]
[0141] where V target represents the target speed, ΔV represents the discrete value of speed reward, and v lim (X d+1 ) represents the maximum speed limit in the braking phase.
[0142] Train control force index R f :
[0143]
[0144] where λ1, λ2 represent the control force reward coefficients, and pos d represents the position of the control handle at the d-th step, and pos d-1 represents the position of the control handle at the (d - 1)-th step, and v lim (X d+1 ) represents X d+1 the maximum allowable driving speed at.
[0145] Air braking index R b :
[0146] R b = -500 t d <0&Ab d = 1 (13)
[0147] where t d <0&Ab d = 1 indicates the use of air braking or the release time is less than 0.
[0148] The problem of constructing a reinforcement learning Markov decision process is as follows:
[0149] State space: It refers to the partial train operation information observed by the agent, which represents the current train operation state. The present invention defines the state space as follows:
[0150] s d =[x d ,v d ,t d ,pos d-1 (14)
[0151] Where s d represents the system state at the d-th step, x d represents the position at the d-th step, v d represents the speed at the d-th step, and pos d-1 represents the traction or electric braking gear at the (d - 1)-th decision-making stage.
[0152] Action space: It is the set of different actions made by the agent according to different environments. The present invention defines the following set of traction and braking forces and air braking sequences in discrete states:
[0153] A d =[F d ,Ab d (15)
[0154] Where A d represents the action at the d-th step, F d represents the traction and braking force at the d-th step, and Ab d represents the air braking state at the d-th step.
[0155] Environment: In the heavy-haul train MDP model designed by the present invention, the simulation and solution model of the train object is a rigid multi-particle model. After converting it into a discrete process, the state transition process of the heavy-haul train MDP model of the present invention is as follows:
[0156]
[0157] Where S d+1 represents the new state of the system at the (d + 1)-th step, represents the brake pad friction coefficient, S d represents the state of the train at the d-th step, g represents the acceleration due to gravity, A d represents the operation performed by the train, acc d () represents the calculation function of the train acceleration, F t represents the train traction force, V d represents the train speed at the d-th step, W0 represents the basic running resistance of the whole vehicle, F eb represents the air braking force of the whole vehicle, W j represents the additional resistance of the whole vehicle, X drepresents the position of the train in the current state, γ represents the train's gyroscopic mass coefficient, V d+1 represents the train speed at the (d + 1)-th step, ΔX represents the discretization step size, Δt represents the time consumed to travel the distance of ΔX, pos d represents the control handle position at the d-th step, pos d-1 represents the control handle position at the (d - 1)-th step, X d+1 represents the position of the train at the (d + 1)-th step, t rmax represents the maximum air brake release time, t d represents the system time at the d-th step, t d+1 represents the system time at the (d + 1)-th step, Ab d represents the air brake state, t bmax represents the maximum air brake application time.
[0158] Policy: The agent selects corresponding actions according to the policy function π(A s |S d ). The form of the policy function designed in the present invention is as follows:
[0159]
[0160] where, A d |S d represents that in the state of S d , the action A d is taken, π(A s |S d ) represents that the current agent selects the optimal action A d * under the current condition according to the policy function, rand(A(S d )) represents a randomly selected action that the agent in the state of S d can execute, δ represents the exploration rate decay rate (δ ∈ [0, 1]), ep represents the number of learning iterations, ep T represents the decay period, and ε0 represents the initial value of the random exploration rate (ε0 ∈ [0, 1]).
[0161] Q-table (state-action expected return table): Records the expected return Q values that can be obtained by selecting different actions in all states.
[0162] The state value of the current state S d can be expressed as:
[0163]
[0164] where, U d represents the state value of the current state S d , represents the expected total reward that can be obtained starting from the state S d , Rd Represents the current state S d The immediate reward obtained after performing an action, γ represents the discount factor, indicating the degree of weight decay of future rewards, Q(S d+1 , A d+1 ) represents the expected total reward in state S d+1 when performing action A d+1 .
[0165] There is a recurrence relationship for the action value function, that is, the current action value function can be represented by the action value function of the next stage. The Q-learning of the present invention adopts a temporal difference update method and uses a recurrence formula under the optimal strategy to update the Q-table:
[0166] Q(S d , A d ) = Q(S d , A d ) + α[U d - Q(S d , A d )] (19)
[0167] Among them, Q(S d , A d ) represents the expected total reward in state S d when performing action A d , and α represents the learning rate.
[0168] During the training process of reinforcement learning Q-learning, in order to determine whether the policy has converged or the Q-table has stabilized, the present invention introduces an early termination criterion based on the change in the total reward. After each training episode ends, the total reward value of this episode will be calculated and compared with the total reward value of the previous episode .
[0169] When the Q-learning algorithm continuously iteratively updates the Q-table, if the state-action value distribution in the Q-table gradually approaches the optimal strategy, the total reward value obtained by the agent under the same or similar operating conditions will tend to be stable. For this reason, the present invention sets a convergence threshold ε tol (for example, ε tol can be taken as 10 or adjusted according to actual operating experience). If the following conditions are met:
[0170]
[0171] then it is considered that the change in the policy performance between this training episode and the previous training is very small, and the Q values in the Q-table have basically converged. At this time, the training process can be terminated and no more unnecessary iterations are required.
[0172] In summary, the present invention proposes a multi-objective optimization method for the combined air and electric control of freight trains with long downhill sections, which applies the reinforcement learning method and considers the refined equivalent model of the brake cylinder. Specifically, it includes:
[0173] 1. Optimization strategy design: Through the Q-learning algorithm in reinforcement learning, the matching strategy of the air braking timing and the electric control power during the cyclic braking process is optimized. An improved Q-table iteration mechanism is introduced to solve the problem of continuous state changes, making the strategy more adaptable.
[0174] 2. Multi-parameter equivalent model of the brake cylinder: Aiming at the inaccurate description of the traditional two-stage empirical model during the cyclic braking process, an equivalent model of the brake cylinder that can be solved quickly is designed. This model can more accurately describe the air braking force and the brake cylinder pressure under different release times and altitude changes, making up for the deficiencies of the traditional model.
[0175] 3. Adaptability to complex operating conditions: Combining the characteristics of long slopes, fully considering the operating disturbances, the discreteness of air braking, and the accuracy requirements of cyclic braking control, the practicality of the model in complex scenarios is improved.
Claims
1. An optimization method for the combined air and electric multi-objective operation of a freight train, characterized in that, It includes the following steps: S1. Set training parameters, vehicle information, and freight train line information, and initialize the reinforcement learning environment and variables; S2. Calculate and select an action using the policy function; S3. Based on the selected action, decompose the action vector to obtain the traction and braking force and the air braking state at the current step, and obtain the new state by calculating the mechanical information. Specifically: S301. Based on the selected action, decompose the action vector to obtain the traction and braking force and the air braking state at the current step; S302. According to the traction and braking force and the air braking state at the current step, calculate the mechanical information by calculating the dynamics of the freight train. Specifically: S3021. Construct a detailed multi-parameter brake cylinder model based on exponential function fitting and correction function; S3022. Calculate the corresponding air braking force according to the detailed multi-parameter brake cylinder model; S3023. Based on the train kinematic equation, calculate the basic running resistance of the freight train and the additional resistance of the whole vehicle, and combine the air braking force to obtain the mechanical information; the expression of the train kinematic equation is as follows: Among them, represents the train running speed, represents the system time, represents the total vehicle weight, represents the train traction force, represents the total vehicle air braking force, represents the total vehicle basic running resistance, represents the total vehicle additional resistance, represents the total vehicle air braking force of the rigid multi - particle model, represents the i air braking force of the n - th carriage, represents the braking duration of air braking, represents the release duration of air braking; represents the train displacement; S303. According to the mechanical information, calculate relevant state variables of the freight train to obtain a new state ; S4. Calculate the state value of the new state; S5. Update the Q value related to the current state in the Q table. Specifically: S501. For continuous states Decompose into surrounding discrete states , and calculate the temporal difference error TD of all neighboring discrete states: Among them, represents the temporal difference error TD, represents the state value at the th step, d represents the action at the th step, represents the decomposed discrete state; S502. Based on the temporal difference error TD of all neighboring discrete states, use the following formula to update the Q table of all discrete states: Among them, represents the state-action value, represents the action, represents the regenerative braking efficiency, represents and the corresponding weight coefficient S6. Determine whether the current step number n exceeds the maximum number of episodes N. If so, go to step S7; otherwise, return to step S3; S7. Use the total reward to determine whether the updated Q table converges. If so, go to step S9; otherwise, go to step S8; S8. Determine whether the current number of iterations reaches the maximum number of iterations. If so, go to step S9; otherwise, reset the train variables and return to step S3; S9. The agent executes the action according to the Q table, calculates the new state, obtains the final sequence of working conditions, and completes the multi-objective operation optimization of the air-electric combination of the freight train.
2. The method for optimizing the combined air and electric multi-objective operation of a freight train according to claim 1, characterized in that The step S2 includes the following steps: S201. Randomly generate a random number , and calculate to obtain the current agent according to the policy function ; S202. Compare the random number with the agent according to the policy function . If the random number , select the action with the largest state value in the current table. If , randomly select an action that can be executed in the current state . Among them, represents taking the action in the state. represents the -th step action. d represents the initial value of the random exploration rate. represents the optimal action under the current condition. represents the learning iteration times. represents the decay period. represents the exploration rate decay rate. represents the exploration rate decay rate.
3. The multi-objective operation optimization method for the air-electric combined system of a freight train according to claim 2, wherein The current agent follows the policy function The expression of which is as follows: Among them, represents an action that can be performed by the randomly selected status agent.
4. The empty and electric combined multi-objective operation optimization method for freight trains according to claim 1, wherein The expression for decomposing the action vector is as follows: Among them, represents the d step movement, represents the d step traction braking force, represents the d step air braking state.
5. The multi-objective operation optimization method for the combined air and electric of a freight train according to claim 1, characterized in that The expression for the detailed multi-parameter brake cylinder model is as follows: Among them, represents the function of the pressure of the brake cylinder changing with time, represents the initial pressure of the brake cylinder, represents the steady-state pressure of the brake cylinder, represents the pressure increase characteristic function of the brake cylinder, represents the system time, represents the time when the pressure of the brake cylinder starts to increase; The expression for the air braking force is as follows: Among them, represents the function of the air braking pressure varying with time, represents the piston area, represents the mechanical structure transmission ratio, represents the braking efficiency, represents the brake pad friction coefficient; The expressions for the basic running resistance of the freight train and the additional resistance of the whole vehicle are as follows: Among them, represents the mass of the th carriage, represents the acceleration due to gravity, represents the basic resistance per unit of the carriage, represents the empirical coefficient, represents the running speed of the train, represents the th carriage's ramp resistance, represents the th carriage's curve resistance when turning, represents the th carriage's air resistance.
6. The method for optimizing the air-electric combined multi-objective operation of a freight train according to claim 1, wherein The new state has the following expression: Among them, represents the new state of the system at the th step, represents the brake pad friction coefficient, represents the state of the train at the th step, represents the acceleration due to gravity, represents the action at the d th step, represents the calculation function of the train acceleration, represents the train traction force, represents the th step train speed, represents the basic running resistance of the whole vehicle, represents the air braking force of the whole vehicle, represents the additional resistance of the whole vehicle, represents the position of the train in the current state, represents the train rotary mass coefficient, represents the th step train speed, represents the discrete step size, represents the time consumed for traveling the distance, represents the th step control handle position, represents the th step control handle position, represents the th step train position, represents the maximum air brake release time, represents the th step system time, represents the th step system time, represents the air brake state, represents the maximum air brake application time.
7. The method for optimizing the combined air and electric multi-objective operation of a freight train according to claim 1, characterized in that The expression for the state value in the step S4 is as follows: wherein, represents the state value of the current state, that is, the expected total reward that can be obtained starting from the current state represents the immediate reward obtained after executing an action in the current state represents the discount factor represents the expected total return when the state executes an action
Citation Information
Patent Citations
Operation curve rolling optimization method for reducing longitudinal impulse of heavy haul train
CN115221717A
Heavy-duty train long and large downhill section operation curve optimization method considering air braking
CN116151025A