Train autonomous decision-making method facing passenger demands

By constructing a dynamic passenger flow model and a safety protection constraint model, combined with a reinforcement learning algorithm, the train's independent decision-making under time-varying passenger flow is realized, the problem of disorderly train operation order is solved, and passenger service level and resource utilization efficiency are improved.

CN120229285APending Publication Date: 2025-07-01SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510516074.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the face of time-varying passenger flow, the existing collaborative research methods cannot effectively adjust the train arrival time, operation interval and full load rate, resulting in disordered train operation order and unable to achieve independent decision-making without pre-design.

Method used

Based on real-time changes in passenger flow, a dynamic passenger flow model is constructed, a multi-train distributed collaboration framework is established, the front and rear vehicle operation relationship is analyzed, a safety protection constraint model is constructed, and the reinforcement learning algorithm is designed to achieve independent decision-making of trains with the goal of balancing train full load rate and minimum passenger waiting time.

Benefits of technology

On the premise of ensuring the safety of train operation, reduce passenger waiting time, achieve balanced full load rate, improve passenger comfort and reduce waste of capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120229285A_ABST
    Figure CN120229285A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of train operation optimization, and particularly discloses a passenger demand-oriented train autonomous decision-making method, which comprises the following steps of: depicting the number of passengers at a platform in a future moment based on passenger flow real-time change, and constructing a dynamic passenger flow model; establishing a multi-train distributed collaborative framework, converting a multi-train problem into a single-train problem of a front train and a rear train, and updating the platform passenger flow state acquired by each train; based on a distributed collaborative framework, analyzing an operation relation between a front vehicle and a rear vehicle, and constructing a front and rear vehicle safety protection constraint model; based on the front and back train safety protection constraint model and the platform passenger flow state obtained by each train, a reinforcement learning algorithm is designed by taking train load factor balance and passenger waiting time minimization as targets, and train autonomous decision making is realized. According to the method, the problems that train arrival and departure time adjustment, operation interval adjustment and load factor indexes under the influence of passenger flow time variation are insufficiently concerned and train autonomous decision-making without a preset plan cannot be realized in an existing collaborative research method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of train operation optimization, and particularly relates to a train autonomous decision-making method oriented to passenger needs. Background Art

[0002] With the network operation of urban rail transit, on the one hand, the total passenger flow is continuously increasing. Taking a certain place as an example, during the double festivals, the average daily passenger volume reaches 5.2352 million person-times, an increase of 4.20% compared with the weekend passenger flow on weekdays. On September 28th, the single-day passenger volume reaches 7.8392 million person-times, breaking the historical record of the line network, and the load factor breaks through 140%. On the other hand, the tidal passenger flow characteristics are obvious. The common phenomenon of "separation of work and residence" makes the subway lines connecting the suburbs and the urban area have obvious tidal characteristics. During the morning peak period, the passenger flow in the inbound direction is huge, while the passenger flow in the outbound direction is small; vice versa during the evening peak.

[0003] Urban rail transit is an organized and planned transportation mode. Its inherent meaning is to reasonably allocate the limited "line resources" to each train, and the trains run longitudinally along the track in sequence according to the established operation diagram. However, during the peak period, the headway is very small, and the initial delay is easily transmitted to other trains, causing associated delays, resulting in multiple trains stopping in the middle to wait, and even affecting the normal departure of oncoming trains, causing the disorder of the train operation order of the whole line. Therefore, it is necessary to adjust the train operation. The existing research on the train operation adjustment and vehicle-to-vehicle cooperation of urban rail transit mainly focuses on adjusting the train timetable and operating with high-density equal intervals under the pre-designed plan to reduce the operation cost and the waiting time of passengers. However, due to the frequent starting and stopping of urban rail trains and the dynamic change of the tracking interval between adjacent trains, the existing cooperation research pays insufficient attention to the adjustment of arrival and departure times, the adjustment of running intervals, and the load factor index under the influence of time-varying passenger flow. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem that the existing cooperation research method pays insufficient attention to the adjustment of train arrival and departure times, the adjustment of running intervals, and the load factor index under the influence of time-varying passenger flow, and cannot achieve the train autonomous decision-making without a pre-designed plan, and proposes a train autonomous decision-making method oriented to passenger needs.

[0005] The technical solution of the present invention is as follows: A train autonomous decision-making method oriented to passenger needs includes the following steps:

[0006] Based on the real-time change of passenger flow, depict the number of passengers on the platform in the future moment, and construct a dynamic passenger flow model;

[0007] Based on the dynamic passenger flow model, establish a multi-train distributed cooperation framework, transform the multi-train problem into a single-train problem of the leading train and the following train, and update the platform passenger flow state obtained by each train;

[0008] Analyze the running relationship between the leading train and the trailing train, and construct a safety protection constraint model for the leading and trailing trains;

[0009] Based on the safety protection constraint model for the leading and trailing trains and the platform passenger flow status obtained by each train, with the goal of balancing the train load factor and minimizing the passenger waiting time, design a reinforcement learning algorithm to achieve autonomous decision-making of the train.

[0010] Preferably, based on the real-time change of passenger flow, characterize the number of platform passengers in the future moment, and construct a dynamic passenger flow model, which specifically includes the following steps:

[0011] According to the time-varying characteristics of passenger flow, establish a platform arrival passenger flow model, expressed as:

[0012]

[0013] Among them, is the platform arrival passenger flow model, j is the platform number, j ∈ N, and N is the set of stations; t is the discrete time constant; V arrive is the rate of passengers arriving at the platform, which is expressed as:

[0014]

[0015] Among them, v a,peak is the rate of passengers arriving at the platform during the peak period, and v a,offpeak is the rate of passengers arriving at the platform during the off-peak period;

[0016] According to the passenger boarding and alighting behavior characteristics, establish a platform boarding and alighting passenger flow model, expressed as:

[0017]

[0018] Among them, is the number of passengers getting off at the j-th platform at time t; M is a constant value; is the number of passengers boarding at the j-th platform at time t; D j is the stop time of the train at the j-th platform; V on and V off are the boarding rate and alighting rate of passengers respectively, which are expressed as:

[0019]

[0020] Among them, v on,peak and v off,peak are the boarding rate and alighting rate of passengers during the peak period respectively; v on,offpeak and v off,offpeak are the boarding rate and alighting rate of passengers during the off-peak period respectively;

[0021] According to the platform arrival passenger model and the platform boarding and alighting passenger flow model, a dynamic passenger flow model is established. The dynamic passenger flow model includes the platform stranded passenger flow model and the train passenger flow model, which can be expressed as:

[0022]

[0023] in, is the number of stranded passengers at platform j at time t; is the number of passengers on the train when it arrives at the jth platform at time t; For m Passengers on the train just after arriving at platform J-1; For m Passengers boarding the bus at platform J-1 at all times; For m Passengers getting off at platform j-1 at any time; m is a discrete moment, and each moment interval is 1s.

[0024] Preferably, the method is based on a dynamic passenger flow model, establishes a multi-train distributed collaborative framework, converts the multi-train problem into a single-train problem of the front and rear trains, and updates the platform passenger flow status obtained by each train, specifically including the following steps:

[0025] According to the principle of vehicle-to-vehicle communication, a multi-train distributed coordination framework is constructed, which can be expressed as:

[0026]

[0027] in, represents a directed graph of order m; V′={1,2,...,m} represents the set of train agents, and m represents the total number of train agents; represents the set of vehicle-to-vehicle communication paths; if from train i x To train y If there is a communication path, then (i x ,i y )∈ε, yes The adjacency matrix of , where R represents the set of real numbers, is an element in the matrix A′, if (i x ,i y )∈ε, then otherwise And there is The adjacent matrix A′ indicates that the current train only has a communication path with the previous adjacent train. If the current train is not the first train, the current train is made to receive the platform passenger flow state calculated by the previous adjacent train. After considering the passengers already or about to be transported away by the previous trains at each platform, the passenger flow state of the current train when arriving at each platform is calculated based on the dynamic passenger flow model. Thus, the platform passenger flow states obtained by each train are as follows:

[0028]

[0029] Among them, is the actual number of platform passengers at the j-th platform received by train i when leaving the k'-th platform. is the estimated number of platform passengers at the j-th platform calculated by train i when leaving the k'-th platform; t i,k′ is the moment when train i leaves the k'-th platform, t i,k′ ∈T; is the number of passengers remaining at the j-th platform after train i leaves the k'-th platform at time t i,k′ ; is the number of passengers arriving at the j-th platform at time t i,k ; is the number of passengers boarding train x at the j-th platform; is the planned stop time of train i at platform w at time t i,k′ ; is the planned running time of train i between platform w and platform w + 1 at time t i,k′ ; V arrive is the rate at which passengers arrive at the platform.

[0030] Preferably, analyzing the running relationship between the preceding train and the following train, and constructing a safety protection constraint model for the preceding and following trains specifically includes the following steps:

[0031] Construct a train kinematic model, expressed as:

[0032]

[0033] Among them, is the derivative of the train's real-time position with respect to time, that is, the running speed v(t) of the train; is the train's real-time acceleration; u i (t) is the traction or braking force per unit mass of the train; g is the acceleration due to gravity; w0 is the resistance during the train's operation, w0 = δ a′ a′ + δ b′ b′v i (t) + δ c′ c′v i (t) 2, where a′, b′, and c′ are the coefficients of the Davis formula, and δ a′ , δ b′ , and δ c′ are the rates of change of a′, b′, and c′ respectively; γ is the coefficient of the train's rotational mass; d(t) is an unknown bounded disturbance, w i is the additional resistance of the line gradient, and w c is the additional resistance of the line curve;

[0034] Based on the train kinematic model, with the goal of punctuality, a dynamic programming algorithm is designed to obtain the speed curves of each train;

[0035] Based on the speed curve of the leading train, a safety protection constraint model for the leading and following trains is constructed, expressed as:

[0036]

[0037] where t is the discrete time constant; is the position of the (i + 1)-th train at time t; is the position of the i-th train at time t; l0 is the safety interval margin between trains.

[0038] Preferably, the step of designing a dynamic programming algorithm based on the train kinematic model with the goal of punctuality to obtain the speed curves of each train specifically includes the following steps:

[0039] Based on the train kinematic model, the inter-station distance of train operation is divided into K + 1 stages at intervals of Δx, and the speed is divided into N + 1 speed intervals at intervals of Δv, thereby obtaining the state grid points of the discretization of the train motion state space. The train operation position at stage k is expressed as x k , k ∈ {1,..., K + 1}, and the train speed state at stage k is expressed as {v k,1 ,..., v k,n ,..., v k,N+1}, where v k,n is the train speed state at stage k;

[0040] With the goal of punctuality, a stage index function is designed, expressed as:

[0041]

[0042] where C k is the cumulative index value from stage k to the last stage K + 1; κ is a certain stage from stage k to the last stage K + 1; Δx is the stage interval distance; v k is the speed at the k-th stage, and v k+1 is the speed at the (k + 1)-th stage;

[0043] Based on the stage index function and the discretized state lattice points, the optimal decision variables for the train speed states at each stage are calculated and expressed as:

[0044]

[0045] where: u k,n is the decision variable corresponding to the train speed state v k,n at stage k; is the optimal decision variable corresponding to the train speed state v k,n at stage k; U is the allowable decision set; c k,n (v k,n , u k,n ) is the index corresponding to the decision made by adopting u k,n when the train speed state is v k,n at stage k; is the optimal index function corresponding to the train speed state v k+1,n′ at stage k + 1. n represents a certain determined train speed state, and n' represents a certain undetermined train speed state, which changes with the change of the optimal objective value. represents the value of the optimal decision variable u k,n corresponding to the calculated optimal index objective function;

[0046] The optimal decision variables of the train speed states at each stage are combined to obtain the train speed curve.

[0047] Preferably, based on the safety protection constraint model of the front and rear trains and the platform passenger flow state obtained by each train, a reinforcement learning algorithm is designed with the goals of balanced train load factor and minimum passenger waiting time to achieve autonomous train decision-making, which specifically includes the following steps:

[0048] Based on the platform passenger flow state obtained by each train and the safety protection constraint model of the front and rear trains, a reinforcement learning algorithm is designed to calculate the optimal running time of the current train in the remaining intervals and the optimal stopping time at the remaining platforms.

[0049] Based on the optimal running time of the current train in the remaining intervals, the train speed curve is obtained.

[0050] At the rolling update moment, the train obtains the speed command at each time step and runs according to the train speed curve. If the train arrives at the platform, the train stops according to the optimal stopping time, and at the same time, the platform passenger information and the remaining platform information are updated until the train arrives at the terminal platform and the operation ends, realizing autonomous train decision-making.

[0051] Preferably, based on the platform passenger flow status obtained from each train and the safety protection constraint model of the front and rear trains, a reinforcement learning algorithm is designed to calculate the optimal running time of the current train in each remaining section and the optimal stopping time at each remaining platform, which specifically includes the following steps:

[0052] Under the condition of satisfying the safety protection constraints of the front and rear trains, with the goal of minimizing the passenger waiting time and balancing the train load factor, a reward function is set, expressed as:

[0053]

[0054] where: s j represents the position state of the train at the j-th platform; is the reward function under the train position state s j ; T wait,i,j is the total passenger waiting time of all subsequent platforms calculated after the i-th train arrives at the j-th platform; R load,i,j is the train load factor balance of all subsequent inter-station sections calculated after the i-th train arrives at the j-th platform; γ1 and γ2 are the weight coefficients of the passenger waiting time and the load factor balance respectively;

[0055] The passenger waiting time is expressed as:

[0056]

[0057] where: t i,j is the time when the i-th train arrives at the j-th platform; is the number of passengers boarding the i-th train at the (χ + 1)-th station; is the number of platform passengers at the (χ + 1)-th station obtained by the i-th train; is the estimated stopping time of the i-th train at the β-th platform; is the estimated running time of the i-th train from the β-th platform to the (β + 1)-th platform, and N is the set of stations;

[0058] The train load factor balance is expressed as:

[0059]

[0060] where: is the number of people on the i-th train when it is at the χ-th platform; C is the maximum capacity of the train; ||·|| represents the second norm;

[0061] Based on the reward function, calculate the optimal Q value for selecting the optimal action in each train position state, and obtain the optimal Q value table:

[0062] Q * ={Q′(s1,a),...,Q′(s j ,a),...,Q′(s n, a)}

[0063] Where: Q * is the optimal Q-value table; Q′(s j , a) is the optimal Q-value corresponding to the optimal action a selected in state s j .

[0064] Calculate the running time of each interval and the stop time corresponding to the optimal Q-value table, expressed as:

[0065] RT obj ={RT1',..., RT′ j ,..., RT′ n}

[0066] ST obj ={ST1',..., ST′ j ,..., ST′ n}

[0067] Where: RT obj and ST obj are the optimal running times of each interval mapped by Q * and the optimal stop times of each platform respectively, RT′ j and ST′ j represent the optimal running time of the train from the j-th platform to the (j + 1)-th platform and the optimal stop time at the j-th platform respectively, Q * →{RT obj , ST obj}.

[0068] Preferably, calculating the optimal Q-value of selecting the optimal action in each train position state based on the reward function specifically includes the following steps:

[0069] Set the maximum iteration step size and randomly initialize the train position state s j ;

[0070] Based on the train position state s j randomly select an action or select the action with the largest Q-value, expressed as:

[0071]

[0072] Where: a is the action selected in the current train position state s j ; s j+1 is the next train position state; a w is the action that the next train position state s j+1 can select; represents the action a corresponding to the maximum Q-value, Q(s w , a j+1 , aw ) is the next position state of the train s j+1 Select action a w The corresponding Q value; a r is a random action; P r is the random probability, P r ∈[0,1];

[0073] Based on the reward function, update the Q value corresponding to the action selected by the current state of the train;

[0074]

[0075] Where: Q′(s j ,a) indicates the updated train position state s j The Q value corresponding to the selected action a, Q(s j ,a) indicates the train is in position state s before updating j The Q value corresponding to the selected action a, α is the learning rate, γ′ is the penalty factor, and α and γ′ are custom constants; is the reward function;

[0076] Update the next position state of the train until it reaches the terminal position, jump out of the current iteration, repeatedly initialize the train position state, select actions, and calculate the Q value of the updated action until the maximum iteration step is reached, and output the optimal Q value table and its corresponding interval running time and stop time.

[0077] The beneficial effects of the present invention are:

[0078] The present invention takes into account the real-time changes in passenger flow and the operation of trains without a preset plan. It can reduce the average waiting time of all passengers on the platform while ensuring the safety of train operation, achieve a balanced full load rate, improve passenger comfort and reduce waste of transportation capacity. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 Shown is a flow chart of a train autonomous decision-making method oriented to passenger demand. DETAILED DESCRIPTION

[0080] Now, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the accompanying drawings are only exemplary and are intended to explain the principles and spirit of the present invention, rather than to limit the scope of the present invention.

[0081] Example:

[0082] like Figure 1 As shown, a train autonomous decision-making method oriented to passenger demand includes the following steps:

[0083] S1. Characterize the number of passengers on the platform at future moments based on the real-time changes in passenger flow, and construct a dynamic passenger flow model;

[0084] S2. Based on the dynamic passenger flow model, establish a multi-train distributed collaborative framework, transform the multi-train problem into a single-train problem of the leading train and the trailing train, and update the platform passenger flow status obtained by each train;

[0085] S3. Analyze the running relationship between the leading train and the trailing train, and construct a safety protection constraint model for the front and rear trains;

[0086] S4. Based on the safety protection constraint model for the front and rear trains and the platform passenger flow status obtained by each train, design a reinforcement learning algorithm with the goals of balanced train load factor and minimum passenger waiting time to achieve autonomous train decision-making.

[0087] In this embodiment, step S1 specifically includes the following steps:

[0088] Establish a platform arrival passenger flow model according to the time-varying characteristics of passenger flow, expressed as:

[0089]

[0090] Among them, is the platform arrival passenger flow model, j is the platform number, j ∈ N, and N is the set of stations; t is the discrete time constant; V arrive is the rate of passengers arriving at the platform, which is expressed as:

[0091]

[0092] Among them, v a,peak is the rate of passengers arriving at the platform during the peak period, and v a,offpeak is the rate of passengers arriving at the platform during the off-peak period;

[0093] Establish a platform boarding and alighting passenger flow model according to the boarding and alighting behavior characteristics of passengers, expressed as:

[0094]

[0095] Among them, is the number of passengers alighting at the j-th platform at time t; M is a constant value; is the number of passengers boarding at the j-th platform at time t; D j is the stopping time of the train at the j-th platform; V on and V off are the boarding rate and alighting rate of passengers respectively, which are expressed as:

[0096]

[0097] Among them, v on,peak and v off,peakare the passenger boarding rate and alighting rate during peak hours; v on,offpeak and v off,offpeak are the passenger boarding rate and alighting rate during off-peak hours respectively;

[0098] According to the platform arrival passenger model and the platform boarding and alighting passenger flow model, a dynamic passenger flow model is established. The dynamic passenger flow model includes the platform stranded passenger flow model and the train passenger flow model, which can be expressed as:

[0099]

[0100] in, is the number of stranded passengers at platform j at time t; is the number of passengers on the train when it arrives at the jth platform at time t; For m Passengers on the train just after arriving at platform J-1; For m Passengers boarding the bus at platform J-1 at all times; For m Passengers getting off at platform j-1 at any time; m is a discrete moment, and each moment interval is 1s.

[0101] In this embodiment, step S2 specifically includes the following steps:

[0102] S21. According to the principle of vehicle-to-vehicle communication, a multi-train distributed coordination framework is constructed, which can be expressed as:

[0103]

[0104] in, represents a directed graph of order m; V′={1,2,...,m} represents the set of train agents, and m represents the total number of train agents; represents the set of vehicle-to-vehicle communication paths; if from train i x To train y If there is a communication path, then (i x ,i y )∈ε, yes The adjacency matrix of , where R represents the set of real numbers, is an element in the matrix A′, if (i x ,i y )∈ε, then otherwise And there is

[0105] S22. Considering the fixed topology structure where a train can only communicate with the previous adjacent train, the adjacency matrix A′ and the Laplace matrix L are obtained as follows:

[0106]

[0107] where \(L = D - A'\) is the Laplacian matrix of the directed graph ; \(D\) is a diagonal matrix, \(diag(\cdot)\) represents a diagonal matrix, is an element in matrix \(D\), and The adjacency matrix \(A'\) indicates that the current train only has a communication path with the previous adjacent train. If the current train is not the first train, the current train is made to receive the platform passenger flow state calculated by the previous adjacent train. After considering the passengers already or about to be transported away by the previous trains at each platform, the passenger flow state of the current train when arriving at each platform is calculated, and thus the platform passenger flow states obtained by each train are as follows:

[0108]

[0109] where, is the actual number of platform passengers at the \(j\)-th platform received by train \(i\) when train \(i\) departs from the \(k'\)-th platform, is the estimated number of platform passengers at the \(j\)-th platform calculated by train \(i\) when train \(i\) departs from the \(k'\)-th platform; \(t\) i,k′ is the moment when train \(i\) departs from the \(k'\)-th platform, \(t\) i,k′ \(\in T\); is the number of passengers remaining at the \(j\)-th platform after train \(i\) departs from the \(k'\)-th platform at time \(t\) i,k′ ; is the number of passengers arriving at the \(j\)-th platform at time \(t\) i,k ; is the number of passengers boarding train \(x\) at the \(j\)-th platform; is the planned stop time of train \(i\) at platform \(w\) at time \(t\) i,k′ ; is the planned running time of train \(i\) in the section between platform \(w\) and platform \(w + 1\) at time \(t\); \(V\) i,k′ ; arrive is the rate at which passengers arrive at the platform.

[0110] In this embodiment, step S3 specifically includes the following steps:

[0111] Construct a train kinematic model, expressed as:

[0112]

[0113] where, is the derivative of the train's real-time position with respect to time, that is, the running speed \(v(t)\) of the train; is the train's real-time acceleration; \(u\) i(t) is the traction or braking force per unit mass of the train; g is the acceleration due to gravity; w0 is the resistance during the train operation, w0 = δ a′ a′ + δ b′ b′v i (t) + δ c′ c′v i (t) 2 , where a′, b′, and c′ are the coefficients of the Davis formula respectively, and δ a′ 、δ b′ and δ c′ are the change rates of a′, b′, and c′ respectively; γ is the coefficient of the train's rotary mass; d(t) is an unknown bounded disturbance, w i is the additional resistance of the line gradient, and w c is the additional resistance of the line curve;

[0114] Based on the train kinematic model and aiming at punctuality, design a dynamic programming algorithm to obtain the speed curves of each train;

[0115] Based on the speed curve of the leading train, construct a safety protection constraint model for the front and rear trains, expressed as:

[0116]

[0117] where t k is the discrete time constant; is the position of the (i + 1)-th train at time t k ; is the position of the i-th train at time t k ; l0 is the safety interval margin between trains.

[0118] In this embodiment, the step of designing a dynamic programming algorithm to obtain the speed curves of each train based on the train kinematic model and aiming at punctuality specifically includes the following steps:

[0119] Based on the train kinematic model, divide the inter-station distance of the train operation into K + 1 stages at intervals of Δx, and divide the speed into N + 1 speed intervals at intervals of Δv, so as to obtain the state grid points of the discretization of the train motion state space. The train operation position at stage k is expressed as x k , k ∈ {1,..., K + 1}, and the train speed state at stage k is expressed as {v k,1 ,..., v k,n ,..., v k,N+1}, where v k,n is the train speed state at stage k;

[0120] Aiming at punctuality, design the stage index function, expressed as:

[0121]

[0122] Among them, C k is the cumulative index value from stage k to the last stage K + 1; κ is a certain stage from stage k to the last stage K + 1; Δx is the stage interval distance; v k is the speed in the k-th stage, v k+1 is the speed in the (k + 1)-th stage;

[0123] The stage index function satisfies the punctuality constraint, which is expressed as:

[0124]

[0125] Among them, RT obj is the given running time between stations; Δt is the time fluctuation within the allowable range, Δt ∈ (-ε, ε); ε is the time error boundary;

[0126] Based on the stage index function and the discretized state lattice points, the optimal decision variables for the train speed states in each stage are calculated, which is expressed as:

[0127]

[0128] Among them: u k,n is the decision variable corresponding to the train speed state v k,n in the k-th stage; is the optimal decision variable corresponding to the train speed state v k,n in the k-th stage; U is the allowable decision set; c k,n (v k,n , u k,n ) is the index corresponding to the train speed state v k,n in the k-th stage when the decision u k,n is adopted; is the optimal index function corresponding to the train speed state v k+1,n′ in the (k + 1)-th stage, represents the optimal decision variable value of u k,n corresponding to the calculated optimal index objective function;

[0129] The optimal decision variables of the train speed states in each stage are combined to obtain the train speed curve.

[0130] In this embodiment, when calculating the optimal running time and the optimal stop time, the safety protection constraints between the front and rear trains need to be considered. If the calculated optimal running time and stop time are too short, then the running speed of the current train will be very fast, which may cause a collision between the current train and the previous train. Therefore, the safety protection constraint model between the front and rear trains is used to avoid the optimal running time and stop time from being too short. Therefore, the step S4 specifically includes the following steps:

[0131] Based on the platform passenger flow status obtained by each train and the safety protection constraint model of the front and rear trains, a reinforcement learning algorithm is designed to calculate the optimal running time of the current train in the remaining intervals and the optimal stop time of the remaining platforms.

[0132] Based on the optimal running time of the current train in the remaining intervals, a train speed curve is obtained.

[0133] At the rolling update moment, the train obtains the speed command at each time step and runs according to the train speed curve. If the train arrives at the platform, the train stops according to the optimal stop time, and at the same time, the platform passenger information and the remaining platform information are updated until the train arrives at the terminal platform and the operation ends, realizing the autonomous decision-making of the train.

[0134] In this embodiment, the design of the reinforcement learning algorithm based on the platform passenger flow status obtained by each train and the safety protection constraint model of the front and rear trains to calculate the optimal running time of the current train in the remaining intervals and the optimal stop time of the remaining platforms specifically includes the following steps:

[0135] Under the condition of satisfying the safety protection constraints of the front and rear trains, with the goal of minimizing the passenger waiting time and balancing the train load factor, a reward function is set, expressed as:

[0136]

[0137] where: s j represents the position state of the train at the j-th platform; is the reward function under the train position state s j ; T wait,i,j is the total waiting time of all subsequent platform passengers calculated after the i-th train arrives at the j-th platform; R load,i,j is the train load factor of all subsequent inter-station trains calculated after the i-th train arrives at the j-th platform; γ1 and γ2 are the weight coefficients of passenger waiting time and load factor balance respectively;

[0138] The passenger waiting time is expressed as:

[0139]

[0140] where: t i,j is the time when the i-th train arrives at the j-th platform; is the number of passengers getting on the i-th train at the (χ + 1)-th station; is the number of platform passengers at the (χ + 1)-th station obtained by the i-th train; is the expected stop time of the i-th train at the β-th platform; is the expected running time of the i-th train from the β-th platform to the (β + 1)-th platform, and N is the set of stations;

[0141] The train full-load rate balance is expressed as:

[0142]

[0143] Where: is the number of people on the i-th train at the χ-th platform; C is the maximum capacity of the train, ||·|| represents the two-norm, and the square root of the sum of the squares of the vectors inside ||·|| is calculated after summing the squares;

[0144] Based on the reward function, calculate the optimal Q value for choosing the optimal action in each train position state, and obtain the optimal Q value table:

[0145] Q * ={Q′(s1,a),...,Q′(s j ,a),...,Q′(s n ,a)}

[0146] Where: Q * is the optimal Q value table; Q′(s j ,a) is the optimal Q value corresponding to choosing the optimal action a in the s j state;

[0147] Calculate the running time of each interval and the stop time corresponding to the optimal Q value table, which is expressed as:

[0148] RT obj ={RT1',...,RT′ j ,...,RT′ n}

[0149] ST obj ={ST1',...,ST′ j ,...,ST′ n}

[0150] Where: RT obj and ST obj are the optimal running time of each interval and the optimal stop time of each platform mapped by Q * respectively, RT′ j and ST′ j represent the optimal running time of the train from the j-th platform to the j + 1-th platform and the optimal stop time at the j-th platform respectively, Q * →{RT obj ,ST obj}.

[0151] In this embodiment, the step of calculating the optimal Q value for choosing the optimal action in each train position state based on the reward function specifically includes the following steps:

[0152] Set the maximum number of iteration steps and randomly initialize the train position state s j ;

[0153] Based on the train position state s j Randomly select an action or select the action with the maximum Q value, which is expressed as:

[0154]

[0155] where: a is the action selected in the current train position state s j ; s j+1 is the next train position state; a w is the action that the train can select for the next position state s j+1 ; Q(s j+1 , a w ) is the Q value corresponding to the action a selected for the next train position state s j+1 ; a w is the random action; P r is the random probability, P r ∈[0, 1]; r ∈[0, 1];

[0156] Based on the reward function, update the Q value corresponding to the action selected for the current train state;

[0157]

[0158] where: Q′(s j , a) represents the Q value corresponding to the action a selected by the train in the position state s j after update, Q(s j , a) represents the Q value corresponding to the action a selected by the train in the position state s j before update, α is the learning rate, γ′ is the penalty factor, and α and γ′ are custom constants; is the reward function;

[0159] Update the next train position state until the end position is reached, jump out of the current iteration, repeat initializing the train position state, selecting an action, and calculating and updating the Q value of the action until the maximum number of iteration steps is reached, and output the optimal Q value table and its corresponding interval running time and stop time.

[0160] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.

Claims

1. A train autonomous decision-making method oriented to passenger demand, characterized in that: The following steps are involved: Based on the real-time changes in passenger flow, the number of passengers on the platform in the future is described to build a dynamic passenger flow model; Based on the dynamic passenger flow model, a multi-train distributed collaborative framework is established to transform the multi-train problem into a single-train problem of the leading and trailing trains, and update the platform passenger flow status obtained by each train; Analyze the operation relationship between the front and rear vehicles, and build a safety protection constraint model for the front and rear vehicles; Based on the safety protection constraint model of the leading and trailing vehicles and the platform passenger flow status obtained by each train, a reinforcement learning algorithm is designed to achieve autonomous train decision-making with the goal of balancing the train load rate and minimizing the passenger waiting time.

2. The train autonomous decision-making method oriented to passenger demand according to claim 1, characterized in that: The method of describing the number of passengers on the platform in the future based on the real-time change of passenger flow and building a dynamic passenger flow model specifically includes the following steps: According to the time-varying characteristics of passenger flow, the platform arrival passenger flow model is established, which can be expressed as: in, is the platform arrival passenger flow model, j is the platform number, j∈N, N is the set of stations; t is the discrete time constant; V arrive is the rate at which passengers arrive at the platform, which is expressed as: Among them, v a,peak is the passenger arrival rate at the platform during peak hours, v a,offpeak The rate of passengers arriving at the platform during off-peak hours; According to the characteristics of passengers getting on and off the bus, a passenger flow model for getting on and off the bus at the platform is established, which can be expressed as: in, is the number of passengers getting off at the jth platform at time t; M is a constant value; is the number of people boarding the bus at platform j at time t; D j is the train’s stop time at the jth platform; V on and V off are the passenger boarding rate and alighting rate, respectively, which are expressed as: Among them, v on,peak and v off,peak are the passenger boarding rate and alighting rate during peak hours; v on,offpeak and v off,offpeak are the passenger boarding rate and alighting rate during off-peak hours respectively; According to the platform arrival passenger model and the platform boarding and alighting passenger flow model, a dynamic passenger flow model is established. The dynamic passenger flow model includes the platform stranded passenger flow model and the train passenger flow model, which can be expressed as: in, is the number of stranded passengers at platform j at time t; is the number of passengers on the train when it arrives at the jth platform at time t; For m Passengers on the train just after arriving at platform J-1; For m Passengers boarding the bus at platform J-1 at all times; For m Passengers getting off at platform j-1 at any time; m is a discrete moment, and each moment interval is 1s.

3. The train autonomous decision-making method oriented to passenger demand according to claim 1, characterized in that: The method is based on a dynamic passenger flow model to establish a multi-train distributed collaborative framework, convert the multi-train problem into a single-train problem of the front and rear trains, and update the platform passenger flow status obtained by each train, which specifically includes the following steps: According to the principle of vehicle-to-vehicle communication, a multi-train distributed coordination framework is constructed, which can be expressed as: in, represents a directed graph of order m; V′={1,2,...,m} represents the set of train agents, and m represents the total number of train agents; represents the set of vehicle-to-vehicle communication paths; if from train i x To train y If there is a communication path, then (i x ,i y )∈ε, yes The adjacency matrix of , where R represents the set of real numbers, is an element in the matrix A′, if (i x ,i y )∈ε, then otherwise And there is The adjacency matrix A′ shows that the current train has a communication path only with the previous adjacent train. If the current train is not the first train, the current train receives the platform passenger flow status calculated by the previous adjacent train, and after taking into account the passengers that have been or will be transported by the previous train at each platform, the passenger flow status of the current train when it arrives at each platform is calculated based on the dynamic passenger flow model, and then the platform passenger flow status obtained by each train is obtained as follows: in, is the actual number of passengers received by train i at platform j when train i leaves platform k′, is the estimated number of passengers at the jth platform calculated by train i when train i leaves the k′th platform; t i,k′ is the time when train i leaves platform k′, t i,k′ ∈T; For train i at t i,k′ The number of passengers who stay at the jth platform after leaving the k′ platform at time t; For i,k The number of passengers arriving at platform j at time ; is the number of passengers boarding the train x at the jth platform; For i,k′ Time, the planned stop time of train i at platform w; For i,k′ At time V, the planned running time of train i between platform w and platform w+1; arrive is the rate at which passengers arrive at the platform.

4. The train autonomous decision-making method oriented to passenger demand according to claim 1, characterized in that: The analysis of the running relationship between the front vehicle and the rear vehicle and the construction of the front and rear vehicle safety protection constraint model specifically includes the following steps: Construct the train kinematic model, expressed as: in, is the derivative of the train's real-time position with respect to time, that is, the train's running speed v(t); is the real-time acceleration of the train; u i (t) is the traction or braking force per unit mass of the train; g is the acceleration of gravity; w0 is the resistance during the train operation, w0 = δ a′ a′+δ b′ b′v i (t)+δ c′ c′v i (t) 2 , a′, b′ and c′ are the coefficients of Davis formula, δ a′ , δ b′ and δ c′ are the rates of change of a′, b′ and c′ respectively; γ is the train rotation mass coefficient; d(t) is the unknown bounded disturbance, w i Add resistance to the line slope, w c Add resistance to line curves; Based on the train kinematic model and with punctuality as the goal, a dynamic programming algorithm is designed to obtain the speed curve of each train; Based on the front vehicle speed curve, the front and rear vehicle safety protection constraint model is constructed, which is expressed as: Where t is the discrete time constant; is the position of the i+1th train at time t; is the position of the i-th train at time t; l0 is the safety interval margin between trains.

5. The train autonomous decision-making method oriented to passenger demand according to claim 4, characterized in that: Based on the train kinematic model, the dynamic programming algorithm is designed with punctuality as the goal to obtain the speed curve of each train, which specifically includes the following steps: Based on the train kinematic model, the distance between train stations is divided into K+1 stages according to the Δx interval, and the speed is divided into N+1 speed intervals according to the Δv interval, so as to obtain the state grid points of the train motion state space discretization. The train running position in stage k is expressed as x k , k∈{1,...,K+1}, the train speed state at stage k is expressed as {v k,1 ,...,v k,n ,...,v k,N+1 },v k,n is the speed state of the train at stage k; Taking punctuality as the goal, the design phase indicator function is expressed as: Among them, C k is the cumulative index value from stage k to the final stage K+1; κ is a stage from stage k to the final stage K+1; Δx is the stage interval distance; v k is the speed of the kth stage, v k+1 is the speed of the k+1th stage; Based on the stage index function and the discretized state grid, the optimal decision variables of the train speed state at each stage are calculated and expressed as: Where: u k,n is the speed state v of the train in stage k k,n The corresponding decision variables; is the speed state v of the train in stage k k,n The corresponding optimal decision variable; U is the set of allowed decisions; c k,n (v k,n ,u k,n ) is the speed state v of the train in stage k k,n Adopt u k,n Indicators corresponding to the decision; is the speed state v of the train in stage k+1 k+1,n′ The corresponding optimal index function, n represents a certain train speed state, n′ represents an uncertain train speed state, which changes with the change of the optimal target value. Indicates the u corresponding to the calculated optimal indicator objective function k,n Optimal decision variable values; The optimal decision variables of the train speed state at each stage are combined to obtain the train speed curve.

6. The train autonomous decision-making method oriented to passenger demand according to claim 1, characterized in that: Based on the safety protection constraint model of the preceding and following vehicles and the platform passenger flow status obtained by each train, a reinforcement learning algorithm is designed to achieve train autonomous decision-making with the goal of balancing the train load rate and minimizing the passenger waiting time, which specifically includes the following steps: Based on the platform passenger flow status obtained by each train and the safety protection constraint model of the front and rear vehicles, a reinforcement learning algorithm is designed to calculate the optimal running time of the current train in the remaining sections and the optimal stop time of the remaining platforms; Based on the optimal running time of the current train in each remaining section, obtain the train speed curve; At the rolling update moment, the train obtains the speed instruction at each time step and runs according to the train speed curve. If the train arrives at the platform, it stops according to the optimal stop time and updates the platform passenger information and the remaining platform information at the same time until the train arrives at the terminal platform and ends the operation, thus realizing autonomous decision-making of the train.

7. The train autonomous decision-making method oriented to passenger demand according to claim 6, characterized in that: The method of designing a reinforcement learning algorithm based on the platform passenger flow status obtained by each train and the safety protection constraint model of the preceding and following vehicles to calculate the optimal running time of the current train in each remaining section and the optimal stop time of each remaining platform specifically includes the following steps: Under the condition of satisfying the safety protection constraints of the front and rear vehicles, the reward function is set with the goal of minimizing the waiting time of passengers and balancing the train load rate, which is expressed as: Where: s j Indicates the position status of the train at the jth platform; is the train position state s j The reward function under T wait,i,j R is the waiting time of all subsequent platform passengers calculated after the i-th train arrives at the j-th platform; load,i,j is the load factor of all subsequent inter-station trains calculated after the i-th train arrives at the j-th platform; γ1 and γ2 are the weight coefficients for balancing the passenger waiting time and load factor respectively; Passenger waiting time is expressed as: Where: t i,j is the time when the i-th train arrives at the j-th platform; is the number of people boarding the i-th train at the x+1th station; The number of platform passengers at the x+1th station obtained for the i-th train; is the estimated stop time of the i-th train at the β-th platform; is the estimated running time of the i-th train from the β-th platform to the β+1-th platform, and N is the set of stations; The train load factor equilibrium is expressed as: in: is the number of people on the i-th train at the x-th platform; C is the maximum capacity of the train; ||·|| represents the second norm; Based on the reward function, the optimal Q value of selecting the optimal action under each train position state is calculated, and the optimal Q value table is obtained: Q * ={Q′(s1,a),...,Q′(s j ,a),...,Q′(s n ,a)} Where: Q * is the optimal Q value table; Q′(s j ,a) is in s j Select the optimal Q value corresponding to the optimal action a in the state; Calculate the interval running time and stop time corresponding to the optimal Q value table, expressed as: RT obj ={RT1',...,RT j ',...,RT n '} ST obj ={ST1',...,ST j ',...,ST n '} Among them: RT obj and ST obj Q * The optimal running time of each section and the optimal stopping time of each station are mapped, RT j ' and ST j 'represents the optimal running time of the train from the jth platform to the j+1th platform and the optimal stopping time at the jth platform, respectively. * →{RT obj ,ST obj }.

8. The train autonomous decision-making method oriented to passenger demand according to claim 7, characterized in that: The method of calculating the optimal Q value of selecting the optimal action under each train position state based on the reward function specifically includes the following steps: Set the maximum iteration step size and randomly initialize the train position state s j ; Based on the train position status s j Randomly select an action or select the action with the largest Q value, expressed as: Where: a is the current position state of the train s j Selected action;s j+1 is the next position state of the train; a w The next position state of the train is s j+1 The actions you can choose; Indicates the action a corresponding to the maximum Q value w , Q(s j+1 ,a w ) is the next position state of the train s j+1 Select action a w The corresponding Q value; a r is a random action; P r is the random probability, P r ∈[0,1]; Based on the reward function, update the Q value corresponding to the action selected by the current state of the train; Where: Q′(s j ,a) indicates the updated train position state s j The Q value corresponding to the selected action a, Q(s j ,a) indicates the train is in position state s before updating j The Q value corresponding to the selected action a, α is the learning rate, γ′ is the penalty factor, and α and γ′ are custom constants; is the reward function; Update the next position state of the train until it reaches the terminal position, jump out of the current iteration, repeatedly initialize the train position state, select actions, and calculate the Q value of the updated action until the maximum iteration step is reached, and output the optimal Q value table and its corresponding interval running time and stop time.