Depth reinforcement learning-based old-age adaptability optimization method for built public transportation network

By optimizing the public transportation network through deep reinforcement learning, the problems of accessibility for the elderly and overall network changes have been solved, resulting in a reduction in walking distance and number of transfers for the elderly, and improving the efficiency and stability of the public transportation system.

CN120911675APending Publication Date: 2025-11-07ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511018015.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing bus network optimization methods fail to effectively address changes in overall network demand and accessibility for the elderly when considering the elderly factor, and are difficult to apply to complex existing bus networks, resulting in unstable optimization results and high costs.

Method used

This study employs a deep reinforcement learning-based approach, constructing an optimization objective function and reward mechanism. It utilizes the DQN algorithm to optimize the public transportation network, combining an elderly walking penalty factor and the K-shortest path algorithm to optimize bus routes and reduce the walking distance and number of transfers for the elderly.

Benefits of technology

It effectively reduces the walking distance and number of transfers for the elderly, improves the overall efficiency and stability of the public transportation system, adapts to the optimization needs of the existing network, and reduces system costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911675A_ABST
    Figure CN120911675A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of public transportation network optimization, and discloses a built public transportation network suitability optimization method based on deep reinforcement learning. Comprising the steps of 1, setting an optimization target; step 2, constructing an optimization objective function according to the optimization objective; and step 3, solving by adopting a DQN-based method. Aiming at the aging development trend of the current society, in the public transit network optimization problem, consideration of aging suitability is added, and reachability, convenience and the like concerned by the elderly are taken as optimization objectives; the invention provides a DQN-based solving method aiming at the problems that a current mainstream search method is weak in result stability and difficult to adapt to an established public transportation network. The invention further provides a DQN-based solving method for solving the problems that the current mainstream search method is weak in result stability and difficult to adapt to the established public transportation network. The full load rate of the supersaturation interval can be effectively reduced, and the average walking distance of old people and the transfer passenger flow ratio can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of public transport network optimization, and in particular to an old-age suitability optimization method for an existing public transport network based on deep reinforcement learning. BACKGROUND

[0002] Transit Route Network Design Problem (TRNDP) has become an important direction in urban traffic research, aiming to build an efficient, economical and user-friendly public transport system

[0003] With the development of aging trend all over the world, especially in developed countries, the elderly factor is increasingly valued in the field of public transport research and application. The public transport research considering the elderly factor is still focused on the analysis of the elderly travel characteristics and demand prediction, the improvement of infrastructure and specific operation strategy, and the design and optimization of the online network. The research on the old-age suitability of the online network design and optimization is relatively basic.

[0004] Currently, there are few studies on the old-age suitability improvement in the process of public transport network design and route optimization. Among them, Chen et al. (2020) made some achievements. In their study, the elderly accessibility was selected as the main optimization target, and two artificially set strategies, increasing the number of stations in the small section and adjusting the position of stations in the small section, were set as the route optimization approach. Finally, a bi-level optimization model was constructed, and a genetic algorithm was used to solve it. This study specifically considers the elderly factor, but there is still much room for improvement in many aspects. First, it takes accessibility as the optimization target, but the calculation of accessibility is simple, and the straight-line distance between the community and the station is simply used as the representation of accessibility, without considering the differences between the elderly passenger flow demand of different stations and the actual walking distance of the elderly group. In terms of optimization strategy, the selection space is relatively small, and only by artificially setting rules to change the number of stations in the small section and the route direction, it is only suitable for single-line small-scale optimization. In terms of solving method and process, it still uses a genetic algorithm-based solving method, which is difficult to apply to the actual situation of large-scale network and long route. The most core problem is that the public transport network is a complex network as a whole, and when a line is changed, the overall network demand distribution will change. In the process of route optimization, it does not consider the impact of route adjustment on the entire network, which greatly reduces the integrity of the method. SUMMARY

[0005] To solve the above technical problems, the present application provides an old-age suitability optimization method for an existing public transport network based on deep reinforcement learning to solve the problems in the prior art. To achieve the above application purposes, the technical solution adopted by the present application is:

[0006] The method for optimizing the established public transport network for the elderly based on deep reinforcement learning comprises the following steps:

[0007] Step 1, setting the optimization goal;

[0008] Step 2, constructing an optimization objective function according to the optimization goal;

[0009] Step 3, using the DQN-based method to solve.

[0010] Further, in step 1, the optimization goal includes better adapting to overall passenger flow changes and better serving the elderly.

[0011] Further, in step 2, for better adaptation to overall passenger flow changes, the following objective function is constructed:

[0012] maxZ c (G * )=1-C(G * )·C(G) -1

[0013] C(G * )=C u (G * )+C o (G * )

[0014]

[0015] C o (G * )=C bus (G * )+C ope (G * )

[0016]

[0017] In the formula, G * is the public transport network after optimization adjustment; Z c (G * ) is the total system cost reduction rate after network optimization, wherein the total system cost C(G * ) includes user time cost G u (G * ) and operator financial cost C o (G * ); w u is the rate of user time cost; n is the number of road network nodes; q ij is the public transport demand between node i and node j; T ij is the travel time between node i and node j, which includes walking time Tw , bus route travel time , station dwell and wait time , and transfer time L ij is the set of links l included in the shortest bus path between node pair ij; DW ij is the set of stations dw on the shortest bus path between node pair ij; TR ij is the set of stations tr where transfer is needed on the shortest bus path between node pair ij; T l is the travel time of the bus on link l; T dw is the dwell time of the bus at station dw; T tr is the transfer time of the traveler at transfer station tr; and are the electric bus vehicle cost and operating cost, respectively; N k is the number of electric bus vehicles needed to be deployed for route k; is the daily operating time of bus route k; LEN k is the route length of bus route k; v b is the average bus speed; f k is the headway of bus route k; Q k,max is the maximum cross-sectional flow of route k during peak hours; N C is the rated passenger capacity of the bus vehicle; is the empty running distance of the operating vehicle of route k in a single charging process; is the empty running time of the operating vehicle of route k in a single charging process; is the charging time of the vehicle; and are the rates of vehicle cost C bus and operating cost C ope , respectively.

[0018] Further, in step 2, for better service for the elderly, the following objective function is constructed:

[0019] min Z d (G * , Q * ) = 1 - D(G * , Q * ) · D(G, Q * ) -1

[0020]

[0021] d ij = d ij,O + d ij,T + dij,D

[0022] where Q * is the demand of public transport for the elderly passengers; min Z d (G * ,Q * ) represents the reduction ratio of the total walking distance of the elderly after optimization; D(G * ,Q * ) is the total walking distance of the elderly in the network G * ; is the elderly passenger flow between nodes i and j; d ij is the walking distance from node i to node j when taking public transport; d ij,O , d ij,T , and d ij,D are the walking distances when boarding, transferring, and alighting, respectively;

[0023] Then, the walking penalty factor β is introduced to update the impedance matrix M * of the superimposed network E and G * :

[0024] M * = [t ij ] n×n ; i = 1, 2, …, n; j = 1, 2, …, n

[0025]

[0026] t ij is the time impedance between nodes i and j; l ij is the length of the road segment between adjacent nodes i and j; v b is the average operating speed of public transport; v w is the walking speed of the elderly; and the values of d ij,O , d ij,T , and d ij,D are calculated by the K-shortest path algorithm based on the updated impedance matrix;

[0027] The arrangement can be obtained as follows:

[0028] min Z f (G * ,Q * ) = 1 - F(G * ,Q * ) · F(G, Q * ) -1

[0029]

[0030] where min Z f (G* ,Q * ) is the reduction ratio of the overall transfer number of the elderly group in the process of public transport; F(G * ,Q * ) is the network G * corresponding to the overall transfer number of the elderly; tf ij is the transfer number of the bus path between node i and node j.

[0031] Further, step 3 comprises the following steps:

[0032] Step 3.1, basic network G analysis, determine the reserved network G r and the network G a that needs to be adjusted;

[0033] Step 3.2, select 1 line that needs to be adjusted in G a ;

[0034] Step 3.3, define state s t , action a t and reward function R(s t+1 |s t , a t );

[0035] Step 3.4, DQN training;

[0036] Step 3.5, determine the optimal adjustment scheme of the line based on ε-greedy;

[0037] Step 3.6, judge whether all lines in G a are adjusted; if yes, continue to step 3.7, otherwise repeat steps 3.2-3.6;

[0038] Step 3.7, end; output the optimized line network.

[0039] Further, in step 3.3, for state s t and action a t , it is represented as:

[0040] s t = (e1, e2, e3, …, e i , …, e N ); e i ∈{0,1,2,…,N}

[0041] a t ∈{0,1,2,…,N}

[0042] The transition probability between the states of the agent is as follows:

[0043]

[0044] where

[0045] where P(s t+1 |s t ,a t ) is the transition probability of the agent from state s t to state s t after taking action a t+1 ; is a fixed vector, m st is the number of non-zero elements in vector s t .

[0046] Further, in step 3.3, the calculation of the reward function R(s t+1 |s t ,a t ) includes the following steps:

[0047] Step 3.31: Determine the generated optimized bus route L t based on the current state s t ;

[0048] Step 3.32: Add L t to G r to form the calculation network G t ;

[0049] Step 3.33: Perform passenger flow distribution on the network G r using the KSP algorithm and the Logit model;

[0050] Step 3.34: Calculate and record Z(G t );

[0051] Step 3.35: Determine if the iteration is complete: if yes, go to step 3.37, otherwise go to step 3.36;

[0052] Step 3.36: If return the reward R(s t+1 |s t ,a t ) = -1; otherwise, update s t , L t , G t , and return to step 3;

[0053] Step 3.37: Calculate the reward R(s t+1 |s t ,a t ) and return;

[0054] where R(s t+1 |st a t ) is represented as:

[0055]

[0056] G t = G r ∪ L t

[0057] E f = E t ∩ (E v ) C

[0058] where R(s t+1 | s t , a t ) is the immediate reward obtained by the agent when taking action a t to reach state s t from state s t+1 ; L t+1 and L t are the generated bus routes corresponding to state s t+1 and state s t , respectively; E t is the set of nodes directly connected to the last node n t of route L t ; E v is the set of nodes already visited by route L t ; (E v ) C is the complement of E v .

[0059] Further, step 3.4, DQN training comprises the following steps:

[0060] Step 3.41 input road network data E, demand matrix D, preserved bus network G r ;

[0061] Step 3.42 set basic training parameters;

[0062] Step 3.43 define Q-network Q(S; Θ), target network T(S; Θ) and experience replay buffer;

[0063] Step 3.44 initialize current node n t , current state s t , L t ;

[0064] Step 3.45 select action a t = argmax Q(s t ; Θ), calculate reward R(s t+1|s t ,a t );

[0065] Step 3.46 stores (n t , L t , s t , a t , R(s t+1 |s t , a t ), n t+1 , L t+1 , s t+1 ) into the experience replay buffer;

[0066] Step 3.47 judges whether the experience replay buffer data is sufficient;Yes: go to step 3.48;No: go to step 3.410;

[0067] Step 3.48 extracts batch data from the experience replay buffer for training and updates the Q(S;Θ) network parameters;

[0068] Step 3.49 updates n t =n t+1 , L t =L t+1 , s t =s t+1 ;

[0069] Step 3.410 judges whether the route is generated according to the length of L t ;Yes: return to step 3.44;No: return to step 3.45;

[0070] Step 3.411 updates the parameters of the target network T(S;Θ) every certain number of training rounds in each iteration;

[0071] Step 3.412 terminates the training when the number of iterations exceeds the number of training rounds.

[0072] The present application has the following beneficial effects:

[0073] Compared with the traditional heuristic search method, the present application can avoid the problems of non-convergence and premature convergence;

[0074] The present application is more suitable for optimizing the existing public transportation network, and the traditional heuristic search method is more suitable for new network construction. The method can avoid the problem of unstable results, and can make greater use of existing bus stations and line facilities, thereby reducing the optimization cost of the public transportation system;

[0075] The present application takes into account the aging of the population, and is more suitable for the current aging of the population. The design effect can effectively reduce the walking distance and the number of transfers of the elderly group when taking public transportation. BRIEF DESCRIPTION OF DRAWINGS

[0076] Figure 1 for the flowchart of the present application;

[0077] Figure 2 for the description of the problem studied by the present application;

[0078] Figure 3 for the description of the state and action. DETAILED DESCRIPTION

[0079] The technical solutions in the embodiments of the present application will be clearly and completely described below. Figures 1-3 The technical solutions in the embodiments of the present application will be clearly and completely described below.

[0080] The present application proposes an optimization method for the adaptability of an existing public transport network based on deep reinforcement learning. Compared with the current mainstream search type line network design method, the method greatly improves the applicability in the optimization problem of the existing network, avoids the problems of non-convergence and premature convergence of traditional algorithms, and provides a new solution for public transport line network design and optimization. Secondly, in view of the current social aging trend, the adaptability of the elderly is considered in the public transport network optimization problem, and the accessibility and convenience of the elderly are taken as the optimization objectives, so that the optimization result can better adapt to the travel structure aging problem that the public transport will inevitably face in the future. The present application proposes a solving method based on DQN (DeepQ-Network) to solve the problem that the current mainstream search method has weak stability and is difficult to adapt to the existing public transport network. The method can effectively reduce the full load rate in the oversaturation interval, reduce the average walking distance of the elderly and the passenger flow ratio of transfer. The general description of the problem studied in the present application is shown in Figure 2 .

[0081] As Figure 1 the optimization method for the adaptability of an existing public transport network based on deep reinforcement learning comprises the following steps:

[0082] Step 1, setting the optimization target;

[0083] Step 2, constructing an optimization objective function according to the optimization target;

[0084] Step 3, solving by using the method based on DQN.

[0085] Further, in step 1, in the current bus line network design and optimization model, the efficiency factor is paid more attention to. And the cost, travel time and the like are selected as the index for expressing system efficiency by more scholars. On the basis of the existing research, the old adaptation factor is further considered in the application. Therefore, the optimization target of the application is set as: adapting to the overall demand change, guaranteeing the system efficiency; and at the same time, considering the accessibility and convenience of the travel of the old age group. Specifically, the optimization target includes better adapting to the overall passenger flow change and better serving the old age group.

[0086] Further, in step 2, for better adapting to the overall passenger flow change, the following objective function is constructed:

[0087] max Z c (G * )=1-C(G * )·C(G) -1

[0088] C(G * )=C u (G * )+C o (G * )

[0089]

[0090] C o (G * )=C bus (G * )+C ope (G * )

[0091]

[0092] In the formula, G * is the bus network after optimization adjustment; Z c (G * ) is the system total cost reduction rate after network optimization, wherein the system total cost C(G * ) includes user time cost C u (G * ) and operator fund cost C o (G * ); w u is the rate of user time cost; n is the number of road network nodes; q ij is the bus demand between node i and node j; T ij is the travel time between node i and node j, which includes walking time T w , bus section travel time , station parking and waiting time and transfer time L ij is the set of links l included in the shortest bus path between node pair ij; DW ij is the set of stops dw on the shortest bus path between node pair ij; TR ij is the set of stops tr where transfer is needed on the shortest bus path between node pair ij; T l is the travel time of the bus on link l; T dw is the dwell time of the bus at stop dw; T tr is the transfer time of the traveler at transfer stop tr; and are the electric bus vehicle cost and operation cost, respectively; N k is the number of electric buses needed to be deployed for route k; is the daily operation time of bus route k; LEN k is the route length of bus route k; v b is the average bus speed; f k is the headway of bus route k; Q k,max is the maximum cross-sectional flow of bus route k in peak hour; N C is the rated passenger capacity of the bus; is the empty running distance of the bus caused by the single charging process of the operation bus of route k; is the empty running time of the bus caused by the single charging process of the operation bus of route k; is the charging time of the bus; and are the rates of bus vehicle cost C bus and operation cost C ope , respectively.

[0093] Further, existing research has shown that in public transportation travel, factors such as walking distance, number of transfers, accessibility, and safety are more concerned by the elderly. Among them, shortening the walking distance of the elderly group and reducing the number of transfers of the elderly group can be effectively improved by optimizing the bus line network. Therefore, for the problem of aging, in step 2, in order to better serve the elderly, the following objective function is constructed:

[0094] min Z d (G * , Q * ) = 1 - D(G * , Q * ) · D(G, Q * ) -1

[0095]

[0096] dij =d ij,O +d ij,T +d ij,D

[0097] In the formula, Q * To meet the public transport needs of elderly passengers; min Z d (G * Q * D(G) represents the percentage reduction in the total walking distance of the elderly after optimization; * Q * ) represents the corresponding network G * Total walking distance for middle-aged and elderly passengers; For elderly passenger flow between nodes i and j; d ij d represents the walking distance from node i to node j when taking the bus; ij,o d ij,T d ij,D These are the walking distances for boarding, transferring, and alighting, respectively.

[0098] Then, a walking penalty factor β is introduced to update the road network E and the public transport network G. * Impedance matrix M of the superimposed network * :

[0099] M * =[t ij ] n×n ; i=1,2,…,n; j=1,2,…,n

[0100]

[0101] t ij The time impedance between nodes i and j; l ij v is the length of the road segment between adjacent nodes i and j; b v represents the average operating speed of public transport. w The walking speed of the elderly is calculated using the K-shortest path algorithm based on the updated impedance matrix. ij,O d ij,T d ij,D The value;

[0102] After sorting, we can obtain:

[0103] minZ f (G * Q * )=1-F(G * Q * )·F(G,Q * ) -1

[0104]

[0105] minZ f (G * ,Q * ) is the overall transfer reduction ratio of the elderly in the process of public transport; F(G * ,Q * ) is the overall transfer number of the elderly in the network G * ; tf ij is the transfer number of the bus route between node i and node j.

[0106] In addition, for the optimization objective function, constraint conditions need to be constructed, as follows:

[0107] (1) Line length constraint:

[0108]

[0109] In the formula, L k is the length of the optimized bus line k; Lmax min are the minimum and maximum length thresholds of the bus line, respectively.

[0110] (2) Adjustment ratio constraint:

[0111] len(G * ) / len(G)≤P max

[0112] In the formula, len(G * ) and len(G) are the number of bus lines included in the network G * and G, respectively; P max is the maximum threshold of the adjusted bus line number ratio.

[0113] (3) Flow conservation constraint:

[0114] The total amount of public transport demand at each node is equal to the total passenger flow on all lines passing through the node.

[0115]

[0116] In the formula, PK i is the set of all bus lines passing through node i; q i,k is the passenger flow on the k line selected at node i.

[0117] The flow of any section l is equal to the total flow of any OD pair passing through the section l:

[0118]

[0119] In the formula, q lis the passenger flow of route k at section l; τ is a binary variable.

[0120] The flow of any route k at section l is equal to the sum of the flow of any OD pair ij passing through section l by route k:

[0121]

[0122] where q k,l is the flow of route k at section l; line i,j,l is the bus route number of the bus route that the bus trip between OD pair ij takes at section l; ρ is a binary variable.

[0123] Further, step 3 comprises the following steps:

[0124] Step 3.1, basic network G analysis, determine the reserved network G r and the network G a that needs to be adjusted;

[0125] Step 3.2, select one route that needs to be adjusted in G a ;

[0126] Step 3.3, define state s t , action a t and reward function R(s t+1 |s t , a t );

[0127] Step 3.4, DQN training;

[0128] Step 3.5, determine the optimal adjustment scheme of the route based on ε-greedy;

[0129] Step 3.6, judge whether all routes in G a have been adjusted; if yes, continue to step 3.7, otherwise repeat steps 3.2-3.6;

[0130] Step 3.7, end; output the optimized line network.

[0131] Further, in step 3.3, the information included in state s t should at least cover two aspects: the current node position and the previously passed nodes and order. Based on this, an integer coding form is adopted, using an N-dimensional vector to represent the state. The index position of the vector represents the node number, and the element size represents the order of the route passing through the node. The corresponding mathematical expressions of state s t and action a t are as follows:

[0132] s t= (e1, e2, e3,..., e i ) N ); e i ∈ {0, 1, 2,..., N}

[0133] a t ∈ {0, 1, 2,..., N}

[0134] In the above encoding form, although the state space S and the action space A are discrete, with the increase of the number of network nodes, the state will present exponential growth, and the state space can be approximated as infinite. It is difficult to handle the infinite state using basic reinforcement learning methods such as Q-learning, therefore, in order to solve this problem, a DQN method is designed, which combines Deep-learning with Q-learning, and uses an MLP network to replace the Q-table in traditional Q-learning, so as to improve the efficiency and optimality of the method. Correspondingly, in view of the above definitions of state s t and action a t and the characteristics of the bus line planning problem, the state s t that the agent can reach by taking action a t in different state s t+1 is determined, so the transition probability between the states of the agent is as follows:

[0135]

[0136] wherein

[0137] In the formula, P(s t+1 |s t , a t ) is the transition probability of the state s t that the agent can reach by taking action a t in state s t+1 ; is a fixed vector, m st is the number of non-zero elements in the vector s t . The descriptions of state s t and action a t are as follows: Figure 3 .

[0138] When the feasible action space E f that the agent can choose in state s t is limited, when the agent chooses an invalid action, a penalty of -1 is given; when the agent takes an effective action a t in state s t to reach state s t+1in the process, the optimized bus line from L t extended to L t+1 , the gain of the process for the optimization target as immediate reward. Unlike the traditional path planning problem, the existing line network has a full consideration of the impact of the optimization line in the process of bus line network optimization adjustment, therefore, in step 3.3, the calculation process of the reward function R(s t+1 |s t ,a t ) includes the following steps:

[0139] Step 3.31: determine the generated optimization bus line L t according to the current state s t ;

[0140] Step 3.32: add L t to G r , form the calculation network G t ;

[0141] Step 3.33: use KSP algorithm and Logit model to perform passenger flow distribution on the network G r ;

[0142] Step 3.34: calculate and record Z(G t );

[0143] Step 3.35: determine whether the iteration is ended: if yes, enter step 3.37, otherwise enter step 3.36;

[0144] Step 3.36: if return reward R(s t+1 |s t ,a t )=-1; otherwise update s t , L t , G t , return to step 3;

[0145] Step 3.37: calculate the reward R(s t+1 |s t ,a t ) and return;

[0146] Wherein, R(s t+1 |s t ,a t ) is expressed as:

[0147]

[0148] G t =G r ∪L t

[0149] E f = E t ∩(E v ) C

[0150] where R(s t+1 |s t ,a t ) is the immediate reward obtained by the agent when taking action a t to reach state s t from state s t+1 ; L t+1 and L t are the generated bus routes corresponding to state s t+1 and state s t , respectively; E t is the set of nodes directly connected to the last node n t of route L t ; E v is the set of nodes already visited by route L t ; (E v ) C is the complement of E v .

[0151] Further, step 3.4, DQN training includes the following steps:

[0152] Step 3.41 input road network data E, demand matrix D, and retained bus network G r .

[0153] Step 3.42 set basic training parameters such as training episodes, learning rate, batch size, etc.

[0154] Step 3.43 define Q network Q(S; Θ), target network T(S; Θ), and experience replay buffer.

[0155] Step 3.44 initialize current node n t , current state s t , L t .

[0156] Step 3.45 select action a t = argmax Q(s t ; Θ), calculate reward R(s t+1 |s t , a t ).

[0157] Step 3.46 (n t , L t , s t , a tR(s t+1 |s t ,a t ),n t+1 ,L t+1 ,s t+1 ) into the experience replay buffer.

[0158] Step 3.47 judges whether the experience replay buffer data is sufficient; yes: go to step 3.48; no: go to step 3.410.

[0159] Step 3.48 extracts batch data from the experience replay buffer for training, and updates the Q(S; Θ) network parameters.

[0160] Step 3.49 updates n t =n t+1 , L t =L t+1 , s t =s t+1 .

[0161] Step 3.410 judges whether the route is generated according to the length of L t ; yes: return to step 3.44; no: return to step 3.45.

[0162] Step 3.411 updates the parameters of the target network T(S; Θ) every certain number of training rounds in each iteration.

[0163] Step 3.412 terminates the training when the number of iterations exceeds the number of training rounds.

[0164] The above-described embodiments are only to describe the preferred modes of the present application, and not to limit the scope of the present application. Without departing from the design spirit of the present application, various modifications, variations, modifications, and replacements of the technical solutions of the present application made by those skilled in the art shall fall within the protection scope determined by the claims of the present application.

Claims

1. A method for optimizing the age-friendliness of an existing public transport network based on deep reinforcement learning, characterized in that, The steps include: Step 1, set optimization target; Step 2, construct optimization objective function according to optimization target; Step 3, use DQN-based method to solve.

2. The method of claim 1, wherein the method is based on deep reinforcement learning. In step 1, the optimization target includes better adaptation to overall passenger flow changes and better service for the elderly. 3.The method of claim 2, wherein, In step 2, for better adaptation to overall passenger flow changes, the following objective function is constructed: max Z c (G * )=1-C(G * )·C(G) -1 C(G * ) = C u (G * )+C o (G * ) C o (G * )=C bus (G * )+C ope (G * ) In the formula, G * To optimize and adjust the public transportation network; Z c (G * ) represents the reduction in total system cost after network optimization, where the total system cost C(G) is the percentage decrease. * This includes user time cost C. u (G * ) and operator's capital costs C o (G * );w u The rate represents the user's time cost; n represents the number of road network nodes; q ij T represents the public transport demand between node i and node j. ij The travel time between node i and node j includes walking time T. w Travel time for bus routes Station parking and waiting time and transfer time L ij Let DW be the set of road segments l included in the shortest bus path between node pairs ij; ij TR is the set of stops dw along the shortest bus route between node pairs ij; ij Let T be the set of stations tr that require transfers on the shortest bus route between node pairs ij; l T represents the travel time of the bus on route segment l. dw The stopping time of the bus at stop dw; T tr This refers to the transfer time for travelers at the transfer station tr. and These are the costs of electric buses and their operating costs; N k The number of electric buses required for route k; For the daily operating hours of bus route k; LEN k v is the length of bus route k; b f represents the average operating speed of the bus. k For bus route k, the departure interval is Q. k,max N represents the maximum cross-sectional flow rate of line k during peak hours; C The rated passenger capacity of a public transport vehicle; The empty driving distance of the operating vehicles on line k during a single charging process; The empty running time caused by a single charging process of the vehicle operating on line k; Charging time for the vehicle; and Vehicle cost C bus and operating expenses C ope The rates. 4.The method of claim 3, wherein, In step 2, for better service for the elderly, the following objective function is constructed: min Z d (G * ,Q * ) = 1 - D(G * ,Q * ) · D(G,Q * ) -1 d ij = d ij,O + d ij,T + d ij,D In the formula, Q * is the demand of the bus for the elderly passengers; min Z d (G * , Q * ) represents the reduction ratio of the total walking distance of the elderly after optimization; D(G * , Q * ) is the total walking distance of the elderly in the network G * ; is the elderly passenger flow between nodes i and j; d ij is the walking distance from node i to node j when taking the bus; d ij,O , d ij,T , d ij,D are the walking distances when boarding, transferring and alighting, respectively; Then introduce the walking penalty factor β to update the road network E and the bus network G * The impedance matrix M of the superimposed network * : M * = [t ij ] n×n ; i = 1, 2,..., n; j = 1, 2,..., n t ij is the time impedance between nodes i and j; l ij is the length of the link between adjacent nodes i and j; v b is the average bus operating speed; v w is the walking speed of the elderly; the values of d ij,O , d ij,T , and d ij,D are calculated by the K-shortest path algorithm based on the updated impedance matrix; The arrangement is as follows: min Z f (G * ,Q * ) = 1 - F(G * ,Q * ) · F(G,Q * ) -1 In the formula, min Z f (G * , Q * ) is the overall transfer reduction ratio of the elderly population in the process of public transport; F(G * , Q * ) is the overall transfer number of the elderly in the network G * ; and tf ij is the transfer number of the bus path between node i and node j. 5.The method of claim 4, wherein, Step 3 includes the following steps: Step 3.1, base network G analysis, determine to retain network G r and adjust network G a ; Step 3.2, in G a selecting one line to adjust; Step 3.3, define state s t , action a t , and reward function R(s t+1 |s t ,a t ). Step 3.4, DQN training; Step 3.5, determine the optimal adjustment scheme for the line based on ε-greedy; Step 3.6, judge G a whether all lines are adjusted; if yes, continue step 3.7, otherwise repeat step 3.2-step 3.6; Step 3.7, end; output the optimized line network. 6.The method of claim 5, wherein, In step 3.3, for state s t and action a t is denoted by: s t = (e1, e2, e3,..., en) ; en∈{0, 1, 2,..., N} i N i = (e1, e2, e3,..., en) ; en∈{0, 1, 2,..., N}​​ a t ∈ {0, 1, 2,..., N} The transition probability between agent states is as follows: wherein where P(s t+1 |s t ,a t ) is the transition probability of the agent reaching state s t from state s t after taking action a t+1 ; is a fixed vector, m st is the number of non-zero elements in vector s t . 7.The method of claim 5, wherein, In step 3.3, for the computation of the reward function R(s t+1 |s t ,a t ), the following steps are included: Step 3.31 : Determine the generated optimized bus route L based on the current state s t determine the generated optimized bus route L t ; Step 3.32: L t G is added r , forming a computing network G t ; Step 3.33: Passenger flow assignment on network G using KSP algorithm and Logit model r is performed. Step 3.34: Calculate and record Z(G t ); Step 3.35: judge whether the iteration is ended: yes, go to step 3.37, otherwise go to step 3.36; Step 3.36: If Return reward R(s t+1 |s t ,a t ) = -1; else update s t ,L t ,G t , return step 3; Step 3.37: Calculate the reward R(s t+1 |s t ,a t ) and return; where R(s t+1 |s t ,a t ) is given by: E f = E t ∩(E v ) C where R(s t+1 |s t ,a t ) is the immediate reward obtained by the agent when taking action a t from state s t to reach state s t+1 ; L t+1 and L t are the generated bus routes corresponding to state s t+1 and state s t , respectively; E t is the set of nodes directly connected to the last node n t of route L t ; E v is the set of nodes already visited by route L t ; (E v ) C is the complement of E v . 8.The method of claim 6, wherein, Step 3.4, DQN training includes the following steps: Step 3.41 Input road network data E, demand matrix D, preserved bus network G r ; Step 3.42 set basic training parameters; Step 3.43 define Q network Q(S; Θ), target network T(S; Θ) and experience replay buffer; Step 3.44 initialize current node n t , current state s t , L t ; Step 3.45 Select action a t = argmax Q(s t ; Θ), compute reward R(s t+1 | s t , a t ); Step 3.46 Store (n t , L t , s t , a t , R(s t+1 |s t , a t ), n t+1 , L t+1 , s t+1 ) in the experience replay buffer; Step 3.47 judge whether the experience replay buffer data is sufficient; yes: go to step 3.48; no: go to step 3.410; Step 3.48 extract batch data from experience replay buffer for training and update Q(S; Θ) network parameters; Step 3.49 update n t = n t+1 = L t = L t+1 = s t = s t+1 ; Step 3.410 According to L t The length of the route determines whether it has been generated; if yes: return to step 3.44; if no: return to step 3.

45. Step 3.411 in each iteration, update the parameters of target network T(S; Θ) every certain number of training rounds; Step 3.412 when the number of iterations exceeds the number of training rounds, terminate training.