Bus dynamic scheduling method and system based on vehicle management

By building a multi-factor scheduling environment simulation system and a deep reinforcement learning network, a dynamic departure schedule is generated, which solves the problems of real-time passenger flow changes and charging restrictions in bus scheduling and realizes efficient and low-cost bus operations.

CN120673615AActive Publication Date: 2025-09-19HANGZHOU DIANZI UNIV +1

Patent Information

Application Number
CN202511173461.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-19
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Existing bus scheduling algorithms are unable to cope with real-time passenger flow changes, resulting in extended waiting times and delays for passengers. Existing reinforcement learning-based methods fail to effectively consider the charging limitations of electric buses, resulting in unstable scheduling strategies and increased costs.

Method used

A multi-factor scheduling environment simulation system was constructed, defining the state space, action space, and reward function. Combined with a deep reinforcement learning network, a dynamic bus scheduling model was trained using Rainbow DQN to generate a dynamic departure schedule, taking into account vehicle management constraints and passenger demand, and optimizing the departure strategy.

Benefits of technology

It has achieved the goal of reducing departure frequency, lowering operating costs, avoiding wrong departures, and improving bus operation efficiency and passenger satisfaction while meeting vehicle management and passenger needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673615A_ABST
    Figure CN120673615A_ABST
Patent Text Reader

Abstract

The invention discloses a bus dynamic scheduling method and system based on vehicle management, and the method comprises the steps: building a scheduling environment simulation system through collecting the historical data and real-time data of the operation of an electric bus, automatically generating passenger travel information based on bus line data and historical passenger getting-on data, carrying out the modeling of the bus operation as a dynamic discrete process, and carrying out the real-time scheduling of the electric bus. Constraints such as vehicle availability and charging limitation are introduced, and a bus operation simulation environment close to reality is constructed; establishing a Markov decision process based on a scheduling simulation environment, and definitely defining vehicle availability management parameters in a state space and a reward function; a Rainbow DQN model is adopted for training, a scheduling strategy is continuously optimized according to environment feedback, a dynamic scheduling scheme is generated by utilizing the model obtained through training according to real-time data reasoning, the travel requirements of passengers and the operation benefits of a bus company are considered on the premise that the vehicle availability is met, and intelligent scheduling decision support is provided for management personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bus scheduling and reinforcement learning, and in particular relates to a bus dynamic scheduling method and system based on vehicle management. Background Art

[0002] Bus scheduling algorithms based on exhaustive methods and genetic algorithms can only generate bus schedules offline based on historical passenger flows, making it difficult to cope with the ever-changing real-time passenger flows during actual operations. In real-world scenarios, passenger flows often experience sudden fluctuations, and offline bus schedules are unable to respond promptly to changes in demand, resulting in extended passenger wait times and even passenger delays. Although reinforcement learning-based scheduling algorithms, such as the DRL-TO method and the dynamic schedule generation method based on DQN, can optimize departure frequency based on real-time passenger flows, they assume that vehicles are always available and ignore constraints such as charging restrictions for electric buses, making them unsuitable for real-world operation. Furthermore, these methods all use DQN for policy learning, which suffers from problems such as overestimation of Q values, insufficient exploration capabilities, and low sample utilization efficiency, making it difficult to obtain stable and efficient scheduling strategies in complex environments. Summary of the Invention

[0003] In order to address the deficiencies of the existing technology and achieve the goal of improving public transportation operation efficiency and reducing public transportation operation costs, the present invention adopts the following technical solutions:

[0004] A bus dynamic scheduling method based on vehicle management includes the following steps:

[0005] Step 1: Obtain vehicle historical operation data and real-time operation status;

[0006] Step 2: Build a multi-factor scheduling environment simulation system that takes into account vehicles, stations, and passengers;

[0007] Step 3: Model the scheduling decision process, define the state space and action space that take into account vehicle constraints, and the reward function that takes into account the interests of public transportation parties;

[0008] Step 4: Construct a bus dynamic scheduling model. The scheduling decision is triggered based on the scheduling environment simulation system. The value distribution of each action is estimated by using the noisy network structure based on the current state of the trigger. The optimal action is selected in combination with the minimum and maximum departure interval constraints. The selected action is applied to the scheduling environment simulation environment to obtain the immediate reward brought by the next state and the current action. The bus dynamic scheduling model is constructed and trained based on the priority experience replay pool constructed based on the current state, action, reward, and next state to generate dynamic departure times.

[0009] Furthermore, the step 3 includes the following steps:

[0010] Step 3.1: Model the dynamic vehicle scheduling problem as a Markov scheduling decision process, and define the state space S, action space A, and reward function R considering vehicle management constraints;

[0011] The state space S contains the different states obtained by the agent from the environment at time m. The characteristics of the state include time information of different scales, demand pressure coefficient, the sum of the waiting time of passengers boarding all vehicles currently running on the route, the carrying capacity utilization rate of all vehicles currently running on the route, and the ratio of the current number of idle buses to the total number of buses.

[0012] The action space contains the action of starting or not starting;

[0013] Step 3.2: Define a reward function that takes into account the interests of both passengers and bus companies, and ensures that bus scheduling does not conflict with bus management; passenger interests are measured by bus congestion and waiting time, while bus company interests are measured by the number of buses dispatched. The reward function Indicates that at the mth moment based on the state Execute an action The reward value obtained provides immediate performance feedback for the deep reinforcement learning model.

[0014] Furthermore, in step 3.1, the demand pressure coefficient is the sum of the number of people waiting at each station at a certain moment, divided by the sum of the load capacities of all vehicles running on the line at that moment, which is obtained based on the number of seats and capacity coefficient of the vehicle; the demand pressure coefficient Defined as:

[0015]

[0016] Among them, SN represents the total number of line stations, represents the number of people waiting at the sth station at time m, Indicates the total number of vehicles currently running on the line, represents the load capacity of the k-th vehicle, represents the number of seats in the k-th car, is the capacity factor, which indicates that the load capacity of a vehicle is a function of its number of seats. times, The smaller the value, the more sufficient the overall capacity of the system is;

[0017] The sum of the waiting time for passengers to board the bus for all vehicles currently operating on the route is calculated by subtracting the time it takes for passengers to arrive at the bus stop from the time it takes for passengers to board the bus at the bus stop, and is calculated based on the total number of vehicles currently operating on the route, the number of bus stops the vehicles have passed, and the number of passengers who boarded the vehicles at the passenger boarding stops.

[0018] The normalized sum of the waiting time for passengers to board all vehicles currently running on the route is defined as:

[0019]

[0020] in, It represents the total waiting time of passengers boarding all vehicles currently running on the line. It represents the parameter that normalizes the total waiting time. The calculation formula is:

[0021]

[0022] in, Indicates the total number of vehicles currently running on the line, represents the number of bus stops that the k-th vehicle has passed, represents the number of passengers boarding the k-th bus at the s-th bus stop, represents the time it takes for the i-th passenger to get on the k-th bus at the s-th bus stop, represents the time when the i-th passenger arrives at the s-th bus stop;

[0023] The carrying capacity utilization rate of all vehicles currently operating on the route is obtained by dividing the carrying capacity actually required by all vehicles on the route by the carrying capacity available to all vehicles;

[0024] The capacity utilization rate of all vehicles currently operating on the route , is an important indicator to measure the degree of matching between actual passenger demand and bus carrying capacity. It can help determine whether it is necessary to dispatch a bus at the current decision time point. It is defined as:

[0025]

[0026] in, Indicates the actual carrying capacity required by all vehicles on the line. 、 and The definition of is the same as before, represents the number of bus stops actually taken by the i-th passenger who gets on the k-th bus at the s-th bus stop, Indicates the carrying capacity that all vehicles can provide. 、 、 Same as above;

[0027] The ratio of the current number of idle buses to the total number of buses is obtained by dividing the number of idle buses at the starting station by the total number of buses on the current route;

[0028] The ratio of the current number of idle buses to the total number of buses , defined as:

[0029]

[0030] in, Indicates the number of idle vehicles on the current route. Indicates the total number of vehicles on the current route, satisfying , Indicates the total number of vehicles currently running on the line, and the number of idle vehicles at the starting station When it reaches zero, the train cannot be started.

[0031] Furthermore, the reward function in step 3.2 includes:

[0032] When passenger flow is low, the agent should prioritize not dispatching a vehicle. Based on the capacity utilization rate, the agent is rewarded for not dispatching a vehicle. The lower the capacity utilization rate, the lower the passenger demand in the current time period. At the same time, a waiting time penalty is introduced.

[0033] When passenger flow is high, the agent should prioritize dispatching. Based on the high passenger demand during the current period, the agent is rewarded for dispatching. However, if the agent still issues a dispatch command when there are no idle vehicles at the starting station, it will be penalized for incorrect dispatch.

[0034] When passenger detention occurs, the agent is penalized for passenger detention based on the number of stranded passengers for both non-departure and departure actions.

[0035] Specifically, the reward function is defined as follows:

[0036]

[0037] When the passenger flow is small, the agent should give priority to not sending the vehicle. , in order to reduce the operating costs of the bus company, therefore, the agent is rewarded for not sending the bus. ,in Indicates the utilization rate of carrying capacity. The smaller the value, the lower the passenger demand in the current time period. However, not sending a bus may increase the waiting time of passengers, so a penalty term is introduced. , is the penalty factor for waiting time;

[0038] When the passenger flow is large, the agent should give priority to the departure action , give the agent a reward for the start action , indicating that the passenger demand is high during the current period, and it is encouraged to depart at the mth minute; in order to meet the bus company's requirements for vehicle management, it is necessary to impose penalties on the agent's wrong departure. When there are no idle vehicles at the starting station, If the agent still issues a start command, it will be punishment, including represents the penalty factor for wrong start;

[0039] In addition, once the passenger is stranded, it will seriously affect the passenger's ride satisfaction. Therefore, regardless of whether the bus is dispatched or not, as long as the passenger is stranded, the agent will be penalized for both the non-dispatch action and the dispatch action. , represents the penalty factor for the number of stranded passengers, It represents the total number of stranded passengers on the current line at time m, and the calculation formula is as follows:

[0040]

[0041] Among them, SN represents the total number of line stations, represents the number of stranded passengers at the s-th station at the m-th time.

[0042] Furthermore, step 4 includes the following steps:

[0043] Step 4.1: Use deep reinforcement learning network to train the bus dynamic scheduling model;

[0044] First, a bus is forced to be dispatched at the first bus time and the environment state is updated. Subsequently, the dispatch environment simulation system triggers a dispatch decision at regular intervals. At each decision moment, the vehicle state is input. Based on the vehicle state, the deep reinforcement learning network uses a noisy network structure to estimate the value distribution of the action of dispatching or not dispatching the bus, and selects the optimal action in combination with the minimum and maximum departure interval constraints. The selected action is applied to the simulation environment to obtain the new vehicle state and the immediate reward brought by the current action. The generated interaction data including state, action, reward, and next state is stored in the priority experience replay pool. After the priority experience replay pool is filled, the deep reinforcement learning network begins training, regularly sampling batch samples from the pool for gradient update to optimize network parameters. After a fixed number of steps, the target network is synchronously updated and the model state is saved, ultimately obtaining a bus dynamic dispatch model.

[0045] Step 4.2: Generate a dynamic departure schedule based on the bus dynamic scheduling model reasoning;

[0046] Based on real-time arrival, departure, and boarding data, the trained bus dynamic scheduling model is used for inference at each decision point. The final action is selected according to the probability distribution of the action in the current state, thereby dynamically generating a departure schedule.

[0047] Furthermore, the step 1 includes the following steps:

[0048] Step 1.1: Obtain vehicle historical operation data and perform preprocessing;

[0049] Collect historical vehicle operation data sets, including route data tables, arrival and departure data tables, and passenger boarding data tables; the route data table records the basic information of all vehicle routes, including route name, route direction, the name, coordinates, and single program number of each station in the current direction. The single program number represents the station number of the station in the current route direction; the arrival and departure data table records the arrival and departure information of the bus during operation, including route name, vehicle ID, route direction, single program number of the station where the vehicle is located, vehicle arrival / departure status, recording time, and remaining battery power; the passenger boarding data table records the boarding information of all passengers, including route name, route direction, vehicle ID, boarding station name, and boarding transaction time;

[0050] Preprocess historical operation data, including data cleaning and deletion of invalid data;

[0051] For each bus route, calculate the average travel time between stops based on the arrival and departure times of the bus; count the frequency of passenger boardings at different times for each route based on the boarding information table; and extract the original bus schedule as a benchmark for subsequent optimization.

[0052] Step 1.2: Get the real-time running status of the vehicle;

[0053] During vehicle operation, real-time arrival and departure data of the vehicle are obtained based on the on-board GPS positioning device and battery management system, and real-time arrival and departure data and real-time passenger boarding data are obtained based on the IC card swiping device and mobile payment terminal. The field settings of these two types of real-time data are consistent with historical data.

[0054] Furthermore, in step 2, the scheduling environment simulation system simulates the operation of the vehicle line based on vehicle management, station management and passenger management, sets the minimum and maximum departure interval scheduling parameters to divide the time period, and sets the decision time point based on each time period, outputs the current simulation status at each decision time point, and receives the scheduling decision instruction of the intelligent agent. After the instruction is executed, it continues to run until the next decision time point, and repeats the cycle until the end of the simulation.

[0055] Furthermore, in terms of vehicle management, the dispatch environment simulation system meticulously models the dispatch status of vehicles. Each vehicle includes status parameters such as the current route direction and station, number of passengers on board, number of available seats, and remaining battery power, which are updated in real time. Vehicles depart from the starting station according to the timetable sequence, stop at each station along the route to complete passenger boarding and disembarking, and finally arrive at the terminal station for the current direction. After a brief stop, the vehicle switches direction according to the dispatch schedule, completes the opposite direction operation task, and returns to the starting station. If the remaining battery power can support the next trip, the vehicle will directly join the idle queue and await subsequent dispatch. Otherwise, it will first join the charging queue and enter the idle queue after completing charging. Parameters such as the number of vehicles, passenger capacity, and battery capacity are set according to the actual vehicle model configuration.

[0056] In terms of station management, each route includes a starting station, a terminal station, and several intermediate stations; the starting station is responsible for managing the idle queue and charging queue of vehicles on this route; each station is responsible for managing the waiting passengers at the station and recording the number of stranded passengers caused by the previous vehicle being full;

[0057] In terms of passenger management, the dispatching environment simulation system dynamically generates passengers based on real passenger flow data patterns; each station continuously generates new passengers during operation, and the number and frequency of new passengers are set according to the characteristics of the route and time period; when the bus arrives at a station, passengers board the bus in order. If the vehicle capacity is full, the remaining passengers will stay and wait for the next vehicle; the dispatching environment simulation system records the number of waiting passengers and the number of stranded passengers at each station in real time.

[0058] Furthermore, in step 2, in order to enhance the authenticity of the simulation, the scheduling environment simulation system introduces random disturbances in the travel time between adjacent stations and simulates various special waiting situations such as bad weather, waiting at traffic lights and congestion.

[0059] A bus dynamic dispatching system based on vehicle management includes a vehicle data and status acquisition module, a dispatching environment simulation module, a dispatching decision module and a bus dynamic dispatching module. According to the bus dynamic dispatching method based on vehicle management, the acquisition of vehicle historical operation data and real-time operation status, the construction of the dispatching environment simulation system, the modeling of the dispatching decision process and the construction of the bus dynamic dispatching model are sequentially performed.

[0060] The advantages and beneficial effects of the present invention are:

[0061] The present invention acquires the historical operating data and real-time operating status of bus routes and constructs a scheduling environment simulation system that takes into account multiple factors such as vehicle management, station management, and passenger management. This system can not only model the historical operating characteristics of the routes and passenger travel needs, but also consider random factors such as weather changes, traffic congestion, and charging restrictions to achieve a true simulation of the operating status of bus routes. On this basis, the dynamic scheduling problem is defined as a Markov decision process that takes into account vehicle management constraints, and the vehicle availability management parameters are clearly defined in the state space and reward function. Finally, the Rainbow DQN is used to train the bus dynamic scheduling model, and a dynamic departure schedule is generated based on the training scheduling model. This can effectively reduce the departure frequency while meeting the vehicle management constraints and passenger travel needs, thereby reducing the operating costs of bus companies. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 is a flow chart of a method in an embodiment of the present invention.

[0063] Figure 2 This is a schematic diagram of the operation of the scheduling environment simulation system in an embodiment of the present invention.

[0064] Figure 3 3. This is a comparison chart of the carrying capacity of the bus schedule for route 7 generated by different methods in the embodiments of the present invention.

[0065] Figure 4 Schematic diagram of the structure of the system in the embodiment of the present invention. DETAILED DESCRIPTION

[0066] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.

[0067] like Figure 1 As shown, a bus dynamic scheduling method based on vehicle management includes the following steps:

[0068] Step 1: Obtain historical operation data and real-time operation status of electric buses.

[0069] Step 1.1: Obtain historical operation data of electric buses and perform preprocessing.

[0070] The historical operation data set of electric buses is collected, including route data tables, arrival and departure data tables, and boarding data tables. The route data table records the basic information of all bus routes, including route name, route direction, the name, coordinates, and single program number of each station in the current direction. The single program number represents the station number of the station in the current route direction. For example, a single program number of 1 indicates that the station is the starting station for the current direction. The arrival and departure data table records the information of the bus arriving at and leaving the station during operation, including route name, vehicle ID, route direction, single program number of the station where the vehicle is located, vehicle arrival / departure status, recording time, and remaining battery power. The boarding data table records the boarding information of all passengers, including route name, route direction, vehicle ID, boarding station name, and boarding transaction time.

[0071] After acquiring historical data, the original data was first cleaned and invalid data was deleted. Then, for each bus route, the average travel time between stops was calculated based on the arrival and departure times of the bus. Based on the boarding information table, the frequency of passenger boardings for each route at different times was counted. The original bus schedule was then extracted as a reference for subsequent optimization.

[0072] Step 1.2: Get the real-time operating status of the electric bus.

[0073] During the operation of electric buses, real-time arrival and departure data are obtained based on the on-board GPS positioning device and battery management system, and real-time arrival and departure data and real-time passenger boarding data are obtained based on IC card swiping devices and mobile payment terminals. The field settings of these two types of real-time data are consistent with historical data.

[0074] Step 2: Construct a dispatching environment simulation system that takes into account multiple factors such as vehicle management, station management, and passenger management, such as Figure 2 shown.

[0075] The dispatch environment simulation system comprehensively considers vehicle management, station management, and passenger management to simulate the daily operation of bus routes. The system sets dispatch parameters such as minimum and maximum departure intervals, divides time into two-minute intervals, and sets the start of each time period as a decision time point. At each decision time point, the system outputs the current state of the simulated environment and receives dispatch decision instructions from the agent. After the instructions are executed, the system continues to operate until the next decision time point, repeating the cycle until the simulation ends.

[0076] In terms of vehicle management, the simulation system meticulously models the dispatch status of electric buses. Each vehicle includes status parameters such as the current route direction and station, the number of passengers on board, the number of available seats, and the remaining battery charge, which are updated in real time. The vehicle departs from the starting station according to the schedule, stops at each station along the route to handle passenger boarding and disembarking, and finally arrives at the terminal for the current direction. After a brief stop, the vehicle switches direction according to the dispatch schedule, completes the opposite direction operation, and returns to the starting station. If the remaining battery charge can support the next trip, the vehicle will directly join the idle queue and await further dispatch; otherwise, it must first join the charging queue and complete charging before joining the idle queue. Parameters such as the number of vehicles, passenger capacity, and battery capacity are set according to the actual vehicle model configuration.

[0077] In terms of station management, each route consists of a starting station, a terminal station, and several intermediate stations. The starting station is responsible for managing the idle and charging queues of vehicles on this route. Each station is responsible for managing waiting passengers at that station and recording the number of stranded passengers caused by the previous vehicle being full.

[0078] In terms of passenger management, the simulation system dynamically generates passengers based on real-world passenger flow data. Each stop continuously generates new passengers during operation, with the number and frequency of new passengers set based on route characteristics and time of day. When a bus arrives at a stop, passengers board in order. If a bus is full, remaining passengers are held back to wait for the next bus. The system records the number of waiting and held passengers at each stop in real time.

[0079] To enhance the realism of the simulation, the system also introduces random disturbances in the travel time between adjacent stations and simulates special situations such as bad weather, waiting at traffic lights and congestion.

[0080] Step 3: Model the Markov decision process and define the state space, action space, and reward function that considers the interests of passengers and bus companies while taking into account vehicle management constraints.

[0081] Step 3.1: Model the Markov decision process and define the state space and action space considering vehicle management constraints.

[0082] The dynamic scheduling problem of urban electric buses is modeled as a Markov decision process, and the corresponding state space S, action space A and reward function R are defined.

[0083] The state space S contains different states , Represents the state characteristics obtained by the agent from the environment at time m, which is specifically defined as follows:

[0084] (1)

[0085] The meaning of each parameter is described in detail below:

[0086] (1) and Represents the normalized time information, where , , Represents the hours in the current environment, Represents the minute in the current environment.

[0087] (2) represents the demand pressure coefficient, which is defined as:

[0088] (2)

[0089] Among them, SN represents the total number of line stations, Indicates the number of people waiting at the sth station at time m. Indicates the total number of vehicles currently operating on the route. represents the load capacity of the k-th vehicle, represents the number of seats in the k-th car, is the capacity factor, which indicates that the load capacity of a vehicle is a function of its number of seats. times. The smaller the value, the more sufficient the overall capacity of the system.

[0090] (3) It represents the normalized sum of the waiting time of passengers boarding all vehicles currently running on the route, and is defined as:

[0091] (3)

[0092] in, is the sum of the waiting time of passengers boarding all vehicles currently running on the line, and the calculation formula is:

[0093] (4)

[0094] in, Indicates the total number of vehicles currently running on the line, represents the number of bus stops that the k-th vehicle has passed, represents the number of passengers boarding the k-th bus at the s-th bus stop, represents the time it takes for the i-th passenger to get on the k-th bus at the s-th bus stop, represents the time when the i-th passenger arrives at the s-th bus stop. is a parameter that normalizes the total waiting time.

[0095] (4) It indicates the capacity utilization rate of all vehicles currently running on the route. It is an important indicator to measure the degree of matching between actual passenger demand and bus capacity, and can help determine whether it is necessary to dispatch a bus at the current decision time. It is defined as:

[0096] (5)

[0097] in, It reflects the actual carrying capacity required by all vehicles on the line. 、 and The definition of is the same as before, It represents the number of bus stops actually taken by the i-th passenger who gets on the k-th bus at the s-th bus stop. Reflects the carrying capacity available to all vehicles. 、 、 The definition of is the same as above.

[0098] (5) It represents the ratio of the current number of idle buses to the total number of buses, which is defined as:

[0099] (6)

[0100] in, Represents the number of idle vehicles on the current route, Represents the total number of vehicles on the current route, satisfying , The total number of vehicles currently running on the line. When it reaches zero, the train cannot be started.

[0101] The action space contains two kinds of actions: Indicates departure, action Indicates no departure.

[0102] Step 3.2: Define a reward function that takes into account the interests of both passengers and bus companies.

[0103] The reward function takes into account the interests of both passengers and bus companies, and ensures that bus scheduling does not conflict with bus fleet management. Passenger interests are measured by bus congestion and waiting time, while bus companies’ interests are measured by the number of buses dispatched. Indicates that at the mth moment based on the state Execute an action The reward value obtained provides immediate performance feedback for the deep reinforcement learning model and is defined as follows:

[0104] (7)

[0105] When the passenger flow is small, the agent should prioritize action 0 (not dispatching the bus) to reduce the bus company's operating costs. Therefore, the agent is rewarded for action 0. ,in Indicates the utilization rate of carrying capacity. The smaller the value, the lower the passenger demand in the current time period. However, not sending a bus may increase the waiting time of passengers, so a penalty term is introduced. , is the penalty factor for waiting time.

[0106] When the passenger flow is large, the agent should prioritize action 1 (departure), so the agent is given a reward for action 1. , indicating that the passenger demand is high during the current period, and the departure of buses at the mth minute should be encouraged. In order to meet the bus company's requirements for vehicle management, it is necessary to impose penalties on the agent's incorrect departure. When there are no idle vehicles at the starting station ( ), if the agent still issues a start command, it will be punishment, including is the penalty factor for wrong start.

[0107] In addition, once the passenger is stranded, it will seriously affect the passenger's satisfaction. Therefore, regardless of whether the bus is dispatched or not, as long as the passenger is stranded, the agent will be penalized for both actions 0 and 1. , is the penalty factor for the number of stranded passengers, It represents the total number of stranded passengers on the current line at time m, and the calculation formula is as follows:

[0108] (8)

[0109] Among them, SN represents the total number of line stations, represents the number of stranded passengers at the s-th station at the m-th time.

[0110] Step 4: Use Rainbow DQN training to obtain a bus dynamic scheduling model and generate a dynamic departure schedule through inference.

[0111] Step 4.1: Use Rainbow DQN training to obtain the bus dynamic scheduling model.

[0112] A dynamic bus scheduling model was constructed using the Rainbow DQN algorithm to learn an optimized dynamic dispatch policy. First, a bus was dispatched at the first bus time, and the environment state was updated. Subsequently, the dispatch environment simulation system triggered a dispatch decision every two minutes. At each decision moment, the model input was a 6-dimensional state vector containing vehicle management information. Based on this state, the Rainbow DQN algorithm used a noisy network structure to estimate the value distribution of each action and selected the optimal action, taking into account minimum and maximum departure interval constraints. The selected action was applied to the simulation environment, resulting in a new state vector and the immediate reward associated with the current action. The interaction data generated by this interactive process (state, action, reward, next state) was stored in a prioritized experience replay pool. Once the prioritized experience replay pool was filled, the Rainbow DQN began training, periodically sampling batches of samples from the pool for gradient updates to optimize the Q network parameters. After a fixed number of steps, the target network was synchronously updated and the model state was saved, ultimately resulting in a dynamic bus scheduling model.

[0113] Step 4.2: Generate a dynamic departure schedule based on the bus scheduling model reasoning.

[0114] Based on real-time arrival and departure data and passenger boarding data, the trained bus scheduling model is used for reasoning at each decision time point, and the final action is selected according to the probability distribution of the action under the current state, thereby dynamically generating a departure schedule.

[0115] In the embodiment of the present invention, as shown in Table 1, the performance of the original bus schedule, the genetic algorithm, the DRL-TO and the bus schedule generated by the method of the present invention are compared.

[0116] Table 1 Evaluation results of different methods

[0117]

[0118] For three bus routes in Hangzhou (Routes 7, 10, and 156), two deep reinforcement learning-based methods (DRL-TO and the proposed method) demonstrated significant advantages over the original bus schedule and the genetic algorithm. While the original bus schedule ensured no incorrect departures, it had a high number of departures and a high number of stranded passengers. Compared to the original schedule, both the genetic algorithm and DRL-TO methods improved on the number of departures, average waiting time, and number of stranded passengers, with the DRL-TO method outperforming the genetic algorithm. However, both the genetic algorithm and DRL-TO generated schedules exhibited incorrect departures. Furthermore, while the genetic algorithm achieved the lowest number of departures for Route 156 in the down direction, the average waiting time and number of stranded passengers were the highest among all methods. In contrast, the proposed method achieved a more balanced performance across all metrics, maintaining low numbers of departures, average waiting time, and number of stranded passengers, while completely avoiding incorrect departures, demonstrating excellent overall scheduling performance.

[0119] like Figure 3 As shown in the figure, the comparison between the carrying capacity provided by the No. 7 bus schedules generated by different methods and the actual passenger demand. Each mark in the figure represents the line carrying capacity within 30 minutes from that time point. As can be seen from the figure, the original bus schedule provided a capacity that far exceeded the actual demand between 9:30 and 13:00, resulting in a large amount of resource waste, and a short-term capacity shortage occurred during the evening peak. The genetic algorithm is not sensitive to changes in passenger demand, and the capacity adjustment lags behind. The DRL-TO method provides redundant capacity during the morning peak period. Although it guarantees the service level, it increases operating costs. In contrast, the method of the present invention better matches the carrying capacity with the actual passenger demand, can dynamically adjust the departure according to real-time passenger demand, and maintain a certain capacity redundancy to deal with emergencies.

[0120] like Figure 4 As shown, a bus dynamic scheduling system based on vehicle management includes a vehicle data and status acquisition module, a scheduling environment simulation module, a scheduling decision module and a bus dynamic scheduling module. According to the bus dynamic scheduling method based on vehicle management, the acquisition of vehicle historical operation data and real-time operation status, the construction of the scheduling environment simulation system, the modeling of the scheduling decision process and the construction of the bus dynamic scheduling model are performed in sequence.

[0121] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A bus dynamic dispatching method based on vehicle management, characterized in that The steps include: Step 1: Obtain vehicle historical operation data and real-time operation status; Step 2: Build a multi-factor scheduling environment simulation system that takes into account vehicles, stations, and passengers; Step 3: Model the scheduling decision process, define the state space and action space that take into account vehicle constraints, and the reward function that takes into account the interests of public transportation parties; Step 4: Construct a bus dynamic scheduling model, trigger the scheduling decision based on the scheduling environment simulation system, and select the optimal action by estimating the value distribution of each action based on the current state of the trigger, and combining the minimum and maximum departure interval constraints. The selected action is applied to the scheduling environment simulation environment to obtain the immediate reward brought by the next state and the current action. The bus dynamic scheduling model is constructed and trained based on the priority experience replay pool constructed based on the current state, action, reward, and next state to generate dynamic departure times.

2. A bus dynamic dispatching method based on vehicle management according to claim 1, characterized in that: The step 3 comprises the following steps: Step 3.1: Model the dynamic vehicle scheduling problem as a Markov scheduling decision process, and define the state space S, action space A, and reward function R considering vehicle management constraints; The state space S contains the different states obtained by the agent from the environment at time m. The characteristics of the state include time information of different scales, demand pressure coefficient, the sum of the waiting time of passengers boarding all vehicles currently running on the route, the carrying capacity utilization rate of all vehicles currently running on the route, and the ratio of the current number of idle buses to the total number of buses. The action space contains the action of starting or not starting; Step 3.2: Define a reward function that balances the interests of passengers and bus companies, ensuring that bus scheduling does not conflict with bus fleet management. Passenger interests are measured by bus congestion and wait times, while bus company interests are measured by the number of bus departures. The reward function represents the reward value obtained by performing an action based on the state at a specific moment.

3. A bus dynamic dispatching method based on vehicle management according to claim 2, characterized in that: In step 3.1, the demand pressure coefficient is the sum of the number of people waiting at each station at a certain moment, divided by the sum of the load capacities of all vehicles operating on the route at that moment, where the load capacity is based on the number of seats and the capacity coefficient of the vehicle; The sum of the waiting time for passengers to board the bus for all vehicles currently operating on the route is calculated by subtracting the time it takes for passengers to arrive at the bus stop from the time it takes for passengers to board the bus at the bus stop, and is calculated based on the total number of vehicles currently operating on the route, the number of bus stops the vehicles have passed, and the number of passengers who boarded the vehicles at the passenger boarding stops. The carrying capacity utilization rate of all vehicles currently operating on the route is obtained by dividing the carrying capacity actually required by all vehicles on the route by the carrying capacity available to all vehicles; The ratio of the current number of idle buses to the total number of buses is obtained by dividing the number of idle buses at the starting station by the total number of buses on the current route.

4. A bus dynamic dispatching method based on vehicle management according to claim 3, characterized in that: The reward function in step 3.2 is as follows: When passenger flow is low, the agent should prioritize not dispatching a vehicle. Based on the capacity utilization rate, the agent is rewarded for not dispatching a vehicle. The lower the capacity utilization rate, the lower the passenger demand in the current time period. At the same time, a waiting time penalty is introduced. When passenger flow is high, the agent should prioritize dispatching. Based on the high passenger demand during the current period, the agent is rewarded for dispatching. However, if the agent still issues a dispatch command when there are no idle vehicles at the starting station, it will be penalized for incorrect dispatch. When passenger detention occurs, the agent is penalized for passenger detention based on the number of stranded passengers for both non-departure and departure actions.

5. The bus dynamic scheduling method based on vehicle management according to claim 1, characterized in that: The step 4 comprises the following steps: Step 4.1: Use deep reinforcement learning network to train the bus dynamic scheduling model; First, a bus is forced to be dispatched at the first bus time and the environment state is updated. Subsequently, the dispatch environment simulation system triggers a dispatch decision at regular intervals. At each decision moment, the vehicle state is input. Based on the vehicle state, the deep reinforcement learning network uses a noisy network structure to estimate the value distribution of the action of dispatching or not dispatching the bus, and selects the optimal action in combination with the minimum and maximum departure interval constraints. The selected action is applied to the simulation environment to obtain the new vehicle state and the immediate reward brought by the current action. The generated interaction data including state, action, reward, and next state is stored in the priority experience replay pool. After the priority experience replay pool is filled, the deep reinforcement learning network begins training, regularly sampling batch samples from the pool for gradient update to optimize network parameters. After a fixed number of steps, the target network is synchronously updated and the model state is saved, ultimately obtaining a bus dynamic dispatch model. Step 4.2: Generate a dynamic departure schedule based on the bus dynamic scheduling model reasoning; Based on real-time arrival, departure, and boarding data, the trained bus dynamic scheduling model is used for inference at each decision point. The final action is selected according to the probability distribution of the action in the current state, thereby dynamically generating a departure schedule.

6. The method for dynamic public transportation scheduling based on vehicle management according to claim 1, characterized in that: The step 1 comprises the following steps: Step 1.1: Obtain vehicle historical operation data and perform preprocessing; Collect historical vehicle operation data sets, including route data tables, arrival and departure data tables, and passenger boarding data tables; The route data table records the basic information of all vehicle routes; The arrival and departure data table records the information of the bus arriving at and leaving the station during operation; The boarding data table records the boarding information of all passengers; For each vehicle route, the average travel time between stations is calculated based on the vehicle arrival and departure times; According to the passenger boarding information table, count the frequency of passengers boarding each line at different times; Extract the original vehicle schedule as a benchmark for subsequent optimization; Step 1.2: Get the real-time running status of the vehicle; During the operation of the vehicle, real-time arrival and departure data and real-time passenger boarding data are obtained.

7. The bus dynamic dispatching method based on vehicle management according to claim 1, characterized in that: In step 2, the scheduling environment simulation system simulates the operation of the vehicle line based on vehicle management, station management and passenger management, sets interval scheduling parameters to divide time periods, and sets decision time points based on each time period. At each decision time point, the current simulation status is output, and the scheduling decision instructions of the intelligent agent are received. After the instructions are executed, the system continues to run until the next decision time point, and repeats the cycle until the end of the simulation.

8. The method for dynamic public transportation scheduling based on vehicle management according to claim 7, characterized in that: In terms of vehicle management, the dispatch environment simulation system meticulously models the dispatch status of vehicles. Each vehicle includes status parameters such as the current route direction and station, the number of passengers on board, the number of available seats, and the remaining battery power, which are updated in real time. Vehicles depart from the starting station according to the schedule, stop at each station along the route to complete passenger boarding and disembarking, and finally arrive at the terminal station for the current direction. After a brief stop, the vehicle switches direction according to the dispatch schedule, completes the opposite direction operation task, and returns to the starting station. If the remaining power is sufficient for the next trip, the vehicle will be directly added to the idle queue and wait for subsequent dispatch. Otherwise, it will first join the charging queue and then enter the idle queue after charging is completed. Vehicle parameters are set according to the actual vehicle model configuration. In terms of station management, each route includes a starting station, a terminal station, and several intermediate stations; the starting station is responsible for managing the idle queue and charging queue of vehicles on this route; each station is responsible for managing the waiting passengers at the station and recording the number of stranded passengers caused by the previous vehicle being full; In terms of passenger management, the dispatching environment simulation system dynamically generates passengers based on real passenger flow data patterns; each station continuously generates new passengers during operation, and the number and frequency of new passengers are set according to the characteristics of the route and time period; when the bus arrives at a station, passengers board the bus in order. If the vehicle capacity is full, the remaining passengers will stay and wait for the next vehicle; the dispatching environment simulation system records the number of waiting passengers and the number of stranded passengers at each station in real time.

9. The method for dynamic public transportation scheduling based on vehicle management according to claim 1, characterized in that: In step 2, the scheduling environment simulation system introduces random disturbances in the travel time between adjacent stations and simulates various special waiting situations.

10. A bus dynamic dispatching system based on vehicle management, comprising a vehicle data and status acquisition module, a dispatching environment simulation module, a dispatching decision module, and a bus dynamic dispatching module, characterized in that: According to a bus dynamic scheduling method based on vehicle management according to any one of claims 1 to 9, the acquisition of vehicle historical operation data and real-time operation status, the construction of a scheduling environment simulation system, the modeling of the scheduling decision process and the construction of a bus dynamic scheduling model are sequentially performed.

Citation Information

Patent Citations

  • Multi-vehicle-type customized bus area scheduling method for real-time demand

    CN113077162A

  • Bus departure timetable dynamic optimization algorithm based on deep reinforcement learning

    CN114240002A

  • Simulation system and method applied to intelligent bus scheduling

    CN115860594A

  • Electric bus dynamic scheduling system and method based on deep reinforcement learning

    CN116895144A

  • DQN-based bus uplink and downlink dynamic equilibrium timetable generation method

    CN118940733A

Cited By

  • Confluence area cooperative control method based on Rainbow DQN and driving risk field

    CN120748204A

  • A Collaborative Control Method for Merging Zones Based on RainbowDQN and Driving Risk Field

    CN120748204B