Vehicle relocation determination for vehicle pools
Through the combination of TM-CNN and LP, the unstable problem of vehicle repositioning strategies in the vehicle pool is solved, more efficient order completion rate and resource utilization are achieved, and the dynamic adaptability and driver income of the system are improved.
Patent Information
- Application Number
- CN202380090940.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-11
- Filing Date
- 2023-11-11
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art is difficult to effectively deal with the problem of space-time mismatch between vehicle supply and demand in real time in the vehicle pool, and the existing methods fail to effectively consider system dynamics and driver controllability, resulting in the vehicle repositioning strategy being not robust enough.
The time memory convolutional neural network (TM-CNN) combined with linear programming (LP) is used to train the value function and estimate the passenger arrival data to generate a vehicle repositioning strategy, considering the nonstationarity of the system and the driver controllable scores, and optimize the vehicle repositioning decision.
It improves the robustness of the vehicle repositioning strategy and order completion rate, optimizes the utilization of vehicle resources, and improves the dynamic adaptability of the system and driver's revenue.
Smart Images

Figure CN120500698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to repositioning or at least determining repositioning information for a pool of vehicles, such as a pool of passenger vehicles. Background Art
[0002] Vehicle relocation information for a vehicle pool (such as a ride-hailing or shared vehicle pool with multiple vehicles) is used to indicate where a vehicle should be relocated during a period when the vehicle is not in use. For example, a ride-hailing service may be provided by a ride-hailing system that includes a ride-hailing vehicle pool with multiple passenger vehicles, each of which can be used to provide a ride to a customer in order to transport the customer from a starting location to an ending location. This type of ride-hailing service is provided by Didi. TM Lyft TM and Uber TM supply.
[0003] The supply and demand for available vehicles and customers fluctuate over time, so there are often instances where a vehicle is "empty," meaning it has no customers. These empty vehicles are considered to be in an idle state while waiting for customers. Conventional systems attempt to determine a repositioning strategy by determining repositioning data indicating where a vehicle should be moved during this idle state when it is empty. Summary of the Invention
[0004] According to one aspect of the present invention, a method for determining vehicle repositioning data for a vehicle pool is provided. The method includes: training a value function using historical ride-sharing data; obtaining estimated passenger arrival data, wherein the estimated passenger arrival data is obtained by generating the estimated passenger arrival data using a temporal memory convolutional neural network (TM-CNN); determining a vehicle repositioning policy based on the trained value function and the estimated passenger arrival data; and determining the vehicle repositioning data for the vehicle pool based on the vehicle repositioning policy.
[0005] According to various embodiments, the method may further include any one of the following features or any technically feasible combination of some or all of these features:
[0006] - the historical ride-sharing data includes trajectory data from the pool of vehicles;
[0007] - the TM-CNN includes a temporal memory (TM) and a convolutional neural network, and wherein the TM includes at least one of a long short-term memory and a gated recurrent unit (GRU);
[0008] - the CNN comprises an encoding layer and a decoding layer, and wherein the TM is inserted in an embedding layer between the encoding layer and the decoding layer;
[0009] The input of TM-CNN includes two-dimensional (2D) passenger arrival data, which represents passenger arrival information for a location in a two-dimensional space and for a given time or time period;
[0010] - the vehicle repositioning strategy is periodically determined according to predetermined time intervals;
[0011] -Vehicle repositioning data for multiple vehicles in a vehicle pool;
[0012] -Determine vehicle repositioning strategies using an optimized look-ahead approach that considers estimated passenger arrival data;
[0013] - the optimized look-ahead method uses linear programming (LP); and / or
[0014] - determining a controllable fraction based on the historical ride-sharing data or other historical ride-sharing data, and wherein the controllable fraction is used to determine the vehicle repositioning strategy.
[0015] According to another aspect of the present invention, a vehicle repositioning system is provided. The vehicle repositioning system includes: at least one processor; and a memory storing computer instructions. The vehicle repositioning system is configured to execute the computer instructions using the at least one processor, such that when the computer instructions are executed by the at least one processor, the vehicle repositioning system: trains a value function using historical ride-sharing data; obtains estimated passenger arrival data, wherein the estimated passenger arrival data is obtained by generating the estimated passenger arrival data using a temporal memory convolutional neural network (TM-CNN); determines a vehicle repositioning strategy based on the trained value function and the estimated passenger arrival data; and determines vehicle repositioning data for a vehicle pool based on the vehicle repositioning strategy.
[0016] According to various embodiments, the vehicle repositioning system may further include any one of the following features or any technically feasible combination of some or all of these features:
[0017] - the historical ride-sharing data includes trajectory data from the pool of vehicles;
[0018] - the TM-CNN comprises a temporal memory (TM) and a convolutional neural network, and wherein the TM comprises at least one of a long short-term memory and a gated recurrent unit (GRU);
[0019] - the CNN comprises an encoding layer and a decoding layer, and wherein the TM is inserted in an embedding layer between the encoding layer and the decoding layer;
[0020] The input of TM-CNN includes two-dimensional (2D) passenger arrival data, which represents passenger arrival information for a location in a two-dimensional space and for a given time or time period;
[0021] - periodically determining a vehicle repositioning strategy according to predetermined time intervals;
[0022] -Vehicle repositioning data for multiple vehicles in a vehicle pool;
[0023] -Determine vehicle repositioning strategies using an optimized look-ahead approach that considers estimated passenger arrival data;
[0024] - the optimized look-ahead method uses linear programming (LP); and / or
[0025] - determining a controllability score based on the historical ride-sharing data or other historical ride-sharing data, and wherein the controllability score is used to determine the vehicle repositioning strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Preferred exemplary embodiments will be described below with reference to the accompanying drawings, wherein like reference numerals represent like elements, and wherein:
[0027] Figure 1 A communication or operational system comprising a vehicle relocation data system, a ride sharing data computer system, and a vehicle pool is depicted according to one embodiment; and
[0028] Figure 2 is a flow chart of a process for determining vehicle repositioning data according to one embodiment;
[0029] Figure 3 is a graphical representation of a spatial grid or map having a set of regions to which vehicle repositioning data belongs according to one embodiment;
[0030] Figure 4 is a diagrammatic representation of a long short-term memory convolutional neural network (LSTM-CNN) according to one embodiment;
[0031] Figure 5 is a flow chart illustrating a method of determining vehicle repositioning data for a vehicle pool according to one embodiment;
[0032] Figure 6 and Figure 7 Shown according to one embodiment without ( Figure 6 ) and have ( Figure 7) a spatial grid of vehicle repositioning data disposed thereon;
[0033] Figure 8 is a graph illustrating completion rates when varying a controllable score according to one embodiment, and illustrating that the disclosed algorithm outperforms other strategies at different controllable scores;
[0034] Figure 9 is a graph showing spatial variance with respect to completion rate according to one embodiment, which shows that as spatial heterogeneity increases, completion rate decreases at different rates for all algorithms;
[0035] Figure 10 is a graph showing the number of vehicles versus completion rate according to one embodiment, illustrating the impact of the total number of vehicles on performance;
[0036] Figures 11A to 11B is a graph showing the completion ratios ( Figure 11A ) and driving (fuel) costs ( Figure 11B ) and a graph of completion rates;
[0037] Figure 12 is a spatial map with a heat map (left) and grouped primary regions of cities (right), according to one embodiment; and
[0038] Figure 13 is a graph illustrating the performance of the disclosed method according to one embodiment relative to other algorithms. DETAILED DESCRIPTION
[0039] According to at least one embodiment, the systems and methods described herein enable determination of vehicle repositioning data based on a vehicle repositioning strategy determined by an optimization process that takes into account the controllability fraction and non-stationarity of the system. The vehicle repositioning data system includes at least one processor and a memory storing computer instructions that, when executed by the at least one processor, cause the system to perform a method for determining vehicle repositioning data (or vehicle repositioning strategy) for a vehicle pool, such as a vehicle pool having a plurality of vehicles participating in a ride-sharing environment (such as a vehicle pool provided by Didi). TM Lyft TM and Uber TMA pool of ride-sharing or ride-hailing vehicles (including those provided by the user) (collectively, the "ride-sharing vehicle pool"). Vehicle repositioning data for the vehicle pool is data indicating a repositioned position of at least one vehicle in the vehicle pool. In an embodiment, the vehicle repositioning data is determined based on a vehicle repositioning strategy, wherein the vehicle repositioning strategy is determined by formulating the ride-sharing system as a vehicle repositioning problem solved using optimization techniques. For example, in one embodiment, the vehicle repositioning problem is formulated as a multi-step look-ahead optimization problem that uses linear programming (LP) according to an embodiment to maximize the order completion rate of the ride-sharing system.
[0040] According to an embodiment, the method includes determining a controllability score for a vehicle pool, and determining a vehicle repositioning strategy for the vehicle pool based on the controllability score for the vehicle pool. As used herein, "controllability score" refers to the percentage or fraction of drivers who follow repositioning recommendations, and this value can be estimated or determined based on historical data. Depending on the implementation, this allows for accounting for the real-world observation that not all drivers will accept repositioning recommendations. Furthermore, the controllability score can be used as part of an optimization problem (such as a linear programming (LP) problem) for formulating the vehicle repositioning strategy.
[0041] In accordance with an embodiment, a vehicle repositioning strategy takes into account the non-stationarity of the vehicle, particularly by formulating the vehicle repositioning problem as a T-step look-ahead optimization problem that seeks to maximize the order completion rate. As used herein, "order completion rate" is a value indicating the number of orders (e.g., ride-hailing or ride-sharing orders) completed by a driver relative to the number or proportion of total orders. In an embodiment, a LP is used to determine the vehicle repositioning strategy, particularly an LP with a look-ahead time horizon T that explicitly models the non-stationarity of the ride-sharing system, such as the vehicle's position within a look-ahead or future time horizon T. In at least some embodiments, the vehicle repositioning strategy is defined based on an objective function or maximization function (collectively referred to as "objective function").
[0042] According to an embodiment, pre-training data is generated using a machine learning (ML) training process and is used to determine a vehicle relocation strategy. In at least some embodiments, the pre-training data is data representing a pre-trained value function that is incorporated into an objective function of the vehicle relocation strategy. According to an embodiment, incorporating the pre-trained value function into the objective function at each time not only optimizes the completion rate (or order completion rate) over T time slots, but also captures future rewards for the system. In an embodiment, the ML training process is a reinforcement learning (RL) process that uses RL to generate a pre-trained value function based on using historical data (such as trajectory data from a vehicle (such as, Global Navigation Satellite System (GNSS) data)) as input.
[0043] According to an embodiment, a neural network (NN) and in some embodiments a convolutional network (CNN) is used to generate estimated passenger arrival data, wherein historical passenger arrival data is used as input to the CNN. In an embodiment, the input includes passenger arrival data specified for two dimensions (e.g., using latitude and longitude) and a given time - this is referred to as two-dimensional (2D) passenger arrival data. In addition, in at least one embodiment, a temporal memory (TM) is used in conjunction with the CNN to provide a TM-CNN comprising a CNN and a TM, and the TM-CNN is configured to generate estimated passenger arrival data that captures the spatiotemporal correlations of the predicted passenger arrival rates. In addition, according to an implementation employing such an embodiment using the TM-CNN, such spatiotemporal prediction of passenger arrival rates works well with LP.
[0044] According to an embodiment, a splitting technique is used to enable consideration of cross-region vehicle occupant orders. The splitting technique can be used to take into account the fact that a driver may accept orders / passengers / customers from other adjacent regions.
[0045] Many existing methods for vehicle repositioning in large-scale ride-hailing platforms either do not consider the spatiotemporal mismatch between supply and demand in real time or do not consider the long-term equilibrium from a system perspective.
[0046] Ride-sharing services have proliferated rapidly over the past decade. Large-scale ride-hailing systems such as Uber TM , Didi TM and Lyft TM ) have fundamentally changed the way people live and move around in cities. The growing popularity of these ride-hailing platforms raises the inevitable question of how to provide reliable, trusted transportation while satisfying most, if not every, passenger request in a highly dynamic environment characterized by an imbalance between supply (drivers) and demand (passengers) in time and space.
[0047] One possible way to address this imbalance is to use a centralized planner that sends relocation suggestions to idle cars (without passengers) to relocate them to locations with future demand shortages at their intended destinations. This can significantly improve drivers’ earnings and passengers’ experience.
[0048] Existing work such as Braverman et al. (Braverman, A.; Dai, JG; Liu, X.; and Ying, L. 2019. Empty-car routing in ridesharing systems; Operations Research, 67(5):1437-1452) formulates the relocation task as a fluid-based optimization problem that can be solved using linear programming (LP) that can dynamically capture the fluctuations of complex systems. However, they assume that the system is already in a steady state and that the model parameters (including the passenger arrival rate at each location) are known, both of which are impossible in real systems. Another related family of forward-looking methods is model predictive control (MPC), which uses short-term demand forecasts based on historical data to obtain relocation strategies by solving a planning problem. However, both existing LP and MPC methods assume that all vehicles are fully controllable, which is obviously impossible in the real world.
[0049] In addition, due to the remarkable success of reinforcement learning (RL) in games and robotics systems, RL has also been applied to the problems of order dispatching and relocation in ride-hailing systems. Jiao et al. (Jiao, Y.; Tang, X.; Qin, ZT; Li, S.; Zhang, F.; Zhu, H.; and Ye, J. 2020. A Deep Value-based Policy Search Approach for Real-world Vehicle Repositioning on Mobility on-Demand Platforms. In NeurIPS Deep RL Workshop) use a spatiotemporal deep value network to learn the driver's perspective state value function and generate relocation actions through value-based policy search. Other works have proposed deep RL methods for vehicle relocation, such as deep Qnetwork (DQN) and neighboring policy optimization (PPO). All of the above methods do not explicitly consider system dynamics, and these methods may be problematic when the controllable fraction of the vehicle becomes larger. For example, crowding a large fleet into a single location with a high value can lead to undesirable imbalance.
[0050] According to an embodiment, in light of the foregoing discussion of existing approaches for vehicle repositioning, a system and method are provided for maximizing order fulfillment rates using a LP-based look-ahead repositioning algorithm combined with RL. The disclosed system and method can be used in real time to provide repositioning recommendations to one or more vehicles in a vehicle pool, such as a complex ride-hailing system.
[0051] refer to Figure 1 , shows a communication or operation system 10 that includes a vehicle relocation data system 12, a ride-sharing data computer system 14, a vehicle pool 16 having a plurality of vehicles, and an interconnected electronic data network 18 for enabling computer systems to communicate with each other, such as for communication between the vehicle relocation data system 12, the ride-sharing data computer system 14, and the vehicle pool 16. Each vehicle of the vehicle pool 16 includes an onboard vehicle computer system that is capable of transmitting vehicle information, such as Global Navigation Satellite System (GNSS) data, such as Global Positioning System (GPS) data, to a remote computer system, such as the vehicle relocation data system 12 and / or the ride-sharing data computer system 14.
[0052] Each computer system discussed herein (including the vehicle relocation data system 12, the ride-sharing data computer system 14, and the onboard vehicle computer system) includes at least one processor and a memory storing computer instructions accessible by the at least one processor. Depending on the embodiment, the hardware of each computer system need not be co-located but may be distributed, and depending on the embodiment, a cloud platform or software service provided by a cloud provider may be used.
[0053] Any one or more processors discussed herein can be implemented as any suitable electronic hardware capable of processing computer instructions and can be selected based on the purpose for which it will be used. Examples of the types of electronic processors that can be used include central processing units (CPUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), microprocessors, microcontrollers, and the like. Any one or more computer-readable memories discussed herein can be implemented as any suitable type of non-transient memory capable of storing data or information in a non-volatile and electronic form so that the stored data or information can be consumed by the electronic processor. The memory can be any of a variety of different electronic memory types and can be selected based on the purpose for which it will be used. Examples of memory types that can be used include magnetic or optical drives, ROM (read-only memory), solid-state drives (SSDs) (including other solid-state storage devices, such as solid-state hybrid drives (SSHDs)), other types of flash memory, hard disk drives (HDDs), non-volatile random access memory (NVRAM), and the like. It should be understood that a computer or computing device can include other memories, such as volatile RAM used by an electronic processor, and / or can include multiple electronic processors.
[0054] In at least some embodiments, the vehicle repositioning data system 12 is configured to determine the vehicle repositioning data using the following method. According to an embodiment, the vehicle repositioning data system 12 includes: training a NN, such as the CNN discussed above, to implement a passenger arrival rate estimation process; performing a passenger arrival estimation process to obtain estimated passenger arrival data for the vehicle pool 16; training (or pre-training) a value function to generate pre-trained data based on historical data; determining a vehicle repositioning strategy based on the pre-trained value function and the estimated passenger arrival data; and determining the vehicle repositioning data based on the vehicle repositioning strategy. For example, referring to Figure 2 , showing an exemplary schematic diagram of a process 100 for determining vehicle repositioning data.
[0055] The process 100 of determining vehicle repositioning data includes obtaining training data (step 110), training a value function based on the obtained training data (step 120) to obtain a pre-trained value function (step 130), determining estimated passenger arrival data for T future time steps or time slots (step 140), combining an LP (step 150) with the pre-trained value function and the estimated passenger arrival data to determine a vehicle repositioning strategy (step 160), and generating vehicle repositioning data according to the vehicle repositioning strategy (step 170). The vehicle repositioning data indicates the repositioned positions of the vehicles of the vehicle pool 16. The vehicle repositioning strategy data is generated at step 160 and is data representing the vehicle repositioning strategy. For a set of regions such as Figure 3 Vehicle repositioning data is determined using a set of regions shown in and discussed below.
[0056] refer to Figure 3 , shows a diagrammatic representation of a spatial grid or map 200 having a set of regions 202, in Figure 3 In the illustrated embodiment, each region is shown as a hexagon. Figure 3 It shows that the vacant vehicle 204 is designated to be removed from the area i ( Figure 3 206) moves to region j ( Figure 3 208 in the vehicle pool 16), as indicated by arrow 210. In an embodiment, the vehicle repositioning data specifies vehicle repositioning data for a plurality of vehicles, such as vehicle repositioning data for two or more vehicles in the vehicle pool 16 or for each vehicle in the vehicle pool 16.
[0057] Return Reference Figure 2 , at each time slot, a neural network that takes into account spatiotemporal correlations and, in some embodiments, the TM-CNN discussed herein is used to predict the passenger arrival rate in the next T time slots. The vehicle repositioning strategy is then generated by solving the LP with the predicted passenger arrival data and other estimated parameters. In an embodiment, a centralized planner (or computer system) sends vehicle repositioning data to one or more drivers or vehicles, such as sending a repositioning recommendation to each idle driver. According to an embodiment, the LP has a look-ahead time horizon T that explicitly models the controllability fraction and non-stationarity of the system. According to an embodiment, the method considers both short-term system dynamics and long-term rewards by incorporating a weighted pre-trained value function (approximately the total benefit of all drivers from time T+1 to the end of the horizon) into the objective function.
[0058] According to an embodiment, a ride-sharing system is modeled using a closed queuing network, such as discussed in Braverman; Afeche et al. (Afeche, P.; Baron, O.; Milner, J.; and Roet-Green, R. 2019. Pricing and prioritizing time-sensitive customers with heterogeneous demand rates, Operations Research, 67(4): 1184-1208). Let N be the number of vehicles and r be the number of regions in the set of regions. In this embodiment, each region is considered as a single server station, where idle vehicles are considered as jobs, and the assignment of a passenger's ride request to an idle vehicle corresponds to the completion of service at the station. The number of passenger arrivals in region i at time slot t is considered to be with parameter Nλ i (t), i∈[r] is a Poisson distribution, where i∈[r] is shorthand for i∈1,2,...,n. In an embodiment, the travel of a vehicle between regions is considered to be an infinite service station, where the vehicle travel is considered to be a job, and the service time of the job corresponds to the travel time of the vehicle. For any i∈[r], j∈[r], let 1 / μ ij (t) represents the average travel time from region i to region j at time slot t. Let P ij (t) represents the order destination distribution at time slot t, i.e., the probability of a passenger traveling to zone j given that the passenger request originates from zone i. Let q(t) = {q ij (t)} represents our relocation strategy at time slot t, so that the idle vehicles in area i are relocated with probability q ij (t) Relocate to region j. According to an embodiment, the goal is to find the optimal solution q * (t) such that the minimum completion rate of ride requests across all zones can be maximized. This maximum objective differs from the objective used in (Braverman et al., 2019) at least because it accounts for the nonstationarity and controllability fraction of the system, which has been found to be more robust in practice.
[0059] Based on the above definitions, the models described below in (1a)-(1f) are LPs with multi-step lookahead and a learned or pre-trained value function.
[0060]
[0061]
[0062] Note that the parameter λ i (t), P ij (t) and μ ik (t) is time-varying. Therefore, a T-slot look-ahead method is provided, which uses parameters averaged over time slots t, t+1, ..., t+T-1 to generate a vehicle relocalization strategy that takes future variations into account.
[0063] In the objective function (1a), the first term is the order fulfillment rate in zone i. Therefore, the first term maximizes the fulfillment rate in the zone with the most severe shortage. Max-min fairness is one of the classic objectives of traffic engineering or congestion control schemes, ensuring fair distribution of resources among competing flows, and is widely used in modern data center networks.
[0064] The second term of (1a) represents the expected discounted total return from time t+T to the end of the time horizon under consideration. In this term, w v is the weight used to strike a balance between the order completion rate over T time slots and the future rewards after T time slots. In particular, V(j,t) is the value function of each vehicle in region j at time slot t, defined as
[0065]
[0066] Where T end is the last time slot of the time horizon, (j, t) is considered as the state of the value function, and γ is the discount factor. t is the reward received by the vehicle at time slot t. It is assumed that the reward of an order is uniformly distributed over the duration of the trip. and are the average number of empty vehicles and full vehicles traveling from area i to area j, respectively.
[0067] In constraint 1(b), the left side is the total inflow of occupied vehicles from region i to region j (infinite server stations), which is equal to the right side That is, the total outflow of occupied vehicles (vehicles with passengers). Constraint (1e) is and The normalized sum of , which is the total number of vehicles. Constraint (1f) ensures that and is a non-negative sum Constraints (1c) and (1d) are discussed below.
[0068] Controllable scores and loose constraints. Ideally, every driver would accept the relocation recommendations sent by the centralized planner. However, in practice, this is not true because drivers have their own personalities and preferences. If each recommendation consumes an incentive reward, it may also be difficult to relocate every idle driver due to budget constraints. The original formulation in (Braverman et al., 2019) does not consider this situation.
[0069] Therefore, in the formula, a new parameter, the controllable fraction η∈[0,1], is introduced, and the controllable fraction η∈[0,1] represents the probability that the driver accepts the relocation suggestion. According to an embodiment, it is assumed that if the driver does not accept the recommendation, the driver will stay at the driver's location (will not move). Therefore, there are the following two equilibrium equations:
[0070]
[0071] Among them, q ij η represents the probability of an event where an empty vehicle is relocated from region i to region j, and the term (1-η+ηq ii ) represents the probability of the event that an empty vehicle stays in the current zone i. Therefore, the right-hand side of constraint (3a) is the total inflow of empty vehicles to the routes from zone i to zone j (by relocation), which is equal to the left-hand side, i.e., the total outflow of empty vehicles. Since q ij ≤1, equation (3a) implies the inequality constraint (1c) in our equation (1). For (3b), the left side is the total outflow of vehicles from area i, and the right side is the total inflow of vehicles into area i.
[0072] Note that the original LP in (Braverman et al. 2019) is formulated under the assumption that the Markov chain is already in a stable state. However, it takes time for the system to converge to a stable state. In practice, when the environment changes rapidly, the system may rarely be in a stable state. Therefore, to capture this non-stationary property, the equation in (3b) is changed to an inequality as follows:
[0073]
[0074] This means that the inflow rate into region i is allowed to be greater than the outflow rate from that region. This may happen when the system is not in a stable state. Since the right-hand side η of (3b) is non-increasing, the right-hand side may be larger, especially when η is small. This idea explains the phenomenon we observed in the simulation: even if the controllable fraction η is small, the proposed algorithm has good performance. Then from (3a) and (4a) we can obtain the relationship
[0075] Exemplary Algorithm. Based on the above LP formulas (1a)-(1f), Algorithm 1 is provided. Note that in (6) is not guaranteed to be non-negative. Therefore, lines 8 to 12 are used to make q * (t) is the effective relocalization strategy (probability matrix).
[0076] Before solving the LP, an estimate of the value function V is obtained; in accordance with an embodiment, this value function V can be trained offline using reinforcement learning (RL) (e.g., table TD learning). According to this embodiment, a discounting technique, such as that used in (Tang et al., 2019), is used to train the value function V offline using historical data, which can be generated by, for example, a simulator or retrieved from a remote data repository. The update rule for the value in region i at time t during training is defined as:
[0077]
[0078] Among them (R t ,η t ,D t ) is a completed order sampled from historical data with starting region i and starting time slot t. i Is the reward for the order, L i is the trip duration, and D i It’s the destination.
[0079]
[0080] The parameters used to train the value function V(j,s) are as follows, where N(j,s) is the number of times the training process visits the state (j,s).
[0081]
[0082] According to an embodiment, the look-ahead length or time horizon T can be set based on an analysis of the look-ahead length with respect to noise in the estimate. In one implementation and according to one embodiment, it was found that as the look-ahead length increases, the completion rate tends to increase first because system information from the near future is taken into account, and then decreases because the noise in the estimate increases when the planning horizon is too long. In one exemplary implementation and according to one embodiment, a look-ahead length of 30 minutes was determined to be optimal.
[0083] After offline training of the value function, at each time slot t, the future passenger arrival rate λ is predicted i (s), s∈{t,t+1,…,t+T-1}. Next, according to the implementation method, the passenger arrival rate {λ i(s)} is divided into its nearest neighbors. Then, given the approximate value function, solve LP to obtain the optimal solution q for T steps * (t) (repositioning matrix). Finally, the centralized planner is based on q * (t) Sample the relocation decisions of each idle driver and send recommendations to the driver. This process can be repeated until the end of the day or some other predetermined time frame.
[0084]
[0085] Prediction of passenger arrival rates. Note that, according to at least some embodiments, the proposed algorithm is a look-ahead strategy that uses future estimates λ i (s), P ij (s) and μ ij (s), s∈{t,t+1,…,t+T-1}. In an embodiment, the travel time μ is estimated by using the average value of historical data ij (s) and the destination distribution P ij =The inverse of (s). Since the passenger arrival rate can vary significantly during a day and also fluctuate on different days, a robust method is needed to estimate the passenger arrival rate (or other passenger arrival data) before using such estimated passenger arrival data as input to the LP. In this section, a neural network (NN)-based prediction method is proposed to estimate the real-time passenger arrival rate.
[0086] It has been found that there are common trends in the number of rides requested during a day; for example, the number of orders has a peak during the rush hour of each day. Therefore, detrending methods, such as (Li, Z.; Li, Y.; and Li, L. 2014. A comparison of detrending models and multi-regime models for traffic flow prediction. IEEE Intelligent Transportation Systems Magazine, 6(4): 34-44), are used to first reduce this common trend before making online predictions. The first step in the detrending method is to determine the trend by taking a simple average of multiple days of historical data. Then, the residual time series can be obtained by subtracting the day-to-day trend. Next, a forecasting method can be used to predict the residual passenger arrival rates in the future. Finally, these residual passenger arrival rates will be added to the current trend to obtain the final predicted future passenger arrival rates.
[0087] In order to utilize the spatiotemporal correlation of passenger arrival rates, a long short-term memory convolutional neural network (LSTM-CNN) is proposed, and according to an embodiment, the LSTM-CNN is combined with a detrending method. The structure of the LSTM-CNN is as follows: Figure 4 As shown in Figure 2, the LSTM component captures temporal correlations, while the CNN captures spatial correlations in passenger arrival rates. Details of the exemplary network structure and parameters are shown below. The network and training parameters for the LSTM-CNN are shown below. Pytorch can be used to implement and train the LSTM-CNN.
[0088]
[0089] Split passenger arrivals. Depending on the implementation, the split passenger arrival method can be used in line 4 in Algorithm 1. The implementation of the split passenger arrival method is defined in Algorithm 2. The reason is that, in practice, a driver may accept order requests from neighboring grids. In previous works, including LP-based or RL-based works, it is assumed that the driver only accepts orders from his current grid (which does not hold true in practice), where orders are dispatched based on a predefined broadcast distance between the customer and the driver. This becomes a problem and cannot be ignored in large-scale systems. For example, for some grids close to hot spots, although these grids may have no passenger arrivals, they do have virtual passenger arrivals because they are allowed to accept orders from their neighbors. After predicting the demand online, the passenger arrival rate of the grid is evenly divided into its neighbors. Note that Algorithm 2 is a weighted splitting algorithm, α c ,α n It is designed based on the platform's order dispatch strategy and grid radius.
[0090] Relocation cost. Depending on the implementation and method, if the LP problem has multiple optimal solutions, it may be desirable to select a solution that relocates to a closer area because the driver wants to reduce travel (fuel) costs or considers that each completed relocation consumes a certain budget as an incentive reward. By further adding a small penalty to the objective function while maintaining its linearity, a new LP is obtained whose solution will have a lower travel cost while maintaining a similar order completion rate. The new objective function is designed as follows:
[0091]
[0092] where d ij is the normalized distance between region i and region j (d ij ≤1), and w tr are parameters that can be adjusted to balance these objectives.
[0093] The above formulas and discussions under “Exemplary Algorithm,” “Prediction of Passenger Arrival Rate,” “Split Passenger Arrivals,” and “Repositioning Costs” describe details of the disclosed system and method according to embodiments and are exemplary in nature.
[0094] refer to Figure 5 , a method 500 for determining vehicle repositioning data for a vehicle pool is shown. In an embodiment, the method 500 is performed by the vehicle repositioning data system 12. Although the method 500 is described below as performing steps 510-540 in a specific order, steps 510-540 may be performed in any technically feasible order, as will be understood by those skilled in the art; for example, in an embodiment, step 520 may be performed before or simultaneously with step 510.
[0095] Method 500 begins at step 510 , where a value function is pre-trained using historical ride-sharing data. In embodiments, the historical ride-sharing data is data indicating trajectory information for one or more vehicles (such as one or more vehicles in a vehicle pool) and / or order completion rate information for a set of regions. In embodiments, it may include more precise data, such as order completion rates for a set of sub-regions within each region within the set of regions. The historical ride-sharing data may be obtained from the ride-sharing data computer system 14 or other remote data repository. As discussed above in conjunction with step 110 , the training data or historical ride-sharing data may be trajectory data, such as GNSS data indicating vehicle positions over time. Furthermore, in some embodiments, order information related to ride-sharing or ride-hailing customer orders may also be used for training. In some embodiments, training the value function includes performing reinforcement learning (RL), particularly temporal difference (TD) learning, such as the tabular TD learning discussed above. In embodiments, pre-training of the value function is performed offline; furthermore, this pre-training may be referred to as offline training or offline pre-training. In embodiments, training the value function (step 510 ) may be performed periodically according to predetermined time intervals. Method 500 continues to step 520 .
[0096] In step 520, estimated passenger arrival data is obtained. In an embodiment, the estimated passenger arrival data is obtained by generating the estimated passenger arrival data using a temporal memory convolutional neural network (TM-CNN), the temporal memory convolutional neural network having a temporal memory (TM) and a convolutional neural network (CNN) with at least two CNN layers, such as the long short-term memory LSTM-CNN discussed above. Thus, according to an embodiment, the TM-CNN is an LSTM-CNN; however, in other embodiments, a gated recurrent unit (GRU) can be used as the TM. In some embodiments, the TM-CNN is configured to receive historical passenger arrival data as input. In an embodiment, the historical passenger arrival data is two-dimensional (2D) passenger arrival data representing passenger arrival information for locations in a two-dimensional space, and the historical passenger arrival data can be generated by a simulator and / or obtained from a remote data repository. Method 500 continues to step 530.
[0097] In step 530, a vehicle repositioning strategy is determined based on the pre-trained value function and the estimated passenger arrival data. In embodiments, the vehicle repositioning strategy is periodically determined over a predetermined period of time and utilizes the estimated passenger arrival data over the predetermined period of time. In at least some embodiments, the vehicle repositioning strategy is determined by solving the linear programming (LP) problem discussed above, which can be expressed as a look-ahead model because future estimated passenger arrival data is considered as part of the LP problem, as described above. Method 500 continues to step 540.
[0098] In step 540, vehicle repositioning data is determined for the vehicle pool based on the vehicle repositioning strategy. In embodiments, the vehicle repositioning data is for a plurality of vehicles in the vehicle pool, such as a plurality of empty vehicles in the vehicle pool. In embodiments, the vehicles in the vehicle pool report location information (e.g., GNSS data) to the vehicle repositioning data system 12, which then uses the vehicle's location and order status information (e.g., vehicle is empty, vehicle is on its way to pick up a passenger / customer) in conjunction with the determined vehicle repositioning strategy to determine vehicle repositioning data, which may indicate a recommended repositioning location. In embodiments, the vehicle repositioning data is for each empty vehicle in the vehicle pool. Method 500 ends.
[0099] Simulation and Evaluation. Extensive experiments using a small-scale simulator with real-world datasets demonstrate that the disclosed method achieves significant improvements over other baseline methods and is robust to prediction errors. The disclosed method is further tuned and evaluated on a more complex and realistic simulation platform similar to the KDD Cup 2020 RL track competition, and simulation results show that the disclosed method achieves state-of-the-art results in improving the completion rate of orders in the system.
[0100] Simulation Results. In this section, we discuss several extensive experiments designed to validate the performance and robustness of our algorithm. First, we introduce the simulation platform and setup.
[0101] Simulation platform. The simulation is performed on a platform built using a public real-world dataset, which is also used by the KDDCup 2020 RL track competition. The platform consists of r = 20 hexagonal grids with a radius of about 600m, which are selected from the hot spots of the dataset (busy traffic conditions). The grids are distributed in Figure 6 As shown in Figure 6 is a diagram of a grid diagram 600 . Figure 7 A diagram of vehicle repositioning data 700 is shown specifying specific vehicle repositioning recommendations for two vehicles; Figure 7 In , the arrows represent the relocation directions and the shading represents the different demand levels of the corresponding grids, where darker represents higher demand. Orders are generated by sampling from the dataset. The passenger arrival rates are synthetic and assumed to follow a Poisson distribution with a time-varying parameter, which changes every 10 minutes. The arrival rate of some grids is defined as the average of the arrival rates of their neighboring grids in order to generate some spatial correlation. The platform dispatches orders to idle drivers and relocates idle drivers every minute. During the simulation, the dispatching algorithm is fixed and unknown to the relocation algorithm. In one of the simulations performed, orders can be dispatched to drivers within a fixed broadcast distance (e.g., 1.2 km). ɑ is then selected in Algorithm 2 c =1,α n =0.5. If a driver does not receive any orders and relocation tasks, the driver will stay at their location. The time interval is from 1:00 pm to 7:59 pm, and the number of drivers (vehicles) during this time period is fixed. The number of drivers is 300, and the controllability fraction is assumed to be 0.2 unless otherwise stated. Drivers during the relocation period can accept orders in the destination grid in advance, and if no order is assigned to the driver, the driver who has completed the relocation needs to wait 5 minutes to accept a new relocation recommendation. These settings were made after discussions with researchers from the ride-hailing company.
[0102] In Algorithm 1, 10 minutes is set as one time slot according to the simulation run. In the simulated embodiment, the look-ahead time T is set to 2 time slots (20 minutes) unless otherwise specified. The relocation route matrix q is recalculated at the beginning of each time slot. * (t); In an embodiment, the vehicle repositioning strategy (and vehicle repositioning strategy data) is recalculated / redetermined periodically according to a predetermined time interval or according to a predetermined schedule.ij (t) and the destination distribution P ij (t), using the average value of historical data. For passenger arrival rate λ i (t), the previous M=6 time slots are used for online prediction. Unless otherwise stated, the weight w v Set to 0.3. The results of each experiment were averaged over 30 trials. It should be understood that the values mentioned above are for the purpose of evaluating embodiments of the present disclosure, and other values may be used depending on the embodiment.
[0103] Baseline algorithms for simulations. In simulations, the disclosed relocalization algorithm is compared with the following strategies: (i) Stay: The stay strategy simply does not relocalize; (ii) Expert: A human expert strategy extracted from historical idle driver transition data that captures collective intelligence; (iii) Nearest Neighbor: This is a heuristic strategy that randomly and uniformly relocates idle drivers to one of the sets of their neighbors and themselves; (iv) LP: This strategy directly uses the solution of LP without a value function as the relocalization strategy. It uses the ground truth of the expected arrival rate or the predicted value of the arrival rate (LP-P); (v) LP-S: This strategy uses the solution of the LP model of (Braverman et al., 2019) under the fixed system assumption. It uses the ground truth of the expected arrival rate; and (vi) Value Based-R (reward-based value function): This random strategy samples the relocalization destination with a probability proportional to the value of the grid of neighbors. (i.e., the method in (Jiao et al., 2020)).
[0104] Controllable Score. In this subsection, the impact of the controllable score on the total order completion rate is studied. Figure 8 The completion rate when the controllable fraction is varied is shown, and it is shown that the disclosed algorithm outperforms other strategies at different controllable fractions. For example, for a controllable fraction η = 0.2, the completion rate of our algorithm increases by 9.3% compared to other non-LP based strategies, and by 12.6% compared to no repositioning (Stay). In addition, the order completion rate increases as the controllable fraction increases until the "saturation" point. The saturation point of the disclosed algorithm is only 0.12, and the disclosed algorithm has achieved a completion rate of up to 0.873 by controlling only 10% of the vehicles, which may be achieved in a real ride-hailing platform. However, LP without considering the non-fixed system and the controllable fraction rate needs to control at least 60% of the vehicles to achieve a similar completion rate. In addition, as Figure 8 As shown, the performance can be further improved by including a value function along with the predictions.
[0105] Spatial Heterogeneity. In this section, we investigate the impact of spatial heterogeneity in arrival rates on the performance of the relocalization algorithm. Some arrivals were manually moved from 10 grids with lower arrival rates to the other 10 grids to increase spatial heterogeneity. The variance of the arrival rates across the 20 grids was used to quantify spatial heterogeneity. Figure 9 It shows that as spatial heterogeneity increases, the completion rate decreases at different rates for all algorithms. For example, for a controllability fraction η = 0.2 and a spatial variance of 70, while the nearest neighbor strategy only has a completion rate of 0.587, the proposed algorithm still achieves a completion rate of 0.786. The proposed algorithm outperforms LP-S, which uses the true arrival rate to obtain the solution, at two different controllability fraction rates.
[0106] The number of vehicles. Figure 10 The effect of the total number of vehicles on performance is shown. Not surprisingly, the completion rate increases when the number of vehicles increases for all algorithms. In fact, the number of vehicles represents the supply-demand gap in the system. There is clearly a supply-demand gap (250-350 vehicles) where relocalization will have a significant gain. When the number of vehicles is large, that is, when the supply is sufficient, no relocalization is required and all relocalization algorithms can perform fairly well. Conversely, when the number of vehicles is small, that is, when the supply is small, no relocalization is required because almost all vehicles are always busy and there are no idle drivers in the system. This shows that there are traffic regimes where relocalization algorithms can play an important role.
[0107] Repositioning costs. Figures 11A-11B Compare the travel (fuel) cost and completion rate between different objective functions. The results show that by choosing a small penalty (w tr ), driving costs can be reduced by up to 23.6% while maintaining similar completion rates.
[0108] Simulation on a more complex and realistic platform. To further verify the practicality of the disclosed algorithm, the disclosed method is modified and evaluated on a more realistic and complex simulation platform similar to the KDD Cup 2020 RL track competition. The simulator contains 1745 grids, which makes solving the LP problem more challenging. Drivers are allowed to accept orders even during relocalization. Since all features introduce new difficulties to solving the LP, the LP is modified with the following changes: (i) since cities usually have several centers (urban and suburban areas), the Louvain method (Blondel et al., 2008) is used for community detection on the map used in the platform; and (ii) the relocalization distance is limited to 2 hops, which significantly reduces the number of variables in the LP problem. As Figure 12As shown, the right figure is the detection result, which shows that there are four main areas in the city, which is consistent with the thermal Figure 1 The city is then divided into three parts (in the box), and each part is modeled with a separate LP. The parallel LP problem and the computational problem can then be solved.
[0109] Simulation results. Figure 13 As shown, the disclosed method still outperforms other algorithms in a complex simulation environment. In simulation, the value function is trained using the RL method from (Tang et al., 2019). Compared to the value-based method ValueBased-R (Tang et al., 2019), the disclosed algorithm increases the completion rate by 1%. The baseline ValueBased-G (target-based value function) is a greedy relocalization strategy that routes empty vehicles to the area with the highest value.
[0110] It should be understood that the foregoing description is of one or more embodiments of the present invention. The present invention is not limited to the specific embodiments disclosed herein, but is limited only by the appended claims. In addition, the statements contained in the foregoing description relate to the disclosed embodiments and should not be interpreted as limitations on the scope of the invention or the definition of terms used in the claims, unless a term or phrase is expressly defined above. Various other embodiments and various changes and modifications to the disclosed embodiments will become apparent to those skilled in the art.
[0111] As used in this specification and claims, when used in connection with a list of one or more components or other items, the terms "e.g.," "for example," "for instance," "such as," and "like," and the verbs "comprising," "having," "including," and their other verb forms, are each to be interpreted as open-ended, meaning that the list should not be viewed to exclude other additional components or items. Other terms are to be interpreted using their broadest reasonable meaning unless they are used in a context that requires a different interpretation. In addition, the term "and / or" should be interpreted as an inclusive "OR." Thus, for example, the phrase "A, B, and / or C" should be interpreted to cover all of the following: "A"; "B"; "C"; "A and B"; "A and C"; "B and C"; and "A, B, and C."
Claims
1. A method for determining vehicle repositioning data for a vehicle pool, the method comprising: Use historical ride-sharing data to train the value function; Obtaining estimated passenger arrival data, wherein the estimated passenger arrival data is obtained by generating the estimated passenger arrival data using a temporal memory convolutional neural network (TM-CNN); determining a vehicle repositioning strategy based on the trained value function and the estimated passenger arrival data; as well as Vehicle repositioning data for a pool of vehicles is determined based on the vehicle repositioning strategy.
2. The method according to claim 1, wherein The historical ride-sharing data includes trajectory data from the pool of vehicles.
3. The method according to claim 1, wherein The TM-CNN includes a temporal memory (TM) and a convolutional neural network, and wherein the TM includes at least one of a long short-term memory (LSTM) and a gated recurrent unit (GRU).
4. The method according to claim 3, wherein: The CNN includes an encoding layer and a decoding layer, and wherein the TM is inserted in an embedding layer between the encoding layer and the decoding layer.
5. The method according to claim 4, wherein The input of the TM-CNN includes two-dimensional (2D) passenger arrival data, which represents passenger arrival information for a location in a two-dimensional space and for a given time or time period.
6. The method according to claim 1, wherein The vehicle repositioning strategy is periodically determined according to predetermined time intervals.
7. The method according to claim 1, wherein The vehicle repositioning data is for a plurality of vehicles in the vehicle pool.
8. The method according to claim 1, wherein The vehicle repositioning strategy is determined using an optimized look-ahead method that takes into account the estimated passenger arrival data.
9. The method according to claim 8, wherein The optimized look-ahead method uses linear programming LP.
10. The method according to claim 1, wherein A controllability score is determined based on the historical ride-sharing data or other historical ride-sharing data, and wherein the controllability score is used to determine the vehicle repositioning strategy.
11. A vehicle positioning system comprising: at least one processor; Memory that stores computer instructions; wherein the vehicle relocation system is configured to execute the computer instructions using the at least one processor, such that when the computer instructions are executed by the at least one processor, the vehicle relocation system: Use historical ride-sharing data to train the value function; Obtaining estimated passenger arrival data, wherein the estimated passenger arrival data is obtained by generating the estimated passenger arrival data using a temporal memory convolutional neural network (TM-CNN); determining a vehicle repositioning strategy based on the trained value function and the estimated passenger arrival data; as well as Vehicle repositioning data for a pool of vehicles is determined based on the vehicle repositioning strategy.
12. The vehicle repositioning system of claim 11, wherein: The historical ride-sharing data includes trajectory data from the pool of vehicles.
13. The vehicle repositioning system of claim 11, wherein: The TM-CNN includes a temporal memory (TM) and a convolutional neural network, and wherein the TM includes at least one of a long short-term memory (LSTM) and a gated recurrent unit (GRU).
14. The vehicle repositioning system of claim 13, wherein: The CNN includes an encoding layer and a decoding layer, and wherein the TM is inserted in an embedding layer between the encoding layer and the decoding layer.
15. The vehicle repositioning system of claim 14, wherein: The input of the TM-CNN includes two-dimensional (2D) passenger arrival data, which represents passenger arrival information for a location in a two-dimensional space and for a given time or time period.
16. The vehicle repositioning system of claim 11, wherein: The vehicle repositioning strategy is periodically determined according to predetermined time intervals.
17. The vehicle repositioning system of claim 11, wherein: The vehicle repositioning data is for a plurality of vehicles in the vehicle pool.
18. The vehicle repositioning system of claim 11, wherein: The vehicle repositioning strategy is determined using an optimized look-ahead method that takes into account the estimated passenger arrival data.
19. The vehicle repositioning system of claim 18, wherein: The optimized look-ahead method uses linear programming LP.
20. The vehicle repositioning system of claim 11, wherein: A controllability score is determined based on the historical ride-sharing data or other historical ride-sharing data, and wherein the controllability score is used to determine the vehicle repositioning strategy.