Vehicle Navigation Simulation for Dispatch Policy Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vehicle dispatch platforms lack the ability to provide drivers with a strategic approach to maximize earnings and passengers with minimized trip time, as they cannot determine the best strategy for accepting requests or routing when carpooling, relying on simple distance-based matching and lacking decision-making algorithms.

Innovation Solution

A vehicle navigation simulation environment is created using reinforcement learning algorithms that train a policy to guide drivers in maximizing rewards and minimizing passenger trip time by simulating scenarios based on historical data, allowing for optimized decision-making in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If simple distance-based matching is used for vehicle-passenger allocation, then the system is easy to operate and quick to implement, but the driver cannot maximize earnings and the passenger cannot minimize trip time

Engineering Contradiction:
Improveease of operationVSAvoiddriver earnings efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent creates a simulation environment that copies real-world vehicle dispatch scenarios into a virtual realm. The simulation replicates passenger requests, locations, and routing decisions, allowing the trained policy to learn optimal strategies without interfering with actual operations. This copying approach enables complex decision-making training while maintaining operational simplicity in the real world.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The trained policy acts as an intermediary between the simple distance-based matching system and the complex optimization goals. It translates real-time dispatch data into optimized routing decisions, serving as a bridge that enhances driver earnings and minimizes passenger trip time without requiring fundamental changes to the existing system architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If a complex decision-making algorithm is implemented to maximize driver earnings and minimize passenger trip time, then productivity and efficiency improve, but the device complexity and computational requirements increase

Engineering Contradiction:
Improvedriver earnings efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary training of the decision-making policy in a simulation environment before deploying it to real-world operations. During training, the system pre-computes optimal routing strategies by evaluating numerous scenarios and outcomes. This preliminary action allows the complex algorithm to be pre-processed and stored as a trained policy, reducing computational complexity during actual vehicle dispatch operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By copying real-world scenarios into a simulation environment, the system can perform complex computational experiments without the overhead of real-time execution. The simulation replicates dispatch scenarios, allowing the trained policy to learn optimal strategies through repeated virtual practice, thereby reducing the computational burden during actual operations.

Inventive Principle:
Principle #26Copying

3Speed

If real-time decision-making is implemented for vehicle routing and passenger allocation, then the system responds quickly to changing conditions, but the computational time and processing requirements increase

Engineering Contradiction:
Improveresponse speedVSAvoidcomputational energy consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary training and policy development during off-peak times or in simulation environments. The trained policy is then deployed for real-time decision-making, which significantly reduces computational requirements during actual operations. This preliminary action allows the system to respond quickly to real-time conditions without the energy consumption of complex on-the-fly computations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The simulation environment provides feedback loops that allow the trained policy to learn from outcomes of virtual dispatch scenarios. This feedback mechanism enables the system to optimize its decision-making strategies through iterative learning, improving response speed while reducing the computational energy required for real-time implementations.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10635764B2Method and device for providing vehicle navigation simulation environment
Publication Date: 2020.04.28 BEIJING DIDI INFINITY TECH & DEV CO LTD
  • US10635764B2 patent drawing
  • US10635764B2 patent drawing
  • US10635764B2 patent drawing

AI summary

A method for providing vehicle navigation simulation environment may comprise recursively performing steps (1)-(4) for a time period: (1) providing one or more states of a simulation environment to a simulated agent, wherein: the simulated agent comprises a simulated vehicle, and the states comprise a first current time and a first current location of the simulated vehicle; (2) obtaining an action by the simulated vehicle when the simulated vehicle has no passenger, wherein the action is selected from: waiting at the first current location of the simulated vehicle, and transporting M passenger groups; (3) determining a reward to the simulated vehicle for the action; and (4) updating the one or more states based on the action to obtain one or more updated states for providing to the simulated vehicle, wherein: the updated states comprise a second current time and a second current location of the simulated vehicle.