Virtual Environment for Ride-Hailing Policy Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Reinforcement learning for driver incentive program recommendation in ride-hailing platforms is impractical due to high trial-and-error costs and the influence of unobservable factors in the ride-hailing environment, such as hidden confounders.

Innovation Solution

A virtual environment is constructed using historical interaction trajectories, with a simulator integrating a platform policy, confounding policy, and driver policy to simulate interactions and approximate the real-world data distribution, incorporating a reward function trained with an uplift inference network to optimize program recommendation policies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If reinforcement learning is applied to train driver incentive program recommendation policies in a real ride-hailing environment, then the policy optimization capability is improved, but the trial-and-error cost becomes prohibitively high

Engineering Contradiction:
Improvepolicy optimization capabilityVSAvoidtrial-and-error cost
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent creates a virtual environment that copies the essential dynamics and characteristics of the real ride-hailing environment. This virtual copy allows reinforcement learning to be performed on simulated data rather than real-world interactions, thereby maintaining policy optimization capability while eliminating the high trial-and-error costs associated with real-world testing.

Inventive Principle:
Principle #26Copying

2Loss of time

If a virtual environment is constructed to simulate real-world interactions, then the trial-and-error cost is reduced, but the complexity of the system increases due to the need to model confounding factors

Engineering Contradiction:
Improvetrial-and-error costVSAvoidsystem complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent extracts and isolates the confounding factors from the complex real-world environment and models them separately within the virtual environment. By taking out these difficult-to-model elements and representing them explicitly through confounding policies and variables, the system can simulate their effects without needing to replicate the entire complexity of the real world, thus reducing overall system complexity while maintaining simulation accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If the simulator models only observable factors, then the system complexity is reduced, but the measurement precision of driver behavior prediction deteriorates due to unobservable confounders

Engineering Contradiction:
Improvesystem complexityVSAvoidbehavior prediction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces confounding policies and confounding variables as intermediary elements that mediate between observable factors and driver behaviors. These intermediaries represent the unobservable confounding factors and allow the simulator to capture their influence on driver decisions without directly observing them, thereby improving behavior prediction accuracy while maintaining manageable system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12198216B2Method and system for constructing virtual environment for ride-hailing platforms
Publication Date: 2025.01.14 BEIJING DIDI INFINITY TECH & DEV CO LTD
  • US12198216B2 patent drawing
  • US12198216B2 patent drawing
  • US12198216B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for constructing a virtual environment for a ride-hailing platform are disclosed. An exemplary method comprises: obtaining a plurality of historical interaction trajectories each comprising one or more interaction records between a driver and a ride-hailing platform, each interaction record comprising a program recommendation of the ride-hailing platform to the driver and a reaction of the driver in response to the program recommendation; training a simulator based on the plurality of historical interaction trajectories; and integrating a reward function with the simulator to construct the virtual environment, wherein the plurality of first program recommendations and the plurality of reactions form a plurality of simulated interactions, and a data distribution of the plurality of simulated interactions approximates a data distribution of a plurality of interaction records in the plurality of historical interaction trajectories.