Generative Adversarial Network for Driver Incentive Policy Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning algorithms for incentivizing drivers in transportation hailing systems face challenges due to reliance on large-scale sampling, inefficiency, and high costs, as well as the inability to directly evaluate and improve incentive policies using static historical data, which is not optimal for preventing 'fading drivers' and is affected by external interference factors.

Innovation Solution

A system utilizing a joint policy model generator and discriminator within a generative adversarial network to construct historical trajectories of driver interactions, generate optimized incentive policies, and simulate driver actions, allowing for reinforcement learning to optimize incentives based on rewards, thereby addressing the inefficiencies and limitations of traditional methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If reinforcement learning algorithms use large-scale sampling to optimize incentive policies, then the accuracy of policy optimization improves, but the cost and time consumption increase significantly

Engineering Contradiction:
Improvepolicy optimization accuracyVSAvoidsampling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates a virtual environment that copies the characteristics of the real driver incentive system. Instead of sampling from real historical data, the system generates synthetic trajectories in a virtual environment that replicate real-world driver behaviors and responses to incentives, enabling efficient reinforcement learning without costly real-world sampling

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system pre-generates a comprehensive virtual environment and driver behavior models before conducting reinforcement learning. By preparing the virtual simulation framework in advance with pre-defined driver characteristics, incentive mechanisms, and interaction protocols, the system eliminates the need for time-consuming real-world sampling during the actual optimization process

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If traditional imitation learning methods are used to learn from historical data, then the implementation simplicity improves, but the ability to optimize beyond historical strategies deteriorates

Engineering Contradiction:
Improveimplementation simplicityVSAvoidpolicy optimization capability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The system transitions from static historical data analysis to dynamic virtual environment simulation. The virtual environment allows drivers and incentive policies to interact dynamically, enabling the discovery of optimized strategies that adapt to various scenarios rather than merely copying past behaviors. The simulation can explore counterfactual scenarios and optimize policies beyond historical constraints

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The virtual environment acts as an intermediary between historical data and policy optimization. It translates static historical trajectories into a dynamic simulation framework where optimized policies can be tested and evaluated, bridging the gap between simple historical analysis and complex optimization while maintaining implementation feasibility

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If static historical data is used for training, then the data availability improves, but the ability to evaluate different incentive policies deteriorates

Engineering Contradiction:
Improvedata availabilityVSAvoidpolicy evaluation capability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The system copies the essential characteristics of historical driver data into a virtual environment where these characteristics can be manipulated and recombined. Instead of being constrained by the fixed nature of historical data, the virtual replication allows flexible exploration of different incentive policies and their potential impacts on driver behaviors

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent adds a temporal and interactive dimension to static historical data by transforming it into a dynamic virtual simulation. This dimensional transformation enables the evaluation of policies across multiple scenarios and time periods, allowing comprehensive policy assessment that static data alone cannot provide

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Reliability

If reinforcement learning is applied directly to real driver data, then the practical relevance improves, but the sampling efficiency and cost worsen

Engineering Contradiction:
Improvepractical relevanceVSAvoidsampling efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system creates a virtual copy of the real driver incentive ecosystem, capturing essential behaviors, preferences, and response patterns. This virtual replica maintains practical relevance by faithfully representing real-world dynamics while enabling high-speed simulation and sampling that would be prohibitively expensive or time-consuming in the real world

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The virtual environment serves as an intermediary layer between theoretical reinforcement learning algorithms and real-world driver data. It enables efficient experimentation and policy optimization in a controlled simulation setting before deploying to real drivers, improving sampling efficiency while maintaining practical relevance through faithful representation of real behaviors

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11861643B2Reinforcement learning method for driver incentives: generative adversarial network for driver-system interactions
Publication Date: 2024.01.02 BEIJING DIDI INFINITY TECH & DEV CO LTD
  • US11861643B2 patent drawing
  • US11861643B2 patent drawing
  • US11861643B2 patent drawing

AI summary

A system and method of determining a policy to prevent fading drivers is described. The system and method creates virtual trajectories of incentives such as coupons offered to drivers in a transportation hailing system and corresponding states of drivers in response to the incentives. A joint policy simulator is created from an incentive policy, a confounding incentive policy, and an incentive object policy to generate the simulated actions of drivers in response to different incentives. The rewards of the simulated actions of the drivers is determined by a discriminator. The incentive policy for preventing fading drivers is optimized by reinforcement learning based on the virtual trajectories generated by the joint policy simulator and discriminator.