Dynamic Pricing via Reinforcement Learning for Ride-Hailing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Online ride-hailing platforms face inefficiencies due to imbalances in supply and demand across time and space, which existing pricing strategies fail to effectively address, leading to suboptimal revenue and operational challenges.

Innovation Solution

A dynamic pricing system based on model-based deep reinforcement learning that updates pricing candidates to minimize cross-entropy with a target policy, maximizing total income by iteratively using a trained reinforcement learning model to optimize pricing actions across the platform.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing pricing strategies are used, then operational simplicity is maintained, but revenue optimization and supply-demand balance are insufficient

Engineering Contradiction:
Improverevenue optimizationVSAvoidpricing system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a reinforcement learning model as an intermediary between the pricing system and the environment. This model learns optimal pricing strategies through continuous interaction, receiving rewards for successful matching and penalties for failures, thereby optimizing revenue without requiring complex manual pricing rules

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The pricing system performs self-optimization through the reinforcement learning agent that automatically learns and adjusts pricing strategies based on real-time feedback from the environment. The system serves itself by continuously improving its pricing policy through trial-and-error learning without external intervention

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If dynamic pricing is implemented to balance supply and demand, then revenue is improved, but system complexity and computational requirements increase

Engineering Contradiction:
Improvesupply-demand balanceVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic pricing where prices continuously adapt to changing supply and demand conditions. The reinforcement learning model updates its pricing policy in real-time based on current platform state, making the system highly adaptable to spatial-temporal variations in ride-hailing demand

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates continuous feedback loops where the reinforcement learning agent receives reward signals based on pricing outcomes (successful matches, cancellations, income generated). This feedback mechanism enables the system to learn from past decisions and continuously improve its supply-demand balancing capability

Inventive Principle:
Principle #23Feedback

3Productivity

If reinforcement learning is used for pricing optimization, then total income is maximized, but training time and computational resources are consumed

Engineering Contradiction:
Improvetotal incomeVSAvoidtraining time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary training of the reinforcement learning model in a simulated environment before deployment. By pre-training the agent offline with synthetic data and simulated ride-hailing scenarios, the system accumulates learning experience in advance, reducing the time needed for real-world adaptation and enabling faster deployment

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11443335B2Model-based deep reinforcement learning for dynamic pricing in an online ride-hailing platform
Publication Date: 2022.09.13 BEIJING DIDI INFINITY TECH & DEV CO LTD
  • US11443335B2 patent drawing
  • US11443335B2 patent drawing
  • US11443335B2 patent drawing

AI summary

Dynamic pricing may be applied in an online ride-hailing platform. Information may be obtained. The information may include a set of pricing candidates and an initial status of a ride-hailing platform. The set of pricing candidates may be updated based on the initial status of the ride-hailing platform to minimize a cross-entropy between the set of pricing candidates and a target pricing policy that maximizes a total income of the ride-hailing platform. A price for at least one current trip request on the ride-hailing platform may be generated based on the updated set of pricing candidates.