Online car-hailing driver reward generation method based on causal inference and operation planning optimization

Through causal inference and operation optimization methods, based on the uplift model and XGBoost regression model, driver populations are dynamically divided and reward strategies are generated, which solves the problem of lack of scientific basis for driver reward design in the existing technology, and realizes personalized distribution of rewards and high ROI.

CN120355463APending Publication Date: 2025-07-22SHANGHAI SAIKE MOBILITY TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510346820.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the existing online ride-hailing system, the design of driver order rewards lacks scientific basis and cannot identify individual driver differences, resulting in inefficient reward distribution, difficulty in dynamically adapting to market changes, and failure to achieve high ROI.

Method used

The causal inference and operation optimization method are used to divide the driver population into responsive, natural and sleepy through the uplift model. The tree model predicts the singular and proportions, and uses XGBoost regression model and linear planning to generate dynamic reward strategies to meet the operation budget and step continuity constraints.

Benefits of technology

The personalized allocation of reward strategies has been realized, the ROI has been improved, the dynamic adaptation to market changes has been improved, and the control efficiency of reward budgets and the rationality of rewards has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355463A_ABST
    Figure CN120355463A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of online car-hailing, and particularly relates to an online car-hailing driver reward generation method based on causal inference and operation planning optimization. The online car-hailing driver reward generation method based on causal inference and operation planning optimization comprises the following specific steps: step 1, crowd division: constructing an uplift model based on historical driver's reward behaviors for order punching to perform crowd division on drivers, and generating three types of crowds of a response type, a natural type and a sleep type; 2, crowd complete order step prediction: constructing a regression model based on a tree model, and predicting a complete order number and a crowd ratio of each order of the crowd in an activity period; and step 3, reward generation: based on step distribution and maximum amount constraint set by operation, a final reward form is obtained through planning and solving. A prediction model is introduced to better control reward budget, a more reasonable reward mode is generated by planning a mathematical solution mode, and a reward with more ROI is generated in combination with the elastic model and the prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online car-hailing, and specifically to a method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization. Background Technique

[0002] In the online car-hailing scenario, in order to stimulate drivers to take orders on the platform, some behaviors to stimulate drivers are often provided, such as the order rush award. It is mainly in the form of agreeing on how much money will be rewarded for the driver to complete a certain number of orders. Generally, according to operational experience, a reward form like [3-8, 5-10, 7-12] will be set. That is, a reward of 8 yuan for completing 3 orders, 10 yuan for completing 5 orders, and 12 yuan for completing 7 orders. However, at present, there is little intervention by algorithms.

[0003] Existing technical solutions:

[0004] Traditional driver order rush rewards adopt a fixed step design (such as corresponding fixed rewards for completing 3 / 6 / 9 orders). Manual experience is used to formulate step targets and reward amounts, and a unified strategy is implemented for all drivers. Technical defects: It is impossible to identify the individual difference response characteristics of drivers, the setting of reward steps and the distribution of amounts lack a scientific basis, the low efficiency of budget allocation leads to ROI losses, and it is difficult to dynamically adapt to the market environment of supply and demand changes. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization to solve the problems raised in the above background technique.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization. The reward generation method is based on a causal inference grouping module, a demand prediction module, and an operations research optimization module. An electrical connection relationship is established among the causal inference grouping module, the demand prediction module, and the operations research optimization module;

[0007] The specific steps of the method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization are as follows:

[0008] Step 1: Population division: Based on the historical drivers' behaviors towards order rush rewards, an uplift model is constructed to divide the drivers into three categories: responsive, natural, and dormant;

[0009] Step 2: Prediction of the order completion steps for the population: Based on a tree model, a regression model is constructed to predict the number of orders completed by the population for each order during the event period and the population proportion;

[0010] Step 3: Reward generation: Based on the step distribution and the maximum amount constraint set by the operation, the final reward form is obtained through solving the programming.

[0011] Preferably, the specific situation of the causal inference grouping module is as follows:

[0012] 1) Model features

[0013] Include weather;

[0014] Driver's individual information;

[0015] The order completion ability includes hundreds of features;

[0016] 2) Uplift model architecture

[0017] Adopt a dual-model difference structure;

[0018] Treatment group prediction model;

[0019] Control group prediction model;

[0020] Uplift value = E[Y|T = 1] - E[Y|T = 0];

[0021] 3) Driver dynamic clustering strategy

[0022] Segment the driver population based on the predicted uplift value obtained from the model;

[0023] Responsive type: uplift value > α;

[0024] Natural type: β < uplift value ≤ α;

[0025] Dormant type: uplift value ≤ β.

[0026] Preferably, the specific situation of the demand prediction module is as follows:

[0027] Model: The regression model of XGBoost;

[0028] Features

[0029] 1) Include weather;

[0030] 2) Driver's individual information;

[0031] 3) The order completion ability includes hundreds of features.

[0032] Preferably, the specific situation of the operations research optimization module is as follows:

[0033] 1) Decision variables:

[0034] x_ik: Whether to set a step for the i-th type of driver population in the k-th order;

[0035] M_ik: The reward amount for the i-th type of driver population in the k-th order;

[0036] 2) Objective function:

[0037] maxΣ(Q_ik*r - M_ik*x_ik);

[0038] ΔQ: The completed order volume of the i - type driver population for the k - th order, r: Average ASP per order;

[0039] 3) Constraints:

[0040] (1) Σc_ik*x_ik*ΔQ_ik ≤ B, total reward budget;

[0041] (2) x_ik ≤ M*z_k, step continuity constraint, the steps need to satisfy an increasing trend;

[0042] (3) ΣM_ik ≤ MT, total step amount constraint;

[0043] (4) Σx_ik ≤ 5, the number of steps cannot exceed 5.

[0044] Preferably, the method for generating rewards for online car - hailing drivers based on causal inference and operations research optimization further includes linear programming solution. After predicting the order - completion steps of the population, linear programming solution is performed, and the linear programming solution is based on the budget subsidy amount and average reward per order.

[0045] Preferably, the method for generating rewards for online car - hailing drivers based on causal inference and operations research optimization further includes effect evaluation. After generating the rewards, based on the operation suggestions and optimizations, the reward activity data is fed back to the constructed uplift model.

[0046] Compared with the prior art, the beneficial effects of the present invention are:

[0047] The uplift model is introduced for population segmentation, solving the difficulty that rewards cannot be refined to different populations. The prediction model is introduced to better control the reward budget, and a more reasonable reward method is generated through the method of planning mathematical solution. Combining the above - mentioned elasticity model and prediction model, a more ROI - oriented reward is generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is the system logic block diagram of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0051] Embodiment 1:

[0052] Please refer to Figure 1 , the present invention provides a technical solution: a method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization. The reward generation method is based on a causal inference grouping module, a demand prediction module, and an operations research optimization module, and an electrical connection relationship is established between the causal inference grouping module, the demand prediction module, and the operations research optimization module;

[0053] The specific steps of the method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization are as follows:

[0054] Step 1: Population division: Based on the historical drivers' behavior of chasing order rewards, an uplift model is constructed to divide the drivers into groups, generating responsive (reward-sensitive), natural (somewhat sensitive), and dormant (insensitive) types;

[0055] Step 2: Prediction of the completion order steps for the population: Based on a tree model, a regression model is constructed to predict the number of completed orders and the population proportion for each order during the activity period;

[0056] Step 3: Reward generation: Based on the step distribution and the maximum amount constraint set by the operation, the final reward form is obtained through programming solution.

[0057] Step 3: Reward generation: Based on the step distribution and the maximum amount constraint set by the operation, the final reward form is obtained through programming solution.

[0058] Effect After the reward is generated, based on the suggestions and optimizations of the operation, the reward activity data is fed back to the constructed uplift model.

[0059] Dynamic response clustering technology:

[0060] Causal forest algorithm integrating counterfactual prediction;

[0061] Driver responsiveness stratification model with real-time update;

[0062] Population step prediction:

[0063] Using population proportion prediction is more stable and controllable;

[0064] Combining multiple prediction models: tree models, linear regression, etc.;

[0065] More reasonable prediction for sparsity;

[0066] Elastic step design mechanism:

[0067] A strategy highly integrated with operations;

[0068] Adaptive step threshold adjustment algorithm.

[0069] Example two:

[0070] Please refer to Figure 1 , the present invention provides a technical solution: the specific situation of the causal inference grouping module is as follows:

[0071] 1) Model features

[0072] Including weather;

[0073] Driver's individual information;

[0074] The order completion ability combines hundreds of features;

[0075] 2) Uplift model architecture

[0076] Adopting a dual-model differential structure;

[0077] Treatment group prediction model;

[0078] Control group prediction model;

[0079] Uplift value = E[Y|T = 1] - E[Y|T = 0];

[0080] 3) Driver dynamic clustering strategy

[0081] Segmenting the driver population based on the predicted uplift value obtained from the model;

[0082] Responsive: uplift value > α;

[0083] Natural: β < uplift value ≤ α;

[0084] Dormant: uplift value ≤ β.

[0085] The specific situation of the demand prediction module is as follows:

[0086] Model: a regression model of XGBoost;

[0087] Features

[0088] 1) Including weather;

[0089] 2) Driver's individual information;

[0090] 3) The comprehensive order-completion ability consists of hundreds of features.

[0091] The specific details of the operation research and optimization module are as follows:

[0092] 1) Decision variables:

[0093] x_ik: Whether to set a step for the i-th type of driver group for the k-th order;

[0094] M_ik: The reward amount for the i-th type of driver group for the k-th order;

[0095] 2) Objective function:

[0096] maxΣ(Q_ik*r - M_ik*x_ik);

[0097] ΔQ: The order-completion volume of the i-th type of driver group for the k-th order, r: Average ASP per order;

[0098] 3) Constraints:

[0099] (1) Σc_ik*x_ik*ΔQ_ik ≤ B, total reward budget;

[0100] (2) x_ik ≤ M*z_k, step continuity constraint, the steps need to satisfy an increasing trend; (3) ΣM_ik ≤ MT, total step amount constraint;

[0101] (4) Σx_ik ≤ 5, the number of steps cannot exceed 5.

[0102] The above shows and describes the basic principles, main features and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0103] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating rewards for online car-hailing drivers based on causal inference and operational research optimization, characterized in that The reward generation method is based on a causal inference grouping module, a demand prediction module, and an operations research optimization module, and an electrical connection relationship is established among the causal inference grouping module, the demand prediction module, and the operations research optimization module; The specific steps of the online car-hailing driver reward generation method based on causal inference and operations research optimization are as follows: Step 1: Population division: Based on the historical drivers' behavior of rushing for order rewards, build an uplift model to divide the drivers into three categories: responsive, natural, and dormant; Step 2: Prediction of the order completion threshold for the population: Based on a tree model, build a regression model to predict the number of completed orders and the population proportion for each order during the event; Step 3: Reward generation: Based on the threshold distribution and the maximum amount constraint set by the operation, obtain the final reward form through solving the optimization problem.

2. The method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization according to claim 1, wherein: The specific situation of the causal inference grouping module is as follows: 1) Model features Include weather; Driver individual information; Comprehensive order completion ability with hundreds of features; 2) Uplift model architecture Adopt a dual-model differential structure; Treatment group prediction model; Control group prediction model; Uplift value = E[Y|T = 1] - E[Y|T = 0]; 3) Driver dynamic clustering strategy Based on the predicted uplift value obtained from the model, divide the driver population; Responsive: uplift value > α; Natural: β < uplift value ≤ α; Dormant: uplift value ≤ β.

3. A method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization according to claim 1, characterized in that: The specific situation of the demand prediction module is as follows: Model: A regression model of XGBoost; Features 1) Include weather; 2) Driver individual information; 3) Comprehensive order completion ability with hundreds of features.

4. A method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization according to claim 1, characterized in that: The specific situation of the operations research optimization module is as follows: 1) Decision variables: x_ik: Whether to set a threshold for the i-th type of driver population in the k-th order; M_ik: The reward amount for the i-th type of driver population in the k-th order; 2) Objective function: maxΣ(Q_ik*r - M_ik*x_ik); ΔQ: The number of completed orders for the i-th type of driver population in the k-th order, r: Average ASP per order; 3) Constraint conditions: (1) Σc_ik*x_ik*ΔQ_ik ≤ B, total reward budget; (2) x_ik ≤ M*z_k, threshold continuity constraint, the threshold needs to be increasing; (3) ΣM_ik ≤ MT, total threshold amount constraint; (4) Σx_ik ≤ 5, the number of thresholds cannot exceed 5.

5. The method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization according to claim 1, characterized in that: The online car-hailing driver reward generation method based on causal inference and operations research optimization also includes linear programming solution. After predicting the order completion threshold for the population, linear programming solution is performed based on the budget subsidy amount and average reward per order.

6. The method for generating rewards for online car-hailing drivers based on causal inference and operations research optimization according to claim 1, characterized in that: The online car-hailing driver reward generation method based on causal inference and operations research optimization also includes effectiveness evaluation. After reward generation, based on the operation's suggestions and optimizations, the reward activity data is fed back to the built uplift model.