Uplift-Tree-based driver marketing sensitivity estimation method and system
Through the Uplift-Tree model, a decision tree is built using the driver data set to quantify marketing sensitivity, and the problems of hypothetical dependence and confounding factors in the online ride-hailing platform are solved, more accurate marketing strategies and resource optimization are achieved, and marketing effectiveness and platform revenue are improved.
Patent Information
- Application Number
- CN202510304022.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technology has uncertainty in hypothetical dependence and complexity in the processing of confounding factors in online ride-hailing platforms, resulting in insufficient accuracy in marketing activity prediction, affecting the scientificity and effectiveness of decision-making.
Using the Uplift-Tree model, the driver data set is collected, the improvement value is calculated, and the KL divergence and decision tree technology is used to build the model, capture key factors, quantify marketing sensitivity, and reduce the impact of hypothesis dependence and confounding factors.
It improves the response rate and conversion rate of marketing activities, optimizes resource allocation, enhances the accuracy and resource utilization efficiency of marketing strategies, reduces the dependence of the model on assumptions, and improves the accuracy and reliability of the estimated results.
Smart Images

Figure CN120259061A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of big data analysis and machine learning, and is directed to calculating the impact of various marketing activities on driver activity indicators, and particularly relates to a method and system for predicting driver marketing sensitivity based on Uplift-Tree. Background Art
[0002] Online reservation taxi service (referred to as online car-hailing) is defined as a new transportation mode that relies on the Internet platform to realize vehicle reservation and rental services. The core of this service mode is to use a smartphone application as a medium. Through this, passengers can initiate a ride request, and drivers can receive orders and provide corresponding services through the same platform. With the rapid development of mobile Internet technology and big data analysis, online car-hailing platforms have gradually become the preferred way for the public to travel. In order to enhance the platform's attractiveness, promote more drivers to join, and improve overall competitiveness, the design and implementation of precise marketing activities are particularly crucial. However, one of the current major challenges is how to accurately predict drivers' responses to various marketing activities.
[0003] The existing technical framework mainly focuses on two major models, S-learn and T-learn. For example, CN202311795965.X discloses an intelligent driver selection algorithm for tenant marketing activities (publication date: April 30, 2024), which is mainly based on the S-learn and T-learn methods; as Figure 6 shown, the basic information of drivers, driving behavior data, and real-time data on vehicle usage are collected from the online car-hailing platform through a data interface. Using feature variables, the S-learn and T-learn algorithms are used to model drivers' marketing responses, and a prediction model for drivers' marketing responses is constructed. The real-time data is input into the trained S-learn and T-learn models to predict drivers' responses to marketing activities. Although these two models have shown significant advantages in causal inference and personalized strategy formulation, their applications still face several limitations and challenges:
[0004] (1) Uncertainty of assumption dependence: The effectiveness and accuracy of the S-learn and T-learn models highly depend on a series of theoretical assumptions, such as the monotonicity assumption of causal relationships and the statistical independence assumption, etc. Unfortunately, in the complex and ever-changing real environment, these assumption conditions may not be fully met, thereby weakening the actual prediction ability of the models.
[0005] (2) Complexity in dealing with confounding factors: At the practical operation level, there are many confounding factors that are difficult to observe or not taken into consideration, and their potential impact on the marketing intervention effect cannot be ignored. The S-learn and T-learn models need to have the ability to effectively handle these confounding factors. Otherwise, it is very easy to lead to bias in the estimation results, affecting the scientific nature and effectiveness of decision-making.
[0006] For this reason, the present invention proposes a method and system for estimating the marketing sensitivity of drivers based on Uplift-Tree. Summary of the Invention
[0007] In view of this, the present invention hopes to provide a method and system for estimating the marketing sensitivity of drivers based on Uplift-Tree to solve or alleviate the technical problems existing in the prior art, that is, to solve the problems of the uncertainty of hypothesis dependence and the complexity of dealing with confounding factors, and to provide at least one beneficial option for this; the technical solution of the present invention is realized as follows:
[0008] In the first aspect, a method for estimating the marketing sensitivity of drivers based on Uplift-Tree:
[0009] (1) Overview:
[0010] The present invention aims to accurately estimate the marketing sensitivity of online car-hailing drivers using the Uplift-Tree model. By collecting a dataset containing driver characteristics, treatment indicators, and response variables, calculate the uplift value for each driver sample, that is, the difference between the potential response value and the actual response value multiplied by the treatment indicator variable. Then, using the dataset as the root node, calculate the distribution difference between the intervention samples and the non-intervention samples based on the KL divergence. Through feature selection and data splitting, recursive growth of the decision tree, and pruning strategies, construct the Uplift-Tree model. This model can capture the key factors affecting the marketing sensitivity of drivers and, based on the input of driver characteristics, output their expected response to marketing activities. By maximizing the uplift value to select the best features and splitting points, ensure that the model can accurately identify the group of drivers with high sensitivity to marketing activities, thereby providing accurate decision-making support for marketing activities on the online car-hailing platform, improving marketing effectiveness and resource utilization. This solution aims to achieve accurate estimation of the marketing sensitivity of online car-hailing drivers through a data-driven method, bringing greater commercial value to the platform.
[0011] (2) Technical solution:
[0012] To achieve the above technical objectives, the present invention selects the following operating steps:
[0013] 2.1 Step S1, data preparation:
[0014] Collect the dataset D containing driver sample i, where each driver sample i contains the following variables:
[0015] T i : The treatment indicator for whether to receive the marketing intervention (T i ∈{0, 1})
[0016] Y i : The response variable of a certain behavior or attitude of the driver;
[0017] X i : The covariate of the driver's characteristics.
[0018] 2.2 Step S2, calculate the uplift value:
[0019] For each driver sample i in the dataset D, calculate the difference between its potential response value (if treated) and the actual response value (if not treated), and multiply it by the treatment indicator variable to obtain the uplift value U i :
[0020] U i =(Y i (1)-Y i (0))×T i ;
[0021] where, Y i (1) represents the response value when driver i is treated, and Y i (0) represents the response value when not treated.
[0022] 2.3 Step S3, calculate the response:
[0023] After normalizing the response degree R i , calculate the ratio of the maximum uplift value U max of the job to obtain the response degree R i value:
[0024] If the uplift value U i is negative, indicating that the marketing activity has a negative impact on the driver, then the response degree R i can be set to 0, or other processing can be performed according to the specific situation.
[0025] 2.4 The training method of the above Uplift-Tree model:
[0026] Use the dataset D as the root node, calculate the distribution difference between the intervention sample T i =1 and the non-intervention sample T i =0 based on the KL divergence, and then perform the following training steps:
[0027] 2.4.1 First step, feature selection and data segmentation:
[0028] Based on the method of ordered clustering, select the key feature X that affects the lift value j , where j represents the index of the feature, find the splitting point a for each feature X j , and divide the data into two child nodes, the left and the right.
[0029] Calculate the distribution differences between the left and right child nodes, and calculate the lift value value(X j ,a) corresponding to the feature X and the splitting point a: j :
[0030] value(X j ,a) = p1×diff1 + p2×diff2 - diff0;
[0031] Among them, p1 and p2 are the proportions of samples in the left and right child nodes respectively, diff1 and diff2 are the distribution differences between the intervened samples and the non - intervened samples in the left and right child nodes respectively, and diff0 is the distribution difference between the intervened samples and the non - intervened samples in the root node.
[0032] 2.4.1.1 Calculate the within - class scatter matrix and the between - class scatter matrix:
[0033] Suppose the dataset D is divided into two subsets D1 and D2 by a splitting point a on a certain feature dimension, where D1 contains all samples less than or equal to a, and D2 contains all samples greater than a.
[0034] The within - class scatter matrix S W and the between - class scatter matrix S B are respectively defined as:
[0035]
[0036] S B = (μ1 - μ2)(μ1 - μ2) T ;
[0037] Among them, μ1 and μ2 are the mean vectors of D1 and D2 respectively, and T represents the transpose operation.
[0038] 2.4.1.2 Calculate the Fisher criterion function:
[0039] The Fisher criterion function J(a) is defined as the ratio of the determinant of the between - class scatter matrix S W to the determinant of the within - class scatter matrix S B :
[0040] Among them, and They are the variances of D1 and D2 respectively.
[0041] 2.4.1.3 Finding the optimal splitting point:
[0042] Traverse all splitting points a, calculate the Fisher criterion function J(a), and select the a that maximizes J(a) as the optimal splitting point.
[0043] In the training of the Uplift - Tree model, the above process can be applied to each feature dimension separately to find the optimal splitting point for each feature. Then, based on these splitting points, a decision tree is constructed, and at each node, the feature and splitting point with the maximum uplift value are selected for splitting.
[0044] 2.4.2 Second step, growing the tree:
[0045] According to the principle of maximizing the uplift value, select the best feature and splitting point, and recursively grow the decision tree. At each step, consider using different splitting points on different features of the current node and select the splitting that maximizes the uplift value.
[0046] 2.4.3 Third step, pruning:
[0047] Adopt pre - pruning or post - pruning strategies to prevent the model from overfitting. Pre - pruning stops the growth of the tree during the tree - growing process when the sample size of a certain node is less than a predetermined threshold or the uplift value gain is less than a predetermined threshold. Post - pruning, on the other hand, deletes those nodes that contribute little to the uplift value from the bottom after the tree is fully grown.
[0048] (III) Mechanism for solving technical problems:
[0049] 3.1 Causal inference - based framework:
[0050] The Uplift - Tree model is essentially a causal inference method. It focuses on the causal effect of an intervention (such as a marketing campaign) on a target variable (such as driver behavior or attitude), rather than simple correlation. This framework enables the model to more accurately estimate the effect of the intervention in the presence of confounding factors. By comparing the changes of the intervention group (drivers who receive marketing intervention) and the control group (drivers who do not receive marketing intervention) before and after the intervention, the model can isolate the influence of other potential confounding factors and focus on evaluating the effect of the marketing campaign itself.
[0051] 3.2 Flexible feature selection and data splitting:
[0052] The Uplift-Tree model recursively performs binary splitting on the training data to maximize the uplift metric for child nodes (such as the distribution difference based on KL divergence). During the feature selection and data splitting process, the model can automatically identify the features that significantly affect the uplift value and their optimal splitting points. This flexibility enables the model to capture key information in complex data environments, reducing the uncertainty of hypothesis dependence. At the same time, by continuously subdividing the data, the model can gradually strip out confounding factors and improve the estimation accuracy of causal effects.
[0053] 3.3 Pruning strategies to prevent overfitting:
[0054] To prevent the model from overfitting, the Uplift-Tree model adopts pre-pruning or post-pruning strategies. Pre-pruning stops the growth of the tree in advance during tree growth, halting the splitting when the sample size of a node is less than a predetermined threshold or the uplift value gain is less than a predetermined threshold. Post-pruning, on the other hand, deletes the nodes that contribute little to the uplift value from the bottom after the tree has fully grown. These pruning strategies help reduce the overfitting of the model to the training data, improve the generalization ability of the model, and thus be more robust when dealing with complex data.
[0055] 3.4 Quantifying the uplift effect to guide precision marketing:
[0056] Finally, the Uplift-Tree model quantifies the sensitivity of different driver groups to marketing activities by calculating the uplift effect value for each leaf node. This quantification index provides valuable decision-making support for the online car-hailing platform, enabling the platform to formulate more precise and effective marketing strategies for driver groups with high sensitivity. In this way, the platform not only improves the utilization efficiency of marketing resources but also enhances the user experience and satisfaction.
[0057] Second aspect, a driver marketing sensitivity prediction system based on Uplift-Tree:
[0058] As Figure 2 shown, this system is used to implement the driver marketing sensitivity prediction method based on Uplift-Tree described above, and it includes:
[0059] (1) A data preparation layer responsible for collecting a dataset containing driver samples, including:
[0060] A dataset collection module: Each driver sample contains a treatment indicator (whether to receive marketing intervention), a response variable (a certain behavior or attitude of the driver), and covariates (driver characteristics).
[0061] (2) A model construction layer that executes the Uplift-Tree model, including:
[0062] Feature Selection and Preprocessing Module for Feature Selection and Preprocessing of the Collected Dataset: Select features that have a significant impact on the uplift value and perform necessary data cleaning and transformation.
[0063] Uplift-Tree Model Training Module for Model Training: Based on the preprocessed dataset, use the Uplift-Tree algorithm for model training. This module includes sub-modules such as feature selection and data splitting, decision tree growth, and pruning. By recursively splitting the data, a decision tree model that maximizes the uplift value is constructed.
[0064] (3) Application Layer for Outputting the Prediction Results of the Uplift-Tree Model:
[0065] Marketing Strategy Formulation Module for Formulating Marketing Strategies for the Online Car-Hailing Platform: This module divides the driver group into different market segments according to the sensitivity of drivers to marketing activities, and formulates corresponding marketing plans and budget allocation schemes.
[0066] Marketing Effect Monitoring Module for Real-Time Monitoring and Feedback of the Implementation Effect of Marketing Activities: Collect actual marketing data and conduct comparative analysis with the model prediction results to verify the practicality and effectiveness of the model, and provide data support for further optimizing the model.
[0067] Compared with the prior art, the beneficial effects of the present invention are:
[0068] I. Improve Marketing Effect: By accurately estimating the marketing sensitivity of drivers, the online car-hailing platform of the present invention can formulate more personalized and effective marketing strategies for high-sensitivity driver groups. This helps to improve the response rate and conversion rate of marketing activities, thereby increasing the platform's revenue and market share.
[0069] II. Optimize Resource Allocation: The present invention can quantify the expected response of different driver groups to marketing activities, enabling the platform to allocate marketing resources more reasonably. By focusing on marketing high-sensitivity groups, the platform can avoid resource waste and improve resource utilization efficiency.
[0070] III. Reduce Dependence on Assumptions: The Uplift-Tree model of the present invention is based on a causal inference framework and can accurately estimate the causal effect of marketing activities in the presence of confounding factors. This reduces the model's dependence on assumptions and improves the accuracy and reliability of the estimation results.
[0071] IV. Handle Complex Data Environments: Through a flexible feature selection and data splitting mechanism, the present invention can capture key information in complex data environments. This enables the model to adapt to the characteristic differences of different driver groups and improve the accuracy and generalization ability of the estimation. Description of the Drawings
[0072] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0073] Figure 1 It is a schematic flowchart of the method of the present invention;
[0074] Figure 2 It is a schematic diagram of the system composition of the present invention;
[0075] Figure 3 It is a schematic diagram of feature selection and data segmentation of the Uplift-Tree model of the present invention;
[0076] Figure 4 It is a schematic diagram of the growth tree of the Uplift-Tree model of the present invention;
[0077] Figure 5 It is a schematic diagram of tree pruning of the Uplift-Tree model of the present invention;
[0078] Figure 6 It is a schematic diagram of the architectures of the two traditional models of S-learn and T-learn. Detailed implementation manners
[0079] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific implementation manners of the present invention with reference to the drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below;
[0080] It should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0081] Embodiment 1: As Figure 1 shown, this embodiment discloses a method for estimating the marketing sensitivity of drivers based on the Uplift-Tree, including a pre-trained Uplift-Tree model, and performing the following operation steps:
[0082] In this embodiment, regarding step S1: Data Preparation: The management personnel need to extract relevant data of the drivers from the database of the online car-hailing platform in order to conduct spot checks on a certain driver and judge his response to the marketing activities.
[0083] Processing Index T i : Whether to accept marketing intervention. For example, Ti = 1 indicates that the driver has received a marketing activity (such as coupons, additional rewards, etc.), and Ti = 0 indicates that he has not.
[0084] Response Variable Y i : The driver's behavior or attitude. Specifically, it may include:
[0085] (1) Driving Duration: The driver's daily driving time.
[0086] (2) Idle Driving Duration: The driver's daily idle driving time without passengers.
[0087] (3) Number of Completed Orders: The number of orders completed by the driver every day.
[0088] (4) Number of Cancelled Orders: The number of orders cancelled by the driver every day.
[0089] (5) Number of Complaints: The number of complaints received by the driver every day.
[0090] (6) Order Revenue: The total daily order revenue of the driver.
[0091] (7) Activity Revenue: The additional revenue obtained by the driver from the marketing activities.
[0092] Covariate X i : The driver's characteristics. Specifically, it includes:
[0093] (1) Registration Days: The length of the driver's registration time on the platform.
[0094] (2) City: The city where the driver is located.
[0095] (3) Gender: The driver's gender.
[0096] (4) Service Score: The service score of the driver on the platform.
[0097] The specific composition of the above variables can be referred to in the following table:
[0098]
[0099]
[0100] In this embodiment, regarding step S2: Calculate the uplift value: For the sampled drivers, we need to calculate the difference in response values between the cases where they receive the marketing campaign and where they do not, that is, the uplift value.
[0101] It is necessary to use the Uplift-Tree model or other causal inference models to train the entire dataset D to estimate the response values of each driver when receiving treatment (T i = 1) and when not receiving treatment (T i = 0). This usually involves steps such as feature selection, data splitting, decision tree growth, etc. For details, see Embodiment 2.
[0102] For the sampled driver i, use the trained model to estimate its response value Y i when receiving treatment (T i = 1) Y i and its response value Y i when not receiving treatment (T i = 0)
[0103] U i = (Y i (1) - Y i (0)) × T i ;
[0104] where Y i (1) represents the response value of driver i when receiving treatment, and Y i (0) represents the response value when not receiving treatment. If the driver receives treatment (T i = 1), then U i is equal to the difference in response values between receiving treatment and not receiving treatment; if the driver does not receive treatment (T i = 0), then Ui is 0.
[0105] It can be understood that the Uplift-Tree model can more accurately estimate the treatment effect by recursively splitting the data, considering multiple features and their interactions, thus reducing the dependence on assumptions. The model will consider confounding factors (i.e., factors that simultaneously affect the treatment indicator and the response variable) during the training process. By adjusting the influence of these factors, it can more accurately estimate the treatment effect, thus solving the problem of dealing with confounding factors.
[0106] In this embodiment, regarding step S3: Calculate the response degree: After obtaining the uplift value, the management personnel need to calculate the response degree of the driver to more intuitively understand the response of the driver to the marketing campaign.
[0107] First, for the uplift value U iPerform normalization to convert it into a value between 0 and 1. Calculate the driver's response degree R based on the normalized uplift value. i . If the uplift value is negative, indicating that the marketing activity has a negative impact on the driver, R i can be set to 0 or other processing can be performed according to specific circumstances:
[0108] If the uplift value U i is negative, indicating that the marketing activity has a negative impact on the driver, then the response degree R i can be set to 0, or other processing can be performed according to specific circumstances.
[0109] Exemplary: Suppose the uplift value Ui of a certain driver is 30 (normalized), and the maximum uplift value of all drivers is 100, then the response degree Ri of this driver = 30 / 100 = 0.3.
[0110] It can be understood that by calculating the response degree, managers can more intuitively understand the drivers' responses to marketing activities, and thus make more accurate decisions. The response degree can be an important indicator for evaluating drivers' sensitivity to marketing activities, helping managers formulate more personalized marketing strategies.
[0111] Furthermore, the following is the Java program for the above execution method:
[0112]
[0113]
[0114]
[0115]
[0116] In the above program, the basic data of the driver is defined in the DriverData class, including the driver ID, processing index Ti, and response variable responseValue. A driver data set is created in the main method. The calculateUpliftValue method performs the calculation of the uplift value. In practical applications, the calculation of the uplift value should be based on the Uplift-Tree model or other causal inference models.
[0117] The calculateResponseDegree method calculates the response degree based on the uplift value. First, the uplift value is normalized, and then it is compared with the maximum uplift value to obtain a value between 0 and 1. Finally, the program outputs the response degree of each driver.
[0118] Example 2: AsFigures 3 to 5 As shown in Figures 3 to 5 , this embodiment further discloses a training method for the Uplift-Tree model as described in Embodiment 1:
[0119] Suppose the dataset D contains the historical data of drivers, such as the number of registration days, service score, city where they are located, gender, historical driving duration, and historical number of completed orders. The intervention sample T i = 1: The driver sample that receives the marketing intervention. The non-intervention sample T i = 0: The driver sample that does not receive the marketing intervention.
[0120] Construct a Uplift-Tree model to predict the response degree of drivers after receiving the marketing intervention and optimize the marketing strategy. Take the dataset D as the root node, and calculate the distribution difference between the intervention sample T i = 1 and the non-intervention sample T i = 0 based on the KL divergence, and then perform the following training steps:
[0121] (The first step) Feature selection and data segmentation:
[0122] Based on the method of ordered clustering, select the key feature X j that affects the uplift value, where j represents the index of the feature, find the splitting point a of each feature X j , and divide the data into two child nodes on the left and right.
[0123] For example, for the feature of the number of registration days, select a number of days threshold as the splitting point. Calculate the distribution difference between the intervention sample and the non-intervention sample in the left and right child nodes (such as using the KL divergence). Calculate the uplift value value(Xj,a) corresponding to the feature Xj and the splitting point a according to the formula. This value reflects the average improvement effect of the intervention on the driver's response at this splitting point, and calculate the uplift value value(X j and the splitting point a corresponding to it: j value(X
[0124] j ,a) = p1 × diff1 + p2 × diff2 - diff0;
[0125] Among them, p1 and p2 are the proportions of samples in the left and right child nodes respectively, diff1 and diff2 are the distribution differences between the intervention sample and the non-intervention sample in the left and right child nodes respectively, and diff0 is the distribution difference between the intervention sample and the non-intervention sample in the root node.
[0126] Calculate the within-class scatter matrix and the between-class scatter matrix: Suppose the dataset D is divided into two subsets D1 and D2 by the splitting point a on a certain feature dimension, where D1 contains all samples less than or equal to a, and D2 contains all samples greater than a.
[0127] Within-class scatter matrix S W and between-class scatter matrix S B are respectively defined as:
[0128]
[0129] S B =(μ1 - μ2)(μ1 - μ2) T ;
[0130] where μ1 and μ2 are the mean vectors of D1 and D2 respectively, and T is the transpose operation.
[0131] The Fisher criterion function J(a) is defined as the ratio of the determinant of the between-class scatter matrix S W to the determinant of the within-class scatter matrix S B : where, and are the variances of D1 and D2 respectively. The Fisher criterion function J(a) measures the ratio of the between-class difference to the within-class difference. The larger J(a) is, the better the splitting point can distinguish the intervention samples and non-intervention samples.
[0132] Traverse all splitting points a, calculate the Fisher criterion function J(a), and select the a that maximizes J(a) as the optimal splitting point, which will be used to partition the data in the decision tree..
[0133] In the training of the Uplift-Tree model, the above process can be applied to each feature dimension separately to find the optimal splitting point for each feature. Then, based on these splitting points, a decision tree is constructed, and at each node, the feature and splitting point with the maximum lift value are selected for splitting.
[0134] (Step 2) Grow the tree:
[0135] Recursively grow the decision tree: Starting from the root node, select the best feature and splitting point for splitting according to the principle of maximizing the lift value. Repeat the above process for each newly generated child node until the stopping condition is met (such as the sample size in the node is less than the predetermined threshold or the lift value gain is less than the predetermined threshold).
[0136] Handle the uncertainty of hypothesis dependence and confounding factors: By recursively splitting the data, the Uplift-Tree model can consider multiple features and their interactions, thereby reducing the dependence and uncertainty on a single feature. At each node, the model will select the best splitting point according to the current data distribution, which helps to handle confounding factors because confounding factors usually affect multiple features.
[0137] (Step 3) Pruning:
[0138] Adopt pre-pruning or post-pruning strategies to prevent model overfitting.
[0139] (1) Pre-pruning strategy: During the tree growth process, when the sample size of a certain node is less than the predetermined threshold or the lift gain is less than the predetermined threshold, stop splitting the node. This helps prevent model overfitting and reduce computational complexity.
[0140] (2) Post-pruning strategy: After the tree is fully grown, traverse the tree from the bottom and delete the nodes that contribute little to the lift value. This can further simplify the model and improve the generalization ability of the model.
[0141] It can be understood that by recursively splitting the data and selecting the best features and split points, the Uplift-Tree model can consider multiple features and their interactions, thereby reducing the dependence and uncertainty on a single feature. The model performs splitting at each node according to the current data distribution, which helps to handle confounding factors because confounding factors usually affect multiple features and their effects may vary in different subgroups. By subdividing the data, the model can better capture these effects and reduce the impact of confounding factors.
[0142] By analyzing the historical data and response situations of drivers, the Uplift-Tree model can predict which drivers are more likely to increase their driving hours or the number of completed orders after receiving marketing interventions. This helps the company formulate more precise marketing strategies, allocate resources to the drivers who are most likely to have a positive response, thereby improving the marketing effect and return on investment.
[0143] Furthermore, the Python execution program of the above training method is as follows:
[0144]
[0145]
[0146]
[0147]
[0148]
[0149] In the above program, load the DataFrame containing driver data from the CSV file. Then use the train_test_split function to split the data into a training set and a test set. Grow the Uplift-Tree into the UpliftTree class, which contains the logic for growing the tree. At each node of the tree, we traverse all features and try different split points to calculate the lift value.
[0150] Select the feature and split point with the maximum uplift value for splitting and recursively grow the tree until the maximum depth is reached or the sample size is less than a predetermined threshold.
[0151] Finally, use the pickle module to save the trained Uplift-Tree model as a file. When needed, the model can be loaded from the file for prediction or other operations.
[0152] All of the above embodiments only express the implementation manners of the relevant practical applications of the present invention. The descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
[0153] For those skilled in the art, it can be further realized that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0154] Meanwhile, those skilled in the art can understand that all or part of the processes in the methods of all the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
Claims
1. A method for estimating the marketing sensitivity of drivers based on Uplift-Tree, characterized in that Including: S1, collecting a dataset D containing driver sample i; S2. For each driver sample i in the dataset D, use the Uplift-Tree model to calculate the difference between its potential response value and the actual response value, and obtain the uplift value U i ; S3. Normalize the response degree R i and calculate the maximum improvement value U of the job max by taking the ratio of the response degree R i value.
2. The driver marketing sensitivity prediction method according to claim 1, wherein: In the S1, each driver sample i includes the following variables: T i : Processing indicator for whether to accept marketing intervention; Y i : Response variable for a certain behavior or attitude of the driver; X i : Covariates of the driver's characteristics.
3. The driver marketing sensitivity estimation method according to claim 1, characterized in that: In the S2, the lifting value U i is calculated as: U i = (Y i (1) - Y i (0)) × T i ; Among them, Y i (1) represents the response value when driver i is being processed, Y i (0) represents the response value when not being processed.
4. The driver marketing sensitivity estimation method according to claim 3, characterized in that: In the S2, the training method of the Uplift-Tree model includes: Step 1, Feature Selection and Data Splitting: Select the key feature X that affects the lift value j , where j represents the index of the feature, find the splitting point a for each feature X j , and divide the data into two left and right child nodes; calculate the distribution differences between the left and right child nodes, and calculate the lift value value(X j , a) corresponding to the feature X j , a); The second step, growing the tree: According to the principle of maximizing the uplift value, select the best feature and splitting point, and recursively grow the decision tree; The third step, pruning: Adopt the post-pruning strategy to prevent the model from overfitting.
5. The driver marketing sensitivity estimation method according to claim 4, characterized in that: In the first step, the promotion value value(X j , a) = p1 × diff1 + p2 × diff2 - diff0; Where, p1 and p2 are respectively the proportions of samples in the left and right child nodes, diff1 and diff2 are respectively the distribution differences between the intervention samples and non-intervention samples in the left and right child nodes, and diff0 is the distribution difference between the intervention samples and non-intervention samples in the root node.
6. The driver marketing sensitivity estimation method according to claim 4, characterized in that: In the first step, the selection method is ordered clustering: Suppose the dataset D is divided into two subsets D1 and D2 by a splitting point a on a certain feature dimension, where D1 contains all samples less than or equal to a, and D2 contains all samples greater than a; Within-class scatter matrix $S$ W and between-class scatter matrix $S_B$ B are respectively defined as: S B =(μ1 - μ2)(μ1 - μ2) T ; Where, μ1 and μ2 are respectively the mean vectors of D1 and D2, and T is the transpose operation; The Fisher criterion function J(a) is the ratio of the determinant of the between-class scatter matrix S W to the determinant of the within-class scatter matrix S B : wherein, and are the variances of D1 and D2 respectively; Traverse all splitting points a, calculate the Fisher criterion function J(a), and select the a that maximizes J(a) as the optimal splitting point.
7. The driver marketing sensitivity prediction method according to claim 4, characterized in that: In the third step, the post-pruning strategy is to delete those nodes with little contribution to the uplift value from the bottom after the tree grows completely.
8. The driver marketing sensitivity estimation method according to claim 1, characterized in that: In the step S3, the response degree R i is calculated as follows:
9. A system for implementing the driver marketing sensitivity prediction method according to any one of claims 1 to 8, characterized in that, The system includes: A data preparation layer responsible for collecting a dataset containing driver samples; A model construction layer that executes the Uplift-Tree model; An application layer that outputs the prediction results of the Uplift-Tree model.
10. The system according to claim 9, characterized in that: The model construction layer includes: A feature selection and preprocessing module for performing feature selection and preprocessing on the collected dataset; A Uplift-Tree model training module for model training The application layer includes: A marketing strategy formulation module for formulating marketing strategies for the online car-hailing platform; A marketing effect monitoring module for real-time monitoring and feedback on the implementation effect of marketing activities.
Citation Information
Patent Citations
Intelligent driver circling algorithm under tenant marketing activity
CN117952729A