Service gain prediction method, model training method, decision-making method and device
By using a deep gain model and a Lagrange dual optimization method, the problem of insufficient causal effect identification in existing business subsidy strategies is solved, achieving optimal resource allocation under budget constraints and improving the scientific nature and economic efficiency of the subsidy strategy.
Patent Information
- Application Number
- CN202511820989.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-02-03
AI Technical Summary
Existing business subsidy strategies typically employ fixed subsidies or simple rule-based allocation methods, which fail to accurately identify the causal effects of subsidies on user behavior. This leads to unscientific resource allocation, potentially resulting in insufficient incentives or losses.
By employing a deep gain model combined with feature selection and a propensity score network, and through multidimensional feature extraction and propensity result correction, we predict the potential outcomes and gains under different intervention methods. We also optimize subsidy allocation by combining the Lagrange dual optimization method.
This approach maximizes overall benefits under budget constraints, improves the scientific and accurate allocation of resources, avoids waste of funds, and enhances the economic efficiency of subsidy strategies.
Smart Images

Figure CN121457743A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to the fields of big data and causal inference. Background Technology
[0002] Business subsidy strategies, such as ride-hailing subsidies, typically employ fixed subsidies or simple rule-based allocation methods. Data-driven personalized incentive delivery models, such as binary classification / regression models, can usually only predict the conversion probability of orders or the overall order volume. Incremental effect modeling has been proposed to directly predict the differentiated response of orders under conditions of intervention versus no intervention. Summary of the Invention
[0003] This disclosure provides a method for predicting business gains, a method for training models, a decision-making method, and an apparatus.
[0004] According to one aspect of this disclosure, a method for predicting business gains is provided, comprising: Use distributional differences to extract multidimensional features related to the estimated gain from business decision units; Based on this multidimensional feature, the deep gain model is used to predict the potential and propensity outcomes under different intervention methods; Based on the potential and propensity outcomes, the gains of the business decision-making unit under different intervention methods are obtained.
[0005] According to another aspect of this disclosure, a method for training a deep gain model is provided, comprising: Use distributional differences to extract multidimensional features related to the estimated gain from business samples; Based on this multidimensional feature, a deep gain model that needs to be trained is used to predict the potential and propensity outcomes of different intervention methods. Based on the labeling results, potential results, and propensity results of this business sample, determine the loss function; Based on this loss function, the deep gain model that needs to be trained is adjusted; Under the condition that the training termination criterion is met, the trained deep gain model is obtained.
[0006] According to another aspect of this disclosure, a decision-making method for selecting the optimal intervention method for a business decision-making unit is provided, comprising: The first operations research optimization model is constructed based on the estimated gains, decision objectives, decision variables, and business constraints of business decision-making units under different intervention methods. The first operations research optimization model is subjected to a Lagrange dual transformation to obtain the second operations research optimization model; Solve the second operations research optimization model to obtain the optimal decision for the business decision-making unit, which includes the optimal intervention method.
[0007] According to another aspect of this disclosure, a business gain prediction apparatus is provided, comprising: The extraction module is used to extract multidimensional features related to the estimated gain from the business decision unit using distribution differences; The prediction module is used to predict the potential and propensity outcomes under different intervention methods based on this multidimensional feature using a deep gain model. The gain module is used to obtain the gain of the business decision-making unit under different intervention methods based on the potential outcome and the propensity outcome.
[0008] According to another aspect of this disclosure, a training apparatus for a deep gain model is provided, comprising: The extraction module is used to extract multidimensional features related to the estimated gain from business samples using distribution differences; The prediction module is used to predict the potential and propensity outcomes under different intervention methods based on the multidimensional features using a deep gain model that needs to be trained. The loss calculation module is used to determine the loss function based on the label results, potential results, and propensity results of the business sample; The training module is used to adjust the deep gain model to be trained based on the loss function; and obtains the trained deep gain model if the training termination criterion is met.
[0009] According to another aspect of this disclosure, a decision-making apparatus for selecting the optimal intervention method for a business decision-making unit is provided, comprising: The module is used to build the first operations research optimization model based on the estimated gains, business objectives, decision variables and business constraints of the business decision-making unit under different intervention methods. The transformation module is used to perform a Lagrange dual transformation on the first operations research optimization model to obtain a second operations research optimization model. The solution module is used to solve the second operations research optimization model to obtain the optimal decision corresponding to the business decision unit, which includes the optimal intervention method.
[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0012] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart illustrating a method for predicting service gain according to an embodiment of the present disclosure. Figure 2 This is a flowchart illustrating a method for predicting service gain according to another embodiment of this disclosure; Figure 3 This is a schematic diagram of the structure of a depth gain model according to an embodiment of the present disclosure; Figure 4 This is a flowchart illustrating a training method for a deep gain model according to an embodiment of the present disclosure. Figure 5 This is a flowchart illustrating a training method for a deep gain model according to another embodiment of the present disclosure; Figure 6 This is a flowchart illustrating a training method for a deep gain model according to an embodiment of the present disclosure. Figure 7 This is a flowchart illustrating a business intervention method according to an embodiment of the present disclosure; Figure 8 This is a flowchart illustrating an optimized ride-hailing subsidy method based on a deep uplift model and a Lagrange dual with post-processing, according to an embodiment of this disclosure. Figure 9 This is a schematic diagram of the structure of a service gain prediction device according to an embodiment of the present disclosure; Figure 10 This is a schematic diagram of the structure of a service gain prediction device according to another embodiment of the present disclosure; Figure 11 This is a schematic diagram of the structure of a training device for a depth gain model according to an embodiment of the present disclosure; Figure 12This is a schematic diagram of the structure of a training device for a depth gain model according to another embodiment of the present disclosure; Figure 13 This is a schematic diagram of the structure of a business intervention device according to an embodiment of the present disclosure; Figure 14 This is a schematic diagram of the structure of a service intervention device according to another embodiment of the present disclosure; Figure 15 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0016] In ride-hailing and food delivery service platforms, the platform typically uses subsidies and incentives to regulate the dynamic balance between passenger demand and driver supply. The goals of subsidies can include: increasing order completion rates, reducing wait times, improving user retention rates, and optimizing overall revenue within budget constraints. However, how to scientifically and accurately allocate subsidies is a problem that urgently needs to be solved.
[0017] Subsidies and incentives are problems of "causal inference + operations optimization." Traditional binary classification / regression models can only predict the conversion probability or order volume of users, but cannot identify the "causal effect" of subsidies on user behavior. Incremental effect modeling (such as uplift modeling) can be used to directly predict the differentiated responses of users under conditions of receiving intervention and not receiving intervention. One type of incremental effect modeling can include models based on meta-learners, such as single-model learners (S-Learner), two-model learners (T-Learner), crossover learners (X-learner), etc.; forest-based causal tree models; and models extended by deep learning, etc. Applying uplift models to travel scenarios, such as pricing, coupon distribution, and user subsidies, can achieve more efficient resource allocation by learning the individualized causal effects of users.
[0018] Subsidy resources are typically constrained by budgets, thus requiring the maximization of overall revenue or platform goals within a limited budget. This type of problem can be formalized as a constrained optimization problem or a variant of the Knapsack problem. Based on linear programming or integer programming, each user's subsidy allocation is treated as a decision variable, maximizing the number of completed orders or revenue under budget constraints.
[0019] Heuristic or approximate optimization algorithms, such as genetic algorithms and Lagrange relaxation algorithms, can be used to solve operational optimization problems involving large-scale decision variables. Reinforcement learning frameworks can be used to model subsidy allocation as a dynamic programming problem, maximizing long-term returns under budget constraints. Tree-based uplift models struggle to handle high-dimensional, dense features, such as user profile embeddings and sequence features. Tree-based uplift models also have limited ability to uncover complex nonlinearities and interactions between data. Furthermore, the greedy splitting strategy in tree-based uplift models makes them prone to overfitting, limiting their generalization ability.
[0020] In statistics, the problem of estimating causal effects can be solved by designing randomized controlled trials (RCTs). A RCT specifically involves randomly assigning samples to experimental and control groups, and then comparing the differences between the two groups to make inferences. However, real-world RCTs have limitations. RCTs require selecting samples that are similar or have little variation in most characteristics, and in the real world, it is difficult to select such samples.
[0021] Regarding incentive and subsidy issues, if the final subsidy scheme or strategy deviates from its intended purpose, it may lead to insufficient incentive effects or even losses. Using multi-head deep neural network (DNN) models, dual-tower models, etc., causal inference can be performed on purely observational data without much consideration for bias correction.
[0022] Therefore, this disclosure proposes a deep gain model by combining feature selection, bias scoring networks and other bias correction methods, which is more suitable for scenarios such as ride-hailing subsidies.
[0023] Figure 1 This is a flowchart illustrating a service gain prediction method 100 according to an embodiment of the present disclosure. In one embodiment, the method may include: S110. Use distributional differences to extract multidimensional features related to the estimated gain from business decision units; S120. Using the deep gain model based on this multidimensional feature, the potential and propensity outcomes under different intervention methods are predicted. S130. Based on the potential and propensity results, the gains of the business decision-making unit under different intervention methods are obtained.
[0024] In this embodiment, the business decision unit may include units that require intervention decisions within the business. The business decision unit may include user information and / or service provider information related to the business. The content of business decision units may differ across different domains. A business decision unit may also be referred to as an online business sample (abbreviated as online sample or business sample), an online business order (abbreviated as order, online order, candidate order), etc. For example, a business decision unit in the ride-hailing field may include various information in a bubble. The intervention methods of the business decision unit may include multiple methods; for example, one discount or promotion method corresponds to one intervention method. Causal effect can represent the influence of one variable (cause, such as intervention, treatment, exposure) on another variable (outcome, such as health, business indicators, etc.). For example, causal effect can represent the influence of an intervention method on an outcome. From the business decision unit, multidimensional features related to causal gain estimation (also known as predicted gain) can be extracted using the distribution differences of business data. Inputting the multidimensional features of the business decision unit into a deep gain model can simultaneously estimate potential outcomes and propensity outcomes. By weighted correction of potential outcomes based on propensity outcomes, the gain of the business decision unit under different intervention methods can be obtained.
[0025] According to embodiments of this disclosure, by extracting multidimensional features related to causal gain estimation, it is beneficial to obtain more accurate gain prediction results under different intervention methods; and by observing the tendency of business decision-making units under different intervention methods, potential results can be corrected, which can further improve the accuracy of gain estimation and facilitate the selection of appropriate intervention methods for business decision-making units in the future.
[0026] Figure 2 This is a flowchart illustrating a business gain prediction method 200 according to another embodiment of the present disclosure. Method 200 can be used to implement S120 in business gain prediction method 100. Method 200 may include: using a deep gain model based on the multidimensional features to predict the potential and propensity results of different intervention methods, and may further include: S210. Each branch of the multi-branch network using the deep gain model predicts the potential results of different intervention methods based on the input features, wherein the potential result of an intervention method represents the possible gain of the business decision unit adopting that intervention method; S220. Using the deep gain model, the bias network predicts the bias outcome of the business decision unit based on the input features.
[0027] In the embodiments disclosed herein, such as Figure 3 As shown, the deep gain model may include an input layer 310, a multi-branch network 320, a bias network 330, and an output layer 340. The input layer can extract multi-dimensional features from the business decision-making unit and use these extracted features as input to the multi-branch network and the bias network. The multi-branch network may include multiple branches, each of which can process the input features separately to predict the potential results of the business decision-making unit under different intervention methods. The potential result of a business decision-making unit under a certain intervention method can represent the possible gain of the business decision-making unit using that intervention method, with values ranging from 0 to 1. Different branches can correspond to different intervention methods. For example, if there are 5 intervention methods for a certain business decision-making unit, using 5 branches can predict the potential results of these 5 intervention methods separately. Each branch may include its own multi-layer computational blocks, each of which can perform nonlinear transformations on the input features to extract deep feature representations. Different branches can share input features, but their parameters are independent.
[0028] In this embodiment of the disclosure, the bias network can correct the selection bias of business decision-making units by estimating the probability that an individual will accept a certain intervention. During the training of the deep gain model, the training loss of the multi-branch layer is weighted using the probability output by the bias network, so that even under "pseudo-random" data, the gain prediction of the decision-making unit under different intervention methods can be closer to the true causal effect.
[0029] According to the embodiments of this disclosure, the potential results of each business decision unit in each branch are weighted according to the tendency of each business decision unit, which can achieve the effect of correction and make the model output results more accurate.
[0030] In one implementation, based on the potential outcome and the propensity outcome, the gain of the business decision unit under different intervention methods is obtained, including: S230, using the propensity outcome to weight the potential outcome of the business decision unit belonging to different branches predicted by the multi-branch network, to obtain the gain of the business decision unit under different intervention methods.
[0031] In this embodiment, the potential outcome of each business decision unit in each branch (a branch can be referred to as a head) is weighted according to the bias result of the bias network to obtain the estimated gain of each business decision unit. The bias network plays a corrective role. Based on the difference between the potential outcome of a business decision unit in a specified intervention method and the potential outcome in the control method, the gain estimate of the business decision unit in that specified intervention method can be obtained. If the samples related to the business decision unit cannot be guaranteed to be unbiased, the gain estimate will also be biased, but by introducing the bias result to correct the sample bias, the gain estimate can be made more accurate.
[0032] In one implementation, gain can represent the improvement in the potential outcome of performing a specified intervention on a business decision-making unit relative to not intervening or performing a control intervention. For example, three intervention methods could include three discounts: 30%, 70%, and no discount, where no discount can also be called the control intervention or control intervention. Based on the potential outcome corresponding to "30% discount" and the potential outcome corresponding to "no discount," the gain for "30% discount" can be obtained, and so on.
[0033] According to embodiments of this disclosure, potential outcomes can be corrected through predisposing results, thereby facilitating the selection of appropriate intervention methods for business decision-making units.
[0034] In one embodiment, the method 100 further includes: using the input layer of the deep gain model to embed a portion of the features of the multidimensional feature to obtain an embedded feature, and concatenating the embedded feature with another portion of the features of the multidimensional feature to obtain the input feature.
[0035] In this embodiment of the disclosure, the input layer includes an embedding layer and other numerical feature input layers. See also Figure 3 The embedding layer maps some high-dimensional sparse features into low-dimensional dense vectors (embedded features), which are then concatenated with the original numerical features from the numerical feature layer to form the final input feature representation. For example, in a 200-dimensional input feature, 80 high-dimensional dense features are processed by the embedding layer to obtain 10-dimensional embedded features; these are then concatenated with 120-dimensional features to obtain 130-dimensional features that are fed into the next layer network.
[0036] According to this embodiment, the input layer divides multidimensional features into two categories for processing: one category is high-dimensional sparse features, which are mapped to a low-dimensional dense vector space through the embedding layer to capture potential complex relationships; the other category is numerical or low-dimensional features, which are directly input and concatenated to avoid unnecessary processing overhead, thereby improving training efficiency and effect while ensuring representation ability.
[0037] Figure 4This is a flowchart illustrating a training method 400 for a depth gain model according to an embodiment of the present disclosure. In one embodiment, the method may include: S410. Use distribution differences to extract multidimensional features related to the estimated gain from business samples; S420. Using the deep gain model that needs to be trained, based on this multidimensional feature, predict the potential and propensity outcomes of different intervention methods; S430. Based on the label results of the business samples, the potential results, and the propensity results, determine the loss function; S440. Based on this loss function, adjust the deep gain model that needs to be trained; S450. Under the condition of satisfying the training termination criterion, the trained deep gain model is obtained.
[0038] In this embodiment, the training samples for the deep gain model can be referred to as business samples, including intervention samples with various specific intervention methods and control group samples without intervention. Business samples during the model training phase can also be referred to as offline business samples (referred to as offline samples), offline business orders (referred to as offline orders), etc. According to this embodiment, the required multi-dimensional features are extracted from the business samples based on data distribution differences as model input. Samples are processed in batches during training: for each training batch, the loss function is calculated based on the batch's true label, the model's prediction results under various intervention conditions, and propensity information, and the model parameters are updated using gradient backpropagation and optimization algorithms. Training termination criteria are set (e.g., loss convergence, early stopping on the validation set, or reaching the maximum number of iterations). When any termination criterion is met, training is terminated and the trained deep gain model is obtained; otherwise, the model parameters or hyperparameters (including but not limited to multi-branch networks and propensity networks) are adjusted and optimized until the termination criterion is met. The trained model can be used to output the potential results of the business decision-making unit under different intervention schemes, and further output the causal gain estimate of the decision-making unit under each intervention method.
[0039] According to embodiments of this disclosure, bias can be used to correct selection bias in business samples, making the output results of the trained model more accurate.
[0040] Figure 5 This is a flowchart illustrating a training method 500 for a deep gain model according to another embodiment of this disclosure. Method 500 can be used to implement step S420 of the training method 400 for the deep gain model. Method 200 may include: determining a loss function based on the label result, the potential result, and the propensity result of the business sample; and may further include: S510. Using the propensity score, the potential results of each sample predicted by the multi-branch network belonging to different branches are weighted and combined with the label results to obtain the loss function.
[0041] In this embodiment of the disclosure, the network structure of the deep gain model can be found in [reference needed]. Figure 3 The calculation methods for the prediction results and trends can be found in the relevant descriptions of the prediction methods in the above embodiments, and will not be repeated here.
[0042] In this embodiment of the disclosure, the specific explanation and examples of weighting the prediction results of different intervention methods for each branch can be found in the relevant description of the prediction method in the above embodiments, and will not be repeated here. In the loss function, for each intervention method of the sample, the prediction results can be weighted using propensity, and then the difference between the weighted prediction results and the label results (the label of the specific intervention method used, and 0 for those not used) can be calculated; thus obtaining the final prediction result.
[0043] One example of a loss function is the binary cross-entropy (BCE) loss function, which measures the difference between the model's predictions and the labeled results, guiding model parameter updates to reduce prediction errors. In the BCE loss function, a bias is introduced to weight the predictions of a sample across different branches; that is, when calculating BCE, a weight related to its corresponding bias is assigned to the prediction results of each branch. Specifically, the bias can be used to weight the prediction results of each branch first, and then the difference between the weighted prediction results and the observed label of the sample can be used as the basis for calculating the binary cross-entropy.
[0044] According to the embodiments of this disclosure, the loss function constructed by combining the prediction results of each business decision unit in the sample with the tendency of each branch as a weight and the label results can more accurately reflect the difference between the prediction results and the label results, improve the model's ability to correct "pseudo-random" data, and make the output results of the trained model more accurate.
[0045] Figure 6 This is a flowchart illustrating a training method 600 for a depth gain model according to an embodiment of the present disclosure. In one embodiment, the method may further include: S610. Select key feature dimensions from the candidate feature set that are related to the intervention methods and outcome variables of the business decision-making unit and may cause confounding biases, and use them as input feature dimensions for the deep gain model that needs to be trained.
[0046] In this embodiment, the candidate feature set may include features from various domains, such as features related to ride-hailing, public transportation, and maps. From these features, key feature dimensions that are relevant to both the intervention and outcome variables in a specific domain (e.g., ride-hailing) and may cause confounding bias can be selected to assist in identifying causal effects for gain estimation. Examples include feature dimensions related to time, space, weather, and order amount. Weather is a key confounding feature; for example, rainy days may reduce user subsidies, so such confounding features need to be added and controlled. The selected feature dimensions can be used as input feature dimensions for training a deep gain model. For example, from 1000 candidate features, 100 features can be selected using feature selection methods such as Kullback-Leibler (KL) divergence and input into the deep gain model.
[0047] According to embodiments of this disclosure, key feature dimensions that are related to both the intervention methods and outcome variables of the business decision-making unit and may cause confounding bias are selected as input feature dimensions of the model, enabling the trained model to more accurately estimate the gain of the intervention methods of the business decision-making unit.
[0048] In one implementation, key feature dimensions that are relevant to both the intervention methods and outcome variables of the business decision-making unit and may cause confounding bias are screened from the candidate feature set. This includes: analyzing the distribution of each candidate feature under different intervention groups, calculating the KL divergence between the feature distribution and the overall or control group distribution to quantify the feature's ability to distinguish intervention gains, and selecting features with significantly larger KL divergences as feature dimensions related to gain estimation. The greater the difference, the higher the ranking, and this can be used to set the final input feature dimension for the deep gain model.
[0049] In one embodiment, the method may further include: using the input layer of the deep gain model to embed a portion of the features of the multidimensional feature to obtain an embedded feature, and concatenating the embedded feature with another portion of the features of the multidimensional feature to obtain the input feature.
[0050] In this embodiment, the process of embedding and concatenating multidimensional features at the input layer can be found in the description of the above embodiments. According to this embodiment, the multidimensional features can be divided into two parts at the input layer for separate processing. For complex features, such as high-dimensional dense features, representation processing can be performed through the embedding layer to more effectively capture their potential information; while for simple features, direct concatenation can reduce unnecessary computational overhead, thereby improving the efficiency of feature extraction and the accuracy of representation.
[0051] Figure 7This is a flowchart illustrating a decision-making method 700 for selecting the optimal intervention method for a business decision-making unit according to an embodiment of the present disclosure. In one embodiment, the method includes: S710. Construct the first operations research optimization model based on the estimated gains, decision objectives, decision variables, and business constraints of the business decision-making unit under different intervention methods; S720. Perform a Lagrange dual transformation on the first operations research optimization model to obtain the second operations research optimization model; S730. Solve the second operations research optimization model to obtain the optimal decision corresponding to the business decision unit. The optimal decision includes the optimal intervention method.
[0052] In this embodiment, the first operations research optimization model is used to characterize maximizing business objectives, such as order completion volume. Specifically, by associating the gain estimates of each business decision-making unit under different intervention methods with their corresponding decision variables and summarizing them, the maximization decision objective is constructed. Adding constraints such as business budget constraints (business constraints) and the values of decision variables, the first operations research optimization model is fully constructed.
[0053] To simplify the solution process of the first operations research optimization model, the primal problem can be transformed into its dual problem using the Lagrange duality method. Specifically, Lagrange multipliers are introduced by imposing constraints on the primal problem to construct a dual function, which is then used as the optimization objective to form the dual problem. This dual problem typically features a simpler structure and higher computational efficiency, and can be used to obtain a lower bound or near-optimal solution to the primal problem, thereby alleviating the complexity of solving the primal problem.
[0054] According to the embodiments of this disclosure, the original problem of the first operations research optimization model is transformed into the dual problem, namely the second operations research optimization model, through Lagrange dual transformation, which can reduce the difficulty of solving the problem and quickly and accurately obtain the appropriate intervention method for the business decision-making unit.
[0055] In one implementation, the decision variable is used to indicate whether the business decision-making unit adopts a specified intervention method; the objective of the first operations optimization model is to maximize the business completion; the first operations optimization model includes decision objectives, decision variables, budget constraints, and other business constraints.
[0056] In one implementation, the decision variable constraint includes: the value of the decision variable is limited to a first state or a second state, wherein the first state indicates that the business decision unit adopts a specified intervention method, and the second state indicates that the business decision unit does not adopt a specified intervention method; and each business decision unit can and can only select one intervention method. For example, if the first... iIf the decision variable for a certain intervention for an order takes the first state, it means that the order adopted this intervention method; if it takes the second state, it means that the order did not adopt this intervention method.
[0057] A first operations research optimization model, aiming to maximize order volume while ensuring subsidies do not exceed a specified budget level, is modeled as follows: in The first predicted by the depth gain model i The gain value of an individual business decision unit (e.g., bubble or order) under different intervention methods (or the individual treatment effect (ITE)). For decision variables, the value can be 0 or 1, representing the first decision variable. i Whether a business decision-making unit adopts the j-th intervention method, and for any business decision-making unit, it can and can only adopt one discount intervention method; For the first i Completion rate of each business decision-making unit without intervention; Let B be the discount rate under the j-th intervention method; B be the total subsidy amount; price i For the first i The price of each business decision unit.
[0058] According to the embodiments of this disclosure, the first operations research optimization model can be used to obtain the intervention method that maximizes the completion of business through the optimization algorithm. However, if the decision dimension exceeds 1 million, the first operations research optimization model needs to be transformed in order to find the optimal solution.
[0059] In one implementation, the first operations research optimization model is subjected to a Lagrange dual transformation to obtain a second operations research optimization model for the business decision-making unit. This includes: constructing the second operations research optimization model using a Lagrange multiplier dual transformation; wherein the second operations research optimization model is solved using a greedy algorithm, transforming the selection of the optimal decision variable into finding the optimal Lagrange multiplier with a fixed budget; the variables of the dual problem are not less than 0; and the Lagrange multiplier is not less than 0.
[0060] In this embodiment of the disclosure, to solve the first operations research optimization model, the Lagrange dual transformation method can be used to transform the original problem into a dual problem. Dual problems typically have the characteristics of simpler structure and higher computational efficiency, and can be used to obtain the lower bound or near-optimal solution of the original problem, thereby alleviating the complexity in solving the original problem.
[0061] An example of a model after Lagrange dual transformation is as follows: in For the variables of the dual problem, These are Lagrange multipliers; the meanings of the other letters are the same as those in the original problem. , indicating the first i The business decision-making unit in the first j The expected cost under each intervention method, then It can be written as i The business decision-making unit in the first j Gain density under various intervention methods (which can be understood as efficiency, or return on investment (ROI)).
[0062] Based on the model after dual transformation, it can be seen that the decision dimension changes from the original n m is reduced to n dimensions.
[0063] According to embodiments of this disclosure, the solution difficulty of the first operations research optimization model can be simplified by using Lagrange dual transformation, and the optimal intervention method of the business decision-making unit can be obtained quickly and accurately.
[0064] In one implementation, solving the second operations research optimization model to obtain the optimal intervention method corresponding to the business decision-making unit includes: using a fixed budget and an iterative approach to solve for the optimal Lagrange multiplier; using the optimal Lagrange multiplier as a fixed value, calculating the gain density of the business decision-making unit under different intervention methods; and selecting the intervention method with the largest gain density to obtain the optimal intervention method corresponding to the business decision-making unit.
[0065] In this embodiment of the disclosure, when When taking a fixed value, we can obtain: That is, the intervention method chosen by the i-th business decision unit is determined from... The intervention method corresponding to the maximum value in (j=0,…,m) is selected. Therefore, in this embodiment of the invention, for each business decision unit to be assigned, by calculating the gain density under different intervention methods and selecting the intervention method with the largest gain density, the optimal solution to the dual problem, which is also the optimal solution to the primal problem, can be obtained. Under the condition of a fixed budget B, by... By searching the value space, the optimal value can be obtained. The value is then used to determine the optimal intervention method.
[0066] According to the embodiments of this disclosure, based on the second operations research optimization model, the intervention method corresponding to the maximum gain density can be used as the optimal intervention method for the business decision-making unit, which reduces the difficulty of solving the original problem and enables the allocation of intervention methods to the business decision-making unit more quickly and accurately.
[0067] In one implementation, the method may further include: calculating the business budget for the next time step based on the business budget for the current time step, the business budget and actual expenditures prior to the current time step, and the total adjustment period; the budget for the next time step is used to update the optimal solution of the second operations optimization model after the Lagrange dual transformation.
[0068] In this embodiment, the subsidy rate is adjusted in real time based on the actual expenditure, business budget, and total adjustment period before the current time step to ensure that expenditure is neither over-spending nor under-spending. Different business budget levels (e.g., subsidy rate) and their optimal Lagrange multipliers are calculated offline based on historical data from the previous period and then directly used in the online business of the current period. If the distribution of business decision units in the two periods is inconsistent, it can lead to unreasonable allocation of online business budget. To more reasonably control budget allocation, the business budget can be updated periodically, and the optimal time slice for the next time slice can be found. This allows for the allocation of optimal interventions to decision-making units.
[0069] An example formula for a business budget update model is as follows: in This is a hyperparameter that can be set according to the magnitude of the error. The business budget amount for the (h+1)th time step (e.g., hour); Let K be the actual budget spent at the i-th time step; K is the total adjustment time step (total adjustment period). Using the above formula, the budget required for the next time step can be calculated. Then, through Lagrange dual optimization, the optimal value for the next time step is obtained. .
[0070] Updating the business budget according to time steps can improve the rationality of budget allocation and reduce or even avoid waste or surplus.
[0071] The business intervention method proposed in this disclosure includes an optimized ride-hailing subsidy scheme based on a deep gain (uplift) model and a Lagrange dual with post-processing. The uplift model has been improved to better suit specific business scenarios. By constructing a deep uplift model, it is possible to more accurately predict the price elasticity of different bubble effects, thereby achieving personalized subsidy strategies. Combined with an operations research optimization model, this disclosure enables dynamic optimization of subsidy allocation under budget constraints, ensuring the reasonable and full use of subsidy funds and avoiding issues of fund waste and budget surplus, thus significantly improving fund utilization and the economic benefits of the subsidy strategy. This disclosure is applicable to ride-hailing platforms, shared mobility services, food delivery platforms, and other service scenarios based on user pricing and subsidy strategies. It can significantly improve the platform's incentive efficiency for users while ensuring the reasonable use of subsidy funds, demonstrating broad application prospects and economic value.
[0072] The subsidy allocation problem under budget constraints can typically be formalized as a constrained optimization problem. Since subsidy decision variables can be discrete variables of type 0-1, solving them using integer programming methods (such as branch and bound) is computationally too complex for real-time decision problems involving millions or even larger scales in ride-hailing scenarios, making it difficult to meet the needs of online applications. This disclosure proposes an optimization method based on Lagrange duality. By introducing budget constraints into Lagrange multipliers and constructing a dual problem, the problem size can be significantly reduced while maintaining approximate optimality, thus supporting efficient solutions for large-scale online subsidy allocation. This method has significant advantages in computational efficiency and scalability. In practical business, when the Lagrange multipliers λ obtained through offline optimization are directly applied to online scenarios, factors such as real-time traffic distribution offsets and user response differences may lead to inaccurate budget execution. This disclosure further proposes a Lagrange duality optimization method with a post-processing mechanism. Through online correction and heuristic adjustment, budget deviations are dynamically corrected, improving the stability and controllability of subsidy allocation.
[0073] This disclosure proposes an optimized ride-hailing subsidy scheme based on a deep uplift model and a Lagrange duality with post-processing. This scheme better captures different price elasticities of bubbling and allocates subsidies based on these elasticities. Furthermore, the combination with the Lagrange duality model with post-processing ensures reasonable budget allocation and full utilization. The specific implementation process is as follows: Figure 8 As shown: S810, Feature Filtering: Feature filtering based on KL divergence.
[0074] The Kohlbek-Leibler (KL) divergence was used to select top features that were important for estimating causal effects. KL divergence (also called relative entropy) is a metric that measures the difference between two probability distributions P and Q. An example calculation is as follows: in, It can represent a probability distribution and The KL divergence between features. Iterate through each feature, applying treatment (or intervention) and not applying treatment for each feature. The greater the difference, the more important the feature. All features are ranked according to the KL divergence with and without the applied treatment, and the top N features are selected as the input features for the deep uplift model.
[0075] S820. Construct an uplift model: Add a multi-head gain (uplift) model based on biased result correction.
[0076] After feature filtering by S810, the selected head features, such as the top 100 features, are input into the deep uplift model. The deep uplift model proposed in this embodiment can be a multi-head uplift model with added bias correction, i.e., an uplift model based on multi-head, multi-treatment bias correction. The goal of this uplift model is to learn the differential causal effects between different treatments and the control group, and to correct bias through biased outcomes, thereby improving the prediction effect and accuracy of causal effects. See the specific model structure below. Figure 8 Different modules have different functions: Input layer 821: The left side is the embedding layer, and the right side is the numerical feature input. Based on input layer 821, the two parts can be concatenated as the model input. The addition of embedding can better capture the information of high-dimensional categorical features.
[0077] Multi-branch layer 822: Each branch can be used to predict the potential outcomes of different treatments. Each branch has its own operational block, such as block 1, block 2, etc., and multiple blocks can form an n-layer network structure (n can be a hyperparameter). Each block is responsible for performing a nonlinear transformation on its own input to extract deep feature representations. Different branches can share inputs, but their parameters are adjusted independently.
[0078] The propensity score network layer 823 (or IPW network) can correct for selection bias by estimating the probability of an individual receiving a particular treatment. The training loss 1 of the multi-branch layer is weighted and adjusted using the output loss 2 of the propensity score network layer, thus making the individual heterogeneity prediction closer to the true causal effect under a "pseudo-random" data distribution.
[0079] Output layer 824: Normalizes the corresponding prediction distribution output from each branch of the multi-branch layer 822 to obtain the final potential prediction results for different treatment and control groups. Then, the prediction results of each sample in each branch are weighted according to propensity score to correct bias and make the model output results more accurate.
[0080] S830. Constructing an operations research optimization model: Calculate the Lagrange multipliers offline using a Lagrange dual optimization model and a greedy algorithm. dictionary.
[0081] This disclosure, in conjunction with the objectives of the ride-hailing business, constructs an operations research optimization model with the goal of maximizing the number of completed rides, the subsidy amount as a constraint, and different discount rates as decision variables. An exemplary calculation method is as follows: in, The first one that can be predicted by the uplift model i Each bubble at different discounts j The uplift value (or ite) below; This can be a decision variable, taking values of 0 or 1, i.e., the first... i Does the bubble take the first one? j There are one type of discount, and for any bubble, only one discount can and must be applied; It can be the initial amount of the original bubble; It can be the first i The completion rate of a bubble-shaped item when there is no discount; discount Discount rates (0.75, 0.8, etc.) can be used. B This can be the total subsidy amount.
[0082] To solve this deep uplift model and apply it to real-time online services (e.g., making instant decisions to apply a discount to a given bubble in a real-world scenario), a Lagrange dual transformation can be used. The original problem is transformed into calculating the "gain density" of each bubble and multiplying it by Lagrange multipliers. By comparing, it can efficiently approximate the optimal solution and meet the requirements of online services for real-time decision-making.
[0083] Bundle Recorded as , can represent the first i The bubble appeared in the first... j The expected cost under this discount. Then... It can be denoted as the first i The bubble appeared in the first... j Gain density under medium discount.
[0084] The following is the model after the Lagrange dual transformation: in, These can be variables in the dual problem; These can be Lagrange multipliers, representing gain density; the other letters have the same meanings as in the original problem. As the model shows, the decision dimension changes from the original n... m is reduced to n dimensions. When the Lagrange multipliers When taking fixed values, an exemplary way to determine the variables of the dual problem is as follows: No. i Which discount applies to each bubble-shaped discount? ( j =0,…, m The discount corresponding to the maximum value is selected from the options. Therefore, in this embodiment of the disclosure, for each decision unit to be assigned (e.g., bubble sort), by calculating the gain density under different discounts and selecting the discount level with the largest gain density, the optimal solution to the dual problem, which is also the optimal solution to the primal problem, can be obtained. Lagrange multipliers The value of can be determined in advance in conjunction with budget constraint B. Specifically, under the condition of a fixed budget B, it can be determined by... The optimal value is obtained by searching the value space. value.
[0085] Furthermore, for the large-scale bubble sort data generated daily, the globally optimal solution can be obtained simultaneously through the above planning and dual solution process. The value, and the optimal discount strategy for each bubble. In actual ride-hailing scenarios, because it is necessary to assign a discount to each bubble in real time, therefore... The value of is typically obtained through offline estimation. Specifically, the optimal value can be obtained offline based on the bubble distribution data from the same period last week. And it can be used directly in the online service of the current cycle. This method also implies the following assumption: if the bubble distribution is consistent between two consecutive periods, then under the same budget constraint, The values are the same.
[0086] S840, Post-processing of the real-time budget control system: online allocation and real-time budget control.
[0087] In S820, it can be assumed that the daily bubble distribution is consistent with the bubble distribution of the same period last week, and the solution is obtained using offline data. To guide the real-time allocation of discounts online.
[0088] For each online sample in the bubble, the optimal discount rate can be selected for each bubble using the optimization formula. An example of an optimization formula is as follows: However, inconsistent bubble distribution can lead to real-time online spending exceeding or falling below the budget. To address this issue, this disclosure presents a time-slice-based budget allocation algorithm to control budget allocation and further correct subsidy deviations. The core solution is that the current cumulative spending error is amortized by the remaining time slices. The specific algorithm model for budget adjustment is as follows: in, It can be a hyperparameter, which can be set to 1.25 or 0.75, depending on the error magnitude. This can be the budget amount for the (h+1)th hour; K can be the budget for the actual expenditure in the i-th hour; K can be the total adjustment period. This can be the budget needed for the next time step; h can be the current time (or time step). This embodiment of the disclosure can calculate the budget needed for the next time step using the above formula, even when the bubble distribution is inconsistent or the budget allocation deviates from the offline optimal solution. Then, through Lagrange dual optimization, the optimal value for the next time step is obtained. This avoids issues such as wasted or surplus subsidies. Here, 'h' can also be other time units, which can be adjusted according to the specific needs of the application scenario.
[0089] The method provided in this disclosure is applicable to ride-hailing platforms, shared mobility services, and other service scenarios based on user pricing and subsidy strategies. It can significantly improve the platform's incentive efficiency for users while ensuring the rational use of subsidy funds. Random data can be collected. Feature filtering is performed on the collected data. A deep uplift model is constructed based on the filtered features. According to this disclosure, the uplift value (or ITE) of each bubble under different discounts can be predicted. An operations research optimization model is built according to the business scenario, and the Lagrange multipliers of this operations research optimization model are solved offline. If the online service budget control is not accurate enough, a post-processing framework, i.e., a real-time budget control model, can be built.
[0090] In this embodiment, by constructing an uplift model, the sensitivity of different bubble formations to price changes can be accurately characterized, thereby enabling personalized subsidy strategies and making the system more user-friendly. Combined with an operations research optimization model, this embodiment can achieve dynamic optimization of subsidy allocation under budget constraints, ensuring that subsidy funds are used reasonably and fully, avoiding waste and budget surpluses, and thus significantly improving fund utilization and the economic benefits of the subsidy strategy.
[0091] Figure 9 This is a schematic diagram of a service gain prediction apparatus 900 according to an embodiment of the present disclosure. In one embodiment, the apparatus may include: Extraction module 910 is used to extract multidimensional features related to the estimated gain from the business decision unit using distribution differences; The prediction module 920 is used to predict the potential and propensity outcomes under different intervention methods based on the multidimensional features using a deep gain model. Gain module 930 is used to obtain the gain of the business decision unit under different intervention methods based on the potential outcome and the propensity outcome.
[0092] Figure 10 This is a schematic diagram of a service gain prediction device 1000 according to another embodiment of the present disclosure. The device may include: an extraction module 1010, a prediction module 1020, and a gain module 1030. The functions and roles of each module can be found in the description of the functions and roles of the modules in the service gain prediction device 1000. In one embodiment, the prediction module 1020 may include: The multi-head prediction submodule 1021 is used to predict the potential results of different intervention methods based on the input features of each branch of the multi-head branch network using the deep gain model, wherein the potential result of an intervention method represents the possible gain of the business decision unit adopting that intervention method; The propensity estimation submodule 1022 is used by the propensity network using the deep gain model to predict the propensity outcome of the business decision unit based on the input features.
[0093] In one implementation, such as Figure 10 As shown, the gain module 1030 includes: The weighting submodule 1031 is used to weight the potential results of the business decision unit belonging to different branches predicted by the multi-branch network using the propensity result, so as to obtain the gain of the business decision unit under different intervention methods.
[0094] In one implementation, such as Figure 10 As shown, the device may further include: The embedding module 1040 is used to embed a portion of the features of the multidimensional feature using the input layer of the deep gain model to obtain the embedded feature, and to concatenate the embedded feature with another portion of the features of the multidimensional feature to obtain the input feature.
[0095] Figure 11 This is a schematic diagram of a training apparatus 1100 for a depth gain model according to an embodiment of the present disclosure. In one embodiment, the apparatus may include: Extraction module 1110 is used to extract multidimensional features related to the estimated gain from business samples using distribution differences; The prediction module 1120 is used to predict the potential and propensity outcomes of different intervention methods based on the multidimensional features using a deep gain model that needs to be trained. The loss calculation module 1130 is used to determine the loss function based on the label results, potential results, and propensity results of the business samples; Training module 1140 is used to adjust the deep gain model to be trained based on the loss function; and to obtain the trained deep gain model when the training termination criterion is met.
[0096] Figure 12 This is a schematic diagram of a training apparatus 1200 for a deep gain model according to another embodiment of the present disclosure. The apparatus may include: an extraction module 1210, a prediction module 1220, a loss calculation module 1230, and a training module 1240. The functions and roles of each module can be found in the aforementioned business gain prediction apparatus 1100. In one embodiment, the loss calculation module 1230 is further used to use the propensity score to weight the potential results of each sample predicted by the multi-branch network as belonging to different branches, and combine this with the label results to obtain the loss function.
[0097] In one implementation, such as Figure 12As shown, the device may further include: The screening module 1250 is used to screen key feature dimensions from the candidate feature set that are related to the intervention methods and outcome variables of the business decision-making unit and may cause confounding biases, as input feature dimensions of the deep gain model that needs to be trained.
[0098] Figure 13 This is a schematic diagram of a decision-making device 1300 for selecting the optimal intervention method for a business decision-making unit according to an embodiment of the present disclosure. In one embodiment, the device may include: Module 1310 is used to construct the first operations research optimization model based on the estimated gains, decision objectives, decision variables and business constraints of the business decision-making unit under different intervention methods. Transformation module 1320 is used to perform a Lagrange dual transformation on the first operations research optimization model to obtain a second operations research optimization model. The solver module 1330 is used to solve the second operations optimization model to obtain the optimal decision corresponding to the business decision unit, which includes the optimal intervention method.
[0099] In one implementation, the decision variable is used to indicate whether the business decision-making unit adopts a specified intervention method; the objective of the first operations optimization model is to maximize the business completion; the first operations optimization model includes decision objectives, decision variables, budget constraints, and other business constraints.
[0100] In one implementation, the transformation module 1320 is further configured to construct the second operations research optimization model by employing dual transformations of Lagrange multipliers. The second operations research optimization model is solved using a greedy algorithm, which transforms the selection of the optimal decision variable into finding the optimal Lagrange multiplier with a fixed budget. The variables in the dual problem are not less than 0; The Lagrange multiplier is not less than 0.
[0101] Figure 14 This is a schematic diagram of a service intervention device 1400 according to another embodiment of the present disclosure. The device may include: a construction module 1410, a transformation module 1420, and a solution module 1430. The functions and roles of each module can be found in the service gain prediction device 1100 described above. In one embodiment, the solution module 1430 includes: Solving submodule 1431 is used to solve for the optimal Lagrange multipliers with a fixed budget using the idea of iteration; The calculation submodule 1432 is used to calculate the gain density of the business decision-making unit under different intervention methods with the optimal Lagrange multiplier as a fixed value. Select submodule 1433 to select the intervention method with the highest gain density, and obtain the optimal intervention method corresponding to the business decision unit.
[0102] In one implementation, such as Figure 14 As shown, the device may further include: The budget calculation module 1440 is used to calculate the business budget for the next time step based on the business budget of the current time step, the business budget and actual expenditure before the current time step, and the total adjustment period; the budget of the next time step is used to update the optimal solution of the second operations optimization model after the Lagrange dual transformation.
[0103] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0104] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0105] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0106] Figure 15 A schematic block diagram of an example electronic device 1500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0107] like Figure 15 As shown, device 1500 includes a computing unit 1501, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1502 or a computer program loaded from storage unit 1508 into random access memory (RAM) 1503. The RAM 1503 may also store various programs and data required for the operation of device 1500. The computing unit 1501, ROM 1502, and RAM 1503 are interconnected via bus 1504. Input / output (I / O) interface 1505 is also connected to bus 1504.
[0108] Multiple components in device 1500 are connected to I / O interface 1505, including: input unit 1506, such as keyboard, mouse, etc.; output unit 1507, such as various types of monitors, speakers, etc.; storage unit 1508, such as disk, optical disk, etc.; and communication unit 1509, such as network card, modem, wireless transceiver, etc. Communication unit 1509 allows device 1500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0109] The computing unit 1501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1501 performs the various methods and processes described above, for example, at least one of a business gain prediction method, a deep gain model training method, and a business intervention method. In some embodiments, at least one of the business gain prediction method, the deep gain model training method, and the business intervention method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1500 via ROM 1502 and / or communication unit 1509. When the computer program is loaded into RAM 1503 and executed by computing unit 1501, it can perform one or more steps of at least one of the service gain prediction method, the deep gain model training method, and the service intervention method described above. Alternatively, in other embodiments, computing unit 1501 can be configured by any other suitable means (e.g., by means of firmware) to perform at least one of the service gain prediction method, the deep gain model training method, and the service intervention method.
[0110] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0111] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0112] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0114] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0115] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0116] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0117] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for predicting business gain, comprising: Use distributional differences to extract multidimensional features related to the estimated gain from business decision units; Based on the aforementioned multidimensional features, the deep gain model is used to predict the potential and propensity outcomes under different intervention methods. Based on the potential and propensity outcomes, the gains of the business decision-making unit under different intervention methods are obtained.
2. The method according to claim 1, wherein, Based on the aforementioned multidimensional features, a deep gain model is used to predict the potential and propensity outcomes of different intervention methods, including: Each branch of the multi-branch network using the deep gain model predicts the potential outcome of different intervention methods based on input features, wherein the potential outcome of an intervention method represents the possible gain of the business decision-making unit adopting the intervention method; The propensity network using the deep gain model predicts the propensity outcome of the business decision unit based on the input features.
3. The method according to claim 1 or 2, wherein, Based on the potential and propensity outcomes, the gains of the business decision-making unit under different intervention methods are obtained, including: The propensity results are used to weight the potential outcomes of the business decision unit belonging to different branches predicted by the multi-branch network, so as to obtain the gain of the business decision unit under different intervention methods.
4. The method according to any one of claims 1 to 3, further comprising: The input layer of the deep gain model is used to embed a portion of the multidimensional features to obtain embedded features, and the embedded features are concatenated with another portion of the multidimensional features to obtain input features.
5. A training method for a deep gain model, comprising: Use distributional differences to extract multidimensional features related to the estimated gain from business samples; Based on the multidimensional features, a deep gain model that needs to be trained is used to predict the potential and propensity outcomes of different intervention methods. Based on the labeling results of the business samples, the potential results, and the propensity results, the loss function is determined; Based on the loss function, the deep gain model that needs to be trained is adjusted; Under the condition that the training termination criterion is met, the trained deep gain model is obtained.
6. The method according to claim 5, wherein, Based on the labeling results, potential results, and propensity results of the business samples, a loss function is determined, including: The loss function is obtained by weighting the potential results of each sample predicted by the multi-branch network to belong to different branches using the propensity results and combining them with the label results.
7. The method according to claim 5 or 6, further comprising: Key feature dimensions that are relevant to the intervention methods and outcome variables of the business decision-making unit and may cause confounding biases are selected from the candidate feature set and used as the input feature dimensions of the deep gain model to be trained.
8. A decision-making method for selecting the optimal intervention method for a business decision-making unit, comprising: The first operations research optimization model is constructed based on the estimated gains, decision objectives, decision variables, and business constraints of business decision-making units under different intervention methods. The first operations research optimization model is subjected to Lagrange dual transformation to obtain the second operations research optimization model; Solve the second operations research optimization model to obtain the optimal decision corresponding to the business decision unit, the optimal decision including selecting the optimal intervention method.
9. The method according to claim 8, wherein, The decision variables are used to indicate whether the business decision-making unit adopts a specified intervention method; the goal of the first operations optimization model is to maximize the business completion; the first operations optimization model includes decision objectives, decision variables, budget constraints, and other business constraints.
10. The method according to claim 8 or 9, wherein, The first operations research optimization model is subjected to a Lagrange dual transformation to obtain the second operations research optimization model for the business decision-making unit, including: The second operations research optimization model is constructed by employing dual transformations of Lagrange multipliers; The second operations research optimization model is solved using a greedy algorithm, which transforms the selection of the optimal decision variable into finding the optimal Lagrange multiplier with a fixed budget. The variables in the dual problem are not less than 0; The Lagrange multipliers are not less than 0.
11. The method according to any one of claims 8 to 10, wherein, Solving the second operations research optimization model yields the optimal intervention method corresponding to the business decision-making unit, including: Using a fixed budget, the optimal Lagrange multiplier is found through iterative methods. Using the optimal Lagrange multiplier as a fixed value, the gain density of the business decision-making unit under different intervention methods is calculated; The intervention method with the highest gain density is selected to obtain the optimal intervention method corresponding to the business decision unit.
12. The method according to any one of claims 8 to 11, further comprising: Based on the business budget of the current time step, the business budget and actual expenditure before the current time step, and the total adjustment period, the business budget of the next time step is calculated; the budget of the next time step is used to update the optimal solution of the second operations research optimization model after the Lagrange dual transformation.
13. A business gain prediction device, comprising: The extraction module is used to extract multidimensional features related to the estimated gain from the business decision unit using distribution differences; The prediction module is used to predict the potential and propensity outcomes under different intervention methods based on the multidimensional features using a deep gain model. The gain module is used to obtain the gain of the business decision-making unit under different intervention methods based on the potential results and the propensity results.
14. A training device for a deep gain model, comprising: The extraction module is used to extract multidimensional features related to the estimated gain from business samples using distribution differences; The prediction module is used to predict the potential and propensity outcomes of different intervention methods based on the multidimensional features using a deep gain model that needs to be trained. The loss calculation module is used to determine the loss function based on the label results of the business samples, the potential results, and the propensity results; The training module is used to adjust the deep gain model to be trained based on the loss function; Under the condition that the training termination criterion is met, the trained deep gain model is obtained.
15. A decision-making device for selecting the optimal intervention method for a business decision-making unit, comprising: The module is used to build the first operations research optimization model based on the estimated gains, decision objectives, decision variables and business constraints of business decision-making units under different intervention methods. The transformation module is used to perform a Lagrange dual transformation on the first operations research optimization model to obtain a second operations research optimization model. The solution module is used to solve the second operations research optimization model to obtain the optimal decision corresponding to the business decision unit, and the optimal decision includes the optimal intervention method.
16. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 12.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 12.
18. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 12.