Offline strategy automatic processing method and device and storage medium
Through the offline strategy automation processing method, predictive models and strategy solution technology are used to solve the problem of incentive difficulties for medium and low-frequency users and new users in the existing technology, and more efficient operations and more accurate incentive strategies are achieved.
Patent Information
- Application Number
- CN202510361725.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
When existing precision marketing solutions deal with low-frequency users and new users, the model learning is difficult, the integer planning solution strategy is inaccurate, and it is impossible to effectively motivate new users.
An offline strategy automated processing method is provided. By obtaining offline characteristic data and user types of target users, using the target prediction model to predict the user's order probability under different incentive amounts, and determining the incentive strategy through policy solutions.
This method can effectively cover any group of new customers, recalls, retentions, etc., improve operational efficiency, save algorithm labor costs, avoid information synchronization errors, and improve the performance of constraint parameters such as additional ROI.
Smart Images

Figure CN120218991A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and particularly to an offline policy automated processing method, device, and storage medium. Background Art
[0002] In the context of intelligent marketing scenarios, determining which group of people to target for precision marketing is a challenging task. Currently, the mainstream precision marketing solutions use causal inference techniques to perform intelligent coupon distribution. Common causal inference models include: single model (S-Learner), dual model (T-Learner), and causal forest model (Causal Forest). By learning the elasticity (price sensitivity) of users through the model, distribution strategies are formulated according to the different elasticities of users. The distribution strategy is implemented using operations research optimization techniques (such as integer programming solution).
[0003] However, due to different customer groups in different marketing scenarios and relatively low overall elasticity of the population, taking low-frequency users as an example, such users have a long dormant time and less historical order data. Therefore, it is very difficult for the model to learn. Coupled with the large difference between the historical average order price of users and the actual average order price at the time of placing an order, the result of the integer programming solution strategy is not completely accurate and may even be biased. And it cannot be used at all for new users without historical orders. Summary of the Invention
[0004] Based on this, it is necessary to provide an offline policy automated processing method, device, and storage medium for the above technical problems to solve at least one of the problems existing in the above prior art.
[0005] In a first aspect, an offline policy automated processing method is provided, including:
[0006] Obtain the offline feature data corresponding to the target user and the user type;
[0007] Based on the user type and the offline feature data, perform prediction through the corresponding target prediction model to obtain the user order probabilities of the target user under different incentive amounts;
[0008] Based on the user order probabilities and the user type, perform policy solution through the corresponding policy solution method to obtain the incentive policy corresponding to the target user.
[0009] In an embodiment, the user type is a new user, and the performing policy solution based on the user order probabilities and the user type includes:
[0010] Perform equal-frequency binning on the user order probabilities, and the number of bins is equal to the number of different incentive amounts;
[0011] Determine the corresponding incentive strategy based on the user order placement probability per bucket. Among them, for the bucket with a higher user order placement probability, a larger incentive amount is distributed.
[0012] In one embodiment, the target user is a recalled user or a retained user. Based on the user order placement probability and the user type, perform policy solution through corresponding policy solution methods to obtain the incentive strategy corresponding to the target user, including:
[0013] Based on the user order placement probability, determine user elasticity, where the user elasticity reflects the sensitivity of the user order placement probability to changes in the incentive amount;
[0014] Based on the offline feature data, determine the historical average order price corresponding to the target user;
[0015] Perform equal-frequency binning on the historical average order price, user elasticity, and user order placement probability. The number of bins is equal to the number of different incentive amounts;
[0016] Based on the binning results, determine the incentive strategy corresponding to each bucket.
[0017] In one embodiment, the target user is a recalled user. Based on the binning results, determine the incentive strategy corresponding to each bucket, including:
[0018] Perform weighted summation on the historical order average price, user elasticity, and user order placement probability respectively to obtain the weight coefficient corresponding to each bucket;
[0019] Based on the weight coefficient corresponding to each bucket, determine the incentive strategy corresponding to each bucket. Among them, for the bucket with a larger weight coefficient, the incentive amount is larger.
[0020] In one embodiment, the target user is a retained user. Based on the binning results, determine the incentive strategy corresponding to each bucket, including:
[0021] Determine the optimization objective, incentive parameters, and constraint parameters;
[0022] Solve the optimization objective for each bucket through an integer programming solution algorithm, incentive parameters, and constraint parameters to obtain the incentive strategy corresponding to each bucket.
[0023] In one embodiment, the solution of the optimization objective for each bucket through an integer programming solution algorithm, incentive parameters, and constraint parameters includes:
[0024] Perform two-way search on the incentive parameters within the first preset search range to obtain the optimal incentive parameters;
[0025] Perform two-way search on the constraint parameters within the second preset search range to obtain the optimal constraint parameters;
[0026] Solve the optimization objective based on the optimal incentive parameters and the optimal constraint parameters.
[0027] In one embodiment, after the two-way search for the incentive parameters according to the preset search range, it includes:
[0028] If the solution fails and the proportion of incentivized users is equal to the preset threshold, select the minimum incentive amount of the first preset proportion from the list of different incentive amounts, and randomly select the incentive amounts for the users of the first remaining proportion from the list of different incentive amounts, where the sum of the first preset proportion and the first remaining proportion is 1;
[0029] If the solution fails and the proportion of incentivized users is less than the preset threshold, randomly select the incentive amounts of the second preset proportion from the list of different incentive amounts, and no incentive amounts are distributed to the users of the second remaining proportion, where the sum of the second preset proportion and the second remaining proportion is 1.
[0030] In one embodiment, after obtaining the incentive strategy corresponding to the target user, it further includes:
[0031] Evaluate the incentive strategy to obtain evaluation information;
[0032] Determine whether there is an abnormality in the evaluation information;
[0033] If not, review the incentive strategy, and execute the incentive strategy after passing the review.
[0034] In a second aspect, an offline policy automated processing device is provided, including:
[0035] An offline data acquisition unit, configured to acquire the offline feature data and user type corresponding to the target user;
[0036] A prediction unit, configured to perform prediction based on the user type and the offline feature data through the corresponding target prediction model to obtain the user order placement probabilities of the target user under different incentive amounts;
[0037] A policy solving unit, configured to perform policy solving based on the user order placement probabilities and the user type through the corresponding policy solving method to obtain the incentive strategy corresponding to the target user.
[0038] In a third aspect, a readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the offline policy automated processing method as described above is implemented.
[0039] The above offline policy automation processing method, device, and storage medium. The implementation of the method includes: obtaining offline feature data corresponding to a target user and the user type; based on the user type and the offline feature data, performing prediction through a corresponding target prediction model to obtain the user order placement probability of the target user under different incentive amounts; based on the user order placement probability and the user type, performing policy solution through a corresponding policy solution method to obtain the incentive policy corresponding to the target user. In the embodiments of the present application, if the user is a non-retained user, such as a new user or a recalled user, a corresponding model can be selected based on the user type to predict the order placement probability under different incentive amounts, and a corresponding solution method can be selected based on the order placement probability and the user type to perform policy solution, which can cover any population such as new customers, recalled users, and retained users, and comprehensively meet the operation requirements. And by adopting an offline policy automation processing process, the labor cost of the algorithm is saved, and at the same time, the frequent information synchronization between the operation and the algorithm side is avoided, and unnecessary errors are avoided. Moreover, the online time of the operation activity is also shortened at the same time, greatly improving the operation efficiency. In addition, after the optimization of the algorithm policy, constraint parameters (such as additional ROI) have also been significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 is a schematic diagram of an application environment of the offline policy automation processing method in an embodiment of the present invention;
[0042] Figure 2 is a schematic flowchart of the offline policy automation processing method in an embodiment of the present invention;
[0043] Figure 3 is a schematic structural diagram of the offline policy automation processing device in an embodiment of the present invention;
[0044] Figure 4 is a schematic diagram of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0046] In one embodiment, as Figure 1 , Figure 2 shown, a method for automatically processing offline policies is provided, including the following steps:
[0047] In step S110, obtain the offline feature data corresponding to the target user and the user type;
[0048] Among them, the offline feature data may include the user's historical order data, such as consumption habits, preferences, etc.
[0049] Optionally, user information can be collected, and a user portrait can be constructed. Then, through the operation staff in the user portrait platform, user groups of different user types can be selected according to user tags, such as historical completed order volume, dormant days, etc. Then, the interface parameters can be configured in the process canvas system, as shown in Table 1 below. Then, the configured interface parameters can be sent to the algorithm platform in the form of an interface for order placement probability prediction and policy solution.
[0050] Table 1: Interface Parameters
[0051] Parameter Name Type Description canvas_id int Canvas ID task_creator_modifier string Activity Creator / Modifier (separated by commas if multiple) tca_id int Strategy ID persons_id int Population Portrait ID target_subsidy_rate float Target Subsidy Rate 0.1 - 10% Subsidy Rate discount_user_rate float Proportion of Subsidized Users 0 - 100 packet_details list <map> < / map> Coupon Packet List input_file_path string Input File Path ( / home / data.csv) evaluate_res_file string Estimation Result File ( / home / evaluation.json) output_results_file string Coupon Sending Result
[0052] Among them, the interface parameters may include incentive amounts, such as different coupon amounts, subsidy amounts, discount ratios, full reduction amounts, etc., core parameters such as validity periods, subsidy rates, budgets, etc.
[0053] It should be noted that users can be divided into user types such as new users, recalled users, and retained users according to their attributes. Among them, new users refer to users who have not placed orders before. Recalled users refer to users who have been dormant for a long time or are low-frequency users. Retained users refer to users who often place orders or often use application programs.
[0054] In step S120, based on the user type and the offline feature data, perform prediction through the corresponding target prediction model to obtain the user order placement probability of the target user under different incentive amounts;
[0055] Optionally, for different business lines, such as logistics, moving, taxi-hailing, e-commerce, etc., and different user groups, such as new users, recalled users, retained users, etc., their data characteristics and data distributions often vary. Therefore, different prediction models can be configured separately, and these prediction models are trained based on the characteristic data of different user groups in different business lines. For example, operation personnel may be more concerned about the likelihood of new users' first purchase, while for old users, they may be more concerned about their repeat purchase behavior. The corresponding prediction models will be trained according to different business requirements and user characteristics to obtain different models. These models will use the data of the corresponding business line and population as input when making predictions and output corresponding prediction results.
[0056] It should be noted that different prediction models are trained using machine learning libraries such as lightgbm or xgboost, and these models are stored in corresponding files according to different business lines and population types. These models cover the model parameters, structures, and algorithm logics obtained from training. Since an important problem will be encountered when using Python for parallel computing, that is, Python model objects cannot be shared among multiple processes. In the multi-process environment of Python, different processes have their own independent memory spaces, and prediction models usually exist in the form of objects, such as the models trained using libraries like lightgbm or xgboost. When using multiple processes for parallel computing, directly sharing these model objects among different processes is not feasible, which will lead to various problems, such as data inconsistency between processes, memory access conflicts, etc., and the efficiency and accuracy of parallel computing cannot be achieved.
[0057] Therefore, the open-source framework m2cgen (Model 2Code Generator) can be used to convert each LightGBM / XGBoost model file into a Python prediction function code file. m2cgen will parse the internal structure and parameters of the model and convert them into equivalent Python code. This conversion is based on the algorithm and structure of the model, converting the model information originally stored in the object into a series of Python function codes, which implement the same prediction logic as the original model. For example, if the original LightGBM model is a gradient boosting decision tree model, m2cgen will express the node information, splitting conditions, predicted values of the leaf nodes, etc. of the decision tree in the form of Python code to form a complete prediction function. This function receives input data, processes the data according to the decision rules and logic of the model, and finally outputs the prediction result. After the conversion by m2cgen, the finally generated is an independent Python prediction function code file, which contains the function to implement the prediction function. This function can be conveniently called in the multi-process environment of Python without the problem of multi-process shared objects. Because the code can be copied and executed in different processes, and the prediction function of the original model can be correctly implemented in different processes, thus avoiding the need to share the model object in the multi-process environment.
[0058] Optionally, if the target user is a non-retained user, then it can be a new user or a recalled user. At this time, based on the user type of the target user, the corresponding prediction model can be selected to predict the order placement probability. Exemplarily, different incentive amount ranges can be set, such as the coupon amount, which can be set to [0, 3, 5, 8, 10, 15, 20, 25]. Then, the prediction model trained by the training data of the user type to which the target user belongs can be selected, such as LightGBM / XGBoost. Then, based on the Python prediction function code file corresponding to the model, the order placement probability of the user under different incentive amounts is predicted.
[0059] In step S130, based on the user order placement probability and the user type, the corresponding policy solution method is used to solve the policy to obtain the incentive policy corresponding to the target user.
[0060] It should be noted that for different user types, different solution strategies can be selected for solving. For example, for new users, the order placement probability strategy can be selected for solving. For recalled users, the mixed weighting strategy can be selected for solving. For retained users, the integer programming optimization method can be used for policy solution.
[0061] Among them, the order placement probability strategy refers to bucketing the user order placement probability, that is, dividing users with the same or similar order placement probabilities into an interval so that the same strategy can be used for users within the same interval. For the group of people with a higher order placement probability, the incentive amount is higher. The hybrid weighting strategy means that after different weight coefficients are weighted for the three dimensions of elasticity, order placement probability, and historical average order price, the final weight coefficient is obtained, and then bucketing can be performed based on the weight coefficient, that is, dividing users with the same or similar weight coefficients into an interval so that the same strategy can be used for users within the same interval. For the person with a higher weight coefficient, the incentive amount is higher. It should be noted that the number of buckets is equal to the number of incentive amounts.
[0062] It should be noted that the incentive strategy refers to the strategy used to motivate users to place orders, which may include specific preferential amounts, discount ratios, subsidy amounts, subsidy rates, etc. corresponding to each user.
[0063] Optionally, after determining the incentive strategy, the incentive strategy can be evaluated to obtain an evaluation result, which may include the estimated subsidy rate, the magnitude of the sub-population, the estimated written-off amount, and the proportion of estimated subsidized users. After the estimated information is returned to the event side, the operation staff will receive a message indicating that the update of the estimated subsidy rate is successful. At this time, the operation staff can view the specific estimated subsidy rate and other information of this event. If any abnormality is found, the event can be suspended, and after repair, it can be put online again. If no abnormality is found, the operation side can review and release the event online. After the event is launched, according to the event execution time, starting from a preset time in advance, for example, 2 hours, the offline coupons will be distributed to the corresponding users.
[0064] Taking the estimated subsidy rate as an example, the total subsidy amount to be distributed can be calculated. This requires multiplying the incentive amount (such as the preferential amount, subsidy amount, etc.) of each user by the expected number of users to be distributed. For different types of incentives, such as coupons, the subsidy amount needs to be calculated based on its discount ratio and the order amount expected to use this coupon; for the subsidy amount, it can be directly added up. Then divide the total subsidy amount by the expected total sales amount to obtain the estimated subsidy rate. Compare the estimated subsidy rate of this event with the subsidy rates of similar events in the company's history. If the current estimated subsidy rate is much higher than the average subsidy rate of historical events, it may indicate that the cost of this incentive strategy is too high, and it is necessary to consider adjusting the intensity or scope of the incentive to control the cost; on the contrary, if it is much lower than the historical average level, it may not be able to achieve the expected user incentive effect, and it is necessary to consider whether to increase the incentive intensity. The estimated written-off amount can be determined based on the ratio of the written-off amount to the total issued preferential amount, and the estimated written-off amount should be compared with the expected revenue, and based on the comparison result, the evaluation result can be determined.
[0065] In an embodiment of the present application, an offline policy automated processing method is provided, including: obtaining offline feature data corresponding to a target user and the user type; based on the user type and the offline feature data, predicting through a corresponding target prediction model to obtain the user order placement probability of the target user under different incentive amounts; based on the user order placement probability and the user type, performing policy solving through a corresponding policy solving method to obtain an incentive policy corresponding to the target user. In an embodiment of the present application, if the user is a non-retained user, such as a new user or a recalled user, a corresponding model can be selected based on the user type to predict the order placement probability under different incentive amounts, and a corresponding solving method can be selected based on the order placement probability and the user type for policy solving, which can cover any population such as new customers, recalled users, and retained users, and comprehensively meet the operation requirements. And by adopting an offline policy automated processing process, the labor cost of the algorithm is saved, and at the same time, the frequent information synchronization between the operation side and the algorithm side is avoided, and unnecessary errors are avoided. Moreover, the online time of the operation activity is also shortened, greatly improving the operation efficiency. In addition, after the optimization of the algorithm policy, constraint parameters (such as additional ROI) have also been significantly improved.
[0066] In an embodiment of the present application, the user type is a new user, and the performing policy solving based on the user order placement probability and the user type includes:
[0067] Performing equal-frequency binning on the user order placement probability, and the number of bins is equal to the number of different incentive amounts;
[0068] Based on the user order placement probability corresponding to each bin, determining a corresponding incentive policy, where the larger the user order placement probability of the bin, the larger the incentive amount distributed.
[0069] Optionally, the user order placement probability corresponding to each predicted incentive amount can be collected. For example, the user order placement probability list is [0.1, 0.3, 0.2, 0.55, 0.4, 0.7, 0.6, 0.85, 0.9, 0.25, 0.45]. The incentive amount list is [0, 5, 10, 15, 20]. Since the number of bins is equal to the number of different incentive amounts, the user order placement probability can be equally divided into 5 bins. The above order placement probabilities can be sorted to obtain [0.1, 0.2, 0.25, 0.3, 0.4, 0.45, 0.55, 0.6, 0.7, 0.85, 0.9], and then the quantile positions are calculated, and binning operations are performed on the sorted data based on the quantiles.
[0070] Then, a one-to-one correspondence can be established between the index of each bucket and the elements in the incentive amount list. For example, bucket 0 corresponds to incentive amount 0, bucket 1 corresponds to incentive amount 5, bucket 2 corresponds to incentive amount 10, bucket 3 corresponds to incentive amount 15, and bucket 4 corresponds to incentive amount 20. Based on the user's order placement probability, find the bucket index where the user is located, and then select the corresponding incentive amount from the incentive amount list according to this index. For each user's order placement probability, the corresponding incentive strategy can be obtained. Among them, users with a lower order placement probability will be assigned to buckets with a lower index and thus receive a lower incentive amount; users with a higher order placement probability will be assigned to buckets with a higher index and thus receive a higher incentive amount.
[0071] In an embodiment of the present application, the target user is a recalled user or a retained user. Based on the user's order placement probability and the user type, policy solving is performed through a corresponding policy solving method to obtain the incentive strategy corresponding to the target user, including:
[0072] Based on the user's order placement probability, determine the user elasticity, where the user elasticity reflects the sensitivity of the user's order placement probability to changes in the incentive amount;
[0073] Based on the offline feature data, determine the historical average order price corresponding to the target user;
[0074] Perform equal-frequency bucketing on the historical average order price, user elasticity, and user's order placement probability. The number of buckets is equal to the number of different incentive amounts;
[0075] Based on the bucketing result, determine the incentive strategy corresponding to each bucket.
[0076] Optionally, for both recalled users and retained users, it is necessary to first predict the user's order placement probability, then determine the user elasticity based on the user's order placement probability, and further determine the historical average order price of the user based on the offline feature data. Then, perform bucketing processing on each dimension based on the number of incentive amounts. The number of buckets is the same as the number of incentive amounts. Then, determine the incentive strategy corresponding to each bucket respectively.
[0077] Exemplarily, as Figure 1 shown, taking recalled users as an example, the specific solving process is as follows:
[0078] First, the operation staff can pre-select the recall population A on the portrait platform based on user tags, such as historical completed order volume, dormant days, etc. Then, after configuring the core parameters such as the incentive amount B, its validity period, the target subsidy rate C, the budget D, etc. in the process canvas system, after filling in the activity information and clicking save, at this time, the relevant configuration parameter information on the activity page can be as shown in Table 1 above, and is transmitted to the algorithm engineering in the form of an interface. After receiving the parameters such as the canvas ID, strategy ID, population ID, activity creator, target subsidy rate, subsidy user ratio, coupon package details, etc. passed from the upstream, the algorithm engineering side conducts parameter verification, then obtains the user ID offline according to the population portrait ID, and obtains the corresponding offline feature data of the user from the ES system according to the user ID, and stores all the user offline feature data in a local csv file. Finally, all the parameters in Table 1 above are transmitted to the python integer programming solver through the command line for policy solving and the generation of incentive policy results.
[0079] Optionally, the online configuration file information can be read first, including the objectives of optimization solving (such as maximizing GMV or order volume), constraint conditions (such as additional ROI), and then the user order placement probabilities of the recalled users under different incentive amounts are predicted according to the selected target prediction model corresponding to the recalled users. For example, if the incentive amount list is [0, 3, 5, 8, 10, 15, 20, 25], the corresponding user order placement probability list can be [0.3, 0.5, 0.6, 0.67, 0.7, 0.75, 0.8, 0.8], and then the slope of the straight line is obtained by linear fitting as the elasticity. For example, the means of the incentive amount list and the user order placement probability list can be calculated respectively, that is, all their elements are added and divided by the number of elements, and then based on this mean, the slope is calculated.
[0080] Then, the recalled population can be bucketed according to the three dimensions of order placement probability - historical order average price - elasticity. The number of buckets can be the same as the number of elements in the incentive amount list. Among them, the user order placement probability can be divided into 5 buckets, the elasticity can be divided into 10 buckets, and the historical order average price can be divided into 5 buckets, with a total of 250 groups. Then, the actual value of each incentive amount in the incentive amount list for each group is calculated according to the average historical order average price of each group, for example, by methods such as proportional adjustment method, difference adjustment method or user value contribution adjustment method.
[0081] The obtained solution data format can be as shown in Table 2 below:
[0082] Table 2: Solution data format
[0083]
[0084] Based on the above-mentioned solution data format, the incentive strategy for each group can be obtained through optimized solution. For example, the coupon issuance result of distributing coupons, and the coupon issuance result can be as shown in Table 3 below.
[0085] Table 3: Coupon Issuance Result:
[0086] User Ordering Probability Binning Elastic Binning Historical Average Order Price Binning Coupon Sending Result prob_level1 elastic_level1 price_level1 3 prob_level1 elastic_level1 price_level3 8 prob_level1 elastic_level2 price_level1 10
[0087] After the strategy solution, the obtained incentive strategy can be evaluated to obtain evaluation information, which may include the estimated subsidy rate, the magnitude of the sub-population, the estimated redemption amount, and the proportion of estimated subsidized users. After obtaining the evaluation information, the evaluation information can be sent to the incentive activity party, and the operation staff will receive a message indicating that the estimated subsidy rate has been successfully updated. At this time, the operation staff can view the specific estimated subsidy rate and other information of the activity. If any abnormality is found, the activity can be suspended and then launched again after repair.
[0088] If no abnormality is found in the evaluation information, the incentive strategy can be audited and released. After the incentive activity is launched, it can be calculated 2 hours in advance according to the execution time, and the offline coupons can be distributed to users.
[0089] In an embodiment of the present application, the target user is a recalled user, and determining the incentive strategy corresponding to each bucket based on the bucketing result includes:
[0090] Weighted summation is respectively performed on the historical average order price, user elasticity, and user order probability to obtain the weight coefficient corresponding to each bucket;
[0091] Based on the weight coefficient corresponding to each bucket, determine the incentive strategy corresponding to each bucket, where the larger the weight coefficient of the bucket, the greater the incentive amount.
[0092] Optionally, equal-frequency bucketing can be respectively performed on the historical average order price, user elasticity, and user order probability based on the number of elements in the incentive amount list. For example, the weight of the historical average order price is 0.2, the weight of user elasticity is 0.4, and the weight of user order probability is 0.2. The historical average order price can be divided into A1, A2, and A3, user elasticity can be divided into B1, B2, and B3, and user order probability can be divided into C1, C2, and C3. Then the weight coefficient of each bucket can be calculated respectively. For example, bucket 1 (weight coefficient) = A1 * 0.2 + B1 * 0.4 + C1 * 0.2. Bucket 2 (weight coefficient) = A2 * 0.2 + B2 * 0.4 + C2 * 0.2. Bucket 3 (weight coefficient) = A3 * 0.2 + B3 * 0.4 + C3 * 0.2. Thus, the weight coefficient corresponding to each bucket can be obtained, and then the incentive amount is allocated based on the weight coefficient. The larger the weight coefficient, the greater the allocated incentive amount.
[0093] In an embodiment of the present application, the target user is a retained user. Determining an incentive strategy corresponding to each bucket based on the bucketing result includes:
[0094] Determine the optimization objective, incentive parameters, and constraint parameters;
[0095] Solve the optimization objective for each bucket through an integer programming solution algorithm, incentive parameters, and constraint parameters to obtain the incentive strategy corresponding to each bucket.
[0096] Among them, the incentive parameter can be the subsidy rate, the constraint parameter can be the additional ROI, generally 2, and the optimization objective is GMV (Gross Merchandise Volume, total transaction amount). Then, the dataset D, the incentive amount list C, and the subsidy user proportion P can be obtained. The dataset D can include the user's order placement probability, user elasticity, historical single order average price, and offline feature data. Input the above data into the integer programming optimization model. First, according to the number of elements in the incentive amount list, perform equal-frequency bucketing operations on the user's order placement probability, user elasticity, and historical single order average price respectively. Then, perform automatic parameter search for each bucket, and solve the optimization objective through the integer programming solution algorithm to obtain the incentive strategy corresponding to each bucket.
[0097] In an embodiment of the present application, solving the optimization objective for each bucket through the integer programming solution algorithm, incentive parameters, and constraint parameters includes:
[0098] Perform a two-way search on the incentive parameters within the first preset search range to obtain the optimal incentive parameters;
[0099] Perform a two-way search on the constraint parameters within the second preset search range to obtain the optimal constraint parameters;
[0100] Solve the optimization objective based on the optimal incentive parameters and the optimal constraint parameters.
[0101] Optionally, the automatic parameter search process is specifically: taking Grid Search as an example, the first-layer loop: "For the incentive parameter (i.e., the subsidy rate B), its search range is set to start from the initial value B, increment by 1 to 99, and start from B - 1, decrement by -1 to 1. The second-layer loop: the search range for the additional ROI is from R to 0 (step size -0.5), thus forming a parameter grid. Traverse the parameter grid, substitute each set of parameter combinations into the optimization objective for solution, and set the result as S. If S is not empty, return S; otherwise, continue the loop. Among them, the input parameters of the integer programming solution function are (b, r, gmv, D, C).; otherwise, continue the loop. Evaluate the objective function value under each combination, and find the parameter combination that makes the objective function optimal, that is, the optimal incentive parameter and the optimal constraint parameter.
[0102] It is understandable that random search or heuristic search can also be used to optimize the objective solution.
[0103] In an embodiment of the present application, after the two-way search for the incentive parameters according to the preset search range, it includes:
[0104] If the solution fails and the proportion of incentivized users is equal to the preset threshold, the minimum incentive amount of the first preset proportion is selected from different incentive amount lists, and the incentive amounts for the users of the first remaining proportion are randomly selected from the different incentive amount lists, where the sum of the first preset proportion and the first remaining proportion is 1;
[0105] If the solution fails and the proportion of incentivized users is less than the preset threshold, the incentive amounts of the second preset proportion are randomly selected from the different incentive amount lists, and no incentive amounts are distributed to the users of the second remaining proportion, where the sum of the second preset proportion and the second remaining proportion is 1.
[0106] Optionally, in extreme cases, it may happen that the result is empty after the automatic search. At this time, in order to be able to normally evaluate the subsidy rate, a default result needs to be generated. The specific generation rule is as follows: If the proportion P of subsidized users is equal to 100, 80% of the grids select the minimum incentive amount from the incentive amount list C, and the remaining 20% are randomly selected from the remaining incentive amounts in C. If the proportion P of subsidized users is less than 100, 80% of the users are not incentivized, and the remaining 20% are randomly selected from the incentive amount list in C where the incentive amount is greater than 0.
[0107] In an embodiment of the present application, after obtaining the incentive strategy corresponding to the target user, it further includes:
[0108] Evaluate the incentive strategy to obtain evaluation information;
[0109] Determine whether there is an abnormality in the evaluation information;
[0110] If not, audit the incentive strategy, and execute the incentive strategy after passing the audit.
[0111] Optionally, after the strategy solution, the obtained incentive strategy can be evaluated to obtain evaluation information, which may include the estimated subsidy rate, the magnitude of the sub-population, the estimated write-off amount, and the estimated proportion of subsidized users. After obtaining the evaluation information, the evaluation information can be sent to the incentive activity party, and the operation personnel will receive a message indicating that the estimated subsidy rate has been updated successfully. At this time, the operation personnel can view the specific estimated subsidy rate and other information of the activity, as shown in Table 4 below. If any abnormality is found, the activity can be suspended and then launched again after repair.
[0112] Table 4: Estimated subsidy information
[0113]
[0114]
[0115] If no abnormalities are found in the evaluation information, the incentive strategy can be reviewed and released. After the incentive activity goes live, it can be calculated 2 hours in advance according to the execution time, and the offline coupons can be distributed to users.
[0116] In the embodiments of the present application, when the user is a non-retained user, such as a new user or a recalled user, a corresponding model can be selected based on the user type to predict the order placement probability under different incentive amounts, and a corresponding solution method can be selected based on the order placement probability and the user type for strategy solution. It can cover any population such as new customers, recalled customers, and retained customers, and fully meets the operation requirements. And an offline strategy automated processing process is adopted, saving more than 80% of the labor cost of the algorithm. At the same time, it also avoids the frequent information synchronization between the operation and the algorithm side, and avoids unnecessary errors. And the online time of the operation activity is also shortened from 1-2 days to 0.5-1.5 hours, greatly improving the operation efficiency. In addition, after the optimization of the algorithm strategy, the constraint parameters (such as additional ROI) have also been significantly improved. By automatically adjusting the parameters to generate the corresponding incentive strategy and automatically launching the relevant incentive activity, the manual intervention on the algorithm side is reduced, and the cost of manually adjusting the parameters and manually launching the strategy offline is reduced; the strategy automated process and the process canvas system are connected, so that after the operation configures the activity on the process canvas system, clicking the estimation button can automatically generate the algorithm strategy of the activity. This greatly reduces the synchronization of experimental information between the algorithm and the operation.
[0117] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0118] In one embodiment, an offline strategy automated processing device is provided, and the offline strategy automated processing device corresponds one-to-one to the offline strategy automated processing method in the above embodiment. As Figure 3 shown, the offline strategy automated processing device includes an offline data acquisition unit 10, a prediction unit 20, and a strategy solution unit 30. The detailed description of each functional module is as follows:
[0119] The offline data acquisition unit 10 is used to acquire the offline feature data corresponding to the target user and the user type;
[0120] A prediction unit 20, configured to perform prediction through a corresponding target prediction model based on the user type and the offline feature data, so as to obtain the user order placement probabilities of the target user under different incentive amounts;
[0121] A policy solving unit 30, configured to perform policy solving through a corresponding policy solving method based on the user order placement probabilities and the user type, so as to obtain the incentive policy corresponding to the target user.
[0122] In an embodiment of the present application, when the target user is a new user, the policy solving unit 30 is further configured to:
[0123] Perform equal-frequency binning on the user order placement probabilities, where the number of bins is equal to the number of different incentive amounts;
[0124] Determine the corresponding incentive policy based on the user order placement probability corresponding to each bin, where the larger the user order placement probability of a bin, the larger the incentive amount distributed.
[0125] In an embodiment of the present application, when the target user is a recalled user or a retained user, the policy solving unit 30 is further configured to:
[0126] Determine user elasticity based on the user order placement probabilities, where the user elasticity reflects the sensitivity of the user order placement probability to changes in the incentive amount;
[0127] Determine the historical average order price corresponding to the target user based on the offline feature data;
[0128] Perform equal-frequency binning on the historical average order price, user elasticity, and user order placement probabilities, where the number of bins is equal to the number of different incentive amounts;
[0129] Determine the incentive policy corresponding to each bin based on the binning result.
[0130] In an embodiment of the present application, when the target user is a recalled user, the policy solving unit 30 is further configured to:
[0131] Perform weighted summation on the historical average order price, user elasticity, and user order placement probabilities respectively to obtain the weight coefficient corresponding to each bin;
[0132] Determine the incentive policy corresponding to each bin based on the weight coefficient corresponding to each bin, where the larger the weight coefficient of a bin, the larger the incentive amount.
[0133] In an embodiment of the present application, when the target user is a retained user, the policy solving unit 30 is further configured to:
[0134] Determine an optimization objective, incentive parameters, and constraint parameters;
[0135] Solve the optimization objective per barrel through an integer programming solution algorithm, incentive parameters, and constraint parameters to obtain the corresponding incentive strategy per barrel.
[0136] In an embodiment of the present application, the policy solving unit 30 is further configured to:
[0137] Perform a two-way search on the incentive parameters according to a first preset search range to obtain the optimal incentive parameters;
[0138] Perform a two-way search on the constraint parameters according to a second preset search range to obtain the optimal constraint parameters;
[0139] Solve the optimization objective based on the optimal incentive parameters and the optimal constraint parameters.
[0140] In an embodiment of the present application, the policy solving unit 30 is further configured to:
[0141] If the solution fails and the proportion of incentivized users is equal to the preset threshold, select the minimum incentive amount of a first preset proportion from different incentive amount lists, and randomly select the incentive amounts for the users of the first remaining proportion from the different incentive amount lists, where the sum of the first preset proportion and the first remaining proportion is 1;
[0142] If the solution fails and the proportion of incentivized users is less than the preset threshold, randomly select the incentive amounts of a second preset proportion from the different incentive amount lists, and do not issue incentive amounts to the users of the second remaining proportion, where the sum of the second preset proportion and the second remaining proportion is 1.
[0143] In an embodiment of the present application, the device further includes a policy evaluation unit, which is used for:
[0144] Evaluate the incentive strategy to obtain evaluation information;
[0145] Determine whether there is an abnormality in the evaluation information;
[0146] If not, audit the incentive strategy, and execute the incentive strategy after passing the audit.
[0147] In the embodiments of the present application, when the user is a non-retained user, such as a new user or a recalled user, a corresponding model can be selected based on the user type to predict the order placement probability under different incentive amounts, and a corresponding solution method can be selected based on the order placement probability and the user type for strategy solution. It can cover any population such as new customers, recalled users, and retained users, fully meeting the operation requirements. And by adopting an offline strategy automated processing process, the labor cost of the algorithm is saved by more than 80%, and at the same time, the frequent information synchronization between the operation side and the algorithm side is avoided, preventing unnecessary errors. Also, the online time of the operation activity is shortened from 1 - 2 days to 0.5 - 1.5 hours, greatly improving the operation efficiency. In addition, after the optimization of the algorithm strategy, constraint parameters (such as additional ROI) have also been significantly improved. By automatically adjusting parameters to generate corresponding incentive strategies and automatically launching relevant incentive activities, the manual intervention on the algorithm side is reduced, and the cost of manually adjusting parameters offline and manually launching strategies is lowered; the automated strategy process and the process canvas system are integrated, so that after the operation configures an activity on the process canvas system and clicks the estimation button, the algorithm strategy of the activity can be automatically generated. This greatly reduces the synchronization of experimental information between the algorithm and the operation.
[0148] For the specific limitations of the offline strategy automated processing device, reference can be made to the limitations of the offline strategy automated processing method described above, and details will not be elaborated here. Each module in the above offline strategy automated processing device can be implemented in whole or in part through software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0149] In one embodiment, a computer device is provided. This computer device can be a terminal device, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer-readable instructions are executed by the processor, an offline strategy automated processing method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0150] In the embodiments of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the offline strategy automated processing method as described above are implemented.
[0151] In an application embodiment, a readable storage medium is provided. The readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the offline policy automation processing method as described above are implemented.
[0152] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the above-described method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above-described division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0154] The above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An offline strategy automation processing method, characterized in that: The method comprises: Obtain offline feature data and user type corresponding to the target user; Based on the user type and offline feature data, a corresponding target prediction model is used to predict the probability of the target user placing an order under different incentive amounts; Based on the user order probability and the user type, a strategy solution is performed through a corresponding strategy solution method to obtain an incentive strategy corresponding to the target user.
2. The offline strategy automation processing method according to claim 1, characterized in that: The user type is a new user, and the strategy solving is performed based on the user order probability and the user type, including: Divide the user's order probability into equal-frequency buckets, where the number of buckets is equal to the number of different incentive amounts; Based on the probability of user orders corresponding to each bucket, the corresponding incentive strategy is determined, where the bucket with a greater probability of user orders has a greater incentive amount.
3. The offline strategy automation processing method according to claim 1, characterized in that: The target user is a recalled user or a retained user, and the strategy solving is performed by a corresponding strategy solving method based on the user order probability and the user type, including: Determine user elasticity based on the user order probability, where the user elasticity reflects the sensitivity of the user order probability to changes in the incentive amount; Determine the historical average price of the target user based on the offline feature data; The historical average order price, user elasticity, and user order probability are divided into equal-frequency buckets, and the number of buckets is equal to the number of different incentive amounts; Based on the bucketing results, determine the incentive strategy corresponding to each bucket.
4. The offline strategy automation processing method according to claim 3, characterized in that: The target user is the recalled user, and the incentive strategy corresponding to each bucket is determined based on the bucketing result, including: The weighted sum of the historical average order price, user elasticity, and user order probability is respectively performed to obtain a weight coefficient corresponding to each bucket; Based on the weight coefficient corresponding to each bucket, the incentive strategy corresponding to each bucket is determined, wherein the larger the weight coefficient of the bucket, the greater the incentive amount.
5. The offline strategy automation processing method according to claim 3, characterized in that: The target user is a retained user, and the incentive strategy corresponding to each bucket is determined based on the bucketing result, including: Determine the optimization objectives, incentive parameters, and constraint parameters; The optimization target under each bucket is solved through integer programming solving algorithm, incentive parameters and constraint parameters to obtain the incentive strategy corresponding to each bucket.
6. The offline strategy automation processing method according to claim 5, characterized in that: The optimization target under each bucket is solved by the integer programming solution algorithm, incentive parameters and constraint parameters, including: Performing a bidirectional search on the excitation parameters according to a first preset search range to obtain an optimal excitation parameter; Performing a bidirectional search on the constraint parameter according to a second preset search range to obtain an optimal constraint parameter; The optimization objective is solved based on the optimal excitation parameters and the optimal constraint parameters.
7. The offline strategy automation processing method according to claim 5, characterized in that: After the bidirectional search of the excitation parameter is performed according to the preset search range, the method includes: If the solution fails and the proportion of incentivized users is equal to the preset threshold, the minimum incentive amount of the first preset proportion is selected from the list of different incentive amounts, and the incentive amount of users of the first remaining proportion is randomly selected from the list of different incentive amounts, and the sum of the first preset proportion and the first remaining proportion is 1; If the solution fails and the proportion of incentivized users is less than the preset threshold, a second preset proportion of incentive amounts is randomly selected from the list of different incentive amounts, and no incentive amounts are issued to users of the second remaining proportion, and the sum of the second preset proportion and the second remaining proportion is 1.
8. The offline strategy automation processing method according to any one of claims 1 to 7, characterized in that: After obtaining the incentive strategy corresponding to the target user, the method further includes: Evaluating the incentive strategy to obtain evaluation information; Determining whether the evaluation information is abnormal; If not, the incentive strategy is reviewed and executed after passing the review.
9. An offline strategy automation processing device, characterized in that: The device comprises: An offline data acquisition unit, used to acquire offline feature data and user type corresponding to the target user; A prediction unit, configured to perform prediction based on the user type and offline feature data through a corresponding target prediction model to obtain the probability of the target user placing an order under different incentive amounts; A strategy solving unit is used to solve the strategy through a corresponding strategy solving method based on the user's order probability and the user type, so as to obtain an incentive strategy corresponding to the target user.
10. A readable storage medium having computer readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, the offline strategy automation processing method according to any one of claims 1 to 8 is implemented.