A strategy recommendation method and device

By analyzing historical data and predictive models, we screened out the most optimized merchant discount strategies, solved the problem of poor merchant profits, and achieved an increase in order conversion rate and revenue.

CN113704619BActive Publication Date: 2025-09-26BEIJING SANKUAI ONLINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111011336.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-09-26
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

In the existing technology, when merchants set up preferential policies, they are unable to maximize their profits, resulting in an irrational increase in order conversion rates.

Method used

Through historical data analysis, we screen effective strategy types and parameters, use pre-trained prediction models to predict merchants' transfer rewards and transfer probabilities under different strategies, and recommend the most optimized strategy.

Benefits of technology

The order conversion rate and revenue of merchants on e-commerce platforms or food delivery platforms have been improved, and the recommended strategies are more reasonable and feasible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113704619B_ABST
    Figure CN113704619B_ABST
Patent Text Reader

Abstract

This specification discloses a strategy recommendation method and device, which can determine a number of candidate strategy types based on the executed strategy types and corresponding rewards of each merchant in history. Afterwards, based on the merchant characteristics of each merchant to be optimized, the strategy parameters under each candidate strategy type, each candidate parameter, and the historical transition probability and historical transfer reward of each merchant for parameter transfer in history, the transition probability and transfer reward of the merchant to be optimized from the strategy parameter to each candidate parameter are predicted. Finally, based on at least one of the transition probability and transfer reward corresponding to each candidate parameter under each candidate strategy type, the optimization strategy recommended to each merchant to be optimized is determined. By determining the transition probability and transfer reward of each merchant to be optimized from the strategy parameter to each candidate parameter under each candidate strategy type, and based on at least one of the transition probability and transfer reward, an accurate or easy-to-implement optimization strategy is recommended to the merchant, thereby promoting the perfect optimization of the merchant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a strategy recommendation method and device. Background Art

[0002] With the rapid development of the Internet, online shopping has gradually become one of people's mainstream consumption methods, so more and more merchants sell goods through e-commerce platforms / takeaway platforms.

[0003] To increase user conversion rates, businesses often implement preferential policies, such as free shipping, reduced delivery fees, or discounts on purchases above a certain amount. Therefore, how to set preferential policies to maximize business profits is an urgent issue that needs to be addressed. Summary of the Invention

[0004] The embodiments of this specification provide a strategy recommendation method and device for partially solving the problems in the prior art.

[0005] The embodiments of this specification adopt the following technical solutions:

[0006] This manual provides a strategy recommendation method, including:

[0007] Determine several candidate strategy types based on the historically executed strategy types and corresponding rewards of each merchant;

[0008] For each merchant to be optimized, determine the strategy parameters of the merchant to be optimized under each strategy type to be selected;

[0009] For each candidate parameter within a specified range under each candidate strategy type, predict the transfer reward of the merchant to be optimized from the strategy parameter to the candidate parameter using a pre-trained first prediction model based on a first feature set, as the transfer reward corresponding to the candidate parameter, wherein the first feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameter of the merchant to be optimized under the candidate strategy type, and the candidate parameter; and / or

[0010] For each candidate parameter within a specified range under each candidate strategy type, a transition probability of the merchant to be optimized from the strategy parameter to the candidate parameter is predicted using a pre-trained second prediction model based on a second feature set, as the transition probability corresponding to the candidate parameter. The second feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the candidate strategy type, the candidate parameters, and historical transition probabilities and historical transfer rewards for each merchant from the strategy parameter to the candidate parameter.

[0011] An optimization strategy is determined based on at least one of the transfer probabilities and transfer rewards corresponding to each candidate parameter under each candidate strategy type, and the optimization strategy is recommended to the merchant to be optimized so that the merchant to be optimized can be optimized.

[0012] Optionally, several candidate strategy types are determined based on the historically executed strategy types and corresponding rewards of each merchant, including:

[0013] Based on the historical executed strategy types of each merchant, select several executed strategy types whose proportion of executed merchants exceeds a first preset threshold;

[0014] According to the determined corresponding rewards of each executed strategy type, several executed strategy types whose rewards exceed the second preset threshold are screened as candidate strategy types.

[0015] Optionally, determining a specified range under the candidate policy type specifically includes:

[0016] Determine the strategy parameters of the merchant to be optimized under the selected strategy type;

[0017] For each parameter selection range preset under the strategy type to be selected, based on the historical transition probability of each merchant transferring from the strategy parameter to the parameter selection range under the strategy type to be selected, each parameter selection range whose historical transition probability exceeds the third preset threshold is determined as the designated range.

[0018] Optionally, the merchant characteristics include at least one of merchant category, location, and historical transaction information, and the first prediction model includes a first sub-model and a second sub-model;

[0019] Based on the first feature set, using a pre-trained first prediction model, predicting the transfer reward of the merchant to be optimized from the strategy parameter to the candidate parameter as the transfer reward corresponding to the candidate parameter specifically includes:

[0020] Based on the merchant features of the merchant to be optimized in the first feature set and the strategy parameters of the merchant to be optimized under the selected strategy type, using the pre-trained first sub-model, predict the future reward of the merchant to be optimized under the selected strategy type and the strategy parameters;

[0021] Based on the merchant characteristics of the merchant to be optimized in the first feature set, the strategy parameters of the merchant to be optimized under the selected strategy type, and the selected parameters, using the pre-trained second sub-model, predict the future reward of the merchant to be optimized under the selected strategy type when the strategy parameters are transferred to the selected parameters;

[0022] Based on the future rewards of the merchant to be optimized under the selected strategy type and the strategy parameters, and the future rewards of the merchant to be optimized transferred from the strategy parameters to the selected parameters under the selected strategy type, the transfer rewards of the merchant to be optimized from the strategy parameters to the selected parameters are determined.

[0023] Optionally, the optimization strategy is determined based on at least one of the transition probability and the transfer reward corresponding to each candidate parameter under each candidate strategy type, specifically including:

[0024] According to the transition probabilities corresponding to the various candidate parameters under the various candidate strategy types, determining the candidate parameter with the largest transition probability as the optimization parameter, determining the candidate strategy type corresponding to the optimization parameter as the optimization strategy type, and determining the optimization strategy based on the optimization parameters and the optimization strategy type; and / or

[0025] According to the transfer rewards corresponding to each candidate parameter under each candidate strategy type, determine the candidate parameter with the largest transfer reward as the optimization parameter, determine the candidate strategy type corresponding to the optimization parameter as the optimization strategy type, and determine the optimization strategy based on the optimization parameter and the optimization strategy type; and / or

[0026] For each candidate parameter, the comprehensive score of the candidate parameter is determined based on the transfer probability and transfer reward corresponding to the candidate parameter, and based on the comprehensive scores of each candidate parameter, the candidate parameter with the highest score is determined as the optimization parameter, and the candidate strategy type corresponding to the optimization parameter is determined as the optimization strategy type, and the optimization strategy is determined based on the optimization parameter and the optimization strategy type.

[0027] Optionally, training the first prediction model specifically includes:

[0028] Identify each merchant that has historically performed parameter transfers under the candidate strategy type;

[0029] Obtaining merchant characteristics of each merchant, strategy parameters of each merchant before parameter transfer under the candidate strategy type, and strategy parameters of each merchant after parameter transfer under the candidate strategy type as training samples;

[0030] Label each training sample based on the transfer reward generated by each merchant's parameter transfer under the selected strategy type;

[0031] For each training sample, input the training sample into a first prediction model to be trained, and determine a transfer reward output by the first prediction model;

[0032] The model parameters in the first prediction model are adjusted with the goal of minimizing the difference between the transfer reward output by the first prediction model and the labels of each training sample.

[0033] Optionally, training the second prediction model specifically includes:

[0034] Identify merchants that have historically been recommended for parameter transfer under the candidate strategy type;

[0035] Determine the training samples based on the merchant characteristics of each merchant, the parameters of each merchant before and after parameter transfer, and the historical transfer probability and historical transfer rewards of each merchant performing the same parameter transfer.

[0036] Label each training sample based on whether each merchant performs parameter transfer under the selected strategy type;

[0037] For each training sample, input the training sample into the second prediction model to be trained, and determine the transition probability output by the second prediction model;

[0038] The model parameters in the second prediction model are adjusted with the goal of minimizing the difference between the transition probability output by the second prediction model and the labels of each training sample.

[0039] This specification provides a strategy recommendation device, including:

[0040] The first determination module is configured to determine a number of candidate strategy types based on the historically executed strategy types and corresponding rewards of each merchant;

[0041] The second determination module is configured to determine, for each merchant to be optimized, the strategy parameters of the merchant to be optimized under each strategy type to be selected;

[0042] A first prediction module is configured to, for each candidate parameter within a specified range under each candidate strategy type, predict, based on a first feature set and using a pre-trained first prediction model, a transfer reward for the merchant to be optimized from the strategy parameter to the candidate parameter, as the transfer reward corresponding to the candidate parameter, wherein the first feature set includes merchant characteristics of the merchant to be optimized, the strategy parameter of the merchant to be optimized under the candidate strategy type, and the candidate parameter; and / or

[0043] A second prediction module is configured to, for each candidate parameter within a specified range under each candidate strategy type, predict, based on a second feature set and using a pre-trained second prediction model, a transition probability of the merchant to be optimized from the strategy parameter to the candidate parameter, as the transition probability corresponding to the candidate parameter, wherein the second feature set includes merchant characteristics of the merchant to be optimized, the strategy parameter of the merchant to be optimized under the candidate strategy type, the candidate parameter, and historical transition probabilities and historical transfer rewards of each merchant from the strategy parameter to the candidate parameter;

[0044] The recommendation module is configured to determine the optimization strategy based on at least one of the transfer probability and transfer reward corresponding to each candidate parameter under each candidate strategy type, and recommend the optimization strategy to the merchant to be optimized so that the merchant to be optimized can be optimized.

[0045] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned policy recommendation method is implemented.

[0046] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the above-mentioned policy recommendation method is implemented.

[0047] At least one of the above technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:

[0048] In this specification, a number of candidate strategy types can be determined based on the historically executed strategy types and corresponding rewards of each merchant. Afterwards, based on the merchant characteristics of each merchant to be optimized, the strategy parameters under each candidate strategy type, each candidate parameter, and the historical transition probability and historical transfer reward of each merchant for parameter transfer in the past, the transition probability and transfer reward of the merchant to be optimized from the strategy parameter to each candidate parameter are predicted. Finally, based on at least one of the transition probabilities and transfer rewards corresponding to each candidate parameter under each candidate strategy type, the optimization strategy recommended to each merchant to be optimized is determined. By determining the transition probability and transfer reward of each merchant to be optimized from the strategy parameter to each candidate parameter under each candidate strategy type, and based on at least one of the transition probabilities and transfer rewards, an accurate or easy-to-implement optimization strategy is recommended to the merchant, thereby promoting the perfect optimization of the merchant. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0050] Figure 1 A flowchart of a strategy recommendation method provided in an embodiment of this specification;

[0051] Figure 2 A schematic diagram of the structure of a strategy recommendation device provided in an embodiment of this specification;

[0052] Figure 3 Schematic diagram of an electronic device for implementing the strategy recommendation method provided in an embodiment of this specification. DETAILED DESCRIPTION

[0053] To make the purpose, technical solutions, and advantages of this specification more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0054] Currently, to improve order conversion rates on e-commerce and food delivery platforms, merchants often independently implement promotional strategies, or platform operators provide strategic recommendations based on the merchant's current status to boost order volume. For example, on food delivery platforms, offering reduced or waived delivery fees or heavily discounted dishes increases the likelihood of users placing orders, leading to faster order growth.

[0055] However, this kind of discount strategy set artificially based on experience is often not reasonable enough and cannot maximize the profits of merchants.

[0056] This specification provides a strategy recommendation method. The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0057] Figure 1 A flow chart of a strategy recommendation method provided in an embodiment of this specification may include the following steps:

[0058] S100: Determine several candidate strategy types based on the historically executed strategy types and corresponding rewards of each merchant.

[0059] In order to recommend effective execution strategies to merchants and ensure that they achieve significant order growth after implementing the recommended strategies, effective strategies can be screened for subsequent recommendations based on the historical order growth generated by merchants on the business platform after implementing various strategies. The strategy recommendation algorithm provided in this specification can be executed by a server on a business platform, which can be an e-commerce platform, a food delivery platform, or other platform with merchant operation services.

[0060] Specifically, the server may first obtain the historically executed policy types for each merchant and, based on these historically executed policy types, select a number of executed policy types whose percentage of executed merchants exceeds a first preset threshold. Then, based on the corresponding rewards determined for each executed policy type, select a number of executed policy types whose rewards exceed a second preset threshold as candidate policy types. Both the first and second preset thresholds can be set as needed.

[0061] Taking the food delivery platform as an example, the types of strategies that have been implemented by merchants include but are not limited to reducing or exempting delivery amounts, setting discounted dishes and setting discount levels, adding instant discounts for new customers and setting instant discount amounts, and increasing business hours. The corresponding rewards for each implemented strategy type include but are not limited to the increase in order volume, order conversion rate, total sales amount, and order revenue after the merchant implements the strategy type. Among them, if the proportion of merchants that reduce or exempt delivery amounts accounts for 20% of all merchants on the platform, it means that the proportion of merchants implementing the strategy type of reducing or exempting delivery amounts is 20%. If the order volume increases by 12% after the merchant adopts the strategy of reducing or exempting delivery amounts, it means that the reward corresponding to the strategy type of reducing or exempting delivery amounts is a 12% increase in order volume.

[0062] Furthermore, in order to ensure the effectiveness of the strategy, the strategy types implemented by each merchant in the recent period of history are usually obtained.

[0063] S102: For each merchant to be optimized, determining the strategy parameters of the merchant to be optimized under each strategy type to be selected.

[0064] In one or more embodiments of this specification, in order to recommend more refined optimization strategies to merchants to be optimized, parameter settings under each strategy type to be selected may also be recommended to each merchant to be optimized based on its own situation.

[0065] For example, if the selected policy type is to reduce or waive the delivery amount, the corresponding parameter setting is the amount setting for the reduced delivery amount. For example, if the selected policy type is to set up a new customer instant discount activity, the corresponding parameter setting is the amount setting for the new customer instant discount.

[0066] Therefore, in this specification, for each merchant to be optimized, the current strategy parameters of the merchant to be optimized under each strategy type to be selected can be determined, so that the parameters can be adjusted according to the current strategy parameters of the merchant to be optimized.

[0067] S104: For each candidate parameter within a specified range under each candidate strategy type, using a pre-trained first prediction model based on a first feature set, predict the transfer reward for the merchant to be optimized from the strategy parameter to the candidate parameter, as the transfer reward corresponding to the candidate parameter. And / or for each candidate parameter within a specified range under each candidate strategy type, using a pre-trained second prediction model based on a second feature set, predict the transition probability of the merchant to be optimized from the strategy parameter to the candidate parameter, as the transition probability corresponding to the candidate parameter.

[0068] In one or more embodiments of the present specification, in order to enable the merchant to achieve maximum profit after executing the optimization strategy, the profit that can be brought by adjusting the current strategy parameters to other parameters under each strategy type to be selected can be predicted to determine the parameter value with the maximum profit, which is the optimal parameter setting for the merchant to be optimized under the strategy type to be selected.

[0069] Specifically, for each policy type to be selected, the specified range set under the policy type to be selected can be determined first. For example, if the policy type to be selected is to reduce or exempt the delivery amount, the corresponding specified range is the range of the delivery amount, which can be set to (0, 8].

[0070] Afterwards, for each candidate parameter within the specified range, the first feature set is used as input to the pre-trained first prediction model to predict the transfer reward generated by the merchant to be optimized from the current strategy parameters to the candidate parameters, as the transfer reward corresponding to the candidate parameters. The first feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the candidate strategy type, and the candidate parameters. The merchant characteristics include at least one of the merchant category, the region where it is located, and historical transaction information. The transfer reward can be the merchant's order growth, conversion rate growth, etc. after the parameter change.

[0071] Assume that the strategy type to be selected is to increase the instant discount for new customers, and the corresponding specified range of the instant discount amount for new customers is (0, 5], then the parameters to be selected within the specified range are instant discount amounts of 1 to 5 yuan. If the current instant discount amount for new customers of the merchant to be optimized is set to 2 yuan, it means that the strategy parameter of the merchant to be optimized under the instant discount for new customers strategy is 2 yuan. Therefore, for the parameter 3 to be selected, the merchant characteristics of the merchant to be optimized, the current strategy parameter 2 yuan, and the parameter 3 to be selected are used as inputs and input into the first prediction model to obtain the order growth generated by the merchant to be optimized after the instant discount for new customers is increased from 2 yuan to 3 yuan.

[0072] Furthermore, when training the first preset model for the candidate strategy type, the first step is to identify all merchants on the business platform that have historically performed parameter transfers under the candidate strategy type, where parameter transfer refers to parameter adjustments made by the merchant. Subsequently, the merchant characteristics of each merchant, the strategy parameters of each merchant before the parameter transfer under the candidate strategy type, and the strategy parameters of each merchant after the parameter transfer under the candidate strategy type are obtained as training samples. Each training sample is then labeled based on the transfer reward generated by each merchant performing the parameter transfer under the candidate strategy type.

[0073] Then, for each training sample, the training sample is input into the first prediction model to be trained, the transfer reward output by the first prediction model is determined, and the model parameters in the first prediction model are adjusted with the goal of minimizing the difference between the transfer reward output by the first prediction model and the labels of each training sample.

[0074] In another embodiment of the present specification, the first prediction model can be divided into a first sub-model and a second sub-model, wherein the first sub-model is used to predict the rewards that the merchant to be optimized will generate in the future with the current strategic parameters, and the second sub-model is used to predict the rewards that the merchant to be optimized will generate in the future after the current strategic parameters are transferred to the parameters to be selected. The transfer from the current strategic parameters to the parameters to be selected means adjusting the current strategic parameters of the merchant to be optimized to the parameters to be selected.

[0075] Therefore, the merchant characteristics of the merchant to be optimized and the strategy parameters of the merchant to be optimized under the selected strategy type can be used as inputs into the first sub-model of the first prediction model to predict the future rewards of the merchant to be optimized under the selected strategy type and strategy parameters. Furthermore, the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the selected strategy type, and the selected parameters can be used as inputs into the second sub-model of the first preset model to predict the future rewards of the merchant to be optimized when the current strategy parameters are transferred to the selected parameters under the selected strategy type.

[0076] Finally, based on the future rewards of the merchant to be optimized under the selected strategy type and the current strategy parameters, and the future rewards of the merchant to be optimized when transferring from the current strategy parameters to the selected parameters under the selected strategy type, the difference between the two rewards is determined as the transfer reward for the merchant to be optimized when transferring from the strategy parameters to the selected parameters.

[0077] Furthermore, when training the first prediction model under the candidate strategy type, the first sub-model and the second sub-model therein may be trained separately.

[0078] Among them, when training the first sub-model, first determine the merchants in the business platform in history that have not undergone parameter transfer under the candidate strategy type within a continuous period of time. Afterwards, obtain the merchant characteristics of each merchant and the strategy parameters of each merchant under the candidate strategy type as training samples. And label each training sample according to the rewards (order volume, conversion rate, etc.) generated by each merchant in the specified historical period. Then, for each training sample, input the training sample into the first sub-model to be trained, determine the predicted reward output by the first sub-model, and adjust the model parameters in the first sub-model with the goal of minimizing the difference between the predicted reward output by the first sub-model and the labeling of each training sample.

[0079] When training the second sub-model, first determine the merchants in the business platform that have undergone parameter transfer under the candidate strategy type within a continuous period of time in history. Then, use the merchant characteristics of each merchant, the strategy parameters of each merchant under the candidate strategy type, and the parameters after the transfer as training samples. And label each training sample based on the rewards (order volume, conversion rate, etc.) generated by each merchant in a specified historical period. Then, for each training sample, input the training sample into the second sub-model to be trained, determine the predicted reward output by the second sub-model, and adjust the model parameters in the second sub-model with the goal of minimizing the difference between the predicted reward output by the second sub-model and the labeling of each training sample.

[0080] In one or more embodiments of this specification, in order to facilitate implementation by merchants, considering the difficulty of implementing each optimization strategy, the probability of executing each optimization strategy can be predicted first to recommend the most easily implemented optimization strategy to the merchant to be optimized.

[0081] Specifically, for each strategy type to be selected, the specified range set under the strategy type to be selected can be determined first. Afterwards, for each parameter to be selected within the specified range, the second feature set is used as input, and the pre-trained second prediction model is input to predict the transition probability of the merchant to be optimized from the current strategy parameter to the parameter to be selected, as the transition probability corresponding to the parameter to be selected. Among them, the second feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the strategy type to be selected, the parameters to be selected, and the historical transition probabilities and historical transfer rewards of each merchant from the strategy parameter to the parameter to be selected. The merchant characteristics include at least one of the merchant category, the region where it is located, and historical transaction information. The higher the transition probability corresponding to the parameter to be selected, the stronger the feasibility of the merchant to perform parameter transfer under the strategy type to be selected.

[0082] Furthermore, when determining the historical transition probability and historical transfer rewards for each merchant under the candidate strategy type, the percentage of merchants on the business platform that have historically transitioned from the strategy parameters to the candidate parameters under the candidate strategy type can be calculated as the historical transition probability. The incremental rewards generated by each merchant before and after the parameter transition under the candidate strategy type, such as the increase in orders or order conversion rate, can also be determined as the historical transfer reward.

[0083] Continuing with the example of adding a new customer discount, if a merchant's current new customer discount amount is 2 yuan and the new customer discount amount to be selected is 5 yuan, the percentage of merchants that have historically increased their new customer discount amount from 2 yuan to 5 yuan will be counted as the transfer probability for the 5 yuan new customer discount amount range. The incremental orders from the merchant after the new customer discount amount is increased will be used as the transfer reward for the 5 yuan new customer discount amount range.

[0084] Furthermore, when training the second prediction model under the candidate strategy type, the first step is to identify the merchants historically recommended for parameter transfer under the candidate strategy type. Next, training samples are determined based on each merchant's merchant characteristics, the parameters before and after parameter transfer, and the historical transfer probabilities and rewards for each merchant performing the same parameter transfer. Each training sample is labeled based on whether each merchant performed parameter transfer under the candidate strategy type. Then, for each training sample, the training sample is input into the second prediction model to be trained to determine the transfer probability output by the second prediction model. Finally, the model parameters in the second prediction model are adjusted with the goal of minimizing the difference between the transfer probability output by the second prediction model and the labels of each training sample.

[0085] When determining training samples, for each identified merchant, the merchant's parameters before and after parameter transfer are determined. The historical percentage of merchants that have undergone the same parameter transfer is determined as the historical transfer probability, and the reward increase generated by each merchant's parameter transfer is determined as the historical transfer reward. Training samples are determined based on the merchant's merchant characteristics, the merchant's parameters before and after parameter transfer, and the historical transfer probabilities and historical transfer rewards of each merchant that has historically undergone the same parameter transfer as the merchant.

[0086] S106: Determine an optimization strategy based on at least one of the transfer probabilities and transfer rewards corresponding to the parameters to be selected under each strategy type to be selected, and recommend the optimization strategy to the merchant to be optimized so that the merchant to be optimized performs optimization.

[0087] In one embodiment of this specification, to facilitate merchant implementation and taking into account the difficulty of implementing each strategy, the server may, based on the determined transition probabilities corresponding to each candidate parameter under each candidate strategy type, determine the candidate parameter with the highest transition probability as the optimization parameter, and determine the candidate strategy type corresponding to the optimization parameter as the optimization strategy type. Finally, based on the determined optimization strategy type and optimization parameters, an optimization strategy is determined and recommended to the merchant to be optimized, so that the merchant to be optimized can improve based on the optimization strategy.

[0088] For example, assuming that the transition probability from the merchant's current delivery reduction amount to the target delivery reduction amount of 3 yuan is 70%, the transition probability from the merchant's current delivery reduction amount to the delivery reduction amount of 4 yuan is 50%, the transition probability from the merchant's current new customer instant discount amount to the new customer instant discount of 4 yuan is 50%, and the transition probability from the merchant's current new customer instant discount amount to the new customer instant discount of 5 yuan is 40%, then the target delivery reduction amount of 3 yuan with the highest transition probability can be determined as the final optimization parameter, and the corresponding delivery reduction amount is the optimization strategy type.

[0089] In another embodiment of the present disclosure, to maximize merchant revenue, the server may determine, based on the transfer rewards corresponding to each candidate parameter under each candidate strategy, the candidate parameter with the highest transfer reward as the optimization parameter, and determine the candidate strategy type corresponding to the optimization parameter as the optimization strategy type. Finally, based on the determined optimization strategy type and optimization parameters, an optimization strategy is determined and recommended to the merchant to be optimized, so that the merchant can improve based on the optimization strategy.

[0090] For example, assuming that the increase in orders brought about by the merchant adjusting the current instant discount for new customers from 0 yuan to 5 yuan is 5,000 orders, and the increase in orders brought about by the merchant reducing the delivery amount from 0 yuan to 3 yuan is 10,000 orders, then the reduction in delivery amount of 3 yuan can be determined as the optimization parameter, and the corresponding reduction in delivery amount can be determined as the optimization strategy type.

[0091] In other embodiments of this specification, in order to combine merchant revenue and strategy implementation feasibility, the server may also weight each candidate parameter based on the transfer probability and transfer reward corresponding to the candidate parameter to determine a comprehensive score for the candidate parameter. Based on the comprehensive scores of the candidate parameters, the candidate parameter with the highest score is determined as the optimized parameter, and the candidate strategy type corresponding to the optimized parameter is determined as the optimized strategy type. Finally, the optimization strategy is determined based on the optimized parameter and the optimization strategy type.

[0092] based on Figure 1The strategy recommendation method shown can first determine several candidate strategy types based on the historically executed strategy types and corresponding rewards of each merchant. Then, based on the merchant characteristics of each merchant to be optimized, the strategy parameters under each candidate strategy type, each candidate parameter, and the historical transition probability and historical transfer reward of each merchant for parameter transfer, the transition probability and transfer reward of the merchant to be optimized from the strategy parameter to each candidate parameter are predicted. Finally, based on at least one of the transition probability and transfer reward corresponding to each candidate parameter under each candidate strategy type, the optimization strategy recommended to each merchant to be optimized is determined. By determining the transition probability and transfer reward of each merchant to be optimized from the strategy parameter to each candidate parameter under each candidate strategy type, and based on at least one of the transition probability and transfer reward, accurate or easy-to-implement optimization measures are recommended to the merchant to promote the perfect optimization of the merchant.

[0093] When determining the designated range for the candidate strategy type in step S104 of this specification, to increase the feasibility of the recommended strategy, facilitate implementation by merchants, and reduce computational complexity, several parameter selection ranges may be pre-set for each candidate strategy type. Subsequently, the strategy parameters for the merchant to be optimized under that candidate strategy type are determined. For each parameter selection range pre-set under that candidate strategy type, the historical transition probability for each merchant under that candidate strategy type from that strategy parameter to that parameter selection range is determined, based on the historical transition probability for each merchant under that candidate strategy type. Each parameter selection range whose historical transition probability exceeds a third preset threshold is determined as the designated range. The third preset threshold can be set as needed.

[0094] Among them, taking the strategy type to be selected as the reduction or exemption of delivery amount as an example, the parameter selection range preset under the strategy type to be selected refers to the selection range of the reduction or exemption of delivery amount. Assume that the parameter selection ranges of the reduction or exemption of delivery amount are (0, 3], (3, 5], (5, 8] respectively, the current reduction or exemption of delivery amount of the merchant to be optimized is 0 yuan, and the proportion of merchants whose reduction or exemption of delivery amount has been adjusted from 0 yuan to (0, 3] yuan in history is 20%, then the historical transition probability of adjusting the reduction or exemption of delivery amount from 0 yuan to (0, 3] yuan is determined to be 20%.

[0095] Of course, if the strategy type to be selected is to increase the business hours of merchants, the preset parameter selection range under this strategy type can be the extension range of business hours, such as extending the business hours (0, 2], (2, 4].

[0096] Furthermore, when determining the historical transition probability of each merchant under the candidate strategy type shifting from the strategy parameters to the parameter selection range, the parameter selection range to which the strategy parameters of the merchant to be optimized belong under the candidate strategy type can be first determined as the current parameter range. Subsequently, each merchant that historically fell within the current parameter range under the candidate strategy type is identified, and the historical transition probability of each merchant shifting from the current parameter range to each parameter selection range is determined.

[0097] For example, assuming that the current delivery exemption amount of the merchant to be optimized is 2 yuan, and the preset parameter selection ranges for the delivery exemption amount are (0, 3], (3, 5], and (5, 8], respectively, then the current parameter range of the delivery exemption amount of the merchant to be optimized is (0, 3]. Therefore, the merchants whose delivery exemption amounts were in (0, 3] in history can be determined, and the proportion of merchants whose delivery exemption amounts were transferred from (0, 3] to (3, 5] and (5, 8] respectively can be determined as the historical transfer probability of each parameter selection range.

[0098] Assuming that the third preset threshold is 50%, the historical transfer probability of each merchant's delivery reduction amount shifting from (0, 3] to (3, 5] is 60%, and the historical transfer probability of shifting from (0, 3] to (5, 8] is 30%, which means that the feasibility of merchants shifting the delivery reduction amount from (0, 3] to (3, 5] is high, so (3, 5] is the specified range.

[0099] In this specification, when the policy type to be selected is to set a discount activity, such as a 5 yuan discount for orders over 20 yuan, the preset parameter selection range for the discount activity can be 3 yuan off for orders over 15 yuan, 5 yuan off for orders over 20 yuan, 8 yuan off for orders over 30 yuan, 10 yuan off for orders over 40 yuan, etc. If the merchant's current discount parameter is 3 yuan off for orders over 15 yuan, the historical transition probability of each merchant changing from 3 yuan off for orders over 15 yuan to 5 yuan off for orders over 20 yuan, 8 yuan off for orders over 30 yuan, and 10 yuan off for orders over 40 yuan can be determined respectively, and based on each historical transition probability, the designated range for the discount activity can be determined.

[0100] Furthermore, usually only by setting more favorable strategies can merchants see an increase in revenue. Therefore, when determining the transition probability of transferring from the merchant's current strategy parameters to the parameter selection ranges to determine the specified range, only the probability of transferring from the merchant's current strategy parameters to the more favorable parameter selection ranges can be determined. For example, assuming that the merchant's current new customer discount amount is 5 yuan, only the transition probability to the parameter selection ranges where the new customer discount amount is greater than 5 yuan can be determined.

[0101] Furthermore, there may be multiple designated ranges under the candidate strategy type determined by the above steps. For each designated range, the transfer probability and transfer reward corresponding to each candidate parameter under each designated range are determined by the method shown in step S104.

[0102] In addition, in step S106 of this manual, in order to facilitate merchants to choose execution strategies according to their needs, the ranking of each candidate strategy and candidate parameters can be displayed to merchants according to at least one of the ranking of each transfer probability, the ranking of each transfer reward, and the ranking of each comprehensive score, so that merchants can make independent choices based on their own circumstances.

[0103] based on Figure 1 A strategy recommendation method is shown in FIG. , and this specification also provides a structural diagram of a strategy recommendation device. Figure 2 shown.

[0104] Figure 2 A schematic diagram of the structure of a strategy recommendation device provided in an embodiment of this specification includes:

[0105] The first determination module 200 is configured to determine a number of candidate strategy types based on the historically executed strategy types and corresponding rewards of each merchant;

[0106] The second determination module 202 is configured to determine, for each merchant to be optimized, the strategy parameters of the merchant to be optimized under each strategy type to be selected;

[0107] The prediction module 204 is configured to, for each parameter to be selected within a specified range under each type of strategy to be selected, predict, based on a first feature set and a pre-trained first prediction model, a transfer reward for the merchant to be optimized from the strategy parameter to the parameter to be selected, as the transfer reward corresponding to the parameter to be selected, wherein the first feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the type of strategy to be selected, and the parameter to be selected; and / or, for each parameter to be selected within a specified range under each type of strategy to be selected, predict, based on a second feature set and a pre-trained second prediction model, a transfer probability of the merchant to be optimized from the strategy parameter to the parameter to be selected, as the transfer probability corresponding to the parameter to be selected, wherein the second feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the type of strategy to be selected, the parameter to be selected, and historical transfer probabilities and historical transfer rewards of each merchant in history from the strategy parameter to the parameter to be selected;

[0108] The recommendation module 206 is configured to determine an optimization strategy based on at least one of the transfer probability and transfer reward corresponding to each candidate parameter under each candidate strategy type, and recommend the optimization strategy to the merchant to be optimized so that the merchant to be optimized can be optimized.

[0109] Optionally, the first determination module 200 is specifically used to screen, based on the historical executed policy types of each merchant, several executed policy types whose proportion of executing merchants exceeds a first preset threshold, and based on the corresponding rewards of each determined executed policy type, screen several executed policy types whose rewards exceed a second preset threshold as candidate policy types.

[0110] Optionally, the prediction module 204 is specifically used to determine the strategy parameters of the merchant to be optimized under the strategy type to be selected, and for each parameter selection range preset under the strategy type to be selected, based on the historical transition probability of each merchant under the strategy type to transfer from the strategy parameters to the parameter selection range, determine each parameter selection range whose historical transition probability exceeds a third preset threshold as the designated range.

[0111] Optionally, the merchant characteristics include at least one of merchant categories, geographical locations, and historical transaction information. The first prediction model includes a first sub-model and a second sub-model. The prediction module 204 is specifically used to predict the future rewards of the merchant to be optimized under the selected strategy type and the strategy parameters based on the merchant characteristics of the merchant to be optimized in the first feature set and the strategy parameters of the merchant to be optimized under the selected strategy type through the pre-trained first sub-model; predict the future rewards of the merchant to be optimized when transferred from the strategy parameters to the selected parameters under the selected strategy type based on the merchant characteristics of the merchant to be optimized in the first feature set and the strategy parameters of the merchant to be optimized under the selected strategy type and the selected parameters through the pre-trained second sub-model; determine the transfer rewards of the merchant to be optimized when transferred from the strategy parameters to the selected parameters based on the future rewards of the merchant to be optimized under the selected strategy type and the strategy parameters, and the future rewards of the merchant to be optimized when transferred from the strategy parameters to the selected parameters under the selected strategy type.

[0112] Optionally, the recommendation module 206 is specifically used to, based on the transition probability corresponding to each candidate parameter under each candidate strategy type, determine the candidate parameter with the largest transfer probability as the optimization parameter, and determine the candidate strategy type corresponding to the optimization parameter as the optimization strategy type, and determine the optimization strategy based on the optimization parameter and the optimization strategy type; and / or based on the transfer reward corresponding to each candidate parameter under each candidate strategy type, determine the candidate parameter with the largest transfer reward as the optimization parameter, and determine the candidate strategy type corresponding to the optimization parameter as the optimization strategy type, and determine the optimization strategy based on the optimization parameter and the optimization strategy type; and / or for each candidate parameter, determine the comprehensive score of the candidate parameter based on the transition probability and transfer reward corresponding to the candidate parameter, and based on the comprehensive score of each candidate parameter, determine the candidate parameter with the highest score as the optimization parameter, and determine the candidate strategy type corresponding to the optimization parameter as the optimization strategy type, and determine the optimization strategy based on the optimization parameter and the optimization strategy type.

[0113] Optionally, the strategy recommendation device further includes a model training module 208, and the model training module 208 is specifically used to determine each merchant that has historically performed parameter transfer under the candidate strategy type, obtain the merchant characteristics of each merchant, the strategy parameters of each merchant before the parameter transfer under the candidate strategy type, and the strategy parameters of each merchant after the parameter transfer under the candidate strategy type as training samples, and label each training sample according to the transfer reward generated by each merchant performing parameter transfer under the candidate strategy type. For each training sample, the training sample is input into the first prediction model to be trained, and the transfer reward output by the first prediction model is determined. With the goal of minimizing the difference between the transfer reward output by the first prediction model and the labeling of each training sample, the model parameters in the first prediction model are adjusted.

[0114] Optionally, the model training module 208 is specifically used to determine the merchants that have been historically recommended to perform parameter transfer under the strategy type to be selected, determine the training samples based on the merchant characteristics of each merchant, the parameters before and after the parameter transfer of each merchant, and the historical transfer probability and historical transfer rewards of each merchant performing the same parameter transfer in history, label each training sample according to whether each merchant performs parameter transfer under the strategy type to be selected, input each training sample into the second prediction model to be trained, determine the transfer probability output by the second prediction model, and adjust the model parameters in the second prediction model with the goal of minimizing the difference between the transfer probability output by the second prediction model and the labels of each training sample.

[0115] The embodiment of this specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1Provides a strategy recommendation method.

[0116] according to Figure 1 A strategy recommendation method is shown in the embodiment of this specification. Figure 3 The schematic structure diagram of the electronic device shown in FIG. Figure 3 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The strategy recommendation method shown.

[0117] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0118] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD through their own programming, without having to hire a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually generating integrated circuit chips, this programming is mostly performed using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages ​​and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.

[0119] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.

[0120] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0121] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0122] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0123] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0124] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0126] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0127] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0128] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0129] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0130] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Thus, this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0132] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0133] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.

Claims

1. A strategy recommendation method, characterized in that: include: Determine several candidate strategy types based on the historically executed strategy types and corresponding rewards of each merchant; For each merchant to be optimized, determine the strategy parameters of the merchant to be optimized under each strategy type to be selected; For each candidate parameter within a specified range under each candidate strategy type, predict the transfer reward of the merchant to be optimized from the strategy parameter to the candidate parameter using a pre-trained first prediction model based on a first feature set, as the transfer reward corresponding to the candidate parameter, wherein the first feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameter of the merchant to be optimized under the candidate strategy type, and the candidate parameter; and / or For each candidate parameter within a specified range under each candidate strategy type, a transition probability of the merchant to be optimized from the strategy parameter to the candidate parameter is predicted using a pre-trained second prediction model based on a second feature set, as the transition probability corresponding to the candidate parameter. The second feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the candidate strategy type, the candidate parameters, and historical transition probabilities and historical transfer rewards for each merchant from the strategy parameter to the candidate parameter. An optimization strategy is determined based on at least one of the transfer probabilities and transfer rewards corresponding to each candidate parameter under each candidate strategy type, and the optimization strategy is recommended to the merchant to be optimized so that the merchant to be optimized can be optimized.

2. The method according to claim 1, wherein Based on the historically executed strategies and corresponding rewards for each merchant, several candidate strategy types are determined, including: Based on the historical executed strategy types of each merchant, select several executed strategy types whose proportion of executed merchants exceeds a first preset threshold; According to the determined corresponding rewards of each executed strategy type, several executed strategy types whose rewards exceed the second preset threshold are screened as candidate strategy types.

3. The method according to claim 1, wherein Determine the specified range under the candidate policy type, specifically including: Determine the strategy parameters of the merchant to be optimized under the selected strategy type; For each parameter selection range preset under the strategy type to be selected, based on the historical transition probability of each merchant transferring from the strategy parameter to the parameter selection range under the strategy type to be selected, each parameter selection range whose historical transition probability exceeds the third preset threshold is determined as the designated range.

4. The method according to claim 1, wherein The merchant characteristics include at least one of merchant category, location, and historical transaction information, and the first prediction model includes a first sub-model and a second sub-model; Based on the first feature set, using a pre-trained first prediction model, predicting the transfer reward of the merchant to be optimized from the strategy parameter to the candidate parameter as the transfer reward corresponding to the candidate parameter specifically includes: Based on the merchant features of the merchant to be optimized in the first feature set and the strategy parameters of the merchant to be optimized under the selected strategy type, using the pre-trained first sub-model, predict the future reward of the merchant to be optimized under the selected strategy type and the strategy parameters; Based on the merchant characteristics of the merchant to be optimized in the first feature set, the strategy parameters of the merchant to be optimized under the selected strategy type, and the selected parameters, using the pre-trained second sub-model, predict the future reward of the merchant to be optimized under the selected strategy type when the strategy parameters are transferred to the selected parameters; Based on the future rewards of the merchant to be optimized under the selected strategy type and the strategy parameters, and the future rewards of the merchant to be optimized transferred from the strategy parameters to the selected parameters under the selected strategy type, the transfer rewards of the merchant to be optimized from the strategy parameters to the selected parameters are determined.

5. The method according to claim 1, wherein Determine the optimization strategy based on at least one of the transition probabilities and transfer rewards corresponding to each candidate parameter under each candidate strategy type, specifically including: According to the transition probabilities corresponding to the various candidate parameters under the various candidate strategy types, determining the candidate parameter with the largest transition probability as the optimization parameter, determining the candidate strategy type corresponding to the optimization parameter as the optimization strategy type, and determining the optimization strategy based on the optimization parameters and the optimization strategy type; and / or According to the transfer rewards corresponding to each candidate parameter under each candidate strategy type, determine the candidate parameter with the largest transfer reward as the optimization parameter, determine the candidate strategy type corresponding to the optimization parameter as the optimization strategy type, and determine the optimization strategy based on the optimization parameter and the optimization strategy type; and / or For each candidate parameter, the comprehensive score of the candidate parameter is determined based on the transfer probability and transfer reward corresponding to the candidate parameter, and based on the comprehensive scores of each candidate parameter, the candidate parameter with the highest score is determined as the optimization parameter, and the candidate strategy type corresponding to the optimization parameter is determined as the optimization strategy type, and the optimization strategy is determined based on the optimization parameter and the optimization strategy type.

6. The method according to claim 1, wherein Training the first prediction model includes: Identify each merchant that has historically performed parameter transfers under the candidate strategy type; Obtaining merchant characteristics of each merchant, strategy parameters of each merchant before parameter transfer under the candidate strategy type, and strategy parameters of each merchant after parameter transfer under the candidate strategy type as training samples; Label each training sample based on the transfer reward generated by each merchant's parameter transfer under the selected strategy type; For each training sample, input the training sample into a first prediction model to be trained, and determine a transfer reward output by the first prediction model; The model parameters in the first prediction model are adjusted with the goal of minimizing the difference between the transfer reward output by the first prediction model and the labels of each training sample.

7. The method according to claim 1, wherein Training the second prediction model includes: Identify merchants that have historically been recommended for parameter transfer under the candidate strategy type; Determine the training samples based on the merchant characteristics of each merchant, the parameters of each merchant before and after parameter transfer, and the historical transfer probability and historical transfer rewards of each merchant performing the same parameter transfer. Label each training sample based on whether each merchant performs parameter transfer under the selected strategy type; For each training sample, input the training sample into the second prediction model to be trained, and determine the transition probability output by the second prediction model; The model parameters in the second prediction model are adjusted with the goal of minimizing the difference between the transition probability output by the second prediction model and the labels of each training sample.

8. A strategy recommendation device, characterized in that: include: The first determination module is configured to determine a number of candidate strategy types based on the historically executed strategy types and corresponding rewards of each merchant; The second determination module is configured to determine, for each merchant to be optimized, the strategy parameters of the merchant to be optimized under each strategy type to be selected; The prediction module is configured to, for each parameter to be selected within a specified range under each strategy type to be selected, predict, based on a first feature set and a pre-trained first prediction model, a transfer reward for the merchant to be optimized from the strategy parameter to the parameter to be selected, as the transfer reward corresponding to the parameter to be selected, wherein the first feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the strategy type to be selected, and the parameter to be selected; and / or, for each parameter to be selected within a specified range under each strategy type to be selected, predict, based on a second feature set and a pre-trained second prediction model, a transfer probability of the merchant to be optimized from the strategy parameter to the parameter to be selected, as the transfer probability corresponding to the parameter to be selected, wherein the second feature set includes the merchant characteristics of the merchant to be optimized, the strategy parameters of the merchant to be optimized under the strategy type to be selected, the parameter to be selected, and historical transfer probabilities and historical transfer rewards of each merchant in history from the strategy parameter to the parameter to be selected; The recommendation module is configured to determine the optimization strategy based on at least one of the transfer probability and transfer reward corresponding to each candidate parameter under each candidate strategy type, and recommend the optimization strategy to the merchant to be optimized so that the merchant to be optimized can be optimized.

9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method, device and equipment for automatically generating data analysis result and storage medium

    CN111340455A

  • Commodity recommendation model training method, commodity recommendation method and device

    CN112598467A