Training method of check strategy recommendation model, check strategy recommendation method and device
By constructing a multi-intervention group experimental sample system and training a check-in strategy recommendation model, the problems of high time cost and low conversion rate of check-in strategy recommendation in e-commerce platforms were solved. This resulted in a payment tool check-in strategy recommendation that better meets user needs, thereby improving the conversion rate and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-27
AI Technical Summary
In existing e-commerce platforms, when determining payment tool selection strategies based on manual methods, the limited features involved result in high time costs and low conversion rates for recommendation strategies, and the recommended results do not meet user needs.
By collecting users' historical behavior data, a multi-intervention group experimental sample system was constructed, a selection strategy recommendation model was trained, and a multi-branch neural network and multivariate joint loss function were used to optimize the model, identify the difference in order completion between users selecting and not selecting, and recommend the optimal selection strategy.
It reduces the time cost of the selection strategy recommendation, improves the accuracy of recommendation results and conversion rate, and enhances the user experience.
Smart Images

Figure CN121745192A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular to a training method of a check strategy recommendation model, a check strategy recommendation method and device. BACKGROUND
[0002] In the transaction scenario of an e-commerce platform, the payment tool check strategy is currently mainly determined based on an artificial strategy, that is, a user group is screened for payment tool checking according to a small number of features, but this method often only refers to a limited number of features, and the behavior habits of a user are often complex and diverse, and it is difficult to accurately predict the behavior of the user through a few features, resulting in a high time cost of check strategy recommendation, and the recommended result does not meet the user demand, and the order completion rate is low. SUMMARY
[0003] Therefore, the embodiments of the present application provide a training method of a check strategy recommendation model, a check strategy recommendation method and device, by collecting user historical behavior data, constructing a multi-intervention group experimental sample system to identify the order completion difference of the user check relative to the non-check, and taking this as a training sample set to train a check strategy recommendation model, which can reduce the time cost of check strategy recommendation, and the recommended result is more in line with the user demand, and improves the order completion conversion rate of the user.
[0004] To achieve the above-mentioned purpose, according to an aspect of an embodiment of the present application, a training method of a check strategy recommendation model is provided, comprising: collecting user historical behavior data in a first preset time period; generating a control group sample and a plurality of intervention group samples based on the user historical behavior data; wherein the control group sample is a user sample that does not check the payment tool, and the intervention group sample is a user sample that checks the corresponding payment tool; constructing a training sample set according to the control group sample and the plurality of intervention group samples, and training a check strategy recommendation model based on the training sample set.
[0005] Optionally, generating a control group sample and a plurality of intervention group samples based on user historical behavior data comprises: determining the payment check result of the user according to the user historical behavior data; grouping the user based on the payment check result to obtain a grouping result; generating a control group sample and a plurality of intervention group samples according to the grouping result.
[0006] Optionally, constructing a training sample set according to the control group sample and the plurality of intervention group samples comprises: determining the payment result of the user of the control group sample and the intervention group sample; in response to the payment result of the user being payment success, marking the sample as a positive sample; In response to the user payment result being a payment failure, the sample is marked as a negative sample; Based on the positive sample and the negative sample, a training sample set is constructed.
[0007] Optionally, the check strategy recommendation model is trained based on the training sample set, and the training includes: A multi-branch neural network structure is constructed, and the multi-branch neural network model is used to evaluate the payment conversion probability of the user under the check strategy of the plurality of payment tools. A multi-element joint loss function is set based on the neural network structure. The neural network structure is trained to obtain the check strategy recommendation model according to the training sample set and the multi-element joint loss function.
[0008] Optionally, after the check strategy recommendation model is trained based on the training sample set, the method further includes: User historical behavior data in a second preset time period is collected, and a test sample set is constructed. The test sample set is input into the check strategy recommendation model to obtain a test recommendation result. The check strategy recommendation model is optimized based on the test recommendation result.
[0009] According to a second aspect of an embodiment of the present application, a check strategy recommendation method is provided, including: Purchase behavior data of a target user is obtained. The purchase behavior data is input into a check strategy recommendation model, and the check strategy recommendation model outputs a check strategy recommendation result of the target user; wherein the check strategy recommendation model is obtained by any method of the first aspect of the present application. An intervention check behavior is performed on the target user based on the check strategy recommendation result.
[0010] According to a third aspect of an embodiment of the present application, a training device of a check strategy recommendation model is provided, including: A collection module is configured to collect user historical behavior data in a preset time period. A generation module is configured to generate a control group sample and a plurality of intervention group samples based on the user historical behavior data; wherein the control group sample is a user sample that does not check a payment tool, and the intervention group sample is a user sample that checks a corresponding payment tool. A training module is configured to construct a training sample set according to the control group sample and the plurality of intervention group samples, and to train a check strategy recommendation model based on the training sample set.
[0011] According to a fourth aspect of an embodiment of the present application, a check strategy recommendation device is provided, including: A collection module is configured to collect user historical behavior data in a preset time period. The prediction module is configured to input the purchase behavior data into a check strategy recommendation model, and the check strategy recommendation model outputs a check strategy recommendation result of the target user; wherein the check strategy recommendation model is obtained by any method of the first aspect of the embodiments of the present application. The recommendation module is configured to perform an intervention check behavior on the target user based on the check strategy recommendation result.
[0012] According to a fifth aspect of the embodiments of the present application, an electronic device is provided, comprising: one or more processors; a storage device configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method of any of the above embodiments.
[0013] According to a sixth aspect of the embodiments of the present application, a computer readable medium is provided, which stores a computer program, and the program is executed by a processor to implement the method of any of the above embodiments.
[0014] According to a seventh aspect of the embodiments of the present application, a computer program product is provided, which comprises a computer program, and the computer program is executed by a processor to implement the method of any of the above embodiments.
[0015] An embodiment of the above application has the following advantages or beneficial effects: by collecting user historical behavior data in a first preset time period; based on the user historical behavior data, a control group sample and a plurality of intervention group samples are generated; wherein the control group sample is a user sample who has not checked the payment tool, and the intervention group sample is a user sample who has checked the corresponding payment tool; a training sample set is constructed according to the control group sample and the plurality of intervention group samples, and a check strategy recommendation model is trained based on the training sample set. This embodiment collects user historical behavior data, constructs a multi-intervention group experimental sample system to identify the single conversion difference of the user check relative to the non-check, and uses it as a training sample set to train the check strategy recommendation model, which can reduce the time cost of the check strategy recommendation, and the recommendation result is more in line with the user demand, and the single conversion rate of the user is improved. By obtaining the purchase behavior data of the target user; the purchase behavior data is input into the check strategy recommendation model, and the check strategy recommendation model outputs the check strategy recommendation result of the target user; wherein the check strategy recommendation model is obtained by any method of the first aspect of the embodiments of the present application; based on the check strategy recommendation result, the intervention check behavior on the target user is performed. This embodiment can make the check strategy recommendation result more in line with the user demand, and improve the user experience.
[0016] The further effects of the above-mentioned non-conventional optional mode will be described in the following combined with the specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings are used to better understand the present application and do not constitute undue limitations on the present application. Among them: Figure 1 is a schematic diagram of the main process of the training method of the check strategy recommendation model according to an embodiment of the present application; Figure 2 is a schematic diagram of the check strategy recommendation method according to an embodiment of the present application; Figure 3 is a schematic diagram of the training method of the check strategy recommendation model according to a preferred embodiment of the present application; Figure 4 is a schematic diagram of the main modules of the training device of the check strategy recommendation model according to an embodiment of the present application; Figure 5 is a schematic diagram of the main modules of the check strategy recommendation device according to an embodiment of the present application; Figure 6 is an exemplary system architecture diagram to which embodiments of the present application can be applied; Figure 7 is a structural schematic diagram of a computer system of a terminal device or a server suitable for implementing embodiments of the present application. DETAILED DESCRIPTION
[0018] Exemplary embodiments of the present application are described below with reference to the accompanying drawings, which include various details of the embodiments of the present application to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Also, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0019] It should be noted that the acquisition, storage and application of personal information and the like involved in the embodiments of the present application comply with relevant laws and regulations and do not violate public order and good customs.
[0020] In the transaction scenario of an e-commerce platform, the payment tool check strategy is currently mainly determined based on artificial strategy, that is, a user group is selected for payment tool checking according to a small number of features, but this method often only refers to a limited number of features, while the behavior habits of users are often complex and diverse, and it is difficult to accurately predict the behavior of users through a few features, resulting in high time cost of check strategy recommendation, and the recommended result does not meet the user demand, and the single rate is low.
[0021] Therefore, according to an aspect of an embodiment of the present application, a training method of a check strategy recommendation model is provided.
[0022] Figure 1is a schematic diagram of the main process of the training method of the check strategy recommendation model according to an embodiment of the present application. As shown in Figure 1 The training method of the check strategy recommendation model according to an embodiment of the present application mainly includes the following steps S101 to S103.
[0023] Step S101, collect user historical behavior data in a first preset time period.
[0024] The first preset time period is a historical data window required for model training, for example, the last thirty days, the last sixty days, or the last ninety days, etc. The specific time length can be flexibly set according to factors such as user activity, check behavior distribution, promotion rhythm, etc. in the business scenario. The user historical behavior data is the behavior log of the user on the platform in the time period, mainly including whether the user participates in shopping, the frequency and duration of browsing goods, the behavior of adding goods to the shopping cart, the record of obtaining and using coupons, whether the payment transaction is successfully completed, whether there is a refund or return behavior, etc. At the same time, it also includes whether the user has intervened in the check behavior, the type of checked goods, whether the check has produced conversion, and other key behavior information.
[0025] Collecting the above-mentioned user historical behavior data can be achieved by real-time collection through an online burying point system. The system reports the behavior data to the log system when the user triggers a behavior event (such as clicking, adding to the shopping cart, ordering, etc.) each time, and the log data is cleaned, aggregated and archived by the data center in the future, forming structured user behavior data. In addition, the existing behavior data table in the platform data warehouse can also be called to regularly extract the data that meets the conditions within the time window through a data processing script.
[0026] Step S102, generate a control group sample and a plurality of intervention group samples based on the user historical behavior data; wherein the control group sample is a user sample that has not checked the payment tool, and the intervention group sample is a user sample that has checked the corresponding payment tool.
[0027] The control group sample is a user sample that has not checked the payment tool option in the first preset time period, that is, these users have not actively selected the payment tool option provided by the platform in the specific time window when browsing, clicking or completing payment and other behaviors. The intervention group sample is a user sample that has checked and used a specific payment tool in the same time period. The behavior of such users can reflect their actual acceptance and use preference for the payment tool.
[0028] Specifically, the operation records about the payment process link in the user behavior data can be marked and screened, specifically including the action field of whether the user clicks and selects the specified payment tool on the order page and payment page. Whether the field is empty or matches a payment tool identifier can distinguish between unchecked and checked users, thereby forming two sample sets. Alternatively, by analyzing the event triggering situation related to the payment tool in the complete path of browsing, adding to cart, ordering, and final payment, if the user never triggers the operation event related to the payment tool in the entire link, the user is classified into the control group sample; if the user has triggered and completed the check event of a specific payment tool, the user is classified into the corresponding intervention group sample, and the intervention group can be subdivided into multiple groups according to different types of payment tools. The above two methods can combine the behavior field and event sequence to accurately construct the sample base required for the control test from different dimensions.
[0029] In step S103, a training sample set is constructed according to the control group sample and the plurality of intervention group samples, and a check strategy recommendation model is trained based on the training sample set.
[0030] In this embodiment, the training sample set for model training can be constructed by aligning the features and normalizing the labels of the control group sample and the plurality of intervention group samples. Specifically, according to the actual click rate or conversion behavior of the user under different strategies, the strategy response results of the control group sample and the intervention group sample are respectively counted, the common behavior features of the samples in the same check strategy dimension are extracted, and a unified strategy label is set according to the features, so that the control group sample and each intervention group sample remain consistent in the feature space. Alternatively, a paired sample construction mechanism is introduced. For each control group sample, a sample with similar user attributes and behavior trajectory in the intervention group is found, forming a sample pair before and after the strategy, and the intervention effect is used as a supervision signal to construct the training sample set, and the effect modeling before and after the strategy intervention is realized.
[0031] When the check strategy recommendation model is trained based on the training sample set, a supervised learning algorithm (such as a gradient boosting tree model, a deep neural network model, etc.) is used, the user features, behavior features, and strategy features extracted in the training sample set are used as inputs, the model is trained to fit the strategy response label or conversion rate label, so as to predict the check tendency or recommendation effect of the user under different strategies. In addition, a structural model can also be constructed using a causal inference framework, for example, a multi-task learning structure is used to model the check behavior of the control group and the intervention group at the same time, and a strategy processing variable is introduced to learn the causal effect of the strategy on the user behavior, thereby realizing the personalized strategy recommendation capability.
[0032] The embodiment collects user historical behavior data in a first preset time period, generates a control group sample and a plurality of intervention group samples based on the user historical behavior data, wherein the control group sample is a user sample that does not check a payment tool, and the intervention group sample is a user sample that checks a corresponding payment tool, constructs a training sample set according to the control group sample and the plurality of intervention group samples, and trains a check strategy recommendation model based on the training sample set. The embodiment collects user historical behavior data, constructs a multi-intervention group experimental sample system to identify the single conversion difference of user check relative to no check, and uses the same as a training sample set to train a check strategy recommendation model, which can reduce the time cost of check strategy recommendation, and the recommended result is more in line with user demand, thereby improving the single conversion rate of users.
[0033] Optionally, generating the control group sample and the plurality of intervention group samples based on the user historical behavior data comprises: determining a payment check result of the user according to the user historical behavior data; grouping the user based on the payment check result to obtain a grouping result; and generating the control group sample and the plurality of intervention group samples according to the grouping result.
[0034] In the embodiment, first, the payment check result in the behavior data of the user in the first preset time period is extracted, the payment check result refers to whether the user performs a check or confirmation operation when the payment option is displayed, and the result can reflect the acceptance tendency or use preference of the user for a certain type of payment method. Based on the payment check result, the user is grouped, such as based on the user check frequency, check time distribution, whether the payment is completed, and the like, the user is divided into a plurality of subgroups with similar payment behavior characteristics by using grouping logic or algorithm. From each subgroup, the control group sample and the intervention group sample are respectively divided according to a certain random sampling or stratified sampling manner. In addition to the user grouping manner based on the payment check result, the user clustering model can also be established by combining the user portrait features and real-time context data, such as combining device types, regions, historical payment methods, and the like, so as to divide the control group sample and the plurality of intervention group samples based on clustering.
[0035] The control group sample and the intervention group sample are generated in the above manner, which can realize fine comparison and analysis of the influence effect of different payment strategies in different user groups, is conducive to improving the controllability and reliability of the payment conversion rate optimization experiment, and finally realizes dynamic optimization and accurate pushing of the payment conversion strategy.
[0036] Optionally, constructing the training sample set according to the control group sample and the plurality of intervention group samples comprises: determining a payment result of the user of the control group sample and the intervention group sample; in response to the payment result of the user being payment success, marking the sample as a positive sample; in response to the payment result of the user being payment failure, marking the sample as a negative sample; and constructing the training sample set based on the positive sample and the negative sample.
[0037] In this embodiment, for each user sample in each group, the actual payment result in the payment scenario is extracted as the labeling basis. Specifically, if the user ultimately completes the payment, the user sample is marked as a positive sample, indicating that the selection strategy has a positive effect on this type of user; if the user does not complete the payment operation, the sample is marked as a negative sample, indicating that the selection strategy failed to convert into successful payment for this type of user. Based on the collected positive and negative samples, a training sample set is constructed. Each sample in the training sample set contains the user's behavioral characteristics before payment, contextual information (such as device type, payment path, time information, etc.), and its corresponding payment result label. This training sample set can be used to model the payment conversion effect of users under different selection strategy interventions, thereby realizing the learning and inference of the optimal strategy. In addition, intermediate behavioral signals such as user dwell time, click behavior, cancellation rate, and payment delay can be used as auxiliary supervision signals to improve the robustness of model training in scenarios with no payment results or sparse payment data; pseudo-labels can also be generated by scoring based on expert rules or system strategies.
[0038] The training sample set constructed through this implementation method can cover user behavior responses under different selection strategy configurations, providing rich and effective data support for the subsequent training of the selection strategy recommendation model. The trained model can identify the optimal strategy configuration under different user characteristics and environmental contexts, thereby improving the personalized adaptability of the selection strategy, optimizing the payment conversion rate, and improving the overall business performance.
[0039] Optionally, a selection strategy recommendation model is trained based on a training sample set, including: constructing a multi-branch neural network structure; wherein the multi-branch neural network model is used to evaluate the payment conversion probability of users under multiple payment tool selection strategies; setting a multivariate joint loss function based on the neural network structure; and training the neural network structure to obtain the selection strategy recommendation model according to the training sample set and the multivariate joint loss function.
[0040] In this embodiment, the multi-branch neural network structure takes user features, context features, and order features as inputs. Each branch corresponds to a selection strategy with different payment tool combinations, used to simulate user payment behavior under that strategy, thereby evaluating the corresponding payment conversion probability. The design of this neural network structure fully considers the impact of different payment tool combinations on user payment decisions, enabling the model to model and evaluate multiple selection strategies in parallel. To optimize the learning effect of the neural network model, a multivariate joint loss function is further defined. The multivariate joint loss function includes, but is not limited to, the main task loss function (such as the cross-entropy loss for click conversion prediction) and the auxiliary task loss function (such as the payment success rate or payment path preference prediction loss). By setting the multivariate joint loss function, the model can simultaneously focus on multiple indicators during training, thereby enhancing its ability to model user payment conversion behavior and improving the model's generalization ability and stability. After the above neural network structure and loss function are set, the neural network structure is trained using a training sample set to finally obtain the selection strategy recommendation model. During training, standard backpropagation and stochastic gradient descent algorithms can be used, combined with batch normalization and other methods for regularization to avoid overfitting and accelerate convergence. In addition to the methods mentioned above, reinforcement learning strategies can also be used to build a selection strategy recommendation model. For example, user payment behavior can be used as an environmental feedback signal to construct a state-action-reward function relationship, and the recommendation model can be trained using methods such as policy gradient or deep Q-network, so that the model can optimize the selection strategy from the perspective of long-term benefits.
[0041] The selection strategy recommendation model trained in the above manner can predictively evaluate the selection combinations of different payment tools in practical applications, recommend the optimal selection strategy combination, and make users more inclined to complete the payment operation.
[0042] Optionally, after training the selection strategy recommendation model based on the training sample set, the method further includes: collecting historical user behavior data within a second preset time period and constructing a test sample set; inputting the test sample set into the selection strategy recommendation model to obtain test recommendation results; and optimizing the selection strategy recommendation model based on the test recommendation results.
[0043] To further improve the accuracy and adaptability of the selection strategy recommendation model in practical use, it can be tested and optimized. Specifically, historical user behavior data can be collected within a second preset time period. This second preset time period can be dynamically configured based on factors such as model update frequency and user activity levels, for example, set to the past week or the past month. Based on the collected historical user behavior data, a corresponding test sample set is constructed. Each sample in the test sample set can include selection context information, candidate option features, user features, and the final user selection result, used to measure the degree of matching between the model's recommended strategy and the user's actual behavior. The test sample set is input into the aforementioned trained selection strategy recommendation model, and the model generates corresponding test recommendation results based on the context and feature information in the test samples. The test recommendation results can be compared with actual user selection behavior to evaluate the model's prediction accuracy, coverage, and ranking effect in the current time period. Based on the evaluation results, the selection strategy recommendation model can be optimized using methods such as reinforcement learning optimization, online gradient updates, and dynamic sample weight adjustments to better adapt the model to current user preference changes and behavioral feature updates. In addition, model compression techniques such as knowledge distillation can be introduced to reduce model complexity and improve inference efficiency while maintaining model performance.
[0044] Through the above optimization process, the adaptive ability of the selection strategy recommendation model can be improved when facing different user groups and behavioral differences at different time periods, thereby enhancing the accuracy of the model recommendation strategy and user satisfaction, and improving the overall system's intelligence level and order completion conversion efficiency.
[0045] According to a second aspect of the present invention, a method for recommending selection strategies is provided.
[0046] Figure 2 This is a schematic diagram of the main flow of the selection strategy recommendation method according to an embodiment of the present invention; as shown below. Figure 2 As shown, the selection strategy recommendation method according to an embodiment of the present invention mainly includes the following steps S201 to S203.
[0047] Step S201: Obtain the purchase behavior data of the target user.
[0048] Step S202: Input the purchase behavior data into the selection strategy recommendation model, and the selection strategy recommendation model outputs the selection strategy recommendation result for the target user; wherein, the selection strategy recommendation model is obtained by any of the methods in the first aspect of the present invention.
[0049] Step S203: Based on the recommendation results of the selection strategy, perform intervention selection behavior on the target user.
[0050] This embodiment can be applied to intervention and recommendation scenarios when users select payment tools on the payment page. First, the system acquires the target user's purchase behavior data, including historical order information, payment success / failure records, the type of payment tool used, payment amount range, types and quantities of purchased goods, and payment time distribution. In practical applications, this purchase behavior data can be extracted and organized from the user's e-commerce platform behavior logs, payment transaction systems, or account management platforms. For users with different dimensions, the system can construct structured feature vectors through preprocessing for subsequent modeling input. The acquired purchase behavior data is input into the selection strategy recommendation model, which is a pre-trained model that can be built using supervised learning. The model's goal is to predict the optimal payment tool selection scheme for the target user in the current payment scenario, such as whether to default to a particular payment tool when multiple payment tools are available, or whether to prompt the user with the optimal payment method. After obtaining the target user's selection strategy recommendation result, intervention and control are implemented based on this recommendation result for the target user's selection behavior on the payment page. Specific intervention methods could include directly selecting the model-recommended payment tools as the default payment tools; setting prominent style prompts for recommended payment tools in the payment tool list to guide users to click; or prompting users to switch to the model-recommended payment tools in the form of a dialog pop-up window, in order to maximize payment success rate and platform incentive utilization.
[0051] In addition to the modeling method based on purchase behavior data described above, this embodiment can also employ other forms of strategy recommendation methods. For example, a time-series decision model can be constructed by combining real-time user behavior signals; or a cold-start recommendation strategy for new users can be adopted, generating a preliminary payment tool selection strategy based on demographic information and basic user profile information.
[0052] By using the above-mentioned selection strategy recommendation method, the intelligence level of payment tool selection can be improved, the user guidance capability in the payment process can be enhanced, and the most suitable payment tool can be dynamically recommended according to the user's personalized characteristics, thereby improving the convenience of the payment process and the payment success rate. At the same time, it provides intelligent support for the platform to realize payment tool diversion and marketing campaign promotion.
[0053] Figure 3 This is a schematic diagram of a training method for a selection strategy recommendation model according to a preferred embodiment of the present invention. Figure 3 As shown, this embodiment provides an intelligent selection recommendation method based on a multi-treatment strategy. It aims to identify the gains that users can gain by selecting specific payment tools compared to not selecting them through modeling, thereby providing differentiated and personalized payment tool selection interventions for target users.
[0054] First, the training sample preparation phase was conducted, constructing unbiased (control group) and multiple intervention group (treatment group) sample sets. The control group samples consisted of users who did not select "EnjoyPay" in the EnjoyPay preference experiment, and an equal number of users who did not select "EnjoyPay" in the EnjoyPay non-preference experiment. The intervention group samples were divided into three subclasses: BaiTiao (a payment platform) selection intervention group, Quick Pay selection intervention group, and EnjoyPay selection intervention group. Each group was selected from users who actually selected the corresponding payment tool, ensuring the sample size was consistent with the control group sample size, thus forming the data structure required for the control experiment.
[0055] The model training phase then begins. This embodiment builds upon the classic DESCN (Deep Estimation for Subgroup Causal Effect Network) algorithm, constructing an improved Multi-DESCN model (a multi-intervention extended version of DESCN) to support multiple intervention strategy inputs and meet the needs of multi-payment tool strategy recommendation scenarios. Specifically, the output of the propensity network is transformed from the original sigmoid function (binary classification function) to a softmax function (multi-class classification function), thus supporting multi-class outputs and deriving a new estimated policy return branch (estr) for each intervention strategy. Simultaneously, to support multiple intervention paths, t2 and u2 branches are added to facilitate separate modeling of the potential effects under each intervention strategy. It is important to emphasize that the control group still uses a unified branch for modeling to avoid overly complex structures. In terms of loss function design, this embodiment configures corresponding loss terms for newly added branches; secondly, the original binary cross-entropy is changed to cross-entropy suitable for multi-class classification to handle the scoring estimation problem under multi-intervention strategies; in addition, to avoid the loss function of the estimation strategy reward branch dominating the overall optimization direction, the loss terms of the two estr branches are each weighted by 0.5, thereby maintaining the stability and balance of the overall training.
[0056] After model training is complete, the model inference and scoring phase begins. The trained model is used to score historical users, predicting their payment conversion probability under different payment tool selection strategies. Subsequently, based on the scoring results, the gain difference between the selected and unselected states for each user is calculated. This gain value represents the increase in the user's payment conversion probability after performing a specific payment tool selection operation.
[0057] During the user selection strategy screening phase, users are ranked by gain score and then filtered based on a set of scoring thresholds that comprehensively consider multiple factors such as payment tool conversion rate, order completion rate, and business GMV (Gross Merchandise Volume) target, thereby identifying the group most worthy of intervention. It is worth noting that the threshold settings are not static constants but dynamically adjusted according to different payment tools and real-time business metrics to maximize overall effectiveness.
[0058] Finally, for users who meet the threshold screening criteria, the system will implement an online selection intervention strategy when they make another purchase and enter the payment process. These intervention methods include, but are not limited to: automatically selecting recommended payment tools, defaulting to high-conversion tools, and pop-up prompts for recommended payment methods, aiming to improve user payment success rates and the platform's payment tool diversion capabilities.
[0059] This embodiment constructs an intelligent selection gain model based on multiple intervention strategies, which not only improves the accuracy and flexibility of payment tool strategy interventions but also breaks through the limitation of the traditional DESCN method, which only supports a single intervention, enabling optimized selection recommendations in a multi-payment tool environment. Furthermore, the model gain calculation mechanism effectively enhances the ability to assess user conversion potential, enabling the system to achieve more benefit-oriented intervention behavior recommendations.
[0060] According to a third aspect of the present invention, a training apparatus for a check-in strategy recommendation model is provided.
[0061] Figure 4 This is a schematic diagram of the main modules of the training device for the selection strategy recommendation model according to an embodiment of the present invention, as shown below. Figure 4 As shown, a training device 400 for a check-out strategy recommendation model includes: The data acquisition module 401 is used to collect historical user behavior data within a preset time period; The generation module 402 is used to generate control group samples and multiple intervention group samples based on users' historical behavior data; wherein, the control group samples are user samples that have not selected a payment tool, and the intervention group samples are user samples that have selected the corresponding payment tool; Training module 403 is used to construct a training sample set based on control group samples and multiple intervention group samples, and to train a check-off strategy recommendation model based on the training sample set.
[0062] Optionally, the generation module 402 is also used for: Determine the user's payment selection result based on the user's historical behavior data; Users are grouped based on payment selection results to obtain grouping results; Based on the grouping results, control group samples and multiple intervention group samples are generated.
[0063] Optionally, the training module 403 is also used for: Determine the payment outcomes for users in the control group and intervention group samples; If the user's payment result is successful, the sample is marked as a positive sample; If a user's payment result is a payment failure, the sample is marked as a negative sample. A training sample set is constructed based on positive and negative samples.
[0064] Optionally, the training module 403 is also used for: Construct a multi-branch neural network structure; wherein, the multi-branch neural network model is used to evaluate the payment conversion probability of users under multiple payment tool selection strategies; A multivariate joint loss function is defined based on the neural network structure; Based on the training sample set and the multivariate joint loss function, a neural network structure is trained to obtain a checklist recommendation model.
[0065] Optionally, the training device 400 also includes a testing module, which is used for: Collect historical user behavior data within a second preset time period and construct a test sample set; Input the test sample set into the selected strategy recommendation model to obtain the test recommendation results; The selection strategy recommendation model is optimized based on the test recommendation results.
[0066] It should be noted that the specific implementation details of the training device for the selection strategy recommendation model in this embodiment of the invention have been described in detail in the training method of the selection strategy recommendation model above, so the details will not be repeated here.
[0067] According to a fourth aspect of the present invention, a selection strategy recommendation device is provided.
[0068] Figure 5 This is a schematic diagram of the main modules of the selection strategy recommendation device according to an embodiment of the present invention, as shown below. Figure 5 As shown, a selection strategy recommendation device 500 includes: Module 501 is used to acquire purchase behavior data of target users; Prediction module 502 is used to input purchase behavior data into the selection strategy recommendation model, and the selection strategy recommendation model outputs the selection strategy recommendation result for the target user; wherein, the selection strategy recommendation model is obtained by any of the methods in the first aspect of the present invention; Recommendation module 503 is used to perform intervention selection behavior on target users based on the recommendation results of the selection strategy.
[0069] It should be noted that the specific implementation details of the selection strategy recommendation device in this embodiment of the invention have been described in detail in the selection strategy recommendation method above, so the details will not be repeated here.
[0070] According to a fifth aspect of the present invention, an electronic device is provided, comprising: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the methods provided by the first aspect and / or the second aspect of the embodiments of the present invention.
[0071] According to a sixth aspect of the present invention, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods provided in the first aspect and / or the second aspect of the present invention.
[0072] According to a seventh aspect of the present invention, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods provided by the first aspect and / or the second aspect of the present invention.
[0073] Figure 6 An exemplary system architecture 600 is shown, which can be used to train a selection strategy recommendation model according to embodiments of the present invention, or to train a selection strategy recommendation model.
[0074] Figure 6 An exemplary system architecture 600 is shown that can be applied to the selection strategy recommendation method or selection strategy recommendation device of the present invention.
[0075] like Figure 6 As shown, system architecture 600 may include terminal devices 601, 602, and 603, a network 604, and a server 605. Network 604 serves as the medium for providing communication links between terminal devices 601, 602, and 603 and server 605. Network 604 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0076] Users can use terminal devices 601, 602, and 603 to interact with server 605 via network 604 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 601, 602, and 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0077] Terminal devices 601, 602, and 603 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0078] Server 605 can be a server providing various services, such as a backend management server supporting shopping websites browsed by users using terminal devices 601, 602, and 603 (for example only). The backend management server can analyze and process data such as received model training requests, and feed back the processing results (such as selecting a strategy to recommend a model - for example only) to the terminal device.
[0079] It should be noted that the training method of the selection strategy recommendation model provided in the embodiments of the present invention is generally run by the server 605, and correspondingly, the training device of the selection strategy recommendation model is generally set in the server 605.
[0080] It should be noted that the selection strategy recommendation method provided in this embodiment of the invention is generally run by server 605, and correspondingly, the selection strategy recommendation device is generally set in server 605.
[0081] It should be understood that Figure 6 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0082] The following is for reference. Figure 7 It shows a schematic diagram of the structure of a computer system 700 suitable for implementing a terminal device of the present invention. Figure 7 The terminal device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0083] like Figure 7 As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 702 or programs loaded from storage section 708 into random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the system 700. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0084] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0085] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is run by central processing unit (CPU) 701, it performs the functions defined above in the system of this invention.
[0086] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0087] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more operable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually operate substantially in parallel, and they may sometimes operate in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0088] The modules described in the embodiments of the present invention can be implemented in software or hardware. The described modules can also be housed in a processor; for example, a processor may include a data acquisition module, a generation module, and a training module. The names of these modules are not necessarily limiting in certain circumstances; for example, a data acquisition module may be described as "a module for collecting historical user behavior data within a preset time period." Alternatively, a processor may include an acquisition module, a prediction module, and a recommendation module. The names of these modules are not necessarily limiting in certain circumstances; for example, a recommendation module may be described as "a module for performing intervention selection behavior on target users based on the recommendation results of a selection strategy."
[0089] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to include: collecting historical user behavior data within a first preset time period; generating a control group sample and multiple intervention group samples based on the historical user behavior data; wherein the control group sample consists of user samples that have not selected a payment tool, and the intervention group sample consists of user samples that have selected the corresponding payment tool; constructing a training sample set based on the control group sample and the multiple intervention group samples; and training a selection strategy recommendation model based on the training sample set. Alternatively, the device implements the following method: acquiring purchase behavior data of a target user; inputting the purchase behavior data into the selection strategy recommendation model, and the selection strategy recommendation model outputting a selection strategy recommendation result for the target user; wherein the selection strategy recommendation model is obtained by any of the methods in the first aspect of the present invention; and performing intervention selection behavior on the target user based on the selection strategy recommendation result.
[0090] The computer program product provided in the embodiments of the present invention includes a computer program that, when executed by a processor, implements the training method of the selection strategy recommendation model in the first aspect of the present invention and / or the selection strategy recommendation method in the second aspect of the present invention.
[0091] According to the technical solution of the present invention, the following advantages or beneficial effects are achieved: By collecting historical user behavior data within a first preset time period; generating a control group sample and multiple intervention group samples based on the historical user behavior data; wherein, the control group sample consists of user samples who have not selected a payment tool, and the intervention group sample consists of user samples who have selected the corresponding payment tool; constructing a training sample set based on the control group sample and multiple intervention group samples; and training a selection strategy recommendation model based on the training sample set. This embodiment, by collecting historical user behavior data and constructing a multi-intervention group experimental sample system to identify the difference in order completion between users selecting and not selecting payment tools, and using this as a training sample set to train a selection strategy recommendation model, can reduce the time cost of selection strategy recommendation, and the recommendation results are more in line with user needs, thereby improving the user's order conversion rate. By obtaining the target user's purchase behavior data; inputting the purchase behavior data into the selection strategy recommendation model, the selection strategy recommendation model outputs the target user's selection strategy recommendation result; wherein, the selection strategy recommendation model is obtained by any method in the first aspect of the present invention; and performing intervention selection behavior on the target user based on the selection strategy recommendation result. This embodiment can make the selection strategy recommendation result more in line with user needs, improving the user experience.
[0092] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
[0093] It should be noted that the acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
Claims
1. A training method for a check-in strategy recommendation model, characterized in that, include: Collect historical user behavior data within the first preset time period; Based on the user's historical behavior data, a control group sample and multiple intervention group samples are generated; wherein, the control group sample consists of user samples who have not selected a payment tool, and the intervention group sample consists of user samples who have selected the corresponding payment tool; A training sample set is constructed based on the control group samples and multiple intervention group samples, and a check-off strategy recommendation model is trained based on the training sample set.
2. The method according to claim 1, characterized in that, Based on the user's historical behavior data, a control group sample and multiple intervention group samples are generated, including: The user's payment selection result is determined based on the user's historical behavior data; Based on the payment selection results, the users are grouped to obtain grouping results; Based on the grouping results, control group samples and multiple intervention group samples are generated.
3. The method according to claim 1, characterized in that, A training sample set was constructed based on samples from the control group and multiple intervention groups, including: Determine the user payment results for the control group sample and the intervention group sample; If the user's payment result is successful, then the sample is marked as a positive sample; If the user's payment result is a payment failure, then the sample is marked as a negative sample; A training sample set is constructed based on the positive and negative samples.
4. The method according to claim 1, characterized in that, A selection strategy recommendation model is trained based on the aforementioned training sample set, including: Construct a multi-branch neural network structure; wherein the multi-branch neural network model is used to evaluate the payment conversion probability of a user under multiple payment tool selection strategies; A multivariate joint loss function is defined based on the aforementioned neural network structure; Based on the training sample set and the multivariate joint loss function, the neural network structure is trained to obtain the check-in strategy recommendation model.
5. The method according to claim 1, characterized in that, After training the check-in strategy recommendation model based on the aforementioned training sample set, the model further includes: Collect historical user behavior data within a second preset time period and construct a test sample set; The test sample set is input into the checkpoint strategy recommendation model to obtain the test recommendation results; The selection strategy recommendation model is optimized based on the test recommendation results.
6. A selection strategy recommendation method, characterized in that, include: Obtain purchase behavior data from target users; The purchase behavior data is input into the selection strategy recommendation model, and the selection strategy recommendation model outputs the selection strategy recommendation result for the target user; wherein, the selection strategy recommendation model is obtained by the method described in any one of claims 1 to 5; Based on the recommendation results of the selection strategy, intervention selection behavior is performed on the target user.
7. A training device for a check-in strategy recommendation model, characterized in that, include: The data collection module is used to collect historical user behavior data within a preset time period. The generation module is used to generate control group samples and multiple intervention group samples based on the user's historical behavior data; wherein, the control group samples are user samples that have not selected a payment tool, and the intervention group samples are user samples that have selected the corresponding payment tool; The training module is used to construct a training sample set based on the control group samples and multiple intervention group samples, and to train a check-off strategy recommendation model based on the training sample set.
8. A selection strategy recommendation device, characterized in that, include: The acquisition module is used to acquire purchase behavior data of target users; A prediction module is used to input the purchase behavior data into a selection strategy recommendation model, wherein the selection strategy recommendation model outputs the selection strategy recommendation result for the target user; wherein the selection strategy recommendation model is obtained by the method described in any one of claims 1 to 5; The recommendation module is used to perform intervention selection behavior on the target user based on the recommendation results of the selection strategy.
9. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-6.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.