Method and device for starting cold start product, electronic equipment and storage medium

By employing reinforcement learning algorithms to iteratively control parameters in cold-start products, the contradiction between overall benefits and efficiency during the startup process is resolved, thereby improving cold-start efficiency while ensuring overall benefits.

CN116578779BActive Publication Date: 2026-04-21MICRO INSURANCE AGENCY LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MICRO INSURANCE AGENCY LTD
Filing Date
2023-04-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing cold start products cannot balance the relationship between the overall benefits of product recommendations and the startup efficiency of cold start products during the startup process, resulting in a decrease in overall benefits.

Method used

A predefined reinforcement learning algorithm is used to iterate the control parameters at preset intervals to obtain the control parameters after each iteration. These parameters are used to determine the number of users flowing into the cold start model within the preset interval. The number of users is updated based on the control parameters after each iteration until the number of iterations reaches the preset number or the number of samples of the cold start product reaches the preset threshold.

Benefits of technology

This approach maximizes overall revenue while improving the startup efficiency of cold-start products, avoiding the improper allocation of exposure opportunities caused by the cold-start model occupying most of the traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116578779B_ABST
    Figure CN116578779B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, electronic device, and storage medium for launching a cold-start product. The method includes: upon detecting a cold-start product, iterating control parameters at preset time intervals based on a predefined reinforcement learning algorithm to obtain control parameters after each iteration. The reinforcement learning algorithm is used to maximize the overall expected return of system users in both the cold-start and warm-start models, as well as the expected return of users flowing into the cold-start model. Based on the control parameters after each iteration, the number of users flowing into the cold-start model within each preset time interval is updated. When the number of iterations reaches a preset number or the number of samples of the cold-start product reaches a preset threshold, the cold-start product launch is determined to be complete. In this way, the reinforcement learning algorithm can simultaneously consider both the overall return of product recommendations and the launch efficiency of the cold-start product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent recommendation technology, and in particular to a startup method, device, electronic device and storage medium for a cold start product. Background Technology

[0002] In recommendation systems, new users and products are constantly being added. Because these new users and products lack sufficient historical behavioral data, they often cannot receive accurate recommendations or be accurately recommended to the right users. This is known as the cold start problem in recommendation systems.

[0003] Currently, existing cold-start products often overlook a crucial issue during the launch process: on current online platforms, such as e-commerce, short video, advertising, and insurance platforms, cold starts and hot starts occur simultaneously. Typically, vendors prepare two models for recommendation ranking: a cold-start model and a hot-start model. The cold-start model handles the distribution of products launched initially, and once a sufficient sample size is accumulated, the cold-start product is then incorporated into the hot-start model for prediction. In this scenario, prioritizing the launch efficiency of cold-start products and allocating as much exposure as possible to them can potentially lead to a decrease in the overall revenue of both cold-start and hot-start products. Therefore, existing cold-start product launch methods do not adequately balance the relationship between overall product recommendation revenue and the launch efficiency of cold-start products.

[0004] Therefore, how to simultaneously consider the overall benefits of product recommendations and the startup efficiency of cold-start products has become an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides a startup method, apparatus, electronic device, and storage medium for a cold-start product, to address the problem that existing startup methods for cold-start products do not adequately balance the relationship between the overall benefits of product recommendation and the startup efficiency of the cold-start product.

[0006] On one hand, embodiments of this application provide a startup method for a cold-start product, the method comprising:

[0007] When a cold start product is detected, the control parameters are iterated at preset intervals based on a predefined reinforcement learning algorithm to obtain the control parameters after each iteration. The control parameters are used to determine the number of users flowing into the cold start model within the preset interval. The reinforcement learning algorithm is used to maximize the overall expected revenue generated by system users in the cold start model and the warm start model, as well as maximize the expected revenue generated by users flowing into the cold start model in the cold start model.

[0008] Based on the control parameters after each iteration, the number of users flowing into the cold start model within each preset time period is updated;

[0009] When the number of iterations reaches a preset number or the number of samples of the cold start product reaches a preset threshold, the cold start product is determined to have started successfully. The number of samples of the cold start product is determined by the relevant operations of the users flowing into the cold start model on the cold start product.

[0010] Optionally, the control parameters based on a predefined reinforcement learning algorithm are iterated at preset intervals to obtain control parameters after each iteration, including:

[0011] Obtain the predefined objective function and related parameters in the reinforcement learning algorithm, wherein the objective function is used to characterize the overall expected benefit generated by the system user in the cold start model and the warm start model, and the related parameters are used to characterize the state, reward and action required to participate in the computation process of the reinforcement learning algorithm;

[0012] Based on the objective function and the relevant parameters, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

[0013] Optionally, the objective function is calculated using the following formula:

[0014]

[0015] Among them, R base t represents the expected revenue generated by users flowing into the hot-start model within the hot-start model. i U represents the i-th preset duration. C R represents the set of users flowing into the cold start model. C (u) represents the expected revenue generated by any user flowing into the cold start model, R. M () represents the expected revenue generated by any user flowing into the hot start model in the hot start model.

[0016] Optionally, the step of iterating the control parameters at preset intervals based on the objective function and the relevant parameters to obtain the control parameters after each iteration includes:

[0017] Based on the state, the reward, and the action, a strategy function is determined to achieve the optimal objective of the objective function. The state is obtained by dividing the expected revenue generated by the cold start model and the warm start model. The reward is used to characterize the expected revenue generated when performing different actions on system users. The action is used to characterize the probability of system users flowing into the cold start model.

[0018] Based on the objective function and the strategy function, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

[0019] Optionally, the formula for determining the control parameters after each iteration is as follows:

[0020]

[0021] in, 'a' represents the action, and 's' represents the state. Let R represent the set of users with state s. C (u) represents the expected revenue generated by any user in state s in the cold start model, R M () represents the expected revenue generated by any user in state s in the hot-start model. This represents the control parameter corresponding to the (i+1)th preset duration. This represents the control parameter corresponding to the i-th preset duration, where i is any positive integer, and γ represents the update step size. Let (a|a,θ) represent the objective function, and (a|a,θ) represent the policy function.

[0022] This represents the gradient of the policy function with respect to the model parameters.

[0023] Optionally, before iterating the control parameters at preset intervals based on the objective function and the policy function to obtain the control parameters after each iteration, the method further includes:

[0024] The expected returns generated by system users in the cold start model are updated at preset intervals to determine the change in expected returns of the cold start model after each iteration.

[0025] The objective function is updated based on the change in expected return of the cold start model after each iteration;

[0026] The step of iterating the control parameters at preset intervals based on the objective function and the strategy function to obtain the control parameters after each iteration includes:

[0027] Based on the updated objective function and the strategy function, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

[0028] Optionally, the formula for determining the change in expected return after each iteration is as follows:

[0029]

[0030] in, This represents the expected change in revenue at the (i+1)th preset time interval. This represents the change in expected revenue corresponding to the i-th preset time period. This represents the total number of users who flow into the cold start model within the i-th preset time period, and β represents the forgetting parameter. ′ C (u) represents the actual benefit generated by any user in the cold start model within the i-th preset time period, R. C (u) represents the expected revenue generated by any user who flows into the cold start model within the i-th preset time period.

[0031] Optionally, the method is applied to a product recommendation system, wherein the products recommended by the product recommendation system include at least one of insurance products, video programs, online goods and advertisements, and the cold start product is a newly added product to be recommended in the product recommendation system.

[0032] On the other hand, embodiments of this application provide a starting device for a cold-start product, the device comprising:

[0033] An iterative module is used to iterate the control parameters at preset intervals based on a predefined reinforcement learning algorithm when a cold start product is detected, so as to obtain the control parameters after each iteration. The control parameters are used to determine the number of users flowing into the cold start model within the preset interval. The reinforcement learning algorithm is used to maximize the overall expected revenue generated by system users in the cold start model and the warm start model, as well as the expected revenue generated by users flowing into the cold start model in the cold start model.

[0034] The update module is used to update the number of users flowing into the cold start model within each preset time period based on the control parameters after each iteration;

[0035] The determination module is used to determine that the cold start product has been launched when the number of iterations reaches a preset number or the number of samples of the cold start product reaches a preset threshold. The number of samples of the cold start product is determined by the relevant operations of the cold start product by the users flowing into the cold start model.

[0036] On the other hand, embodiments of this application provide an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0037] Memory, used to store computer programs;

[0038] When a processor executes a program stored in memory, it implements the steps of the startup method for a cold-start product as described in any embodiment of the first aspect.

[0039] On the other hand, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the startup method for a cold-start product as described in any embodiment of the first aspect.

[0040] In this embodiment, upon detecting a cold-start product, a predefined reinforcement learning algorithm is used to iterate control parameters at preset intervals to obtain control parameters after each iteration. These control parameters determine the number of users flowing into the cold-start model within the preset interval. The reinforcement learning algorithm maximizes the overall expected return of system users in both the cold-start and warm-start models, as well as the expected return of users flowing into the cold-start model. Based on the control parameters after each iteration, the number of users flowing into the cold-start model within each preset interval is updated. When the number of iterations reaches a preset number or the number of samples of the cold-start product reaches a preset threshold, the cold-start product is determined to have completed its startup. The number of samples of the cold-start product is determined by the relevant operations performed by users flowing into the cold-start model. Through this method, the reinforcement learning algorithm can be used to control the number of users flowing into the cold-start model within each preset interval, thereby maximizing the overall expected return of system users in both the cold-start and warm-start models, as well as the expected return of users flowing into the cold-start model. In other words, this reinforcement learning algorithm can be used to balance the overall benefits of product recommendation with the launch efficiency of cold-start products, avoiding the extreme situation where the cold-start model occupies most of the traffic and allocates exposure opportunities to cold-start products. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 A schematic flowchart illustrating a startup method for a cold-start product provided in an embodiment of this application;

[0044] Figure 2 This is a schematic diagram of the structure of a product recommendation system provided in an embodiment of this application;

[0045] Figure 3 A schematic diagram of a test result provided in an embodiment of this application;

[0046] Figure 4 A schematic diagram illustrating another test result provided in an embodiment of this application;

[0047] Figure 5 This is a schematic diagram of the structure of a starting device for a cold start product provided in an embodiment of this application;

[0048] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0050] See Figure 1 , Figure 1 This is a flowchart illustrating a startup method for a cold-start product provided in an embodiment of this application. Figure 1 As shown, the startup method of this cold start product includes:

[0051] Step 101: When a cold start product is detected, the control parameters are iterated at preset intervals based on a predefined reinforcement learning algorithm to obtain the control parameters after each iteration. The control parameters are used to determine the number of users flowing into the cold start model within the preset interval. The reinforcement learning algorithm is used to maximize the overall expected revenue generated by system users in the cold start model and the warm start model, as well as maximize the expected revenue generated by users flowing into the cold start model in the cold start model.

[0052] It should be noted that the cold start product launch method provided in this application embodiment can be applied to any product recommendation system, such as an insurance policy recommendation system, a product recommendation system, a short video recommendation system, etc. This product recommendation system includes a cold start model and a hot start model. The cold start model is responsible for distributing cold start products. When a certain number of samples are accumulated, the cold start product will enter the hot start model for prediction. Here, we use one cold start product a.C As an example, let's assume that the product to be sorted by the hot-start model is I. M Then, the cold start model is responsible for sorting the products I. M +a C The hot-start model is used for the unordered product I that it is responsible for. M The system sorts the products, generates a corresponding recommendation list based on the sorting results, and finally recommends products to the corresponding users according to the sorting order of this recommendation list. The cold start model is used to sort the products I that it is responsible for. M +a C The system sorts the products and generates a corresponding recommendation list based on the sorting results. Finally, products are recommended to the corresponding users according to the sorting order of this recommendation list. Of course, the number of cold-start products can also be multiple; this embodiment does not impose a specific limitation. The product recommendation system also includes a predefined reinforcement learning algorithm. This predefined reinforcement learning algorithm can estimate the expected revenue generated by each of the recommendation lists of the hot-start model and the cold-start model, update the number of users flowing into the hot-start model and the cold-start model respectively, and then update the control parameters in the reinforcement learning algorithm based on the new overall expected revenue generated by the updated hot-start model and the cold-start model. Through continuous iterative learning and updating of the control parameters, the overall expected revenue generated by the hot-start model and the cold-start model during the entire launch process of the cold-start product is maximized. As an optional embodiment, the structure of this product recommendation system can be as follows: Figure 2 As shown, the cold start model coverage traffic refers to the number of users flowing into the cold start model, the hot start model coverage traffic refers to the number of users flowing into the hot start model, and the T+X strategy update refers to updating the control parameters in the reinforcement learning algorithm every preset time interval X.

[0053] Specifically, the aforementioned cold-start product refers to a product that lacks user interaction features, such as a product that has not been clicked, converted, or viewed by users. The purpose of the aforementioned reinforcement learning algorithm is to execute a series of appropriate actions based on a policy to maximize cumulative reward. The goal of this reinforcement learning algorithm is to maximize the overall expected return generated by system users in both the cold-start and warm-start models, while simultaneously maximizing the expected return generated by users flowing into the cold-start model. This reinforcement learning algorithm can be a value function-based algorithm, a policy-based algorithm, an actor-critic algorithm, etc., and this application does not impose any specific limitations.

[0054] The preset duration can be set according to actual needs. As an option, the preset duration can be determined based on the startup time of the cold start product. For example, assuming the estimated startup time for a cold start product is 24 hours, this 24 hours can be divided into several consecutive segments, each of which is a preset duration. The above control parameters are parameters in the reinforcement learning algorithm, used to determine the number of users flowing into the cold start model for each preset duration.

[0055] Step 102: Based on the control parameters after each iteration, update the number of users flowing into the cold start model within each preset time period.

[0056] In this step, after obtaining the control parameters after each iteration, the number of users flowing into the cold start model within each preset time period can be updated based on the control parameters after each iteration. Since the number of users flowing into the cold start model affects the expected revenue generated in the cold start model and the startup efficiency of the cold start product, while ensuring that the overall expected revenue generated by the warm start model and the cold start model is maximized, we can control the inflow of users into the cold start model as much as possible, so as to maximize the expected revenue generated by the users flowing into the cold start model.

[0057] Step 103: When the number of iterations reaches the preset number or the number of samples of the cold start product reaches the preset threshold, the cold start product is determined to be launched successfully. The number of samples of the cold start product is determined by the relevant operations of the users flowing into the cold start model on the cold start product.

[0058] Specifically, the preset number of times and preset threshold can be set according to actual needs, and this application embodiment does not impose specific limitations.

[0059] In this step, the model parameters are updated at preset intervals during reinforcement learning. This means the number of users flowing into the cold start model is updated at preset intervals. This ensures that each iteration maximizes the overall expected return for users in both the cold start and warm start models, as well as the expected return for users flowing into the cold start model. When the preset number of iterations is reached or the number of samples in the cold start product reaches a preset threshold, the cold start product is considered to have successfully started, and iteration stops.

[0060] In this embodiment, a reinforcement learning algorithm can be used to control the number of users flowing into the cold start model within each preset time period, thereby maximizing the overall expected revenue generated by system users in both the cold start and warm start models, as well as maximizing the expected revenue generated by users flowing into the cold start model. In other words, this reinforcement learning algorithm can simultaneously consider the overall revenue of product recommendations and the startup efficiency of cold start products, avoiding the extreme situation where the cold start model occupies most of the traffic and allocates exposure opportunities to cold start products.

[0061] Further, in step 101 above, based on a predefined reinforcement learning algorithm, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration, including:

[0062] Obtain the predefined objective function and related parameters in the reinforcement learning algorithm. The objective function is used to characterize the overall expected benefit generated by the system user in the cold start model and the warm start model. The related parameters are used to characterize the state, reward and action required to participate in the computation process of the reinforcement learning algorithm.

[0063] Based on the objective function and relevant parameters, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

[0064] Specifically, the objective function described above can be used to characterize the overall expected revenue generated by system users in the cold start model and the warm start model, and is expressed by the following formula:

[0065]

[0066] Among them, R M R represents the total expected return generated by the hot-start model during the entire startup process. C t represents the total expected return generated by the cold start model during the entire startup process. i U represents the i-th preset duration, where i is any positive integer. M U represents the set of users flowing into the warm-start model. C R represents the set of users flowing into the cold start model. C (u) represents the expected revenue generated by any user flowing into the cold start model in the cold start model, R. M () represents the expected revenue generated by any user flowing into the hot start model.

[0067] As an alternative implementation, the expected revenue R generated by any user flowing into the hot start model can be defined. M The calculation formula for () and the expected return R generated by any user flowing into the cold start model in the cold start model. CThe formulas for calculating (u) are as follows:

[0068]

[0069]

[0070] Among them, I M Let a represent the set of products to be sorted in the hot start model. C This indicates a cold start product, u represents a specific user in the system, and Imp(a) represents a user who is cold-started. i (u) represents the binary exposure function, i.e., product a i Given user u, whether the user will be exposed is determined by a value of 0 / 1 (1 indicates exposure, 0 indicates no exposure), CR(a i ,u) represents the corresponding user product pair (a i The estimated conversion rate of u), V(a) i ) represents the price of product ai, α represents the value measurement factor, and N C (a C ) represents the collected cold start product a C The number of samples.

[0071] Therefore, the difference between the cold start model and the hot start model lies in the fact that the cold start model requires the collection of cold start product a. C The number of samples is considered as part of the overall contribution value. Of course, as another alternative implementation, when defining the objective function, a simplified version of the above formula (1) can also be used (such as R in the above formula (1)). M () and R C (u) Increase the corresponding weight coefficients respectively, or adjust R in the above formula (1). M () and R C (u) Calculations are performed by adding a certain constant, etc., and the embodiments of this application are not specifically limited.

[0072] Specifically, the aforementioned parameters are used to characterize the state, reward, and action parameters required for the computation process of the reinforcement learning algorithm. Here, the state represents the system state, and as an optional implementation, it can be represented by the formula s(k,j)=[(R C,k ,R C,k+1 ),(R M,j ,R M,j+1 Let R be a state, where R is a variable. C,k <R C,k+1 ,R M,j <R M,j+1 Both k and j represent user-defined serial numbers, which can be understood as starting with R. C and R MConstruct two coordinate axes, R C,k and R M,j Let R represent points on the two coordinate axes respectively, hence the requirement to R. C,k <R C,k+1 ,R M,j <R M,j+1 This similar requirement essentially represents the coordinate axis values ​​gradually increasing with the subscript. Here, the reward (r(s,a)) represents the gain generated by taking action a in a specific state s. Based on the aforementioned objective function, the reward can be defined as follows: if the current state is led to a warm-start model, then the current reward is 0; if it is led to a cold-start model, then the current reward is R. C -R M A good action should generate a positive reward; conversely, if an action generates a negative reward, then the actions corresponding to the entire state need to be updated. Here, action (a) represents the probability of guiding the user in state s to the cold start model. If action (a) is 1, it means that all users in state s will be guided to the cold start model.

[0073] In one embodiment, a predefined objective function and related parameters in the reinforcement learning algorithm can be obtained first. Then, based on the objective function and related parameters, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration. This ensures that each iteration approaches the optimal objective function, thereby maximizing the overall expected return of system users in both the cold start and warm start models, as well as maximizing the expected return of users flowing into the cold start model.

[0074] Furthermore, the formula for calculating the objective function is as follows:

[0075]

[0076] Among them, R base t represents the expected revenue generated by users flowing into the warm-start model. i U represents the i-th preset duration. C R represents the set of users flowing into the cold start model. C (u) represents the expected revenue generated by any user flowing into the cold start model in the cold start model, R. M (u) represents the expected revenue generated by any user flowing into the hot-start model. It should be noted that R... base This can be understood as R corresponding to each user flowing into the warm-start model. M The sum of (u) can approximate R. base Treat it as a constant.

[0077] In one embodiment, the objective function is transformed into a function that solves for... The value of , i.e., the goal of the cold start phase, can be transformed into: the expected revenue generated by the users allocated to the cold start model should be as high as possible. Based on the aforementioned analysis of R... C As the definition suggests, this requires generating as many samples as possible for users assigned to the cold start model, such as clicking on new products, converting to new products, and viewing new products. Simultaneously, the expected return generated by the cold start model for the current user should also be as high as possible. In this way, through the converted objective function, the launch efficiency of cold start products can be better improved while ensuring the overall return of product recommendations.

[0078] Furthermore, based on the objective function and relevant parameters, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration, including:

[0079] Based on state, reward, and action, a strategy function is determined to achieve the optimal objective function. The state is determined by dividing the expected revenue generated by the cold start model and the warm start model. The reward is used to characterize the expected revenue generated when performing different actions on system users. The action is used to characterize the probability of system users flowing into the cold start model.

[0080] Based on the objective function and the policy function, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

[0081] Specifically, the strategy function π(a|s,θ) can be expressed by the following formula:

[0082]

[0083] Where 'a' represents the action, 's' represents the state, and 'θ' represents the control parameter. Let R represent the set of users with state s. C (u) represents the expected revenue generated by any user flowing into the cold start model in the cold start model, R. M () represents the expected revenue generated by any user flowing into the hot start model.

[0084] In one embodiment, a strategy function that achieves the optimal objective function can be determined based on predefined states, rewards, and actions. Then, based on the objective function and the strategy function, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration. This allows for subsequent updates to the number of users flowing into the cold start model within each preset interval based on the control parameters after each iteration.

[0085] Furthermore, the formula for determining the control parameters after each iteration is as follows:

[0086]

[0087] in, 'a' represents the action, and 's' represents the state. Let R represent the set of users with state s. C (u) represents the expected revenue generated by any user in state s in the cold start model, R M () represents the expected revenue generated by any user in state s in the hot-start model. This represents the control parameter corresponding to the (i+1)th preset duration. This represents the control parameter corresponding to the i-th preset duration, where i is any positive integer, and γ represents the update step size. Let (a|s,θ) represent the objective function (i.e., Equation 4 above), and let (a|s,θ) represent the policy function. This represents the gradient of the policy function with respect to the model parameters.

[0088] According to the above formula (6), the control parameter θ can be updated at preset intervals, thereby continuously updating action a so that action a is updated in the direction that produces the optimal expected return (i.e., maximizing the objective function). When the value of i (i.e., the number of iterations) reaches a certain set value, or N C (a C When the value of ) (i.e. the number of samples of cold start products) reaches a certain set value, the iteration can be stopped.

[0089] Furthermore, before the above steps, which iterate the control parameters at preset intervals based on the objective function and the policy function to obtain the control parameters after each iteration, the method further includes:

[0090] The expected revenue generated by system users in the cold start model is updated every preset time interval to determine the change in expected revenue of the cold start model after each iteration.

[0091] The objective function is updated based on the change in expected return after each iteration of the cold start model;

[0092] The above steps, based on the objective function and policy function, iterate the control parameters at preset intervals to obtain the control parameters after each iteration, including:

[0093] Based on the updated objective function and policy function, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

[0094] In one embodiment, considering the lack of user interaction features in the cold start model, the originally estimated expected return R CThere may be inaccuracies. To address this issue, while updating the control parameter θ, the expected return R of the original estimate of the cold start model can be updated. C Also update it. This way, it can be done by updating R. C The continuous updating of the objective function makes the calculation results more accurate, thereby improving the accuracy of the control parameter θ.

[0095] Furthermore, the formula for determining the change in expected return after each iteration is as follows:

[0096]

[0097] in, This represents the expected change in revenue at the (i+1)th preset time interval. This represents the change in expected revenue corresponding to the i-th preset time period. This represents the total number of users who enter the cold start model within the i-th preset time period, and β represents the forgetting parameter. ′ C (u) represents the actual revenue generated by any user in the cold start model within the i-th preset time period, R. C (u) represents the expected revenue generated by any user who flows into the cold start model within the i-th preset time period.

[0098] Using the formula (7) above, the change in expected return after the i-th iteration can be calculated. Subsequently, this change in expected return can be used as a basis. The expected return for the (i+1)th iteration The correction is made, and the calculation formula is as follows:

[0099] This effectively solves the problem of the expected return R estimated in the original cold start model. C There may be inaccuracies, which could lead to low accuracy in the calculation of the control parameter θ, thus improving the accuracy of the control parameter θ.

[0100] Furthermore, the cold start product launch method is applied to a product recommendation system. The products recommended by the product recommendation system include at least one of insurance products, video programs, online goods, and advertisements. The cold start product is a newly added product to be recommended in the product recommendation system.

[0101] In one embodiment, the cold-start product launch method provided in this application can be applied to a product recommendation system. This system can recommend any product to be recommended, such as one or more combinations of insurance products, video programs, online goods, and advertisements. When the product recommendation system detects a cold-start product, it can iterate the control parameters at preset intervals based on a predefined reinforcement learning algorithm to obtain the control parameters after each iteration. Then, based on the control parameters after each iteration, it updates the number of users flowing into the cold-start model within each preset interval. This accelerates the conversion efficiency of cold-start samples while ensuring the overall system benefit. When the number of iterations reaches a preset number or the number of cold-start product samples reaches a preset threshold, the iteration stops, and it is determined that the cold-start product has been successfully launched. The cold-start product can then be added to the warm-start model for prediction.

[0102] In one example, the product recommendation system can be applied to an online insurance recommendation scenario. In this scenario, in order to address the potential loss of overall revenue due to independently optimizing the cold start model, the cold start product launch method provided in this application embodiment (i.e., a scheduling algorithm based on delayed reinforcement learning oriented towards overall revenue) can be used to recommend new insurance products in an online insurance recommendation scenario where revenue is relatively sensitive.

[0103] Specifically, the cold-start product launch method described above maximizes the conversion of the original cold-start objective (i.e., collecting user samples for new insurance products) into equivalent revenue (expected return), while simultaneously maximizing the overall return generated by integrating the original cold-start objective with the revenue (expected return) from the warm-start and cold-start models. After the cold-start insurance product is deployed online, this improves the efficiency of new sample collection while minimizing overall revenue loss. The cold-start product launch method provided in this embodiment was deployed on a platform and underwent a week of online testing. The test results are as follows: Figure 3 and Figure 4 As shown. According to Figure 3 and Figure 4 As can be seen, compared to the baseline cold start scheme (i.e., random traffic grouping), the cold start product launch method provided in this application embodiment can achieve an efficiency improvement of 1.21x (for a single cold start product) and 1.31x (for multiple cold start products), while achieving an overall revenue improvement of 1.05x (for a single cold start product) and 1.02x (for multiple cold start products). This demonstrates that the cold start product launch method provided in this application embodiment ensures both efficient cold start and overall revenue.

[0104] In addition, see Figure 5 , Figure 5This is a schematic diagram of the structure of a starting device for a cold start product provided in an embodiment of this application. Figure 5 As shown, the starting device 500 of the cold start product includes:

[0105] The iteration module 501 is used to iterate the control parameters at preset intervals based on a predefined reinforcement learning algorithm when a cold start product is detected, so as to obtain the control parameters after each iteration. The control parameters are used to determine the number of users flowing into the cold start model within the preset interval. The reinforcement learning algorithm is used to maximize the overall expected revenue generated by system users in the cold start model and the warm start model, as well as the expected revenue generated by users flowing into the cold start model.

[0106] The update module 502 is used to update the number of users flowing into the cold start model within each preset time period based on the control parameters after each iteration.

[0107] The determination module 503 is used to determine that the cold start product has been launched when the number of iterations reaches a preset number or the number of samples of the cold start product reaches a preset threshold. The number of samples of the cold start product is determined by the relevant operations of the users flowing into the cold start model on the cold start product.

[0108] Furthermore, the iteration module 501 includes:

[0109] The acquisition submodule is used to acquire the predefined objective function and related parameters in the reinforcement learning algorithm. The objective function is used to characterize the overall expected benefit generated by the system user in the cold start model and the warm start model, and the related parameters are used to characterize the state, reward and action required to participate in the computation process of the reinforcement learning algorithm.

[0110] The iteration submodule is used to iterate the control parameters at preset intervals based on the objective function and related parameters to obtain the control parameters after each iteration.

[0111] Furthermore, the formula for calculating the objective function is as follows:

[0112]

[0113] Among them, R base t represents the expected revenue generated by users flowing into the warm-start model. i U represents the i-th preset duration. C R represents the set of users flowing into the cold start model. C (u) represents the expected revenue generated by any user flowing into the cold start model in the cold start model, R. M() represents the expected revenue generated by any user flowing into the hot start model.

[0114] Furthermore, the iterative submodule includes:

[0115] The determination unit is used to determine the strategy function that achieves the optimal objective function based on the state, reward, and action. The state is obtained by dividing the expected revenue generated by the cold start model and the warm start model. The reward is used to characterize the expected revenue generated when performing different actions on system users. The action is used to characterize the probability of system users flowing into the cold start model.

[0116] The iterative unit is used to iterate the control parameters at preset intervals based on the objective function and the policy function to obtain the control parameters after each iteration.

[0117] Furthermore, the formula for determining the control parameters after each iteration is as follows:

[0118]

[0119] in, 'a' represents the action, and 's' represents the state. Let R represent the set of users with state s. C (u) represents the expected revenue generated by any user in state s in the cold start model, R M () represents the expected revenue generated by any user in state s in the hot-start model. This represents the control parameter corresponding to the (i+1)th preset duration. This represents the control parameter corresponding to the i-th preset duration, where i is any positive integer, and γ represents the update step size. Let (a|s,θ) represent the objective function and (a|s,θ) represent the policy function. This represents the gradient of the policy function with respect to the model parameters.

[0120] Furthermore, the iteration module 501 includes:

[0121] The first update submodule is used to update the expected revenue generated by system users in the cold start model at preset intervals, and to determine the change in expected revenue of the cold start model after each iteration.

[0122] The second update submodule is used to update the objective function based on the change in expected return of the cold start model after each iteration.

[0123] The iteration submodule is also used to iterate the control parameters at preset intervals based on the updated objective function and policy function to obtain the control parameters after each iteration.

[0124] Furthermore, the formula for determining the change in expected return after each iteration is as follows:

[0125]

[0126] in, This represents the expected change in revenue at the (i+1)th preset time interval. This represents the change in expected revenue corresponding to the i-th preset time period. This represents the total number of users who enter the cold start model within the i-th preset time period, and β represents the forgetting parameter. ′ C (u) represents the actual revenue generated by any user in the cold start model within the i-th preset time period, R. C (u) represents the expected revenue generated by any user who flows into the cold start model within the i-th preset time period.

[0127] It should be noted that the starting device 500 of the cold start product can implement the steps of the starting method of the cold start product provided in any of the aforementioned method embodiments, and can achieve the same technical effect, which will not be described in detail here.

[0128] like Figure 6 As shown in the illustration, this application also provides an electronic device, including a processor 611, a communication interface 612, a memory 613, and a communication bus 614, wherein the processor 611, the communication interface 612, and the memory 613 communicate with each other via the communication bus 614.

[0129] Memory 613 is used to store computer programs;

[0130] In one embodiment of this application, when the processor 611 executes the program stored in the memory 613, it implements the startup method of the cold start product provided in any of the foregoing method embodiments, including:

[0131] When a cold start product is detected, the control parameters are iterated at preset intervals based on a predefined reinforcement learning algorithm to obtain the control parameters after each iteration. The control parameters are used to determine the number of users flowing into the cold start model within the preset interval. The reinforcement learning algorithm is used to maximize the overall expected revenue generated by system users in the cold start model and the warm start model, as well as maximize the expected revenue generated by users flowing into the cold start model in the cold start model.

[0132] Based on the control parameters after each iteration, the number of users flowing into the cold start model within each preset time period is updated;

[0133] When the number of iterations reaches a preset number or the number of samples of the cold start product reaches a preset threshold, the cold start product is determined to have started successfully. The number of samples of the cold start product is determined by the relevant operations of the users flowing into the cold start model on the cold start product.

[0134] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the startup method for a cold-start product as provided in any of the foregoing method embodiments.

[0135] In addition, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of the startup method for the cold-start product provided in any of the foregoing method embodiments.

[0136] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0137] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for starting a cold-start product, characterized in that, The method includes: When a cold start product is detected, the control parameters are iterated at preset intervals based on a predefined reinforcement learning algorithm to obtain the control parameters after each iteration. The control parameters are used to determine the number of users flowing into the cold start model within the preset interval. The reinforcement learning algorithm is used to maximize the overall expected revenue generated by system users in the cold start model and the warm start model, as well as maximize the expected revenue generated by users flowing into the cold start model in the cold start model. Based on the control parameters after each iteration, the number of users flowing into the cold start model within each preset time period is updated; When the number of iterations reaches a preset number or the number of samples of the cold start product reaches a preset threshold, it is determined that the cold start product has been successfully launched. The number of samples of the cold start product is determined by the relevant operations of the users flowing into the cold start model on the cold start product. The aforementioned reinforcement learning algorithm, which iterates the control parameters at preset intervals to obtain the control parameters after each iteration, includes: Obtain the predefined objective function and related parameters in the reinforcement learning algorithm, wherein the objective function is used to characterize the overall expected benefit generated by the system user in the cold start model and the warm start model, and the related parameters are used to characterize the state, reward and action required to participate in the computation process of the reinforcement learning algorithm; Based on the objective function and the relevant parameters, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

2. The method according to claim 1, characterized in that, The objective function is calculated using the following formula: in, This represents the expected revenue generated by users flowing into the hot-start model within the hot-start model. This represents the i-th preset duration. This represents the set of users flowing into the cold start model. This represents the expected revenue generated by any user flowing into the cold start model within the cold start model. This represents the expected revenue generated by any user flowing into the hot start model within the hot start model.

3. The method according to claim 1, characterized in that, The step of iterating the control parameters at preset intervals based on the objective function and the relevant parameters to obtain the control parameters after each iteration includes: Based on the state, the reward, and the action, a strategy function is determined to achieve the optimal objective of the objective function. The state is obtained by dividing the expected revenue generated by the cold start model and the warm start model. The reward is used to characterize the expected revenue generated when performing different actions on system users. The action is used to characterize the probability of system users flowing into the cold start model. Based on the objective function and the strategy function, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

4. The method according to claim 3, characterized in that, The formula for determining the control parameters after each iteration is as follows: in, 'a' represents the action. Indicates the state, Indicates the state is The user set, This represents the expected revenue generated by any user flowing into the cold start model within the cold start model. This represents the expected revenue generated by any user flowing into the hot start model within the hot start model. This represents the control parameter corresponding to the (i+1)th preset duration. This represents the control parameter corresponding to the i-th preset duration, where i is any positive integer. Indicates the update step size. Denotes the objective function, This represents the policy function. This represents the gradient of the policy function with respect to the model parameters.

5. The method according to claim 3, characterized in that, Before iterating the control parameters at preset intervals based on the objective function and the policy function to obtain the control parameters after each iteration, the method further includes: The expected returns generated by system users in the cold start model are updated at preset intervals to determine the change in expected returns of the cold start model after each iteration. The objective function is updated based on the change in expected return of the cold start model after each iteration; The step of iterating the control parameters at preset intervals based on the objective function and the strategy function to obtain the control parameters after each iteration includes: Based on the updated objective function and the strategy function, the control parameters are iterated at preset intervals to obtain the control parameters after each iteration.

6. The method according to claim 5, characterized in that, The formula for determining the change in expected return after each iteration is as follows: in, This represents the expected change in revenue at the (i+1)th preset time interval. This represents the change in expected revenue corresponding to the i-th preset time period. This represents the total number of users who flow into the cold start model within the i-th preset time period. Indicates the forgotten parameter, This represents the actual benefit generated by any user in the cold start model. This represents the expected revenue generated by any user flowing into the cold start model within the cold start model.

7. The method according to claim 1, characterized in that, The method is applied to a product recommendation system, wherein the products recommended by the product recommendation system include at least one of insurance products, video programs, online goods and advertisements, and the cold start products are newly added products to be recommended in the product recommendation system.

8. A starting device for a cold-start product, characterized in that, The device includes: The iteration module is used to iterate the control parameters at preset intervals based on a predefined reinforcement learning algorithm when a cold start product is detected, so as to obtain the control parameters after each iteration. The control parameters are used to determine the number of users flowing into the cold start model within the preset interval. The reinforcement learning algorithm is used to maximize the overall expected revenue generated by system users in the cold start model and the warm start model, as well as the expected revenue generated by users flowing into the cold start model in the cold start model. The update module is used to update the number of users flowing into the cold start model within each preset time period based on the control parameters after each iteration; The determination module is used to determine that the cold start product has been launched when the number of iterations reaches a preset number or the number of samples of the cold start product reaches a preset threshold. The number of samples of the cold start product is determined by the relevant operations of the users flowing into the cold start model on the cold start product. The iterative module includes: The acquisition submodule is used to acquire the predefined objective function and related parameters in the reinforcement learning algorithm, wherein the objective function is used to characterize the overall expected benefit generated by the system user in the cold start model and the warm start model, and the related parameters are used to characterize the state, reward and action required to participate in the computation process of the reinforcement learning algorithm; The iteration submodule is used to iterate the control parameters at preset intervals based on the objective function and the relevant parameters to obtain the control parameters after each iteration.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When a processor executes a program stored in a memory, it implements the steps of the startup method of the cold-start product according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the startup method of the cold start product according to any one of claims 1-7.

Citation Information

Patent Citations

  • Conversation management method and system of outbound system, electronic equipment and storage medium

    CN111104502A

  • Music recommendation method, device and equipment and storage medium

    CN111506762A