Recommended Model Training Method, Electronic Device, and Storage Medium

By dynamically adjusting the click weight and conversion weight in the recommendation model, the problem of difficulty in adjusting the weight of the loss function in the existing technology is solved, and the accuracy of predicting click rate and conversion rate is improved.

CN113868523BActive Publication Date: 2025-06-24TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111132414.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2025-06-24
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

When predicting click-through rate and conversion rate, it is difficult for existing recommended models to adjust the weight parameters of the loss function, resulting in poor training results and low prediction accuracy.

Method used

By obtaining the sample feature set, input the initial prediction model to obtain the predicted click-through rate and conversion rate, generate a click loss function and a conversion loss function, calculate the weight loss function based on these loss functions and model parameters, dynamically adjust the click weight and conversion weight, generate a target loss function to correct the model parameters, and obtain the target prediction model.

Benefits of technology

By dynamically adjusting the click weight and conversion weight, the balance between the two loss functions in the target loss function is maintained, and the training effect of the target prediction model is improved, making the prediction click rate and conversion rate more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113868523B_ABST
    Figure CN113868523B_ABST
Patent Text Reader

Abstract

Embodiments of this application disclose a method for training a recommendation model, an electronic device, and a storage medium, which are applied to the technical field of machine learning. The method includes: obtaining a sample feature set and inputting it into an initial prediction model to obtain the predicted click-through rate and predicted conversion rate of a sample user for a sample recommendation object, generating a click loss function based on a click label and the predicted click-through rate, generating a click conversion loss function based on the click label, conversion label, predicted click-through rate, and predicted conversion rate, obtaining a weight loss function according to the click loss function, click conversion loss function, and model parameters, generating a click weight and a click conversion weight based on the weight loss function, obtaining an objective loss function based on the click weight, click conversion weight, click loss function, and click conversion loss function, and correcting the model parameters based on the objective loss function to obtain an objective prediction model. By using the embodiments of this application, the prediction accuracy of the objective prediction model for the click-through rate and conversion rate can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of machine learning, and particularly relates to a method for training a recommendation model, an electronic device, and a storage medium. Background Art

[0002] Currently, in a recommendation scenario, it is necessary to predict the degree of interest of a user in a recommended object, so as to push the recommended objects that the user is interested in to the user terminal. The accuracy of the prediction is directly reflected in the click-through rate and conversion rate of the user for the recommended object. Correspondingly, accurate push in the recommendation scenario can be achieved by predicting the click-through rate and conversion rate. Existing prediction methods usually jointly model the click-through rate prediction task and the conversion rate prediction task, and train a target prediction model for predicting the click-through rate and conversion rate by combining the loss function of the click-through rate prediction task and the loss function of the conversion rate prediction task. However, since it is difficult to adjust the corresponding weight parameters when the two loss functions are combined, it will affect the training effect of the model, resulting in low prediction accuracy. Therefore, how to improve the prediction accuracy of the click-through rate and conversion rate has become an urgent problem to be solved. Summary of the Invention

[0003] Embodiments of the present application provide a method for training a recommendation model, an electronic device, and a storage medium, which can improve the prediction accuracy of a target prediction model for the click-through rate and conversion rate.

[0004] On the one hand, an embodiment of the present application provides a method for training a recommendation model, and the method includes:

[0005] Obtain a sample feature set; the sample feature set includes user attribute information of a sample user, object attribute information of a sample recommended object, a click label of the sample user for the sample recommended object, and a conversion label;

[0006] Input the sample feature set into an initial prediction model, and obtain a predicted click-through rate and a predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model;

[0007] Generate a click loss function based on the click label and the predicted click-through rate, and generate a click-conversion loss function based on the conversion label, the predicted click-through rate, and the predicted conversion rate;

[0008] Obtain a weight loss function according to the click loss function, the click-conversion loss function, and the model parameters of the initial prediction model, and generate a click weight corresponding to the click loss function and a click-conversion weight corresponding to the click-conversion loss function based on the weight loss function;

[0009] Obtain a target loss function based on the click weight, the click-conversion weight, the click loss function, and the click-conversion loss function, and correct the model parameters of the initial prediction model based on the target loss function to obtain a target prediction model.

[0010] On the one hand, an embodiment of the present application provides a data recommendation method based on a recommendation model, and the method includes:

[0011] Obtain the target user attribute information of the predicted user and the target object attribute information of the object to be recommended;

[0012] Input the target user attribute information and the target object attribute information into the target prediction model;

[0013] Generate the target predicted click-through rate and the target predicted conversion rate of the predicted user for the object to be recommended in the target prediction model;

[0014] Based on the target predicted click-through rate and the target predicted conversion rate, obtain the interest score of the predicted user for the object to be recommended;

[0015] If the interest score is greater than the interest score threshold, push the object to be recommended to the user terminal corresponding to the predicted user.

[0016] On the one hand, an embodiment of the present application provides a recommendation model training device, and the device includes:

[0017] An acquisition module, configured to acquire a sample feature set; the sample feature set includes the user attribute information of the sample user, the object attribute information of the sample recommended object, the click label and the conversion label of the sample user for the sample recommended object;

[0018] The acquisition module is further configured to input the sample feature set into an initial prediction model, and obtain the predicted click-through rate and the predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model;

[0019] A generation module, configured to generate a click loss function based on the click label and the predicted click-through rate, and generate a click conversion loss function based on the conversion label, the predicted click-through rate and the predicted conversion rate;

[0020] The generation module is further configured to obtain a weight loss function according to the click loss function, the click conversion loss function and the model parameters of the initial prediction model, and generate a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function based on the weight loss function;

[0021] A correction module, configured to obtain a target loss function based on the click weight, the click conversion weight, the click loss function and the click conversion loss function, and correct the model parameters of the initial prediction model based on the target loss function to obtain a target prediction model.

[0022] On the one hand, an embodiment of the present application provides a data recommendation device based on a recommendation model, and the device includes:

[0023] An acquisition module, configured to acquire target user attribute information of a predicted user and target object attribute information of a to-be-recommended object;

[0024] An input module, configured to input the target user attribute information and the target object attribute information into a target prediction model;

[0025] A generation module, configured to generate a target predicted click-through rate and a target predicted conversion rate of the predicted user for the to-be-recommended object in the target prediction model;

[0026] The acquisition module is further configured to acquire an interest score of the predicted user for the to-be-recommended object based on the target predicted click-through rate and the target predicted conversion rate;

[0027] A push module, configured to push the to-be-recommended object to a user terminal corresponding to the predicted user if the interest score is greater than an interest score threshold.

[0028] On the one hand, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute some or all of the steps in the above method.

[0029] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, they are used to execute some or all of the steps in the above method.

[0030] Correspondingly, according to an aspect of the present application, there is provided a computer program product or a computer program, which includes program instructions. The program instructions are stored in a computer-readable storage medium. A processor of a computer device reads the program instructions from the computer-readable storage medium, and the processor executes the program instructions, so that the computer device executes the above-provided recommendation model training method and / or the data recommendation method based on the recommendation model.

[0031] In the embodiments of the present application, a sample feature set can be obtained, the sample feature set is input into an initial prediction model, the predicted click-through rate and predicted conversion rate of the sample user for the sample recommended object are obtained based on the initial prediction model, a click loss function is generated based on the click label and the predicted click-through rate, a click conversion loss function is generated based on the conversion label, the predicted click-through rate, and the predicted conversion rate, a weight loss function is obtained according to the click loss function, the click conversion loss function, and the model parameters of the initial prediction model, a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function are generated based on the weight loss function, a target loss function is obtained based on the click weight, the click conversion weight, the click loss function, and the click conversion loss function, and the model parameters of the initial prediction model are corrected based on the target loss function to obtain a target prediction model. By implementing the method proposed above, the click weight and the click conversion weight can be dynamically adjusted during the model training process to maintain the balance between the two loss functions in the target loss function, thereby making the training effect of the target prediction model better, and further improving the accuracy of the target prediction model for the predicted click-through rate and conversion rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0033] Figure 1 It is a schematic diagram of an application architecture provided by an embodiment of the present application;

[0034] Figure 2 It is a schematic flowchart of a recommended model training method provided by an embodiment of the present application;

[0035] Figure 3 It is a schematic flowchart of a recommended model training method provided by an embodiment of the present application;

[0036] Figure 4 It is a schematic diagram of a scenario for obtaining attribute features provided by an embodiment of the present application;

[0037] Figure 5 It is a schematic diagram of a scenario for obtaining fused attribute features provided by an embodiment of the present application;

[0038] Figure 6 It is a schematic diagram of a scenario for model training provided by an embodiment of the present application;

[0039] Figure 7 It is a schematic flowchart of a data recommendation method based on a recommended model provided by an embodiment of the present application;

[0040] Figure 8a A schematic diagram of a push scenario for predicting users provided by an embodiment of the present application;

[0041] Figure 8b A schematic diagram of a push scenario for predicting users provided by an embodiment of the present application;

[0042] Figure 9 A schematic structural diagram of a recommendation model training device provided by an embodiment of the present application;

[0043] Figure 10 A schematic structural diagram of a data recommendation device based on a recommendation model provided by an embodiment of the present application;

[0044] Figure 11 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0046] The recommendation model training method proposed in the embodiments of the present application can be implemented on an electronic device, which can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto.

[0047] In some embodiments, please refer to Figure 1 , Figure 1 An application architecture schematic diagram provided by an embodiment of the present application, and the recommendation model training method proposed by the present application can be executed through this application architecture. As Figure 1 shown, Figure 1 it may include an electronic device, and an initial prediction model is deployed in the electronic device.

[0048] Among them, the initial prediction model may include a shared layer, a prediction layer, and a weight adjustment layer; specifically, (1) the shared layer may include a feature processing layer and a prediction feature generation layer. The feature processing layer can be used to process the sample feature set input by the electronic device to obtain fused attribute features. The prediction feature generation layer may include a click weight prediction network, a conversion weight prediction network, and N prediction feature generation networks. The electronic device can use the click weight prediction network, the conversion weight prediction network, and the N prediction feature generation networks in the prediction feature generation layer and obtain click prediction features and conversion prediction features based on the fused attribute features;

[0049] (2) The prediction layer may include a click prediction network and a conversion prediction network. The electronic device may use the click prediction network in this prediction layer and obtain a predicted click-through rate based on click prediction features, and use the conversion prediction network and obtain a predicted conversion rate based on conversion prediction features. Furthermore, a click loss function may be generated based on the obtained predicted click-through rate and the click label in the sample feature set, and a click conversion loss function may be generated based on the obtained predicted conversion rate and the conversion label in the sample feature set.

[0050] (3) The weight adjustment layer may be used to obtain a weight loss function by using the click loss function, the click conversion loss function, and the model parameters of the initial prediction model during the training process of the initial prediction model, and generate a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function based on the weight loss function. That is, the initial click weight corresponding to the click loss function (the click weight used during the previous model training) and the initial click conversion weight corresponding to the click conversion loss function (the click conversion weight used during the previous model training (such as the i-th time)) may be dynamically adjusted based on the weight loss function to obtain the click weight and the click conversion weight respectively. The click loss function is weighted by the click weight to obtain a weighted click loss function, and the click conversion loss function is weighted by the click conversion weight to obtain a weighted click conversion loss function. The weighted click loss function and the weighted click conversion loss function are summed to obtain the target loss function for this model training (such as the (i + 1)-th time). Subsequently, the model parameters of the initial prediction model may be corrected based on this target loss function to train a target prediction model. The electronic device may execute the data recommendation method based on the recommendation model proposed in this application through this target prediction model. Subsequently, the electronic device may use this target prediction model to perform prediction tasks in the recommendation scenario and perform recommendation tasks in the recommendation scenario based on the prediction results obtained from this prediction task. It can be understood that the target prediction model used in the application process includes the shared layer and the prediction layer part of the initial prediction model. In addition, the click-through rate is also called CTR (Click-through Rate), and the conversion rate is also called CVR (Conversion Rates).

[0051] It can be understood that Figure 1 it only exemplarily represents the possible application architecture of the technical solution of this application and does not limit the specific architecture of the technical solution of this application. That is, the technical solution of this application may also provide other forms of application architectures.

[0052] Optionally, in some embodiments, the electronic device may execute the recommended model training method according to actual service requirements to improve the prediction accuracy of click-through rate and conversion rate. The technical solution of this application can be applied to the prediction tasks in any recommendation scenario, that is, the electronic device can utilize the relevant information of the sample user and the sample recommended object (such as the user attribute information of the sample user, the object attribute information of the sample recommended object, the click label and conversion label of the sample user for the sample recommended object, etc.), and based on the model training method included in the technical solution of this application, train the initial prediction model. Subsequently, a data recommendation method based on the recommended model can be executed to obtain the target user attribute information of the predicted user and the target object attribute information of the object to be recommended, and use the trained target prediction model to predict the target user attribute information of the predicted user and the target object attribute information of the object to be recommended, so as to obtain the target predicted click-through rate and target predicted conversion rate of the predicted user for the object to be recommended. Furthermore, the interest score of the predicted user for the object to be recommended can be determined in combination with the target predicted click-through rate and target predicted conversion rate, and accurate push can be realized based on the interest score.

[0053] For example, it can be applied to the recommendation scenario of live products. At this time, the sample recommended object (object to be recommended) can be an online anchor, and the sample user (predicted user) can be a user who clicks on the online anchor in the live list interface. The electronic device can use the relevant information of the sample user and the sample recommended object to train the model, and subsequently use the prediction result of the target prediction model to determine the interest score of the predicted user for the online anchor, and push the online anchor to the user terminal of the predicted user based on the interest score. Another example is that it can also be applied to the recommendation scenario of news products. At this time, the sample recommended object (object to be recommended) can be the published news article, and the sample user (predicted user) can be a user who clicks on and reads the article in the news list interface. The electronic device can use the relevant information of the sample user and the sample recommended object to train the model, and subsequently use the prediction result of the target prediction model to determine the interest score of the predicted user for the news article, and push the news article to the user terminal of the predicted user based on the interest score. Another example is that it can also be applied to the recommendation scenario of e-commerce products. At this time, the sample recommended object can be the purchased commodity, and the sample user (predicted user) can be a user who clicks on the commodity in the commodity list interface. Subsequently, use the prediction result of the target prediction model to determine the interest score of the predicted user for the commodity, and push the commodity to the user terminal of the predicted user based on the interest score.

[0054] Optionally, the data involved in this application, such as predicted click-through rate and predicted conversion rate, etc., can be stored in a database, or can be stored in a blockchain, such as stored through a blockchain distributed system. This application does not make a limitation.

[0055] It can be understood that the above scenarios are only examples and do not constitute a limitation on the application scenarios of the technical solutions provided by the embodiments of the present application. The technical solutions of the present application can also be applied to other scenarios. For example, as is known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0056] Based on the above description, an embodiment of the present application proposes a method for training a recommendation model, and this method can be executed by the above-mentioned electronic device. Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for training a recommendation model provided by an embodiment of the present application. As Figure 2 shown, the process of the method for training a recommendation model according to an embodiment of the present application may include the following:

[0057] S201. Obtain a sample feature set.

[0058] Among them, the sample feature set includes user attribute information of the sample user, object attribute information of the sample recommendation object, click labels and conversion labels of the sample user for the sample recommendation object.

[0059] Optionally, the user attribute information of the sample user may have one or more attributes, which can be used to characterize user features. The object attribute information of the sample recommendation object may have one or more attributes, which can be used to characterize the object features of the recommendation object. And the user attribute information of the sample user and the object attribute information of the sample recommendation object can be set by relevant business personnel according to the actual application scenario, and there is no limitation here. For example, taking the recommendation scenario of a live broadcast product as an example, the user attribute information of the sample user may be basic portrait attributes (including age, gender, education level, etc.), statistical attributes (the duration of live broadcast watched by the sample user within one day or three days, the number of flowers sent by the sample user to the anchor within three days or one week, the number of rewards given by the sample user to the anchor within three days or one week, etc.), and so on; the object attribute information of the online anchor, that is, the sample recommendation object, may be basic portrait attributes (including age, gender, education level, etc.), real-time attributes (the total number of viewing users or the maximum number of viewing users in the live broadcast room within one day or three days, the type of the live broadcast room, the playing duration of the live broadcast room within one day, etc.), and so on.

[0060] Optionally, the click label and conversion label of the sample user for the sample recommended object are determined according to whether the sample user has a click behavior and a conversion behavior for the sample recommended object. That is, if the sample user has a click behavior on the sample recommended object but no conversion behavior, the corresponding click label is 1 and the conversion label is 0; if the sample user has a click behavior and a conversion behavior on the sample recommended object, the corresponding click label is 1 and the conversion label is 1; if the sample user has no click behavior on the sample recommended object, the corresponding click label is 0 and the conversion label is 0. For example, in the recommendation scenario of a live broadcast product, the click behavior refers to the user entering the live broadcast room of an online anchor, and the conversion behavior refers to the user watching for a certain duration (e.g., 30s) in the live broadcast room of the online live broadcast.

[0061] S202. Input the sample feature set into the initial prediction model, and obtain the predicted click-through rate and predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model.

[0062] In a possible implementation manner, the electronic device can input the sample feature set into the initial prediction model, and in the initial prediction model, use the user attribute information of the sample user in the sample feature set to generate user attribute features corresponding to the user attribute information, and use the object attribute information of the sample recommended object in the sample feature set to generate object attribute features corresponding to the object attribute information. Among them, the user attribute features and object attribute features can be generated by using the feature processing layer in the shared layer of the initial prediction model.

[0063] Optionally, taking the generation of user attribute features as an example, the specific way for the electronic device to generate user attribute features by using the feature processing layer in the shared layer of the initial prediction model can be to perform one-hot encoding on the user attribute information by using the feature processing layer to obtain user attribute features corresponding to the user attribute information. For example, the user attribute information includes an age attribute, and it is assumed that the age attribute is divided into [<18, 19 - 30, 31 - 40, 41 - 50, 51 - 60, >60]. If the age of the sample user is 24, the attribute features corresponding to the age attribute obtained by one-hot encoding can be represented as [0, 1, 0, 0, 0, 0]. Therefore, one or more attributes included in the user attribute information can generate corresponding user attribute features, and the user attribute features can be a feature matrix composed of one or more attribute features obtained from the attributes included in the user attribute information. Optionally, the specific way to generate object attribute features can be the same as the specific way to generate user attribute features, which will not be elaborated here.

[0064] In some embodiments, the electronic device obtains the predicted click-through rate and the predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model. Specifically, the feature processing layer in the shared layer of the initial prediction model is used to obtain the fused attribute features of the user attribute features and the object attribute features, and the predicted click-through rate and the predicted conversion rate of the sample user for the sample recommended object are obtained based on the fused attribute features. Among them, the specific process of the electronic device using the feature processing layer in the shared layer of the initial prediction model to obtain the fused attribute features of the user attribute features and the object attribute features is as follows: the user attribute features and the object attribute features are concatenated through the feature processing layer to obtain the fused attribute features, and the fused attribute features can be the feature vectors obtained by concatenating the feature matrix of the user attribute features and the feature matrix of the object attribute features. For example, the user attribute features are represented as [0,1,0,0,0,0], [0,0,0,1,0,0], and the object attribute features are represented as [0,0,0,1,0,0], [0,1,0,0,0,0], so the obtained fused attribute features can be represented as [0,1,0,0,0,0,0,0,0,1,0,0,0,0,0,1,0,0,0,1,0,0,0,0].

[0065] In a possible implementation manner, the specific process of the electronic device obtaining the predicted click-through rate and the predicted conversion rate of the sample user for the sample recommended object based on the fused attribute features is as follows: the click prediction features and the conversion prediction features are obtained by using the prediction feature generation layer in the shared layer of the initial prediction model and based on the fused attribute features, and the predicted click-through rate is obtained based on the click prediction features and the predicted conversion rate is obtained based on the conversion prediction features in the prediction layer of the initial prediction model.

[0066] Optionally, the prediction feature generation layer in the shared layer of the initial prediction model may include a click weight prediction network, a conversion weight prediction network, and N prediction feature generation networks; the specific process of the electronic device using the prediction feature generation layer in the shared layer of the initial prediction model and obtaining the click prediction features and the conversion prediction features based on the fused attribute features is as follows: the fused attribute features are input into the prediction feature generation layer, and the click prediction features are generated by using the click weight prediction network and the N prediction feature generation networks in the prediction feature generation layer, and the conversion prediction features are generated by using the conversion weight prediction network and the N prediction feature generation networks. Among them, the specific implementation process of generating the click prediction features and the conversion prediction features can refer to the relevant description in step S302 below.

[0067] In addition, the prediction layer in the initial prediction model may include a click prediction network constructed for predicting the click-through rate and a conversion prediction network constructed for predicting the conversion rate. When the electronic device obtains the predicted click-through rate based on the click prediction features and the predicted conversion rate based on the conversion prediction features in the prediction layer of the initial prediction model, specifically, the click prediction features are input into the click prediction network in the prediction layer to obtain the predicted click-through rate, and the conversion prediction features are input into the conversion prediction network in the prediction layer to obtain the predicted conversion rate. Among them, the specific implementation process of obtaining the predicted click-through rate and the predicted conversion rate can refer to the relevant description in step S302 below.

[0068] S203. Generate a click loss function based on the click label and the predicted click-through rate, and generate a click-conversion loss function based on the click label, the conversion label, the predicted click-through rate, and the predicted conversion rate.

[0069] In a possible implementation manner, the electronic device may use multiple sample feature sets as a sample data set, and determine the click loss function based on the predicted click-through rate obtained for each sample feature set in the sample data set, and determine the click-conversion loss function based on the predicted click-through rate and the predicted conversion rate obtained for each sample feature set in the sample data set. The electronic device may generate a click loss function for the initial prediction model based on the click label and the predicted click-through rate. The click loss function L ctr (t) may be:

[0070]

[0071] where N is the number of sample feature sets in the sample data set, x i(ctr) represents the click prediction features obtained according to the i-th sample feature set in the sample data set, y i represents the click label of the i-th sample feature set, θ ctr represents the network parameters of the click prediction network, l represents the cross-entropy loss function, f1(v1,u1) represents generating the corresponding predicted click-through rate according to the input u1 and v1, t may represent the t-th training process of the initial prediction model, that is, t may be the number of times or the moment at the t-th training (this moment is the moment compared with the 0-th model (assuming the indicated moment is 0) training). The following takes t as the number of training times as an example for illustration.

[0072] In addition, the electronic device may generate a click-conversion loss function for the initial prediction model based on the click label, the conversion label, the predicted click-through rate, and the predicted conversion rate. The click-conversion loss function L ctcvr (t) may be:

[0073]

[0074] where N is the number of sample feature sets in the sample data set, and x i(ctr) represents the click prediction feature obtained from the i-th sample feature set in the sample data set, and x i(cvr) represents the conversion prediction feature obtained from the i-th sample feature set in the sample data set, y i represents the click label of the i-th sample feature set, z i represents the conversion label of the i-th sample feature set, θ ctr represents the network parameters of the click prediction network, and θ cvr represents the network parameters of the conversion prediction network. l represents the cross-entropy loss function. f1(v1, u1) represents generating the corresponding predicted click-through rate according to the input u1 and v1. f2(v2, u2) represents generating the corresponding predicted conversion rate according to the input u2 and v2. t can represent the t-th training process of the initial prediction model, that is, t can be the number or moment at the t-th training (this moment is the moment compared to the 0-th model (assuming the indicated moment is 0) training).

[0075] In addition, during the training process, the training objective is a multi-task objective, that is, training the click-through rate prediction task and the conversion rate prediction task. Since the number of click labels with a value of 1 for the sample users' sample recommended objects is large, and the number of conversion labels with a value of 1 is small, there will be a difference between the click sample space and the conversion sample space, resulting in poor training effects. By introducing the click-conversion loss function, that is, CTCVR = CTR * CVR, the training objective of the model can be understood as training the click-through rate prediction task and the click-conversion rate prediction task, which can connect the click sample space and the conversion sample space and reduce the sample space difference, that is, it can solve the problem of sample selection bias in model training.

[0076] Optionally, both the network parameters of the above click prediction network and the network parameters of the conversion prediction network can be the model parameters in the initial prediction model.

[0077] S204. Obtain a weight loss function according to the click loss function, the click-conversion loss function, and the model parameters of the initial prediction model, and generate a click weight corresponding to the click loss function and a click-conversion weight corresponding to the click-conversion loss function based on the weight loss function.

[0078] In a possible implementation, the electronic device obtains a weight loss function based on a click loss function, a click conversion loss function, and the model parameters of an initial prediction model. Specifically, it can obtain a first weight function corresponding to the click loss function and a second weight function corresponding to the click conversion loss function, determine a click gradient function based on the first weight function, the click loss function, and the model parameters of the initial prediction model, determine a click conversion gradient function based on the second weight function, the click conversion loss function, and the model parameters of the initial prediction model, and generate a weight loss function based on the click gradient function and the click conversion gradient function.

[0079] Among them, the first weight function can be used to determine the click weight, and the second weight function can be used to determine the click conversion weight. The electronic device can determine a model loss function for training the initial prediction model based on the first weight function, the click loss function, the second weight function, and the click conversion loss function. The first weight function and the second weight function can be set by relevant business personnel according to empirical values.

[0080] For example, let the first weight function be W ctr (t), and the second weight function be W cvr (t). Therefore, the determined model loss function L task (t) can be expressed as:

[0081] L task (t) = W ctr (t)L ctr (t) + W cvr (t)L ctcvr (t)

[0082] Among them, t can represent the t-th training process of the initial prediction model; L ctr (t) represents the click loss function, and L ctcvr (t) represents the click conversion loss function.

[0083] In a possible implementation, the electronic device generates a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function based on the weight loss function. Specifically, based on the weight adjustment layer of the initial prediction model and using the weight loss function, it obtains a first weight adjustment function corresponding to the first weight function and a second weight adjustment function corresponding to the second weight function, obtains the click weight according to the first weight function and the first weight adjustment function, and obtains the click conversion weight according to the second weight function and the second weight adjustment function.

[0084] S205. Obtain a target loss function based on the click weight, the click conversion weight, the click loss function, and the click conversion loss function, and correct the model parameters of the initial prediction model based on the target loss function to obtain a target prediction model.

[0085] In a possible implementation, during one round of training for the initial prediction model, let the click weight be W ctr , and the click conversion weight be W ctcvr , the click loss function be L ctr , and the click conversion loss function be L ctcvr , so the obtained target loss function L task is:

[0086] L task = W ctr L ctr + W ctcvr L ctcvr

[0087] It can be understood that when obtaining the click weight according to the first weight function and the click conversion weight according to the second weight function, the electronic device can replace the first weight function with the click weight and the second weight function with the click conversion weight in the model loss function. At this time, the model loss function is the target loss function.

[0088] Therefore, the electronic device can correct the model parameters of the initial prediction model based on the target loss function at this time. The model obtained by correcting the model parameters in this round will participate in the next round of model training until the model converges, and thus the corresponding target prediction model can be obtained. Subsequently, the target prediction model can be applied to the recommendation scenario for the object to be recommended. For example, it can be specifically used to predict the target click-through rate and target conversion rate according to the target user attribute information of the predicted user and the target object attribute information of the object to be recommended (such as an online anchor or an information article). Furthermore, the electronic device can implement the recommendation task for the predicted user and the object to be recommended based on the predicted target click-through rate and target conversion rate of the predicted user for the object to be recommended. For example, the recommendation task can be that the electronic device determines the object to be pushed based on the predicted target click-through rate and target conversion rate and pushes the object to be pushed to the user terminal of the predicted user. The predicted user can click to view the relevant information of the pushed object to be recommended (such as the live broadcast room corresponding to the online anchor or the detailed content corresponding to the information article).

[0089] In the embodiments of the present application, a sample feature set can be obtained, the sample feature set is input into an initial prediction model, and based on the initial prediction model, the predicted click-through rate and predicted conversion rate of the sample user for the sample recommended object are obtained. A click loss function for the initial prediction model is generated based on the click label and the predicted click-through rate, and a click conversion loss function for the initial prediction model is generated based on the click label, conversion label, predicted click-through rate, and predicted conversion rate. A weight loss function is obtained according to the click loss function, click conversion loss function, and model parameters of the initial prediction model. A click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function are generated based on the weight loss function. A target loss function is obtained based on the click weight, click conversion weight, click loss function, and click conversion loss function, and the model parameters of the initial prediction model are corrected based on the target loss function to obtain a target prediction model. By implementing the method proposed in the embodiments of the present application, the click weight and click conversion weight can be dynamically adjusted during the model training process to maintain the balance between the two loss functions in the target loss function, thereby making the training effect of the target prediction model better, and further improving the accuracy of the target prediction model for the predicted click-through rate and conversion rate.

[0090] Please refer to Figure 3 , Figure 3 FIG. is a schematic flowchart of a recommendation model training method provided by an embodiment of the present application, and this method can be executed by the above-mentioned electronic device. As Figure 3 shown, the process of the recommendation model training method in the embodiments of the present application may include the following:

[0091] S301. Obtain a sample feature set. Among them, for the specific implementation manner of step S301, reference can be made to the relevant description of step S201 above, and details are not described herein again.

[0092] S302. Input the sample feature set into the initial prediction model, and based on the initial prediction model, obtain the predicted click-through rate and predicted conversion rate of the sample user for the sample recommended object.

[0093] In a possible implementation manner, the electronic device can use the feature processing layer in the shared layer of the initial prediction model to generate user attribute features corresponding to user attribute information and object attribute features corresponding to object attribute information. Among them, the feature processing layer may include an attribute feature generation layer.

[0094] Optionally, the electronic device generates user attribute features corresponding to user attribute information and object attribute features corresponding to object attribute information through the feature processing layer. Specifically, in the attribute feature generation layer, the electronic device generates user attribute features using user attribute information and generates object attribute features using object attribute information. Taking the generation of user attribute features as an example, the electronic device generates user attribute features using user attribute information in the attribute feature generation layer. Specifically, the electronic device performs one-hot encoding on the user attribute information to obtain user attribute codes corresponding to the user attribute information, constructs an embedding matrix for the user attribute, and determines the user attribute features using the user attribute codes and the corresponding embedding matrix.

[0095] In some embodiments, the specific method of performing one-hot encoding on the user attribute information to obtain user attribute codes corresponding to the user attribute information can refer to the relevant description of step S202 above; the specific method of determining the user attribute features using the user attribute codes and the corresponding embedding matrix is as follows: obtain the column numbers where the elements with a value of 1 are located in the user attribute codes, and obtain the corresponding row vectors from the embedding matrix according to the column numbers. The row number of the row vector in the embedding matrix is the same as the column number, and use the row vector as the user attribute feature. For example, the user attribute information includes an age attribute. Suppose the age attribute is divided into [<18, 19 - 30, 31 - 40, 41 - 50, 51 - 60, >60]. If the age of the sample user is 24, the attribute feature corresponding to the age attribute obtained by one-hot encoding can be expressed as [0, 1, 0, 0, 0, 0]. Therefore, the column number where the element with a value of 1 is located is 2, and the row vector with a row number of 2 is obtained from the embedding matrix. Suppose the embedding matrix is 6*8, and the 8 elements included in the row vector are used as the attribute feature corresponding to the age attribute. Therefore, one or more row vectors can be obtained from one or more attributes included in the user attribute information, and a feature matrix can be formed based on the one or more row vectors. This feature matrix is the user attribute feature. Among them, the embedding matrices corresponding to multiple attributes can be the same or different, and the size of the embedding matrices corresponding to multiple attributes and the specific values of each element in the embedding matrix are not limited.

[0096] Optionally, the specific method of generating object attribute features can be the same as the specific method of generating user attribute features, which will not be elaborated here. In addition, the embedding matrix for user attributes and the embedding matrix for object attributes described above can be set by relevant business personnel according to empirical values, or can be obtained through model training as model parameters of the initial prediction model.

[0097] For example, as Figure 4 shown, Figure 4A schematic diagram of a scenario for obtaining attribute features provided by an embodiment of this application. Among them, the user attribute information includes user attribute 1, and the attribute code obtained by performing one-hot encoding on user attribute 1 is [0, 1, 0, 0, 0, 0]. Therefore, the row vector of the second row is obtained from the corresponding embedding matrix 1 as the attribute feature (V1) corresponding to user attribute 1. Based on multiple user attributes (1, 2,..., n) included in the user attribute information, the corresponding attribute features (V1, V2,..., Vn) are obtained. This user attribute feature is the feature matrix 1 composed of the attribute features (V1, V2,..., Vn). Also, based on multiple object attributes (n + 1, n + 2,..., N) included in the object attribute information, the corresponding attribute features (Vn+1, Vn+2,..., V N ) are obtained. This object attribute feature is the feature matrix 2 composed of the attribute features (Vn+1, Vn+2,..., V N ). The target feature matrix can be composed of this feature matrix 1 and feature matrix 2. Optionally, the target feature matrix can be in two forms, and the two forms of the target feature matrix can be as shown in Figure 4 's target feature matrix 1 and target feature matrix 2. Among them, Figure 4 's target feature matrix 1 represents that the same attribute feature is represented by columns, Figure 4 's target feature matrix 2 represents that the same attribute feature is represented by rows.

[0098] In a possible implementation manner, the electronic device obtains the predicted click-through rate and predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model. Specifically, it obtains the user attribute feature and the object attribute feature, and obtains the predicted click-through rate and predicted conversion rate based on the user attribute feature and the object attribute feature. Among them, when the electronic device obtains the predicted click-through rate and predicted conversion rate based on the user attribute feature and the object attribute feature, specifically, it performs feature fusion on the user attribute feature and the object attribute feature to obtain the fused attribute feature; and obtains the predicted click-through rate and predicted conversion rate based on the fused attribute feature.

[0099] In some embodiments, when the electronic device performs feature fusion on the user attribute feature and the object attribute feature to obtain the fused attribute feature, specifically, it can splice the user attribute feature and the object attribute feature to obtain the fused attribute feature. This fused attribute feature can be the feature vector spliced from the feature matrix of the user attribute feature and the feature matrix of the object attribute feature; or, the feature processing layer can also include a feature fusion layer, and the feature fusion layer in the feature processing layer is used to perform feature fusion on the user attribute feature and the object attribute feature to obtain the fused attribute feature.

[0100] Optionally, the electronic device uses the feature fusion layer in the feature processing layer to perform feature fusion on the user attribute features and the object attribute features to obtain the fused attribute features. Specifically, in the feature fusion layer, the user attribute features and the object attribute features are used as a feature set, and feature crossing (i.e., inner product of two vectors) is performed on every two features in the feature set to obtain cross features. The user attribute features and the object attribute features are concatenated to obtain concatenated features, and the fused attribute features are obtained based on the cross features and the concatenated features. Among them, obtaining the fused attribute features based on the cross features and the concatenated features can be to tile each element of the cross features and the concatenated features in sequence to form a target vector, and use this target vector as the fused attribute features. In addition, performing feature crossing on every two features references the idea of the Factorization Machine (FM), that is:

[0101] <V i ,V j >1≤i≤N,1≤j≤N

[0102] Among them, <> represents the inner product of vectors, V i ,V j represents any two features in the feature set, and N represents the number of features in the feature set.

[0103] It can be understood that the fused attribute features contain the correlation information of the features among the user attribute features themselves, among the object attribute features themselves, and between the user attribute features and the object attribute features, reflecting the relationship between features obtained through explicit interaction among multiple features. Furthermore, the relationship between every two features can be combined in the fused attribute features, avoiding the problem that will have a certain negative impact on subsequent model training when the features are relatively sparse.

[0104] Further optionally, after concatenating the user attribute features and the object attribute features to obtain the concatenated features, the first target attribute feature of the user attribute features and / or the second target attribute feature of the object attribute features are obtained, and the first target attribute feature and / or the second target attribute feature are multiplied using the user attribute information in the concatenated features to obtain the processed concatenated features. Then, the fused attribute features are obtained based on the cross features and the processed concatenated features. Among them, the first target attribute feature and the second target attribute feature can be set by relevant business personnel according to the actual situation and empirical values.

[0105] For example, as Figure 5 shown, Figure 5 is a schematic diagram of a scenario for obtaining fused attribute features provided by an embodiment of the present application. Among them, for a feature set containing user attribute features (V1, V2,..., Vn) and object attribute features (Vn+1, Vn+2,..., V N)Perform feature cross of pairwise features to obtain cross features, splice user attribute features and object attribute features to obtain spliced features, and obtain fused attribute features (such as fused attribute feature 1) based on the cross features and the spliced features. Further, after obtaining the spliced features, set the first target attribute feature as V1 (set as [1, 2, 3, 4, 5, 6]), and this first target attribute feature corresponds to the average click-through rate attribute of the sample user (assuming the attribute information is represented as 0.3). Therefore, in the spliced features, use the attribute information corresponding to the first target attribute feature to perform a multiplication process on the first target attribute feature to obtain the processed spliced feature (i.e., V 1: [1, 2, 3, 4, 5, 6] * 0.3 → V′ 1: [0.3, 0.6, 0.9, 1.2, 1.5, 1.8]). Obtain fused attribute features (such as fused attribute feature 2) based on the cross features and the processed spliced features.

[0106] In a possible implementation manner, the electronic device obtaining the predicted click-through rate and the predicted conversion rate based on the fused attribute features can specifically be generating a click prediction feature and a conversion prediction feature based on the fused attribute features, and obtaining the predicted click-through rate based on the click prediction feature and obtaining the predicted conversion rate based on the conversion prediction feature. Among them, the electronic device generating the click prediction feature and the conversion prediction feature based on the fused attribute features can specifically be using a prediction feature generation layer and generating the click prediction feature and the conversion prediction feature based on the fused attribute features.

[0107] In some embodiments, the prediction feature generation layer can be a shared layer structure in a multi-task learning model (such as an MMoE (Multi-gate Mixture-of-Experts) model), and the prediction feature generation layer can include a click weight prediction network, a conversion weight prediction network, and N prediction feature generation networks; the electronic device generating the click prediction feature and the conversion prediction feature based on the fused attribute features and using the prediction feature generation layer can specifically be generating an initial prediction feature corresponding to each prediction feature generation network based on the fused attribute features and the N prediction feature generation networks, predicting a first prediction weight for each prediction feature generation network based on the fused attribute features and the click weight prediction network, predicting a second prediction weight for each prediction feature generation network based on the fused attribute features and the conversion weight prediction network, using the first prediction weight corresponding to each prediction feature generation network to perform a weighted sum on the initial prediction feature corresponding to each prediction feature generation network to obtain the click prediction feature, and using the second prediction weight corresponding to each prediction feature generation network to perform a weighted sum on the initial prediction feature corresponding to each prediction feature generation network to obtain the conversion prediction feature.

[0108] Among them, the electronic device generates a network based on the fusion attribute features and N prediction feature generation networks, and generates the initial prediction features corresponding to each prediction feature generation network. Specifically, the fusion attribute features are input into the N prediction feature generation networks, and the corresponding initial prediction features are generated according to the fusion attribute features in each prediction feature generation network respectively; and the electronic device predicts the first prediction weight for each prediction feature generation network based on the fusion attribute features and the click weight prediction network, and predicts the second prediction weight for each prediction feature generation network based on the fusion attribute features and the conversion weight prediction network. Specifically, the fusion attribute features are input into the click weight prediction network, and the first prediction weight corresponding to each prediction feature generation network is generated in the click weight prediction network. The fusion attribute features are input into the conversion weight prediction network, and the second prediction weight corresponding to each prediction feature generation network is generated in the conversion weight prediction network.

[0109] According to the above description, it can be expressed as:

[0110] The first prediction weight g ctr (x) is: g ctr (x) = softmax(W ctr x)

[0111] The second prediction weight g cvr (x) is: g cvr (x) = softmax(W cvr x)

[0112] The click prediction feature f ctr (x) is:

[0113] The conversion prediction feature f cvr (x) is:

[0114] Among them, W ctr represents the network parameters of the click weight prediction network, W cvr represents the network parameters of the conversion weight prediction network, x represents the fusion attribute features, g ctr (x) i represents the first prediction weight corresponding to the i-th prediction feature generation network among the N prediction feature generation networks, g cvr (x) i represents the second prediction weight corresponding to the i-th prediction feature generation network among the N prediction feature generation networks; f i (x) represents the initial prediction feature corresponding to the i-th prediction feature generation network among the N prediction feature generation networks.

[0115] Optionally, the above click weight prediction network, conversion weight prediction network, and N prediction feature generation networks may all be composed of one or more fully connected layers (FC). The click weight prediction network and the conversion weight prediction network may also be referred to as gate networks, and the N prediction feature generation networks may also be referred to as expert networks. Moreover, the model parameters of the initial prediction model include the network parameters of the click weight prediction network, the conversion weight prediction network, and the N prediction feature generation networks. The value of N can be set by relevant business personnel according to empirical values.

[0116] In a possible implementation manner, for the electronic device to obtain the predicted click-through rate based on the click prediction features and obtain the predicted conversion rate based on the conversion prediction features, specifically, it can utilize the prediction layer in the initial prediction model and obtain the predicted click-through rate based on the click prediction features and obtain the predicted conversion rate based on the conversion prediction features. Among them, the prediction layer may be an ESMM model (Entire Space Multi-Task Model). This prediction layer may include a click prediction network constructed for the click-through rate prediction task (CTR Tower) and a conversion prediction network constructed for the conversion rate prediction task (CVR Tower). Therefore, the predicted click-through rate can be obtained by using the click prediction network and based on the click prediction features, and the conversion click-through rate can be obtained by using the conversion prediction network and based on the conversion prediction features. Among them, both the click prediction network and the conversion prediction network may include one or more fully connected layers.

[0117] S303. Generate a click loss function based on the click label and the predicted click-through rate, and generate a click-conversion loss function based on the click label, the conversion label, the predicted click-through rate, and the predicted conversion rate.

[0118] S304. Obtain a first weight function corresponding to the click loss function and a second weight function corresponding to the click-conversion loss function. The specific implementation manners of steps S303 - S304 can refer to the relevant descriptions of steps S203 - S204 above, and will not be elaborated here.

[0119] S305. Determine a click gradient function according to the first weight function, the click loss function, and the model parameters of the initial prediction model, and determine a click-conversion gradient function according to the second weight function, the click-conversion loss function, and the model parameters of the initial prediction model.

[0120] In a possible implementation, the initial prediction model further includes a weight adjustment layer. The electronic device can utilize the weight adjustment layer and determine a click gradient function based on a first weight function, a click loss function, and the model parameters of the initial prediction model, and determine a click conversion gradient function based on a second weight function, a click conversion loss function, and the model parameters of the initial prediction model, and determine a weight loss function, a click weight, and a click conversion weight in the weight adjustment layer.

[0121] In some embodiments, the electronic device can utilize the weight adjustment layer and determine a click gradient function based on a first weight function, a click loss function, and the model parameters of the initial prediction model, and determine a click conversion gradient function based on a second weight function, a click conversion loss function, and the model parameters of the initial prediction model. Specifically, it can determine the click gradient function by using the network parameters of the shared layer in the first weight function, the click loss function, and the model parameters, and determine the click conversion gradient function by using the second weight function, the click conversion loss function, and the network parameters of the shared layer in the model parameters. That is:

[0122]

[0123] Wherein, G ctr W represents the click gradient function, G ctcvr W represents the click conversion gradient function, θ share represents the network parameters of the shared layer; t can represent the t-th training process of the initial prediction model; L ctr (t) represents the click loss function, L ctcvr (t) represents the click conversion loss function; W ctr (t) represents the first weight function; W ctcvr (t) represents the second weight function; represents the calculated gradient of A with respect to B.

[0124] Optionally, the network parameters of the shared layer here can refer to the network parameters in N prediction feature generation networks, or other network parameters in the shared layer, and there is no limitation here.

[0125] It can be understood that the weight adjustment layer realizes gradient normalization to balance the gradients (or weights) of the two tasks in model training (Gradient Normalization, GradNorm). Therefore, the click weight of the click loss function and the click conversion weight of the click conversion loss function can be dynamically adjusted during model training through this weight adjustment layer, so as to achieve the coefficient balance of the two loss functions in the target loss function, and then obtain better model training results. The click weight is the coefficient of the click loss function in the target loss function, and the click conversion weight is the coefficient of the click conversion loss function in the target loss function.

[0126] S306. Generate a weight loss function according to the click gradient function and the click conversion gradient function.

[0127] In a possible implementation, the electronic device can generate a weight loss function in the weight adjustment layer according to the click gradient function and the click conversion gradient function. Specifically, according to the first weight function, the second weight function, the click loss function, the click conversion loss function, and the model parameters, determine the average gradient function, determine the target click gradient function according to the click gradient function and the average gradient function, determine the target click conversion gradient function according to the click conversion gradient function and the average gradient function, and generate a weight loss function according to the target click gradient function and the target click conversion gradient function. Among them, the model parameters used in the average gradient function can be the network parameters of the click prediction network and the network parameters of the conversion prediction network.

[0128] That is, determine the average gradient Specifically, it can be:

[0129]

[0130] Among them, represents the average gradient function, θ ctr represents the network parameters of the click prediction network, θ cvr represents the network parameters of the conversion prediction network; G 1 W (t) represents the first gradient function determined according to the first weight function, the click loss function, and the network parameters of the click prediction network; G 2 W (t) represents the second gradient function determined according to the second weight function, the click conversion loss function, and the network parameters of the conversion prediction network. Therefore, the average gradient function is obtained by taking the average of the first gradient function and the second gradient function; E[A + B] represents the mean result (average result) of taking the mean of the input A and B; t can represent the t-th training process of the initial prediction model; L ctr (t) represents the click loss function, Lctcvr (t) represents the click conversion loss function; W ctr (t) represents the first weight function; W ctcvr (t) represents the second weight function; W ctr (t) represents the first weight function; W ctcvr (t) represents the second weight function; represents the calculated gradient of A with respect to B.

[0131] In a possible implementation, the process and principle of the electronic device determining the target click gradient function and the target click conversion gradient function are the same. Here, taking the determination of the target click gradient function as an example for illustration. The electronic device determines the target click gradient function according to the click gradient function and the average gradient function. Specifically, according to the click loss function and the click conversion loss function, the target inverse training rate for model training is determined, and according to the click gradient function, the average gradient function, and the target inverse training rate, the target click gradient function is determined.

[0132] Among them, the electronic device determines the target inverse training rate for model training according to the click loss function and the click conversion loss function. Specifically, the initial click loss function and the initial click conversion loss function are obtained. According to the initial click loss function and the click loss function, the inverse training rate for the click-through rate prediction task in model training is determined, and according to the initial click conversion loss function and the click conversion loss function, the inverse training rate for the conversion rate prediction task in model training is determined. The average inverse training rate is determined according to the inverse training rate for the click-through rate prediction task and the inverse training rate for the conversion rate prediction task. The relative inverse training rate for the click-through rate prediction task is determined according to the inverse training rate for the click-through rate prediction task and the average inverse training rate. The relative inverse training rate for the click-through rate prediction task is used as the target inverse training rate. Optionally, when determining the target click conversion gradient function, the average inverse training rate used is the same as that when determining the target click gradient function. That is:

[0133]

[0134] Among them, represents the inverse training rate for the click-through rate prediction task, represents the inverse training rate for the conversion rate prediction task, L ctr (0) represents the initial click loss function, that is, the initial click loss function during model training, L ctcvr (0) represents the initial click conversion loss function, that is, the initial click conversion loss function during model training, represents the average inverse training quantity, r ctr(t) represents the relative inverse training rate for the click-through rate prediction task (i.e., the target inverse training rate when determining the target click gradient function), r cvr (t) represents the relative inverse training rate for the conversion rate prediction task (i.e., the target inverse training rate when determining the target click conversion gradient function), Grad ctr (t) represents the target click gradient function, Grad ctcvr (t) represents the target click conversion gradient function. t can represent the t-th training process of the initial prediction model, and α is a hyperparameter that can be set by relevant business personnel according to empirical values.

[0135] In some embodiments, according to the target click gradient function Grad ctr (t) and the target click conversion gradient function Grad ctcvr (t), the weight loss function Grad(t) generated can be expressed as:

[0136] Grad(t) = Grad ctr (t) + Grad ctcvr (t)

[0137] S307. Generate the click weight corresponding to the click loss function and the click conversion weight corresponding to the click conversion loss function based on the weight loss function.

[0138] In a possible implementation manner, the electronic device can adjust the first weight function based on the weight loss function to obtain the adjusted first weight function, and use the adjusted first weight function to obtain the click weight (i.e., the adjusted click weight), and adjust the second weight function based on the weight loss function to obtain the adjusted second weight function, and use the adjusted second weight function to obtain the click conversion weight (i.e., the adjusted click conversion weight).

[0139] In some embodiments, the electronic device adjusts the first weight function based on the weight loss function to obtain the adjusted first weight function. Specifically, it can be to take the derivative of the first weight function based on the weight loss function to obtain the first weight adjustment function, and then obtain the adjusted first weight function according to the first weight adjustment function and the first weight function. The electronic device can determine the click weight for this round of model training according to the adjusted first weight function, that is, the first weight adjustment function and the first weight function. And adjust the second weight function based on the weight loss function to obtain the adjusted second weight function. Specifically, it can be to take the derivative of the second weight function based on the weight loss function to obtain the second weight adjustment function, and then obtain the adjusted second weight function according to the second weight adjustment function and the second weight function. The electronic device can determine the click conversion weight for this round of model training according to the adjusted second weight function, that is, the second weight adjustment function and the second weight function. That is, the adjusted first weight function W′ ctr (t) can be:

[0140] W′ ctr (t) = W ctr (t) + λ ctr β ctr (t)

[0141] Wherein, W ctr (t) represents the first weight function, β ctr (t) represents the first weight adjustment function, t can represent the t-th training process of the initial prediction model, and λ ctr represents a hyperparameter, which can be set by relevant business personnel according to empirical values.

[0142] And, the adjusted second weight function W′ ctcvr (t) can be:

[0143] W′ ctcvr (t) = W ctcvr (t) + λ ctcvr β ctcvr (t)

[0144] Wherein, W ctcvr (t) represents the second weight function, β ctcvr (t) represents the second weight adjustment function, t can represent the t-th training process of the initial prediction model, and λ ctcvr represents a hyperparameter, which can be set by relevant business personnel according to empirical values, and λ ctr and λ ctcvr can be the same or different.

[0145] It can be understood that the adjusted first weight function and the adjusted second weight function in this round of model training serve as the first weight function and the second weight function in the next round of model training, and continue to adjust the first weight function and the second weight function in the next round to obtain the adjusted first weight function and the adjusted second weight function in the next round, thereby obtaining the click weight (i.e., the adjusted click weight in the next round) and the click conversion weight (i.e., the adjusted click conversion weight in the next round) in the next round, and determining the target loss function in the next round of model training based on the obtained click weight and click conversion weight in the next round, and using the target loss function in the next round to correct the model parameters in the next round until the model converges after multiple rounds of model training to obtain the final target prediction model.

[0146] S308. Obtain the target loss function based on the click weight, click conversion weight, click loss function, and click conversion loss function, and correct the model parameters of the initial prediction model based on the target loss function to obtain the target prediction model.

[0147] In a possible implementation manner, for the electronic device to obtain the target loss function based on the click weight, click conversion weight, click loss function, and click conversion loss function, specifically, it may be to weight the click loss function according to the click weight to obtain the first weighted loss function, weight the click conversion loss function according to the click conversion weight to obtain the second weighted loss function, and generate the target loss function based on the first weighted loss function and the second weighted loss function. The electronic device can correct the model parameters of the initial prediction model based on the target loss function to obtain the target prediction model. That is, only one round of model training is taken as an example here for illustration, and the process and principle of each round of model training are the same. Until the model converges to obtain the target prediction model. It can be understood that in each round of model training process, not only will the model parameters be corrected based on the target loss function, but the target loss functions used for model training in each round can be the same or different. That is, when obtaining the target loss function in each round, the click weight (initial click weight) and click conversion weight (initial click conversion weight) in the target loss function used in the previous model training will also be dynamically adjusted in combination with the determined weight loss function to obtain the click weight and click conversion weight in the target loss function used in this model training.

[0148] For example, please refer to Figure 6 , Figure 6A schematic diagram of a model training scenario provided by an embodiment of the present application. Among them, the initial prediction model includes a shared layer, a prediction layer, and a weight adjustment layer. The shared layer includes a feature processing layer and a prediction feature generation layer. The feature processing layer includes an attribute feature generation layer and a feature fusion layer. (1) The electronic device inputs the sample feature set into the initial prediction model, and generates user attribute features and object attribute features in the attribute feature generation layer; (2) generates fused attribute features according to the user attribute features and the object attribute features in the feature fusion layer; (3) generates click prediction features and conversion prediction features according to the fused attribute features in the prediction feature generation layer; (4) in the prediction layer, uses the click prediction network and generates a predicted click-through rate according to the click prediction features, and uses the conversion prediction network and generates a predicted conversion rate according to the conversion prediction features, and obtains a click loss function according to the predicted click-through rate, and obtains a click conversion loss function according to the predicted click-through rate and the predicted conversion rate; (5) in the weight adjustment layer, obtains a click weight and a click conversion weight according to the click loss function, the click conversion loss function, and the model parameters of the initial prediction model, thereby obtaining an objective loss function, and uses the objective loss function to correct the initial prediction model to obtain a target prediction model for predicting the click-through rate and the conversion rate.

[0149] After a large number of tests on the target prediction model, it is found that the prediction accuracy and efficiency for the click-through rate and the conversion rate have been greatly improved compared with the prior art. That is, taking the recommendation scenario of a live broadcast product as an example, for the recommendation task based on the target prediction model, the test results show that both the click-through rate and the conversion rate (i.e., the effective viewing duration) have been improved. See Table 1 below:

[0150] Dataset Exposure Click Conversion (Effective View) Training Set 42.48 million 14.51 million 1.48 million Test Set 4.35 million 1.56 million 0.16 million

[0151] Table 1

[0152] Among them, the target prediction model is obtained by training the initial prediction model using the training set as the sample feature set, and the target prediction model is tested using the test set. It is known that when the exposure volume (push for online anchors) is about 4.35 million times, the click volume is about 1.56 million times, and the conversion volume is about 160,000 times.

[0153] And by comparing multiple prediction models, it is found that the model evaluation index AUC (area under the curve) of the target prediction model proposed by the technical solution of the present application has also been greatly improved compared with the existing prediction models. See Table 2 below:

[0154] Model CTR AUC CVR AUC ESMM 0.7096 0.7427 ESMM + GradNorm 0.7094 0.7429 FM + ESMM 0.7099 0.7432 FM + MMOE + ESMM 0.7248 0.7447 FM + ESMM + GradNorm 0.7102 0.7433 FM + MMOE + ESMM + GradNorm 0.7473 0.7487

[0155] Table 2

[0156] Among them, CTR AUC represents the AUC corresponding to the predicted click-through rate, and CVR AUC represents the AUC corresponding to the predicted conversion rate. The larger the CTR AUC and CVR AUC are, the better the model's performance in predicting the click-through rate and conversion rate, that is, the higher the accuracy. From the above description, it can be known that the technical solution of this application uses the last model structure (FM+MMOE+ESMM+GradNorm) to train the target prediction model. Using this target prediction model, the CTR AUC can reach 0.7473, which is the highest compared with other models. Using this target prediction model, the CVR AUC can reach 0.7487, which is also the highest compared with other models.

[0157] In the embodiments of this application, a sample feature set can be obtained, the sample feature set is input into the initial prediction model, the predicted click-through rate and predicted conversion rate of the sample user for the sample recommendation object are obtained based on the initial prediction model, a click loss function is generated based on the click label and the predicted click-through rate, a click conversion loss function is generated based on the click label, conversion label, predicted click-through rate, and predicted conversion rate, the first weight function corresponding to the click loss function and the second weight function corresponding to the click conversion loss function are obtained, and the click gradient function is determined according to the first weight function, click loss function, and model parameters of the initial prediction model, and the click conversion gradient function is determined according to the second weight function, click conversion loss function, and model parameters of the initial prediction model. A weight loss function is generated according to the click gradient function and the click conversion gradient function, the click weight corresponding to the click loss function and the click conversion weight corresponding to the click conversion loss function are generated based on the weight loss function, the target loss function is obtained based on the click weight, click conversion weight, click loss function, and click conversion loss function, and the model parameters of the initial prediction model are corrected based on the target loss function to obtain the target prediction model. By implementing the method proposed in the embodiments of this application, the model training process can use gradients to dynamically adjust the click weight and click conversion weight to maintain the balance of the two loss functions in the target loss function, thereby making the training effect of the target prediction model better, and further improving the accuracy of the target prediction model for predicting the click-through rate and conversion rate.

[0158] Please refer to Figure 7 , Figure 7 is a schematic flowchart of a data recommendation method based on a recommendation model provided by an embodiment of this application. This method can be executed by the electronic device mentioned above. As Figure 7 shown, the process of the data recommendation method based on the recommendation model in the embodiments of this application can include the following:

[0159] S701. Obtain the target user attribute information of the predicted user and the target object attribute information of the object to be recommended.

[0160] In a possible implementation, the trained target prediction model can predict the click-through rate and conversion rate of the user for the object to be recommended, so as to achieve precise push of the object to be recommended for the predicted user. Taking the recommendation scenario of a live broadcast product as an example, the predicted user is the user who logs in to the application client of the live broadcast product, and the object to be recommended is the online anchor. When the predicted user clicks on the live broadcast interface, the electronic device obtains one or more current online anchors, and obtains the target user attribute information of the predicted user and the target object attribute information of each object to be recommended; the electronic device can be referred to as the background device of the application client.

[0161] S702. Input the target user attribute information and the target object attribute information into the target prediction model.

[0162] In a possible implementation, the electronic device inputs the target user attribute information of the predicted user and the target object attribute information of the object to be recommended into the target prediction model. When there are multiple objects to be recommended, the target user attribute information of the predicted user is respectively input into the target prediction model together with the target object attribute information of each object to be recommended.

[0163] Among them, the target prediction model can be trained through the relevant descriptions in the above Figure 2 illustrated embodiments and / or Figure 3 illustrated embodiments.

[0164] S703. Generate the target predicted click-through rate and target predicted conversion rate of the predicted user for the object to be recommended in the target prediction model.

[0165] In some embodiments, the electronic device can respectively generate the target predicted click-through rate and target predicted conversion rate of the predicted user for each object to be recommended in the target prediction model. It can be understood that the target prediction model includes a shared layer and a prediction layer, and the shared layer includes a feature processing layer and a predicted feature generation layer, and the feature processing layer includes an attribute feature generation layer and a feature fusion layer.

[0166] S704. Obtain the interest score of the predicted user for the object to be recommended based on the target predicted click-through rate and the target predicted conversion rate.

[0167] In some embodiments, based on the target predicted click-through rate (P ctr ) and the target predicted conversion rate (P cvr ), the interest score (P) of the predicted user for the object to be recommended can be obtained through the following formula:

[0168] P = P ctr * P cvr a

[0169] Among them, a is a hyperparameter that can be set by relevant business personnel according to empirical values.

[0170] For example, a can be set to [0.0, 0.25, 0.5, 0.75, 1], and the percentage of the current group's traffic in the total traffic is used. Each group of traffic is set to 1% of the traffic. By testing with multiple prediction models and the target prediction model of the present application, it is known that the target prediction model of the present application has the best test effect. That is, taking the recommendation scenario of the live product as an application, see Table 3 below:

[0171]

[0172]

[0173] Table 3

[0174] Among them, each column represents the model test result for the traffic proportion when a is a specified value. Taking a = 0.5 as an example, the traffic proportion represents the traffic occupied by the effective viewing duration of the test users within a specified time and the traffic consumed for pushing to the test users. The larger the traffic proportion, the better the model's effect on predicting the click-through rate and conversion rate, that is, the higher the accuracy. From the above description, it can be known that the technical solution of the present application uses the last model structure (FM + MMOE + ESMM + GradNorm) to train the target prediction model. Using this target prediction model, the traffic proportion is the highest compared with other models under the condition that a is the same value.

[0175] S705. If the interest score is greater than the interest score threshold, then push the to-be-recommended object to the user terminal corresponding to the predicted user.

[0176] In some embodiments, if the interest score of the predicted user for the to-be-recommended object is greater than the interest score threshold, then the to-be-recommended object can be pushed to the user terminal of the predicted user object, that is, push the to-be-recommended object (online anchor) to the predicted user in the live broadcast interface of the application client. The interest score threshold can be set by relevant business personnel according to empirical values.

[0177] Optionally, when the interest scores of multiple to-be-recommended objects are greater than the interest score threshold, the multiple to-be-recommended objects can be randomly pushed or pushed in sequence according to the size of the interest scores; or, when there are multiple to-be-recommended objects, they can also be pushed in sequence in descending order of the interest scores of the predicted user for the multiple to-be-recommended objects.

[0178] For example, as Figure 8a - Figure 8b shown, Figure 8a - Figure 8b is a schematic diagram of a push scenario for a predicted user provided by an embodiment of the present application. Among them, as Figure 8a, when it is predicted that the user clicks on the live broadcast interface of the application client of the live product, the electronic device obtains the current online anchors, and forms prediction pairs ((u, i1), (u, i2), (u, i3)......, (u, iN)) between the predicted user and each online anchor respectively. And based on multiple prediction pairs, the target user attribute information of the predicted user and the target user attribute information of the online anchor are extracted from the attribute storage platform (such as a database), and the target user attribute information of the predicted user and the target user attribute information of each online anchor are respectively combined to form attribute pairs ((u, i1), (u, i2)......). Each attribute pair is input into the target prediction model in turn to obtain the target prediction click-through rate and the target prediction conversion rate of the predicted user for the online anchor in each attribute pair. The electronic device can use the target prediction click-through rate and the target prediction conversion rate to obtain the interest score of the predicted user for each online anchor, and sort each online anchor using the interest score of the predicted user for each online anchor to obtain the sorted online anchors, and push the sorted online anchors to the user terminal of the predicted user in turn, that is, push them to the live broadcast interface of the application client, such as Figure 8b .

[0179] In the embodiments of the present application, the target user attribute information of the predicted user and the target object attribute information of the object to be recommended can be obtained, the target user attribute information and the target object attribute information are input into the target prediction model, and the target prediction click-through rate and the target prediction conversion rate of the predicted user for the object to be recommended are generated in the target prediction model. Based on the target prediction click-through rate and the target prediction conversion rate, the interest score of the predicted user for the object to be recommended is obtained. If the interest score is greater than the interest score threshold, the object to be recommended is pushed to the user terminal corresponding to the predicted user. By implementing the method proposed in the embodiments of the present application, the interest score of the predicted user for the object to be recommended can be obtained using the target prediction model, and the object to be recommended is pushed based on the interest score to achieve precise push in the recommendation scenario.

[0180] Please refer to Figure 9 , Figure 9 is a schematic structural diagram of a recommendation model training device provided by the present application. It should be noted that, Figure 9 The shown recommendation model training device is used to execute the methods of the embodiments of the present application Figure 2 and Figure 3 shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown, and the specific technical details are not disclosed. Please refer to the embodiments of the present application Figure 2 and Figure 3 shown. The recommendation model training device 900 may include: an acquisition module 901, a generation module 902, and a correction module 903. Among them:

[0181] An acquisition module 901, configured to acquire a sample feature set; the sample feature set includes user attribute information of a sample user, object attribute information of a sample recommendation object, click tags and conversion tags of the sample user for the sample recommendation object;

[0182] The acquisition module 901 is further configured to input the sample feature set into an initial prediction model, and acquire a predicted click-through rate and a predicted conversion rate of the sample user for the sample recommendation object based on the initial prediction model;

[0183] A generation module 902, configured to generate a click loss function based on the click tags and the predicted click-through rate, and generate a click conversion loss function based on the conversion tags, the predicted click-through rate and the predicted conversion rate;

[0184] The generation module 902 is further configured to obtain a weight loss function according to the click loss function, the click conversion loss function and the model parameters of the initial prediction model, and generate a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function based on the weight loss function;

[0185] A correction module 903, configured to obtain a target loss function based on the click weight, the click conversion weight, the click loss function and the click conversion loss function, and correct the model parameters of the initial prediction model based on the target loss function to obtain a target prediction model.

[0186] In a possible implementation manner, when the generation module 902 is configured to obtain a weight loss function according to the click loss function, the click conversion loss function and the model parameters of the initial prediction model, it is specifically configured to:

[0187] Obtain a first weight function corresponding to the click loss function and a second weight function corresponding to the click conversion loss function;

[0188] Determine a click gradient function according to the first weight function, the click loss function and the model parameters of the initial prediction model;

[0189] Determine a click conversion gradient function according to the second weight function, the click conversion loss function and the model parameters of the initial prediction model;

[0190] Generate a weight loss function according to the click gradient function and the click conversion gradient function.

[0191] In a possible implementation manner, when the generation module 902 is configured to generate a weight loss function according to the click gradient function and the click conversion gradient function, it is specifically configured to:

[0192] Determine an average gradient function according to the first weight function, the second weight function, the click loss function, the click conversion loss function and the model parameters;

[0193] Determine the target click gradient function according to the click gradient function and the average gradient function;

[0194] Determine the target click conversion gradient function according to the click conversion gradient function and the average gradient function;

[0195] Generate a weight loss function according to the target click gradient function and the target click conversion gradient function.

[0196] In a possible implementation manner, when the generation module 902 is used to generate the click weight corresponding to the click loss function and the click conversion weight corresponding to the click conversion loss function based on the weight loss function, it is specifically used for:

[0197] Derive the first weight adjustment function from the first weight function based on the weight loss function, and determine the click weight according to the first weight adjustment function and the first weight function;

[0198] Derive the second weight adjustment function from the second weight function based on the weight loss function, and determine the click conversion weight according to the second weight adjustment function and the second weight function.

[0199] In a possible implementation manner, when the correction module 903 is used to obtain the target loss function based on the click weight, the click conversion weight, the click loss function, and the click conversion loss function, it is specifically used for:

[0200] Weight the click loss function according to the click weight to obtain the first weighted loss function;

[0201] Weight the click conversion loss function according to the click conversion weight to obtain the second weighted loss function;

[0202] Generate the target loss function according to the first weighted loss function and the second weighted loss function.

[0203] In a possible implementation manner, when the acquisition module 901 is used to obtain the predicted click-through rate and the predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model, it is specifically used for:

[0204] In the initial prediction model, generate the user attribute features corresponding to the user attribute information, and generate the object attribute features corresponding to the object attribute information;

[0205] Obtain the predicted click-through rate and the predicted conversion rate based on the user attribute features and the object attribute features.

[0206] In a possible implementation manner, when the acquisition module 901 is used to obtain the predicted click-through rate and the predicted conversion rate based on the user attribute features and the object attribute features, it is specifically used for:

[0207] Fuse the user attribute features and object attribute features to obtain fused attribute features;

[0208] Obtain the predicted click-through rate and predicted conversion rate based on the fused attribute features.

[0209] In a possible implementation, when the obtaining module 901 is used to obtain the predicted click-through rate and predicted conversion rate based on the fused attribute features, it is specifically used for:

[0210] Generate click prediction features and conversion prediction features based on the fused attribute features;

[0211] Obtain the predicted click-through rate based on the click prediction features, and obtain the predicted conversion rate based on the conversion prediction features.

[0212] In a possible implementation, the above initial prediction model includes a click weight prediction network, a conversion weight prediction network, and N prediction feature generation networks, where N is a positive integer;

[0213] When the obtaining module 901 is used to generate click prediction features and conversion prediction features based on the fused attribute features, it is specifically used for:

[0214] Input the fused attribute features into N prediction feature generation networks, and respectively generate corresponding initial prediction features according to the fused attribute features in each prediction feature generation network;

[0215] Input the fused attribute features into the click weight prediction network, and generate a first prediction weight corresponding to each prediction feature generation network in the click weight prediction network;

[0216] Input the fused attribute features into the conversion weight prediction network, and generate a second prediction weight corresponding to each prediction feature generation network in the conversion weight prediction network;

[0217] Use the first prediction weight corresponding to each prediction feature generation network to perform weighted summation on the initial prediction features corresponding to each prediction feature generation network to obtain click prediction features;

[0218] Use the second prediction weight corresponding to each prediction feature generation network to perform weighted summation on the initial prediction features corresponding to each prediction feature generation network to obtain conversion prediction features.

[0219] In the embodiments of the present application, the acquisition module inputs the sample feature set into the initial prediction model, and obtains the predicted click-through rate and predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model; the generation module generates a click loss function based on the click label and the predicted click-through rate, and generates a click conversion loss function for the click conversion loss function based on the conversion label, the predicted click-through rate, and the predicted conversion rate; the generation module obtains a weight loss function according to the click loss function, the click conversion loss function, and the model parameters of the initial prediction model, and generates a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function based on the weight loss function; the correction module obtains a target loss function based on the click weight, the click conversion weight, the click loss function, and the click conversion loss function, and corrects the model parameters of the initial prediction model based on the target loss function to obtain a target prediction model. By implementing the above-mentioned device, the click weight and the click conversion weight can be dynamically adjusted during the model training process to maintain the balance between the two loss functions in the target loss function, thereby making the training effect of the target prediction model better and the accuracy of the predicted click-through rate and conversion rate higher.

[0220] In each embodiment of the present application, the functional modules can be integrated into one module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module, which is not limited in the present application.

[0221] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a data recommendation device based on a recommendation model provided by the present application. It should be noted that Figure 10 the data recommendation device based on the recommendation model shown Figure 7 is used to execute the method of the embodiments shown in the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown, and the specific technical details are not disclosed. Please refer to the embodiments shown in the present application Figure 7 The data recommendation device 1000 based on the recommendation model may include: an acquisition module 1001, an input module 1002, a generation module 1003, and a push module 1004. Among them:

[0222] The acquisition module 1001 is used to acquire the target user attribute information of the predicted user and the target object attribute information of the object to be recommended;

[0223] The input module 1002 is used to input the target user attribute information and the target object attribute information into the target prediction model;

[0224] The generation module 1003 is used to generate the target predicted click-through rate and the target predicted conversion rate of the predicted user for the object to be recommended in the target prediction model;

[0225] An obtaining module 1001 is further configured to obtain an interest score of a predicted user for a to-be-recommended object based on a target predicted click-through rate and a target predicted conversion rate;

[0226] A pushing module 1004 is configured to push the to-be-recommended object to a user terminal corresponding to the predicted user if the interest score is greater than an interest score threshold.

[0227] In a possible implementation manner, the target prediction model may be trained by using the relevant descriptions in the above Figure 2 illustrated embodiments and / or Figure 3 illustrated embodiments.

[0228] In the embodiments of the present application, an obtaining module obtains target user attribute information of a predicted user and target object attribute information of a to-be-recommended object; an input module inputs the target user attribute information and the target object attribute information into a target prediction model; a generating module generates a target predicted click-through rate and a target predicted conversion rate of the predicted user for the to-be-recommended object in the target prediction model; the obtaining module obtains an interest score of the predicted user for the to-be-recommended object based on the target predicted click-through rate and the target predicted conversion rate; if the interest score is greater than the interest score threshold, the pushing module pushes the to-be-recommended object to a user terminal corresponding to the predicted user. By implementing the above-mentioned device, the target prediction model can be used to obtain an interest score of a predicted user for a to-be-recommended object, and the to-be-recommended object can be pushed based on the interest score to achieve accurate pushing in a recommendation scenario.

[0229] In each embodiment of the present application, each functional module may be integrated into one module, may exist separately physically for each module, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module, which is not limited in the present application.

[0230] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 11 shown, the electronic device 1100 includes: at least one processor 1101 and a memory 1102. Optionally, the electronic device may further include a network interface. Among them, data can be exchanged between the processor 1101, the memory 1102, and the network interface. The network interface is controlled by the processor 1101 to send and receive messages. The memory 1102 is used to store a computer program, and the computer program includes program instructions. The processor 1101 is used to execute the program instructions stored in the memory 1102. Among them, the processor 1101 is configured to call the program instructions to execute the above method.

[0231] Among them, the memory 1102 may include volatile memory, such as random-access memory (RAM); the memory 1102 may also include non-volatile memory, such as flash memory, solid-state drive (SSD), etc.; the memory 1102 may further include a combination of the above types of memories.

[0232] Among them, the processor 1101 may be a central processing unit (CPU). In one embodiment, the processor 1101 may also be a Graphics Processing Unit (GPU). The processor 1101 may also be a combination of a CPU and a GPU.

[0233] In a possible implementation, the memory 1102 is used to store program instructions, and the processor 1101 may call the program instructions to perform the following steps:

[0234] Obtain a sample feature set; the sample feature set includes user attribute information of a sample user, object attribute information of a sample recommended object, a click label and a conversion label of the sample user for the sample recommended object;

[0235] Input the sample feature set into an initial prediction model, and obtain a predicted click-through rate and a predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model;

[0236] Generate a click loss function based on the click label and the predicted click-through rate, and generate a click conversion loss function based on the click label, the conversion label, the predicted click-through rate, and the predicted conversion rate;

[0237] Obtain a weight loss function according to the click loss function, the click conversion loss function, and the model parameters of the initial prediction model, and generate a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function based on the weight loss function;

[0238] Obtain a target loss function based on the click weight, the click conversion weight, the click loss function, and the click conversion loss function, and correct the model parameters of the initial prediction model based on the target loss function to obtain a target prediction model.

[0239] In a possible implementation, when the processor 1101 is used to obtain a weight loss function according to the click loss function, the click conversion loss function, and the model parameters of the initial prediction model, it is specifically used for:

[0240] Obtain the first weight function corresponding to the click loss function and the second weight function corresponding to the click conversion loss function;

[0241] Determine the click gradient function according to the first weight function, the click loss function, and the model parameters of the initial prediction model;

[0242] Determine the click conversion gradient function according to the second weight function, the click conversion loss function, and the model parameters of the initial prediction model;

[0243] Generate a weight loss function according to the click gradient function and the click conversion gradient function.

[0244] In a possible implementation, when the processor 1101 is used to generate a weight loss function according to the click gradient function and the click conversion gradient function, it is specifically used for:

[0245] Determine the average gradient function according to the first weight function, the second weight function, the click loss function, the click conversion loss function, and the model parameters;

[0246] Determine the target click gradient function according to the click gradient function and the average gradient function;

[0247] Determine the target click conversion gradient function according to the click conversion gradient function and the average gradient function;

[0248] Generate a weight loss function according to the target click gradient function and the target click conversion gradient function.

[0249] In a possible implementation, when the processor 1101 is used to generate the click weight corresponding to the click loss function and the click conversion weight corresponding to the click conversion loss function based on the weight loss function, it is specifically used for:

[0250] Derive the first weight adjustment function by taking the derivative of the first weight function with respect to the weight loss function, and determine the click weight according to the first weight adjustment function and the first weight function;

[0251] Derive the second weight adjustment function by taking the derivative of the second weight function with respect to the weight loss function, and determine the click conversion weight according to the second weight adjustment function and the second weight function.

[0252] In a possible implementation, when the processor 1101 is used to obtain the target loss function based on the click weight, the click conversion weight, the click loss function, and the click conversion loss function, it is specifically used for:

[0253] Weight the click loss function according to the click weight to obtain the first weighted loss function;

[0254] Weight the click conversion loss function according to the click conversion weight to obtain a second weighted loss function;

[0255] Generate an objective loss function according to the first weighted loss function and the second weighted loss function.

[0256] In a possible implementation, when the processor 1101 is used to obtain the predicted click-through rate and the predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model, it is specifically used for:

[0257] In the initial prediction model, generate user attribute features corresponding to the user attribute information, and generate object attribute features corresponding to the object attribute information;

[0258] Obtain the predicted click-through rate and the predicted conversion rate based on the user attribute features and the object attribute features.

[0259] In a possible implementation, when the processor 1101 is used to obtain the predicted click-through rate and the predicted conversion rate based on the user attribute features and the object attribute features, it is specifically used for:

[0260] Perform feature fusion on the user attribute features and the object attribute features to obtain fused attribute features;

[0261] Obtain the predicted click-through rate and the predicted conversion rate based on the fused attribute features.

[0262] In a possible implementation, when the processor 1101 is used to obtain the predicted click-through rate and the predicted conversion rate based on the fused attribute features, it is specifically used for:

[0263] Generate click prediction features and conversion prediction features based on the fused attribute features;

[0264] Obtain the predicted click-through rate based on the click prediction features, and obtain the predicted conversion rate based on the conversion prediction features.

[0265] In a possible implementation, the above initial prediction model includes a click weight prediction network, a conversion weight prediction network, and N prediction feature generation networks, where N is a positive integer;

[0266] When the processor 1101 is used to generate click prediction features and conversion prediction features based on the fused attribute features, it is specifically used for:

[0267] Input the fused attribute features into the N prediction feature generation networks, and respectively generate corresponding initial prediction features according to the fused attribute features in each prediction feature generation network;

[0268] Input the fused attribute features into the click weight prediction network, and generate a first prediction weight corresponding to each prediction feature generation network in the click weight prediction network;

[0269] Input the fusion attribute features into the transformation weight prediction network, and generate the second prediction weights corresponding to each prediction feature generation network in the transformation weight prediction network;

[0270] Use the first prediction weights corresponding to each prediction feature generation network to perform weighted summation on the initial prediction features corresponding to each prediction feature generation network respectively, to obtain the click prediction features;

[0271] Use the second prediction weights corresponding to each prediction feature generation network to perform weighted summation on the initial prediction features corresponding to each prediction feature generation network respectively, to obtain the conversion prediction features.

[0272] In a possible implementation, the memory 1102 is used to store program instructions, and the processor 1101 can call the program instructions to execute the following steps:

[0273] Obtain the target user attribute information of the predicted user and the target object attribute information of the object to be recommended;

[0274] Input the target user attribute information and the target object attribute information into the target prediction model;

[0275] Generate the target prediction click-through rate and the target prediction conversion rate of the predicted user for the object to be recommended in the target prediction model;

[0276] Based on the target prediction click-through rate and the target prediction conversion rate, obtain the interest score of the predicted user for the object to be recommended;

[0277] If the interest score is greater than the interest score threshold, then push the object to be recommended to the user terminal corresponding to the predicted user.

[0278] Wherein, the target prediction model is trained by using the relevant descriptions in the above Figure 2 illustrated embodiments and / or Figure 3 illustrated embodiments.

[0279] In specific implementation, the above-described device, processor 1101, memory 1102, etc. can execute the implementation manners described in the above method embodiments, and can also execute the implementation manners described in the embodiments of the present application, which will not be elaborated here.

[0280] In an embodiment of the present application, a computer (readable) storage medium is further provided. The computer storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor can execute some or all of the steps executed in the above method embodiment. Optionally, the computer storage medium may be volatile or non-volatile. The computer-readable storage medium mainly includes a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.

[0281] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0282] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The foregoing program can be stored in a computer storage medium, and the computer storage medium can be a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the foregoing storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0283] The foregoing disclosure is only part of the embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments, and the equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A method for training a recommendation model, characterized in that, The method includes: Obtain a sample feature set; the sample feature set includes user attribute information of a sample user, object attribute information of a sample recommended object, click tags and conversion tags of the sample user for the sample recommended object; Input the sample feature set into an initial prediction model, and obtain a predicted click-through rate and a predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model; Generate a click loss function based on the click tags and the predicted click-through rate, and generate a click conversion loss function based on the click tags, the conversion tags, the predicted click-through rate and the predicted conversion rate; Determine a click gradient function according to the click loss function, a first weight function corresponding to the click loss function and model parameters of the initial prediction model, and determine a click conversion gradient function according to the click conversion loss function, a second weight function corresponding to the click conversion loss function and the model parameters. Generate a weight loss function according to the click gradient function and the click conversion gradient function; Generate a click weight corresponding to the click loss function and a click conversion weight corresponding to the click conversion loss function based on the weight loss function; Obtain a target loss function based on the click weight, the click conversion weight, the click loss function and the click conversion loss function, and correct the model parameters of the initial prediction model based on the target loss function to obtain a target prediction model.

2. The method according to claim 1, characterized in that, The generating the weight loss function according to the click gradient function and the click conversion gradient function includes: Determine an average gradient function according to the first weight function, the second weight function, the click loss function, the click conversion loss function and the model parameters; Determine a target click gradient function according to the click gradient function and the average gradient function; Determine a target click conversion gradient function according to the click conversion gradient function and the average gradient function; Generate the weight loss function according to the target click gradient function and the target click conversion gradient function.

3. The method according to claim 1, wherein The generating the click weight corresponding to the click loss function and the click conversion weight corresponding to the click conversion loss function based on the weight loss function includes: Derive the first weight function based on the weight loss function to obtain a first weight adjustment function, and determine the click weight according to the first weight adjustment function and the first weight function; Derive the second weight function based on the weight loss function to obtain a second weight adjustment function, and determine the click conversion weight according to the second weight adjustment function and the second weight function.

4. The method according to claim 1, characterized in that, The obtaining the target loss function based on the click weight, the click conversion weight, the click loss function and the click conversion loss function includes: Weight the click loss function according to the click weight to obtain a first weighted loss function; Weight the click conversion loss function according to the click conversion weight to obtain a second weighted loss function; Generate the target loss function according to the first weighted loss function and the second weighted loss function.

5. The method according to claim 1, wherein The obtaining of the predicted click-through rate and predicted conversion rate of the sample user for the sample recommended object based on the initial prediction model includes: In the initial prediction model, generate user attribute features corresponding to the user attribute information, and generate object attribute features corresponding to the object attribute information; Obtain the predicted click-through rate and the predicted conversion rate based on the user attribute features and the object attribute features.

6. The method according to claim 5, wherein The obtaining of the predicted click-through rate and the predicted conversion rate based on the user attribute features and the object attribute features includes: Perform feature fusion on the user attribute features and the object attribute features to obtain fused attribute features; Obtain the predicted click-through rate and the predicted conversion rate based on the fused attribute features.

7. The method according to claim 6, wherein The obtaining of the predicted click-through rate and the predicted conversion rate based on the fused attribute features includes: Generate a click prediction feature and a conversion prediction feature based on the fused attribute features; Obtain the predicted click-through rate based on the click prediction feature, and obtain the predicted conversion rate based on the conversion prediction feature.

8. The method according to claim 7, wherein The initial prediction model includes a click weight prediction network, a conversion weight prediction network, and N prediction feature generation networks, where N is a positive integer; The generating of the click prediction feature and the conversion prediction feature based on the fused attribute features includes: Input the fused attribute features into the N prediction feature generation networks, and generate corresponding initial prediction features according to the fused attribute features in each prediction feature generation network; Input the fused attribute features into the click weight prediction network, and generate a first prediction weight corresponding to each prediction feature generation network in the click weight prediction network; Input the fused attribute features into the conversion weight prediction network, and generate a second prediction weight corresponding to each prediction feature generation network in the conversion weight prediction network; Use the first prediction weight corresponding to each prediction feature generation network to perform weighted summation on the initial prediction features corresponding to each prediction feature generation network to obtain the click prediction feature; Use the second prediction weight corresponding to each prediction feature generation network to perform weighted summation on the initial prediction features corresponding to each prediction feature generation network to obtain the conversion prediction feature.

9. A data recommendation method based on a recommendation model, characterized in that, The method includes: Obtain the target user attribute information of the predicted user and the target object attribute information of the object to be recommended; Input the target user attribute information and the target object attribute information into a target prediction model; the target prediction model is trained by using the method described in any one of claims 1-8 above; Generate the target predicted click-through rate and target predicted conversion rate of the predicted user for the object to be recommended in the target prediction model; Obtain the interest score of the predicted user for the object to be recommended based on the target predicted click-through rate and the target predicted conversion rate; If the interest score is greater than the interest score threshold, push the object to be recommended to the user terminal corresponding to the predicted user.

10. An electronic device, characterized in that, It includes a processor and a memory. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the method according to any one of claims 1-9.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Training method of behavior prediction model, behavior prediction method and device, and equipment

    CN112784157A