Recommender system conversion rate prediction method and system based on unbiased data and meta-learning
By using unbiased data sets and meta-learning methods to train the conversion rate prediction model in the recommendation system, the selection bias and data sparse problems are solved, and the prediction accuracy and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202410700152.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-05-31
AI Technical Summary
When the existing recommendation system conversion rate prediction model faces the problems of selection bias and data sparseness, the prediction performance decreases, and the deviation, variance and generalization errors are too large.
The conversion rate prediction method of recommendation system based on unbiased data and meta-learning is adopted. The propensity prediction model is assisted in training the propensity prediction model and filling the prediction model through unbiased data sets, reducing the deviation and variance of the model, and using meta-learning methods to train the model on small sample tasks to improve the generalization ability of the model.
It effectively reduces the deviation, variance and generalization error of the conversion rate prediction model of the recommendation system, improves the prediction accuracy and generalization ability of the model, and solves the problems of selection bias and data sparseness.
Smart Images

Figure CN118521380B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of recommendation system conversion rate prediction, and in particular to a recommendation system conversion rate prediction method and system based on unbiased data and meta-learning. Background Art
[0002] In recent years, recommender systems (RS) have been widely used in e-commerce platforms because they can provide personalized recommendations for each user. Among them, post-click conversion rate (CVR) is an important indicator for measuring user interest and increasing the revenue of e-commerce platforms. CVR prediction represents the probability that a user consumes the recommended item after clicking on it. In reality, only items that users are interested in are clicked and consumed by users. This leads to serious data sparsity and selection bias problems in the user feedback data collected by the recommendation system. It is inaccurate to train a CVR prediction model based only on sparse and biased user click data.
[0003] Some debiasing methods can be used to solve the bias problem in CVR prediction: (1) Error imputation based (EIB) estimator, which introduces a filling prediction model to generate pseudo labels for missing user click data to calculate the filling error of non-click events, and use this to estimate the prediction error of all events. However, since the filling error is usually inaccurate, this estimator usually has a large bias. (2) Inverse propensity score (IPS) estimator, which uses the predicted user propensity to inversely weight the prediction error of each click event to increase the weight of low-click samples and reduce the weight of high-click samples, thereby balancing the data distribution differences between high-propensity click samples and low-propensity click samples to reduce the bias of the model. However, using sparse and biased data to predict user propensity is not accurate, so IPS estimators often have high variance problems. (3) Transfer via joint reconstruction (TJR). This method realizes knowledge transfer and sharing between unbiased data and biased data by jointly reconstructing the loss function. It has achieved certain results in effectively integrating unbiased data into biased data. However, in the process of extracting user preference information from biased data and unbiased data in the upper branch of the model, the distribution difference between click samples and non-click samples in biased data is not taken into account, and missing data is not filled. Only a small amount of biased click data combined with unbiased data cannot achieve double robustness. (4) Doubly robust (DR) estimator. This estimator combines the advantages of EIB estimator and IPS estimator. Its core idea is to correct error bias to make the prediction of the filling model more accurate. It also uses the inverse propensity score for weighting to reduce the distribution difference between click samples and non-click samples. When the filling error or the prediction tendency is accurate, it can reduce both bias and variance.
[0004] Among the above methods, the DR estimator shows good performance overall, but there are still some problems. Through theoretical analysis of the DR estimator, it is found that its bias, variance and generalization error all depend on the product of the propensity prediction error bias and the filling prediction error bias weighted by the inverse of the propensity score, which easily leads to excessive bias, variance and generalization error. It is analyzed from two aspects:
[0005] (1) Selection bias: The click data collected by the recommendation system has serious selection bias. These click data only represent the preferences of some users for certain items. The user preferences predicted by such data often cannot represent all users, making the user preferences predicted by the model inaccurate, which in turn leads to excessive bias in the preference prediction error. Secondly, the filling prediction model generates pseudo labels for unclicked data based on click data, and the click data itself has serious selection bias, which will pass the data bias to the generated pseudo labels, resulting in inaccurate filling errors, and thus excessively large error bias in the filling prediction model.
[0006] (2) Data sparsity: Data sparsity causes the sample size of click events to be much smaller than that of non-click events. The model lacks sufficient user-item feature information when estimating the overall item tendency, resulting in inaccurate tendency prediction, which leads to excessive deviation in tendency prediction error. Secondly, since a small number of click events lack sufficient user-item feature information, the pseudo-labels generated by the filling model for non-click events based on a small number of click events are also inaccurate, resulting in excessive deviation in the error of the filling model.
[0007] The above two situations will lead to excessive bias, variance and generalization error of the DR estimator, resulting in a decrease in the overall prediction performance of the DR estimator. Summary of the invention
[0008] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a recommendation system conversion rate prediction method and system based on unbiased data and meta-learning, aiming to solve the problems of selection bias and data sparsity, and to improve the prediction accuracy of the tendency prediction model and the filling prediction model by assisting the training of debiasing parameters of the tendency prediction model and the filling prediction model through unbiased data sets. The model is trained using the meta-learning method in small sample learning, thereby alleviating the problem of poor model prediction accuracy caused by data sparsity.
[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0010] A first aspect of the present invention provides a method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning.
[0011] The recommendation system conversion rate prediction method based on unbiased data and meta-learning includes the following steps:
[0012] Obtain biased data sets from the historical data of interactions between multiple users and multiple items in the recommendation system, and simultaneously obtain unbiased data sets from multiple users and multiple items;
[0013] Build a UMEDR model that includes a filling prediction model, a tendency prediction model, and a conversion rate prediction model, and train the UMEDR model:
[0014] Using an unbiased data set to train a filling prediction model and a tendency prediction model to obtain a trained filling prediction model and a tendency prediction model, and obtaining an unbiased filling error and a user tendency based on the trained filling prediction model and the tendency prediction model respectively;
[0015] Based on biased data sets, unbiased filling errors and user tendencies, a meta-learning method is used to train the conversion rate prediction model to obtain a trained conversion rate prediction model.
[0016] Based on the trained conversion rate prediction model, the conversion rate prediction of users clicking on items in the recommendation system is realized.
[0017] A second aspect of the present invention provides a recommendation system conversion rate prediction system based on unbiased data and meta-learning.
[0018] Recommendation system conversion rate prediction system based on unbiased data and meta-learning, including:
[0019] The data set acquisition module is configured to: acquire biased data sets from historical data of interactions between multiple users and multiple items in the recommendation system, and acquire unbiased data sets between multiple users and multiple items;
[0020] The training module is configured as follows:
[0021] Build a UMEDR model that includes a filling prediction model, a tendency prediction model, and a conversion rate prediction model, and train the UMEDR model:
[0022] Using an unbiased data set to train a filling prediction model and a tendency prediction model to obtain a trained filling prediction model and a tendency prediction model, and obtaining an unbiased filling error and a user tendency based on the trained filling prediction model and the tendency prediction model respectively;
[0023] Based on biased data sets, unbiased filling errors and user tendencies, a meta-learning method is used to train the conversion rate prediction model to obtain a trained conversion rate prediction model.
[0024] The prediction module is configured to: based on the trained conversion rate prediction model, realize the conversion rate prediction of users in the recommendation system who click on items and then make purchases.
[0025] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning as described in the first aspect of the present invention.
[0026] The fourth aspect of the present invention provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning as described in the first aspect of the present invention are implemented.
[0027] One or more of the above technical solutions have the following beneficial effects:
[0028] The present invention provides a method and system for predicting conversion rate of a recommendation system based on unbiased data and meta-learning, and proposes a new meta-learning dual robust debiasing model UMEDR that integrates unbiased data, aiming to solve the problems of selection bias and data sparsity. An unbiased dataset is introduced, which reflects user preferences in an unbiased manner. The unbiased dataset is used to assist the training of debiasing parameters of a tendency prediction model and a filling prediction model, thereby improving the prediction accuracy of the tendency prediction model and the filling prediction model, thereby reducing the overall bias, variance and generalization error of the DR estimator.
[0029] The present invention uses the meta-learning method in small sample learning to train the model. The meta-learning method is good at improving the model prediction performance when the amount of data is small. Since the data set in the recommendation system is relatively sparse, the use of meta-learning alleviates the problem of poor model prediction accuracy caused by data sparsity.
[0030] When the present invention adopts the meta-learning training model, the biased data set is divided into a small sample task set, wherein the data contained in each small sample task is not exactly the same, and each round of model update extracts small sample tasks of different batches for learning. Since the small sample tasks used in each round of update are not exactly the same, the updated UMEDR model can obtain sufficiently strong generalization ability, so that it can better fit when facing new and never-before-seen data.
[0031] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0033] Figure 1 Schematic diagram of the feedback loop and selection bias stages in the recommendation system.
[0034] Figure 2 Use a schematic diagram for an unbiased dataset.
[0035] Figure 3 This is the workflow diagram of the UMEDR model.
[0036] Figure 4 It is the index of each method in Yahoo! R3 in the comparative experiment.
[0037] Figure 5 It is the indicator of each method in Coat Shopping in the comparative experiment.
[0038] Figure 6 are the indicators of various methods in Yahoo! R3 in the ablation experiment.
[0039] Figure 7 It is the indicator of each method in Coat Shopping in the ablation experiment.
[0040] Figure 8 Line graph of the model performance in Yahoo! R3 for different unbiased dataset sizes.
[0041] Fig. 9 The performance line chart of the model in Coat Shopping under different unbiased dataset sizes.
[0042] Fig.10 Line graph of model performance at different embedding sizes in Yahoo! R3.
[0043] Fig.11 Line chart of model performance under different embedding sizes in Coat Shopping. DETAILED DESCRIPTION
[0044] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0045] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0046] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0047] The overall idea proposed by the present invention is:
[0048] Through research, we found that unbiased datasets do not have selection bias problems. Using such datasets to assist model training can effectively improve model prediction performance and reduce model bias. In addition, small sample learning can learn and achieve good results with relatively small amounts of data, which can alleviate the problem of inaccurate predictions caused by data sparsity in recommendation systems. This provides a new motivation for further optimizing the DR estimator, that is, starting with the use of unbiased datasets and small sample learning.
[0049] Based on the above problems and research motivations, on the basis of the dual robust estimator, this paper proposes a new meta-learning doubly robust debiasing model (UMEDR, meta-learning doubly robust debiasing model with unbiased data) that integrates unbiased data, and introduces an unbiased dataset that has better performance. Because there is no selection bias in the unbiased dataset, using an unbiased dataset to assist model training will achieve better results. By using an unbiased dataset to assist the learning of debiasing parameters of the tendency prediction model and the filling prediction model, the prediction accuracy of the tendency prediction model and the filling prediction model can be improved, making the predicted user tendency and filling error more accurate, thereby reducing the overall bias, variance and generalization error of the DR estimator.
[0050] Then, in terms of model training, the present invention adopts the meta-learning method in small sample learning. Meta-learning is based on multiple different small sample tasks, focusing on using a small amount of data to improve the model prediction performance. Because the data set used by the recommendation system is sparse, using the meta-learning method to train the model on a sparse data set has a natural advantage, which can greatly improve the prediction accuracy of the model. Furthermore, since the data sets contained in each task used by meta-learning when training the model are different, this greatly avoids the problem of model overfitting and improves the model generalization ability.
[0051] Compared with the prior art, the present invention has the following differences. First, the research motivation is different. Most of the previous works are based on the perspective of changing the model loss weight and model joint training, while the work studied in the present invention starts from the introduction of new unbiased data sets and model training methods. Secondly, the research focus is different. Previous work focused on changing the model training loss function to balance the distribution of click data and non-click data, thereby reducing model bias, variance and generalization error, while the present invention digs deeper into the deep-seated causes of model bias, variance and generalization error, namely the data sparsity and selection bias problems of the biased data set itself, and then finds corresponding strategies to solve the problem. Finally, the model training method is different. Previous studies have been carried out on a single data set through multiple rounds of iterations to enable the model to learn better parameters, which can easily cause model overfitting. The present invention uses multiple different small sample data sets to train the model, which greatly improves the generalization ability of the model and reduces the risk of model overfitting.
[0052] Embodiment 1
[0053] This embodiment discloses a method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning.
[0054] like Figure 1 As shown in FIG. 1 , the recommendation system conversion rate prediction method based on unbiased data and meta-learning includes the following steps:
[0055] Obtain biased data sets from the historical data of interactions between multiple users and multiple items in the recommendation system, and simultaneously obtain unbiased data sets from multiple users and multiple items;
[0056] Build a UMEDR model that includes a filling prediction model, a tendency prediction model, and a conversion rate prediction model, and train the UMEDR model:
[0057] Using an unbiased data set to train a filling prediction model and a tendency prediction model to obtain a trained filling prediction model and a tendency prediction model, and obtaining an unbiased filling error and a user tendency based on the trained filling prediction model and the tendency prediction model respectively;
[0058] Based on biased data sets, unbiased filling errors and user tendencies, a meta-learning method is used to train the conversion rate prediction model to obtain a trained conversion rate prediction model.
[0059] Based on the trained conversion rate prediction model, the conversion rate prediction of users clicking on items in the recommendation system is realized.
[0060] Next, the relevant technologies of this embodiment will be described in detail:
[0061] 1. Related Work
[0062] The essence of a recommender system is a feedback loop between users, data, and models, such as Figure 1 As shown in Figure 1, it consists of three stages: (1) Collecting data through user behavior. (2) Using the collected data to train the recommendation system. (3) The trained recommendation system recommends items to the user. Selection bias occurs in the first stage, that is, users can freely choose the items to be evaluated, resulting in the data collected by the recommendation system being troubled by the user's self-selection, causing the problem of missing non-random data (MNAR). Training the recommendation system with biased data will further aggravate the bias.
[0063] Therefore, we consider introducing an unbiased dataset and using a meta-learning method to train the recommendation system. Unlike biased datasets, unbiased datasets do not have the problem of selection bias. The recommendation system learned on an unbiased dataset has better recommendation performance.
[0064] 1.1 Estimator based on filling error
[0065] The error imputation based (EIB) estimator uses a filling prediction model to calculate the filling error. Using the filling error of unclicked events and the prediction error of clicked events, the loss function of the EIB estimator is:
[0066]
[0067] Among them, when the filling error When accurate, the EIB estimator is unbiased. However, models trained on biased data often have inaccurate fill errors, so the EIB estimator usually has large bias in practice.
[0068] 1.2 Inverse Propensity Score Estimator
[0069] The inverse propensity score (IPS) estimator calculates the predicted user’s propensity use Each click sample is weighted to calibrate the distribution difference between high-click samples and low-click samples to reduce the bias. The loss function of the IPS estimator is:
[0070]
[0071] Among them, when the predicted user tendency When accurate, the IPS estimator is unbiased, but due to the sparsity of biased data sets, click samples only account for a small part of the data set, so the user tendency predicted with very little data is often inaccurate, and the IPS estimator usually has a large variance.
[0072] 1.3 Transfer via Joint Reconstruction
[0073] The TJR method is divided into two branches. The upper branch uses the union of biased data and unbiased data to extract user preference information and generate predictions. The lower branch uses the biased dataset to extract user bias information and generate predictions Then the unbiased prediction is obtained in a linear manner, i.e. Then a new loss function is designed to jointly reconstruct biased data and unbiased data to make unbiased prediction More accurately, the TJR loss function is:
[0074]
[0075] in is the true label from the biased dataset, are the true labels from the unbiased dataset, It is an additional loss function designed to better learn biased features.
[0076] 1.4 Dual Robust Estimators
[0077] In order to solve the defects of EIB estimator and IPS estimator, a doubly robust estimator (DR) is proposed. The DR estimator combines the EIB estimator and the IPS estimator at the same time, using the filling error To estimate the prediction error e for all events u,i , and through the predicted user tendency Inverse propensity weighting to correct for the bias of no-click events The loss function of the DR estimator is as follows:
[0078]
[0079] Among them, when the filling error Or predicted user trends When accurate, the DR estimator is unbiased.
[0080] Given the predicted user preference and filling error The bias of the DR estimator is:
[0081]
[0082] The variance of the DR estimator is:
[0083]
[0084] The generalization error of the DR estimator is:
[0085]
[0086] From formulas (5), (6), and (7), we can see that the bias, variance, and generalization error of the DR estimator are all related to the predicted user tendency. and filling error About, when or When inaccurate, the bias, variance, and generalization error of the DR estimator will be large.
[0087] 2 Problem Description
[0088] Let U={u1,u2,...,u M} is a set of M users, I = {i1,i2,...,i N} is a set of N items. T represents a biased dataset collected from historical user interaction data, D U This is the unbiased dataset introduced, which has better performance. u,i =1,(u,i)∈D T} represents a collection of click samples, where each item o u,i ∈{0,1} indicates whether user u clicked on item i. If yes, o u,i =1, otherwise o u,i =0.
[0089] R∈{0,1} M×N As the true conversion rate label matrix, each item r u,i ∈{0,1} reflects the real rating of user u on the recommended item i, x u,i is the potential feature vector of the user-item pair (u,i), which is used in the prediction model f(x u,i ,θ) to predict the score r of user u for the recommended item i u,i , in the recommendation system, most of the r u,i is missing and is non-randomly missing, resulting in sparse data and large bias. Represents the predicted conversion rate label matrix, where each item represents the conversion rate predicted by the model. Ideally, if all r u,i are observed, then the prediction model f(x u,i ,θ):
[0090]
[0091] However, since only when ou,i = 1 to observe the true conversion rate label r u,i , so the ideal loss is not computable, and restricting the analysis to non-random missing data will lead to biased conclusions. Therefore, different debiasing methods have been designed to approximate and replace the ideal loss, such as EIB, IPS and DR estimators.
[0092] 3 Meta-learning dual robust debiasing model integrating unbiased data
[0093] In this section, the proposed debiasing solution will be introduced in detail, including the use of unbiased datasets, training on small sample task sets, design details of the proposed model, model unbiasedness analysis, and the model learning process.
[0094] 3.1 Use of unbiased datasets
[0095] As mentioned above, the bias, variance, and generalization error of the DR estimator mainly depend on the tendency prediction model Predicted user trends and fill prediction model g(x u,i ,φ) filling error If the prediction accuracy of the tendency prediction model and the filling prediction model can be improved, the overall prediction performance of the DR estimator will be effectively increased. T There is a large data deviation, resulting in and g(x u,i ,φ) deviates from the real user feedback results, resulting in the overall performance degradation of the DR estimator.
[0096] With D T The difference is that the unbiased dataset D U Contains unbiased user feedback information, which can provide unbiased update guidance signals to make model predictions more accurate. Figure 2 As shown, the unbiased data set D U Used to guide propensity prediction models and fill prediction model g(x u,i ,φ) to obtain the optimal debiasing parameters and provide the conversion rate prediction model f(x u,i ,θ) provides unbiased user preferences and filling error Since the overall performance of the DR estimator depends on and In the unbiased dataset D U The two learned above are close to unbiased, so the prediction performance of the improved UMEDR model is greatly improved.
[0097] 3.2 Training on small sample task sets
[0098] When training the traditional DR estimator, on a large dataset D T Multiple cycles are performed to converge the model. Although this method can obtain the user-item feature information contained in the data set, it also has a problem, that is, it is easy to cause the DR estimator to overfit, resulting in the DR estimator performing well on the training data, but performing poorly in unseen test data or actual applications.
[0099] In order to improve the generalization ability of the DR estimator, the dataset D T Divide into a small sample task set T = {t1, t2, ..., t i}, where each small sample task t i The data contained are different. Each round of model update extracts different batches of small sample tasks from T for learning. Since the small sample tasks used in each round of update are not exactly the same, the updated UMEDR model can obtain strong enough generalization ability, so that it can fit well when facing new and unseen data. In order to show the effectiveness of this method, multiple small sample tasks and a large dataset were used for verification in the experimental part. The experimental results show that the use of multiple small sample tasks significantly improves the model performance.
[0100] 3.3UMEDR model
[0101] The UMEDR model definition process is divided into three steps:
[0102] (1) Step 1: Use an unbiased dataset D U Guide the filling prediction model to update and obtain the filling error
[0103]
[0104] In formula (9), the square loss is used to reduce and e u,i To improve the gap between accuracy.
[0105] (2) Step 2: Use an unbiased dataset D U Guide the update of the tendency prediction model to generate predicted user tendencies
[0106]
[0107] In formula (10), the cross entropy loss is used to make the predicted user tendency Continuously approaching real user clicks u,i , to improve accuracy.
[0108] (3) Step 3: Based on biased dataset D T And the results obtained in steps 1 and 2 and And using the meta-learning method, the proposed UMEDR model is as follows:
[0109]
[0110] In formula (11), user preference is used Inverse propensity weighting is used to balance the difference between clicked and non-clicked samples.
[0111] The main components of the UMEDR model are similar to the traditional DR estimator, but the filling prediction model g(x u,i ,φ) and tendency prediction model Secondly, the meta-learning method enables the UMEDR model to T Multiple small sample tasks T = {t1, t2, ..., t i} is updated using a double-layer optimization approach to alleviate the data sparsity problem. The model update process will be described in detail in Section 3.4.
[0112] 3.4 Analysis of the unbiasedness of the UMEDR model
[0113] In the click space O, when the expected loss of the UMEDR model is equal to the ideal loss, that is, The UMEDR model is considered to be unbiased when . Next, the bias of the UMEDR model is derived and its unbiasedness is proved.
[0114] Theorem 1. When in an unbiased data set D U User tendencies learned from and filling error Accurate, that is , the proposed UMEDR model is unbiased.
[0115] prove.
[0116]
[0117] From the above proof, it can be concluded that when in the unbiased data set D U User tendencies learned from and filling error When accurate, the UMEDR model is unbiased, that is, Bias[L UMEDR (θ)] = 0, so far, the unbiasedness of the UMEDR model has been proved.
[0118] 3.5UMEDR model learning process
[0119] Under the initial parameter state, when training the traditional dual robust model, the filling prediction model g(x u,i ,φ), generate pseudo labels for missing data And based on the generated pseudo labels Calculate the filling error of the filling model Then update the propensity prediction model Get predicted user trends Finally, according to and Update the conversion rate prediction model f(x u,i ,θ), and predict the user conversion rate This process is all done on the biased data set D T The user feedback data in this dataset has a large deviation, which leads to the filling error of the filling prediction model. And the user tendency predicted by the tendency prediction model Inaccurate, resulting in the conversion rate prediction model f(x u,i ,θ) predicted The actual conversion rate r u,i Compared with the traditional DR estimator, there is a large error. In addition, the traditional DR estimator uses a single data set D T The training is performed in multiple rounds of iterations, and D T The data in is very sparse, which further reduces the prediction performance of the DR estimator and leads to its poor generalization performance.
[0120] The update process of the UMEDR model adopts a meta-learning two-layer optimization method, such as Figure 3 As shown, it is divided into two parts: local update and global update. The unbiased data set D is used in the update process. U Guide the filling prediction model g(x u,i ,φ) and tendency prediction model Update to obtain the optimal debiasing parameters, and f(x u,i ,θ) update process in the dataset D T Divided into multiple small sample tasks T = {t1, t2, ..., t i To simplify the model complexity, a common prediction model f(x u,i ,ψ) as the copy model of the filling model, the tendency model, and the conversion rate prediction model, and the local update of the copy model f(x u,i ,ψ) is updated to obtain the initial and The optimal debiasing parameters of the replica model under . The local update process is as follows:
[0121] By filling the prediction model g(x u,i ,φ) and tendency prediction model In the biased dataset D T The small sample tasks divided into i The initial filling error is obtained by and predicted user preferences Based on initial and For the replica prediction model f(x u,i ,ψ) to update and obtain the optimal replica model debiasing parameter ψ*:
[0122]
[0123] Here, the learning rate η1 is used to perform gradient descent to update the replica model f(x u,i ,ψ).
[0124] Since in the biased data set D T What I learned and There is a large deviation, which makes it impossible for the DR estimator to achieve the best prediction effect. Therefore, when filling the prediction model g(x u,i ,φ) and tendency prediction model In the learning process, an unbiased dataset D is introduced U To guide its update to obtain the optimal filling prediction model and tendency prediction model debiasing parameter φ * and
[0125]
[0126] After this, the unbiased filling error is calculated and user preferences based on and For the true prediction model f(x u,i ,θ) to update and obtain the optimal debiasing parameter θ* of the prediction model:
[0127]
[0128] In the above update process, fill the prediction model g(x u,i ,φ) is updated using squared loss and uses the inverse tendency Weighted to balance the distribution difference between high-tendency click samples and low-tendency click samples, the loss function is defined as follows:
[0129]
[0130] in is the prediction error, measured by predicting the conversion rate label and the real conversion rate label ru,i It is calculated by the binary cross entropy between . is the filling error, which is measured by filling the pseudo labels generated by the prediction model and predicted conversion rate tags It is calculated by the binary cross entropy between . λ is a regularization term to prevent the model from overfitting.
[0131] Propensity Model Update using binary cross entropy loss:
[0132]
[0133] o u,i Indicates the real click rate in the click sample, Indicates D U The user tendency predicted under the guidance of the algorithm is reduced through the binary cross entropy loss to improve the accuracy of tendency prediction.
[0134] Based on unbiased data D U The padding error after the update and predicting user preferences Prediction model f(x u,i ,θ) The loss function used in the update process is defined as follows:
[0135]
[0136] UMEDR model filling error and predicting user preferences Compared with the traditional DR estimator, it is more accurate and the overall performance of the model is significantly improved.
[0137]
[0138] 4. Experiment
[0139] In this example, a large number of experiments were conducted to verify the effectiveness of the proposed method. The experiments were designed to answer the following research questions:
[0140] RQ1: Does the proposed UMEDR method improve upon the current state-of-the-art debiasing methods?
[0141] RQ2: What is the impact of the completeness of the UMEDR model on the overall performance of the model?
[0142] RQ3: What is the performance of the UMEDR model under different unbiased dataset sizes?
[0143] RQ4: How do different embedding dimension sizes affect the UMEDR model?
[0144] 4.1 Experimental Setup
[0145] To answer the above questions, this embodiment uses two commonly used data sets, Yahoo! R3 and Coat Shopping, to conduct experiments. The specific details of the data sets are as follows.
[0146] Yahoo! R3: This is a dataset that contains a biased user subset and an unbiased random subset. The user subset includes more than 300,000 ratings from 15,400 users on 1,000 songs. It is collected through the normal interaction process between users and the recommendation system, so the user subset has data bias. The random subset is collected from the first 5,400 users out of 15,400 users. Each user is asked to randomly select 10 songs from these 1,000 songs for rating. The random selection strategy ensures that each song has the same exposure opportunity, so this random subset is unbiased. This example follows previous research and the ratings are binarized by a threshold of 4, that is, observed ratings greater than 4 are marked as positive feedback (r=1), otherwise they are negative feedback (r=-1). The experiment uses the biased user dataset D T The random unbiased dataset is divided into three parts as a training set: 5% is used to assist in training the unbiased dataset D U , 5% is used for the validation set D to tune the hyperparameters V , 90% is used to evaluate the test set D of the model test .
[0147] Coat Shopping: This is a dataset collected from Amazon that contains 290 users’ ratings of 300 coats. Similar to Yahoo! R3, this dataset also contains biased user subsets and unbiased random subsets. The difference is that the unbiased random subset is collected by these 290 users randomly selecting 16 coats from 300 coats for rating. The preprocessing and splitting methods of the two subsets in Coat Shopping are the same as those in Yahoo! R3. The statistics of the dataset are shown in Table 1.
[0148] Table 1 Dataset information
[0149]
[0150] The experiment uses the matrix factorization model (MF) as the underlying model, and uses the SGD optimizer to optimize the fill prediction, tendency prediction, and conversion rate prediction models. MSE, NLL, Recall@5, and NDCG@5 are used as evaluation indicators to evaluate the debiasing performance of the proposed model.
[0151] 4.2 Performance Comparison with Existing Baselines (RQ1)
[0152] This experiment compares both with an improved model based on a dual robust model and with a debiased model using an unbiased dataset. The following will introduce the experimental baseline in detail and analyze the experimental results through specific experimental data.
[0153] 4.2.1 Comparative Experimental Baseline
[0154] (1) MF(D T )、MF(D U )、MF(D U +D T ): respectively in the biased data set D T , unbiased dataset D U And the combined dataset D U +D T Basic matrix factorization model trained on .
[0155] (2) CausE: A method that uses an additional alignment loss function to extract unbiased information.
[0156] (3) KDCRec-Label: A method that trains two models simultaneously using unbiased data and biased data, and transfers unbiased information from the model trained with unbiased data to the model trained with biased data.
[0157] (4) Inverse Propensity Score Estimator (IPS): This method uses the predicted user propensity to inversely weight the prediction error of each click event, thereby balancing the data distribution differences between high-propensity click samples and low-propensity click samples. (5) Dual Robust Estimator (DR-JL): This method combines the inverse propensity score estimator and the filling error estimator.
[0158] (6) More Robust Doubly Robust Estimator (MRDR-JL): A method to reduce variance by redesigning the learning objective of the padding model.
[0159] 4.2.2 Comparative experimental results analysis
[0160] Table 2 and Figure 4-5 The results of the comparative tests show that the proposed UMEDR model has achieved significant improvements over the other eight baseline debiasing methods on the Yahoo! R3 dataset and the CoatShopping dataset.
[0161] Among them, dual-robust models such as UMEDR, DR-JL, and MRDR-JL perform better than the IPS model in NDCG@5 and Recall@5. This is because the dual-robust model combines the advantages of the IPS model and the EIB model and has dual robustness, while IPS does not have this feature.
[0162] The UMEDR model has a significant improvement over the variants of dual robust models such as DR-JL and MRDR-JL. The NDCG@5 index on the two datasets is improved by 6.5% and 1.5% compared with the DR-JL model, and the NDCG@5 index is improved by 5.8% and 1.3% compared with the recently proposed MRDR-JL model. First, this is mainly due to the fact that the proposed UMEDR model uses an unbiased dataset D U Guide the learning of the filling prediction model and the tendency prediction model so that the obtained filling error and user preferences More accurate. Secondly, the model training method of small sample learning and meta-learning double-layer optimization makes the parameters learned by the UMEDR model on two sparse data sets more unbiased.
[0163] Compared with the CausE and KDCRec-Label models that also use unbiased datasets, UMEDR has also made significant improvements. Compared with CausE, the NDCG@5 index on the two datasets has increased by 7% and 2.4%, respectively, and compared with KDCRec-Label, the NDCG@5 index has increased by 4.7% and 2.7%, respectively. This is because UMEDR is a dual-robust model with dual robustness, while the CausE and KDCRec-Label models do not have this feature.
[0164] Compared with the matrix decomposition basic models such as MF, the improvement of the UMEDR model is more obvious. This is because UMEDR uses the MF model as the underlying framework and adds meta-learning, dual robustness, unbiased data and other steps on the basis of MF, making the UMEDR model learn more fully. U The trained model MF(D U ) showed poor results on both datasets because the unbiased dataset D U Smaller in scale, in D U The trained model resulted in severe overfitting.
[0165] Table 2 Comparative experiments on Coat Shopping and Yahoo! R3 datasets
[0166]
[0167] 4.3 Ablation Experiment (RQ2)
[0168] This experiment is mainly to illustrate the necessity of the complete UMEDR model, and mainly compares the performance of the UMEDR model and its five variants on two real data sets. The following will introduce each variant in detail and analyze the characteristics of each variant through specific experimental data.
[0169] 4.3.1 Ablation Experiment Baseline
[0170] (1)UMEDR(NO D U ): Do not use unbiased data set D U Instead of guiding the filling prediction model and the tendency prediction model to update, the traditional biased dataset D is used T .
[0171] (2)UMEDR(NO D T ): The entire update process removes the biased data set D T , only using the unbiased dataset D U Guide model learning.
[0172] (3) UMEDR (NO Meta): does not use meta-learning two-level optimization to train the model, but retains the unbiased dataset D U and small sample learning.
[0173] (4)UMEDR (Single-task): Instead of training the model using multiple small sample tasks, the traditional method of performing multiple rounds of iterations on a large dataset is used for training.
[0174] (5)UMEDR (Only D U +D T ): Remove the two-layer optimization method of small sample learning and meta-learning at the same time.
[0175] 4.3.2 Analysis of ablation experiment results
[0176] Table 3 and Figure 6-7 The UMEDR model is compared with its five variants on the Yahoo! R3 and Coat Shopping datasets.
[0177] Compared with the other five variants, UMEDR improves NDCG@5 by 6.3%, 24.1%, 13.2%, 14.8%, and 14.2% on the Yahoo! R3 dataset, and improves NDCG@5 by 0.5%, 21.7%, 1%, 1.8%, and 5% on the Coat Shopping dataset. U The improvement on the Coat Shopping dataset is not as significant as on the Yahoo! R3 dataset, probably because the Coat Shopping dataset is small and the model is not very efficient. U The termination condition is reached without sufficient learning.
[0178] UMEDR(NO D T ) The results are poor because D U Smaller, in D UTraining the model on this will lead to severe overfitting.
[0179] Table 3 Ablation experiments on Coat Shopping and Yahoo! R3 datasets
[0180]
[0181] 4.4 UMEDR model performance under different unbiased dataset sizes (RQ3)
[0182] To further verify the unbiased dataset D U In the UMEDR model, this example uses unbiased datasets of different sizes D U To test the model performance, set D U The size of is 1% to 6% of the MAR data in the dataset. The experimental results are as follows Figure 8-9 As shown in the figure, it can be seen that the performance of the UMEDR model on unbiased datasets of 1% to 6% scale is better than other methods that also use unbiased datasets.
[0183] 4.5 UMEDR model performance under different embedding dimensions (RQ4)
[0184] The underlying framework of the UMEDR model is the MF algorithm, and the user-item embedding matrix is obtained through MF, so it is necessary to explore the performance of the UMEDR model under different user-item embedding sizes. In this section, the performance of the UMEDR model on the NDCG@5 index is explored when the user-item embedding size is K=2, 4, 6, 8, and 10. The test results on the Yahoo! R3 and Coat Shopping datasets are shown in the figure. Figure 10-11 As shown, the UMEDR model achieves the best performance when the embedding size is K=10.
[0185] This embodiment mainly studies how to improve the debiasing performance of the DR estimator, and proposes a new meta-learning dual robust debiasing model that integrates unbiased data, aiming to solve the selection bias and data sparsity problems. This embodiment starts from two aspects. First, the traditional DR estimator only uses biased data sets to guide the learning of the filling prediction model and the tendency prediction model. The biased data set has a large selection bias and cannot make the DR estimator achieve the best effect. For this reason, an unbiased data set is introduced to guide the update of the filling prediction model and the tendency prediction model, making the filling prediction model and the tendency prediction model more accurate, thereby improving the overall prediction performance of the DR estimator. For the problem of data sparsity, the model is updated by meta-learning. Meta-learning focuses on using a small amount of data to improve model performance, thereby alleviating the impact of data sparsity. At the same time, the learning method of meta-learning double-layer optimization makes the model parameter learning more sufficient, further improving the model performance.
[0186] For future work, we plan to start with an unbiased dataset and find new data augmentation methods to expand the unbiased dataset to avoid overfitting or insufficient learning of the model due to a small unbiased dataset. At the same time, we will continue to explore more methods to optimize the filling prediction model and the tendency prediction model to further improve the overall performance of the DR estimator.
[0187] Embodiment 2
[0188] This embodiment discloses a recommendation system conversion rate prediction system based on unbiased data and meta-learning.
[0189] Recommendation system conversion rate prediction system based on unbiased data and meta-learning, including:
[0190] The data set acquisition module is configured to: acquire biased data sets from historical data of interactions between multiple users and multiple items in the recommendation system, and acquire unbiased data sets between multiple users and multiple items;
[0191] The training module is configured as follows:
[0192] Build a UMEDR model that includes a filling prediction model, a tendency prediction model, and a conversion rate prediction model, and train the UMEDR model:
[0193] Using an unbiased data set to train a filling prediction model and a tendency prediction model to obtain a trained filling prediction model and a tendency prediction model, and obtaining an unbiased filling error and a user tendency based on the trained filling prediction model and the tendency prediction model respectively;
[0194] Based on biased data sets, unbiased filling errors and user tendencies, a meta-learning method is used to train the conversion rate prediction model to obtain a trained conversion rate prediction model.
[0195] The prediction module is configured to: based on the trained conversion rate prediction model, realize the conversion rate prediction of users in the recommendation system who click on items and then make purchases.
[0196] Embodiment 3
[0197] The purpose of this embodiment is to provide a computer-readable storage medium.
[0198] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning as described in Example 1 of the present disclosure.
[0199] Embodiment 4
[0200] The purpose of this embodiment is to provide an electronic device.
[0201] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in the method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning as described in Example 1 of the present disclosure are implemented.
[0202] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0203] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0204] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A method for predicting conversion rate of recommendation system based on unbiased data and meta-learning, characterized in that: The following steps are involved: Obtain biased data sets from the historical data of interactions between multiple users and multiple items in the recommendation system, and simultaneously obtain unbiased data sets from multiple users and multiple items; Build a UMEDR model that includes a filling prediction model, a tendency prediction model, and a conversion rate prediction model, and train the UMEDR model: Using an unbiased data set to train a filling prediction model and a tendency prediction model to obtain a trained filling prediction model and a tendency prediction model, and obtaining an unbiased filling error and a user tendency based on the trained filling prediction model and the tendency prediction model respectively; Based on biased data sets, unbiased filling errors and user tendencies, a meta-learning method is used to train the conversion rate prediction model to obtain a trained conversion rate prediction model. The filling prediction model is trained using an unbiased dataset, and the loss function is: in is the prediction error, measured by predicting the conversion rate label and the true conversion rate label r u,i It is calculated by the binary cross entropy between; is the filling error, which is measured by filling the pseudo labels generated by the prediction model and predicted conversion rate tags It is calculated by the binary cross entropy between; λ is the regularization term; To predict user preferences; The tendency prediction model is trained using an unbiased dataset, and the loss function is: Among them u,i Indicates the real click rate in the click sample, Indicates D U Predicted user preferences under guidance; The loss function used in the conversion rate prediction model update process is defined as follows: Based on the trained conversion rate prediction model, the conversion rate prediction of users clicking items in the recommendation system is realized; The filling prediction model and the tendency prediction model are trained as follows: A common prediction model is defined as a replica model of the filling prediction model, the tendency prediction model, and the conversion rate prediction model, and the biased data set is divided into a small sample task set, where the data contained in each small sample task is not exactly the same; The initial filling error and initial user tendency are obtained on the small sample task divided by the biased dataset through the filling prediction model and the tendency prediction model; Based on the initial filling error and the initial user tendency, the replica model is updated to obtain the optimal debiasing parameters of the replica model, and the training of the replica model is completed; Based on the trained replica model, the filling prediction model and the tendency prediction model are trained respectively using an unbiased data set to obtain the optimal debiasing parameters of the filling prediction model and the tendency prediction model, and complete the training of the filling prediction model and the tendency prediction model; The process of obtaining the biased data set is as follows: In the recommendation system, determine the set of M users U = {u1,u2,...,u M }, a set of N items I = {i1,i2,...,i N }; Obtain the user's clicks on the items of interest and obtain the set of click samples O = {(u,i)|o u,i =1,(u,i)∈D T }, where each item o u,i ∈{0,1} indicates whether user u clicked on item i. If yes, then o u,i =1, otherwise o u,i =0; (u,i) is a user-item pair; Obtain the actual consumption data of users after clicking on samples in the sample set, and generate the actual conversion rate label matrix R∈{0,1} M×N , where each item r u,i ∈{0,1} reflects the real rating of user u on the recommended item i. The historical data of user interaction on the items of interest is taken as the biased data set D T ; The process of obtaining the unbiased data set is as follows: Obtain the clicks and true ratings of some users on some items, and obtain an unbiased dataset D of some users on some items U .
2. The method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning according to claim 1, characterized in that: When using the meta-learning method to train the conversion rate prediction model, each round of model update extracts different batches of small sample tasks from the small sample task set for learning, ensuring that the small sample tasks used in each round of update are not exactly the same.
3. A recommendation system conversion rate prediction system based on unbiased data and meta-learning, characterized in that: include: The data set acquisition module is configured to: acquire biased data sets from historical data of interactions between multiple users and multiple items in the recommendation system, and acquire unbiased data sets between multiple users and multiple items; The training module is configured as follows: Build a UMEDR model that includes a filling prediction model, a tendency prediction model, and a conversion rate prediction model, and train the UMEDR model: Using an unbiased data set to train a filling prediction model and a tendency prediction model to obtain a trained filling prediction model and a tendency prediction model, and obtaining an unbiased filling error and a user tendency based on the trained filling prediction model and the tendency prediction model respectively; Based on biased data sets, unbiased filling errors and user tendencies, a meta-learning method is used to train the conversion rate prediction model to obtain a trained conversion rate prediction model. The prediction module is configured to: predict the conversion rate of consumption after users click on items in the recommendation system based on the trained conversion rate prediction model; The filling prediction model is trained using an unbiased dataset, and the loss function is: in is the prediction error, measured by predicting the conversion rate label and the true conversion rate label r u,i It is calculated by the binary cross entropy between; is the filling error, which is measured by filling the pseudo labels generated by the prediction model and predicted conversion rate tags It is calculated by the binary cross entropy between; λ is the regularization term; To predict user preferences; The tendency prediction model is trained using an unbiased dataset, and the loss function is: Among them u,i Indicates the real click rate in the click sample, Indicates D U Predicted user preferences under guidance; The loss function used in the conversion rate prediction model update process is defined as follows: The filling prediction model and the tendency prediction model are trained as follows: A common prediction model is defined as a replica model of the filling prediction model, the tendency prediction model, and the conversion rate prediction model, and the biased data set is divided into a small sample task set, where the data contained in each small sample task is not exactly the same; The initial filling error and initial user tendency are obtained on the small sample task divided by the biased dataset through the filling prediction model and the tendency prediction model; Based on the initial filling error and the initial user tendency, the replica model is updated to obtain the optimal debiasing parameters of the replica model, and the training of the replica model is completed; Based on the trained replica model, the filling prediction model and the tendency prediction model are trained respectively using an unbiased data set to obtain the optimal debiasing parameters of the filling prediction model and the tendency prediction model, and complete the training of the filling prediction model and the tendency prediction model; The process of obtaining the biased data set is as follows: In the recommendation system, determine the set of M users U = {u1,u2,...,u M }, a set of N items I = {i1,i2,...,i N }; Obtain the user's clicks on the items of interest and obtain the set of click samples O = {(u,i)|o u,i =1,(u,i)∈D T }, where each item o u,i ∈{0,1} indicates whether user u clicked on item i. If yes, then o u,i =1, otherwise o u,i =0; (u,i) is a user-item pair; Obtain the actual consumption data of users after clicking on samples in the sample set, and generate the actual conversion rate label matrix R∈{0,1} M×N , where each item r u,i ∈{0,1} reflects the real rating of user u on the recommended item i. The historical data of user interaction on the items of interest is taken as the biased data set D T ; The process of obtaining the unbiased data set is as follows: Obtain the clicks and true ratings of some users on some items, and obtain an unbiased dataset D of some users on some items U .
4. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning as described in any one of claims 1 to 2 are implemented.
5. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for predicting conversion rate of a recommendation system based on unbiased data and meta-learning are implemented as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Unbiased machine learning method
CN113077057A
Platform-related advertisement click rate prediction method based on deep learning
CN113689234A