A Dual Enhancement Propensity Score Estimation Method in Sequential Recommendation

Through the dual enhanced propensity score estimation method, combining the Transformer layer and the GRU neural network, the propensity score is estimated from the user and item perspectives, which solves the deviation problem in the sequence recommendation model and improves the accuracy and effectiveness of recommendations.

CN115599972BActive Publication Date: 2025-07-25RENMIN UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211270803.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2025-07-25
Estimated Expiration
2042-10-17

AI Technical Summary

Technical Problem

Existing serial recommendation models have serious bias problems, especially when the test environment is related to office products, and the correlation between users and items cannot be accurately captured, resulting in reduced recommendation effectiveness.

Method used

A dual enhanced propensity score estimation method is proposed. By designing a preamble recommendation model composed of the Transformer layer and the Prediction layer, combining the GRU neural network to estimate the propensity score from the user and the item, and weighted learning is performed through unbiased preamble model training method to optimize the recommendation model.

Benefits of technology

In the unbiased test setup, it significantly outperforms existing baseline models, improving the accuracy and effectiveness of recommendations, especially when users are complex in relation to items.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599972B_ABST
    Figure CN115599972B_ABST
Patent Text Reader

Abstract

Through a method in the field of network security, the present invention realizes a method for estimating dual-enhanced propensity scores in sequential recommendation. When a user behavior occurs, a pre-sequential recommendation model composed of a Transformer layer and a Prediction layer is designed as the architecture of the basic model. The real-valued vector e(u) composed of the context information of the target user-item pair (u, i) collected by the system is used as the input to obtain the final predicted recommendation result for the user behavior. And by designing a network structure for learning propensity scores to weight the user behavior, the recommendation model can obtain an accurate predicted recommendation result for the user behavior. The present invention provides a new IPS estimation method to make up for the exposure or selection bias in sequential recommendation. Evaluating propensity scores from the perspectives of items and users provides theoretical reliability and end-to-end learning. A large number of experimental results on four real datasets show that DEPS can significantly outperform the state-of-the-art baselines under unbiased test settings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for estimating dual-enhanced propensity scores in sequential recommendation. Background Art

[0002] In recent years, sequential recommendation has received increasing attention from the industry and academia. Basically, the key advantage of sequential recommendation models lies in the explicit modeling of the temporal correlation of items. To accurately capture such information, many models based on Markov chains or recurrent neural networks have been proposed in recent years. Although these models have achieved remarkable success, existing sequential recommendation systems have very serious bias problems: for example Figure 1 (a) In the given user behavior sequence, the next item observed is a coffee pot cleaner. By building a model based on the observed data, the correlation between the cleaner and the coffee pot can be understood. However, from the perspective of user preferences, the next item can also be an ink cartridge. But the model has no chance to capture the correlation between the printer and the ink cartridge because they have not been recommended in the data and have not been observed in the training environment. This bias will reduce the effectiveness of the recommendation, especially when the test environment is more relevant to office products.

[0003] To alleviate the above problems, most previous models are based on the inverse propensity score (IPS) technique. If a training sample is more likely to appear in the dataset, then its weight should be lower in the optimization process. In previous studies, given the historical information H, that is, P(u, i|H), the probability of observing the user-item pair (u, i) is accurately approximated. To this end, previous methods usually decompose P(u, i|H) into P(i|u, H)P(u|H), and focus on parameterizing P(i|u, H) (that is, estimating P(u, i|H) from the perspective of the item side) in order to predict the item that the user will interact with given the previous item.

[0004] However, we believe that the probability of observing the user-item pair can also be considered from a dual perspective, that is, for an item, predicting the next interacting user given the users who have interacted with it before. In principle, this is equivalent to decomposing P(u, i|H) in another way into P(u|i, H)P(i|H), where P(u|i, H) is accurately aimed at predicting the user of the given item and the historical users (that is, estimating P(u, i|H) from the user's perspective). Intuitively, for the same item, if two users interact with it for a short time, then they should have some similarities at that time. We believe that this user-oriented method can provide complementary information for the previous item-oriented models.

[0005] For example, in Figure 1In (b), from the perspective of item prediction, the tripod can be regarded as the next item in user sequences A and B because of the similar historical information. However, from the perspective of user prediction, we can infer that it should be easier to observe sequence A because female users have a higher interaction frequency with tripods recently. For example, due to reasons such as women's clothing promotions. This example shows that time user-related signals can well compensate for traditional item-oriented IPS methods.

[0006] A common method to correct bias is through inverse propensity score (IPS). Devooght et al., Hu et al. used previous experience as a sample of propensity scores for unified reweighting. UIR and UBPR proposed using models that estimate propensity scores with latent probabilities. Agarwal et al., Fang et al. utilized intervention signals to learn propensity scores. USR proposed a network to estimate propensity scores from the perspective of items in sequential recommendations. And our invention aims to utilize both user and item sequential information to obtain propensity scores.

[0007] The present invention mainly lies in proposing to establish an unbiased sequential recommendation model with dual-enhanced IPS estimation (referred to as DEPS), aiming to solve the following main challenges: (1) First, how to estimate item-oriented and user-oriented IPS from the same set of user feedback data; (2) Second, how to combine these two features; (3) Finally, how to theoretically ensure that the proposed objective is still unbiased. Summary of the Invention

[0008] For this purpose, the present invention first proposes a method for dual-enhanced propensity score estimation in sequential recommendation, designs a pre-sequential recommendation model composed of a Transformer layer and a Prediction layer when a user behavior occurs as the architecture of the basic model, takes the real-valued vector e(u) composed of the context information of the target user-item pair (u, i) collected by the system as input to obtain the final predicted recommendation result for the user behavior, and assigns weights to the user behavior by designing a network structure for learning propensity scores, and trains the architecture of the basic model based on the weighted data through the training method of the unbiased pre-sequential model combined with weighting, so that the recommendation model can obtain accurate predicted recommendation results for user behavior;

[0009] The network structure for learning propensity scores uses dual GRU neural networks to estimate propensity scores, and estimates the above-mentioned pseudo-propensity scores from two perspectives, the user side and the commodity side, when a user behavior occurs, that is, assigns weights to the user behavior;

[0010] The training method of the unbiased pre-sequential model synchronously learns propensity scores and utilizes the learned propensity scores to obtain an accurate pre-sequential recommendation model.

[0011] The Transformer layer consists of two transformers. One converts the item sequence into a representation vector, and the other Transformer also converts the user sequence into another representation vector;

[0012] Specifically, the overall item representation of the input tuple (u, i, t) of the Transformer layer is the concatenation of the item id embedding and the user historical sequence embedding:

[0013]

[0014] where the operator | represents concatenating two vectors, e(i) is the embedding of the item, is the user sequence related to the target user u at timestamp t, is the vector encoding the sequence, defined as the mean of the vectors output by the transformer:

[0015]

[0016] where Mean is the average pooling operation of all input vectors, and Transformer1 is a transformer architecture;

[0017] Meanwhile, the overall user representation of the input tuple (u, i, t) of the Transformer layer is the concatenation of the user id embedding and the item historical sequence embedding:

[0018]

[0019] e(u) is the embedding of the user, is the user sequence related to the target item i at timestamp t, is the vector encoding the sequence, defined as the mean of the vectors output by the transformer:

[0020]

[0021] After that, the obtained representation is input into the MLP through the Prediction layer to obtain the final prediction:

[0022] 3. The method for estimating the dual enhanced propensity score in sequence recommendation according to claim 1, wherein: the method for estimating the propensity score using the dual GRU neural network is as follows: First, two GRU units are respectively used to estimate the propensity scores from two angles. Among them, GRU1 is the GRU unit that processes the sequence from the item angle, and GRU2 is the GRU unit that processes the sequence from the user angle;

[0023] The method for processing the propensity score of sequence estimation from the item perspective is as follows: Given a tuple which represents that at timestamp t, user u accesses the system and interacts with item i, and the item sequence that user u interacted with before t is The user sequence that interacted with item i before time t is where represents the propensity score estimated from the item perspective, that is, it is represented as the embedding e(i) of the item expression and the output of the last layer in the GRU network, that is, the i l(u,t) -th. The sequence is used as the input of the GRU network. Therefore, we can write the estimated propensity score from the item perspective as:

[0024]

[0025] where y(i l(u,t) ) is the output of its last layer of GRU, that is, the output corresponding to the GRU of layer l(u,t), and l(u,t) represents the number of items that user u interacted with before time t. The GRU scans the items in as follows: At the k-th layer, the k-th item e(i k ) of the item embedding is required as the input embedding and outputs y(i k ). The representation of each layer is as follows, k = 1,..., l(u,t):

[0026] y(i k ), z k = GRU1(e(i k ), z k-1 )

[0027] where z k and z k-1 are the hidden vectors of the k-th and k-1 steps;

[0028] where represents the propensity score estimated from the item perspective, that is, it is represented as the dot product of the embedding e(u) of the user expression and the output of the last layer of the i l(u,t) -th. The sequence is used as the input of the GRU network. We write the estimated propensity score from the user perspective as:

[0029]

[0030] where y(i l(i,t) ) is the output of its last layer, that is, the i l(u,t) -th layer of GRU, that is, the output corresponding to the GRU of layer l(i,t). The GRU scans The users in it are as follows: At the k-th layer, k = 1, …, l(i, t), where l(i, t) represents the number of users who have interacted with item i before time t, and the k-th item e(u k ) is required as the input embedding and outputs y(u k ), which represents the subsequence [u 1 , u 2 , …, u k :

[0031] y(u k ), z k = GRU2(e(u k ), z k-1 )

[0032] where z k and z k-1 are the hidden vectors at the k-th and (k - 1)-th steps.

[0033] 4. A method for estimating dual-enhanced propensity scores in sequential recommendation according to claim 1, characterized in that: the training method of the unbiased pre-order model first requires a set of parameters to be learned, denoted as Θ = {θ e , θ p , θ t , θ m}, where the parameter θ e represents the model that outputs user and item embeddings in the embedding, the parameter θ p in GRU1 and GRU2 used to estimate the propensity score, the parameter θ t in the transformer, and the parameter θ m in the MLP are used to make the final recommendation;

[0034] Design a two-stage learning process to learn the model parameters. In the first stage, the parameters of {θ e , θ p , θ t} are trained in an unsupervised learning manner to achieve good initialization. In the second stage, all parameters with the above unbiased learning objectives are learned, and the Adam optimizer is used for optimization;

[0035] The specific algorithm process is as follows: The dataset to be trained is defined as The number of learning rounds is defined as n p , n u , n b , the loss weight is defined as λ p , and Θ is randomly initialized;

[0036] First, perform the pre-training process in the first stage to optimize the pre-training loss function The number of loop rounds is n p , in each loop round, first update the parameter θ p through the optimization function: This function utilizes the Autoregressive model in deep learning. Among them

[0037]

[0038]

[0039] ; then update the parameter θ e , θ t through the optimization function: where

[0040]

[0041] is the user sequence related to the target user u at timestamp t, and respectively represent the set of users and items in the system, l(u, t) represents the number of users who interacted with user u before time t, and l(u, t) represents the number of users who interacted with user u before time t;

[0042] Then perform unbiased learning in the second stage, with the number of loop rounds being n u , and the number of nested loop rounds n b in each loop round, execute to update the parameter θ p through the above optimization function: for obtaining a stable propensity score estimation function, where

[0043]

[0044]

[0045] is the user sequence related to the target user u at timestamp t, and respectively represent the set of users and items in the system, l(u, t) represents the number of users who interacted with user u before time t, and l(u, t) represents the number of users who interacted with user u before time t; then update the parameter θ through the optimization function: Update the parameter θ e , θ t , θ m , where

[0046]

[0047]

[0048] The final output parameter Θ = {θ e , θ p , θ t , θ m}.

[0049] The technical effect to be achieved by the present invention is as follows:

[0050] The present invention proposes a new IPS estimation method called Dual-Enhanced Propensity Score Estimation (DEPS) to address the exposure or selection bias in sequential recommendation. DEPS evaluates the propensity score from both item and user perspectives and offers several advantages: theoretical soundness and end-to-end learning. Extensive experimental results on four real-world datasets show that DEPS can significantly outperform state-of-the-art baselines under unbiased test settings. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 Schematic illustration of the bias problem in sequential recommendation;

[0052] Figure 2 Causal graph based on time series;

[0053] Figure 3 Dual propensity score estimation network;

[0054] Figure 4 Base model for sequential recommendation based on Transformer encoding;

[0055] Figure 5 Dataset statistics;

[0056] Figure 6 Overall experimental results;

[0057] Figure 7 Ablation study on the Amazon-Digital-Music test set;

[0058] Figure 8 Comparison with non-sequential IPS methods;

[0059] Figure 9 Model-Agnostic experiments;

[0060] Figure 10 Experiments on the clipping value M; DETAILED DESCRIPTION OF THE INVENTION

[0061] The following are the preferred embodiments of the present invention in combination with the accompanying drawings, and the technical solutions of the present invention will be further described, but the present invention is not limited to this embodiment.

[0062] Embodiment 1: This embodiment takes the movie recommendation scenario as an example. At the same time, the method for estimating the dual-enhanced score of sequential recommendation can be widely applied to Internet shopping platforms to improve the training effect of product recommendation models and other similar scenarios.

[0063] For the movie recommendation scenario, the present invention proposes a method for estimating the dual-enhanced propensity score in sequential recommendation. First, the user logs in to the platform, and the platform will automatically recommend a list of movies the user likes based on the user's historical behavior. The user watches the movies by clicking on the recommended movies. At the same time, the platform applies a method for estimating the dual-enhanced propensity score in sequential recommendation to record the user's click behavior as training data for further learning the parameters of the recommendation model in the future. Then, using a method for estimating the dual-enhanced propensity score in sequential recommendation, the recommendation model is trained using the user's click behavior to form a more accurate movie recommendation list, which can improve the effective conversion rate of movies on the entire website, increase the viewing rate and viewing duration of movie recommendations, and enhance the revenue of the platform.

[0064] The key to the success of the above process is that the user's click behavior can accurately reflect the user's true preference for movies, that is, the movies that have been clicked are the ones the user really likes, and the movies that have not been clicked represent that the user does not like them. However, in real-world recommendations, the user's click behavior is also affected by many factors. For example, users tend to trust more the movies ranked at the top, and are inclined to click even if they are not relevant; if the recommended movies are placed at the bottom of the sequence or on the next page, even if the movies will be liked by the user, they will not be clicked because the user does not notice them. We call the gap between the above click behavior and the true preference the "bias". How to remove the above "bias" is the key to accurate movie recommendation. An effective method to eliminate the bias is to assign a weight value to the user's click behavior using the inverse propensity score. If a movie is ranked at the top of the recommendation result, the behavior of its being clicked will be assigned a relatively small weight value, while if a movie ranked at the bottom of the sequence is clicked, it will be assigned a relatively large weight value. Training the recommendation model with the re-weighted user click data will achieve better results than directly using the user click data to train the model.

[0065] This patent focuses on how to calculate the above-mentioned pseudo-propensity score in the sequential recommendation scenario (such as: when the user visits the platform to watch movies multiple times, how to recommend movies to him each time) to train a better recommendation model and obtain better recommendation results. Specifically, when observing that a user clicks on a movie, we obtain the recommended movies for the user from the user side and the product side (movie).

[0066] Specifically, a method for estimating the dual-enhanced propensity score in sequential recommendation first needs to define the symbols of the model. Assume that a sequential recommendation system manages a set of user-item (movie) historical interactions Each tuple (u, i, c t ) records that at timestamp t, user accesses the system and interacts with an item (movie) , and the user's feedback is c t ∈ {0, 1}, where and represent the sets of users and items (movies) in the system respectively, and c t = 1 indicates that user u clicks on item (movie) i at time t, otherwise 0. Additionally, the context information of (u, i) collected from the system is typically represented as real-valued vectors (embeddings) e(u), e(i) ∈ R d , where d represents the dimension of the embedding.

[0067] At a specific time t, given a target user-item (movie) pair (u, i), two types of interaction sequences can be derived from : (1) The sequence of items (movies) that user u has interacted with before t: where l(u, t) represents the number of items (movies) that user u has interacted with before time t; (2) The sequence of users who have interacted with item (movie) i before time t: where l(i, t) represents the number of users who have interacted with item (movie) i before time t.

[0068] The task of sequential recommendation becomes, based on the user-item (movie) interactions and user feedback in A sequential recommendation model for item (movie) sequences that dual-enhances the propensity score estimation can be represented as a function t to predict user u's preference for item (movie) i at time t. It is expected that the predicted preference is close to the true but unobservable user preference r t ∈ {0, 1}, where r

[0069] In sequential recommendation, bias occurs when user u is systematically under-exposed / over-exposed to certain items (movies). As Figure 2 shown, in a causal sense, r t = 1 only when item (movie) i is relevant to user, and item (movie) i is exposed to u (o uit = 1) the user will click on item (movie) c t = 1 at time t, formally c t = r t · o uit. Further assume that two interaction sequences and also affect whether user u knows item (movie) i. Since the model predicts user preference r t by observing the clicks c t , the prediction is inevitably affected by the item (movie) exposure o uit . This is because after observing the clicks in the causal graph, o uit becomes a confounding factor.

[0070] For the unbiased training of the recommendation model, for the historical interactions of a given user-item (movie), we define the ideal learning objective of sequential recommendation as evaluating the preference every time the user visits the system:

[0071]

[0072] We prove that Theorem 1 - the ideal learning objective can be represented by two dual loss functions:

[0073]

[0074] where and can be expressed as

[0075]

[0076]

[0077] For the variance problem of the model, we apply the IPS clipping method:

[0078]

[0079]

[0080] Meanwhile, we theoretically prove that the loss function of our model training can be controlled by the hyperparameter M of gradient clipping (where this M is a hyperparameter we defined, and in practice, we usually adjust it to a certain value in 0.05 - 0.2 according to different data):

[0081]

[0082] where We conclude that weighting the loss does not introduce additional variance. Intuitively, the clipping value M is a trade - off between unbiasedness and variance and provides a mechanism to control variance.

[0083] Based on the above definitions, first, the design of the basic model architecture is carried out. The basic model architecture is the sequential recommendation model, which is composed of a Transformer layer and a Prediction layer. The Transformer layer consists of two transformers. One converts the item (movie) sequence into a representation vector, and the other also converts the user sequence into another representation vector. The Prediction layer connects the vectors and uses an MLP for prediction. The overall architecture is shown in Figure 4 .

[0084] For the overall item (movie) representation of the input tuple (u, i, t) of the Transformer layer, it can be expressed as the concatenation of the item (movie) id embedding and the user historical sequence embedding:

[0085]

[0086] where the operator | represents concatenating two vectors, e(i) is the embedding of the item (movie), is the user sequence related to the target user u, is the vector encoding the sequence, which is defined as the mean of the vectors output by the transformer:

[0087]

[0088] where `Mean' is the average pooling operation of all input vectors, and Transformer1 is a transformer architecture.

[0089] Similarly, the overall user representation of the input tuple (u, i, t) of the Transformer layer can be expressed as the concatenation of the user id embedding and the item (movie) historical sequence embedding:

[0090]

[0091] e(u) is the embedding of the user, is the user sequence related to the target item (movie) i, is the vector encoding the sequence, which is defined as the mean of the vectors output by the transformer:

[0092]

[0093] For the Prediction layer, we input the obtained representation into the MLP to get the final prediction:

[0094]

[0095] After that, a network structure for learning the propensity score is designed to weight the user behavior, and the architecture of the basic model is learned based on the weighted data through the training method of combining the weight and the unbiased pre-order model, so that the recommendation model can obtain an accurate prediction recommendation result of the user behavior.

[0096] Specifically, a dual GRU neural network is used to estimate the propensity score. The overall model framework diagram is shown in Figure 3 . Two GRUs are respectively used to estimate the propensity scores from two perspectives. Given a tuple From the perspective of an item (movie), its propensity score is estimated as the maximum value M and the value of, where is proportional to the dot product of the embedding e(i) expressed by the item (movie), and the sequence is used as the input. By applying the clipping technology as described above, we write the estimated propensity score from the item (movie) perspective as:

[0097]

[0098] where y(i l(u,t) ) is the output of the last layer of the GRU (i.e., the output corresponding to the GRU of the l(u,t) layer). The GRU scans the items (movies) in as follows: at the kth (k = 1,..., l(u,t)) layer, it requires the kth item e(i k ) as the input embedding and outputs y(i k ), which represents the subsequence [i 1 , i 2 , …, i k :

[0099] y(i k ), z k =GRU1(e(i k ), z k-1 )

[0100] where z k and z k-1 are the hidden vectors at the kth and k - 1 steps, and GRU1 is the GRU unit that processes the sequence from the item (movie) perspective.

[0101] Similarly, we can estimate the propensity score from the user perspective according to the above symbol definition:

[0102]

[0103] where y(i l(i,t)) is the output of the GRU at its last layer (i.e., the output of the GRU corresponding to layer l(i,t)). The GRU scans the items (movies) in as follows: At the k-th (k = 1, …, l(i,t)) layer, it requires the k-th item e(u k ) as the input embedding and outputs y(u k ), which represents the subsequence [u 1 , u 2 , …, u k :

[0104] y(u k ), z k = GRU2(e(u k ), z k-1 )

[0105] where z k and z k-1 are the hidden vectors at the k-th and k - 1 steps, and GRU2 is the GRU cell that processes the sequence from the user's perspective.

[0106] The specific content of the training method for the unbiased pre - order model is as follows: For the proposed model, there is a set of parameters to be learned, denoted as Θ = {θ e , θ p , θ t , θ m}, where the parameter θ e represents the model that outputs the user and item (movie) embeddings in the embedding, the parameter θ p in GRU1 and GRU2 for estimating the propensity score, the parameter θ t in the transformer, and the parameter θ m in the MLP are used to make the final recommendation.

[0107] Inspired by the pre - training then fine - tuning paradigm, we also designed a two - stage learning process to learn the model parameters. In the first stage, the parameters of {θ e , θ p , θ t} are trained in an unsupervised learning manner to achieve good initialization. Then, in the second stage, all the parameters with the above - mentioned unbiased learning objective are learned. The Adam optimizer is used for optimization.

[0108] For the first stage, we apply n p epochs to optimize For the second stage, we apply n u epochs to alternative training. The entire algorithm flow can be seen in the DEPS algorithm flow.

[0109] The training process of this method:

[0110]

[0111]

[0112] Algorithm Flow: DEPS

[0113] Among them, the loss function of pre-training is defined as follows:

[0114]

[0115] For the propensity score estimation module, we apply the auto-regressive training method:

[0116]

[0117]

[0118] For the transformer, we apply the form of masked language model:

[0119]

[0120] Example 2:

[0121] According to the settings of the unbiased experiment, we maintained the same setting conditions. The experimental settings of Example 2 are as follows:

[0122] These experiments were conducted on four publicly available large-scale sequential recommendation benchmarks:

[0123] MIND: A large-scale news recommendation dataset. Users / items with fewer than 5 item / user interactions have been removed to avoid extremely sparse situations.

[0124] Amazon-Beauty / Amazon-Digital-Music: Two subsets of the Amazon product dataset (in the beauty and digital music fields). Similarly, users / items with fewer than 5 item / user interactions were removed. We regarded users' 4-5 star ratings on the Amazon dataset as positive feedback (labeled as 1), and others as negative feedback (labeled as 0).

[0125] Huawei Dataset: To verify the effectiveness of our method on production data, we collected 1-month traffic logs from the Huawei music service system, and there were approximately 245K interactions after sampling.

[0126] Figure 5Lists the statistics of four datasets. The debiased recommendation model is evaluated based on an unbiased test set, trained using the top 50% interactions sorted by interaction time, and the other 50% of the data is resampled for evaluation and testing. Specifically, assuming that item i is clicked m i times, we use the inverse probability m i / max j m j for sampling. Then we use 20% and 30% of the sorted data for validation and testing respectively.

[0127] For the hyperparameters in all models, the learning rate is in [1e-3, 1e-4] and the clipping coefficient M for propensity score estimation is adjusted in the range of [0.01, 0.2]. The trade-off coefficient in the first stage is set to 0.5. The hidden dimension d of the neural network is adjusted among {64, 128, 256}.

[0128] And we selected the following representative sequential recommendation models as baseline models:

[0129] STAMP simulates users' long-term and short-term preferences; GRU4Rec+ is an improved version of GRU4Rec with data augmentation and considering the changes in the input; BERT4Rec uses an attention module to simulate users' behavior and is trained in an unsupervised manner; FPMC captures users' preferences by combining matrix factorization with a first-order Markov chain; DIN applies an attention module to adaptively learn users' interests from their historical behaviors; BST applies a transformer architecture to adaptively learn users' interests from historical behaviors and side information of users and items; LightSANs is a low-rank factorization-based SANs recommendation model. We also selected the following unbiased recommendation models as baselines: UIR is an unbiased recommendation model that uses heuristic estimation of propensity scores; CPR is a pairwise debiasing method for exposure bias; UBPR is an IPS method for non-negative pairwise loss. DICE: A debiasing model focusing on user communities. USR: A debiasing sequential model aiming to mitigate the bias caused by potential confounding factors.

[0130] Main experimental results:

[0131] Figure 6 Reports the experimental results of DEPS and the baselines on all four datasets, measured by NDCG@K and HR@K for recommendation accuracy. '*' indicates that the improvement relative to the best baseline is statistically significant (t-test and p-value < 0.05). Underline indicates the method with the best performance.

[0132] As can be seen from the reported results, on Huawei commercial data, DEPS significantly outperforms almost all baselines in terms of NDCG and the expected NDCG@5 of HR, verifying the effectiveness of DEPS in improving the accuracy of sequential recommendations. In addition, DEPS significantly outperforms the unbiased model, demonstrating the importance of estimating propensity scores from the perspective of item and user sequential recommendations.

[0133] Ablation experiment

[0134] To further illustrate the importance of estimating the propensity scores of the two types of sequences from the perspectives of users and items, we also studied their unbiased performance in the second-stage training when has been optimized. Specifically, we show the NDCG@K and HR@K of several DEPS variants. These variations include learning a recommendation model without propensity score estimation (denoted as "w / o IPS"), estimating the propensity score only from the perspective of the item sequence ("w / o user-oriented IPS"), and using only the user sequence view ("w / o item-oriented IPS"). From the performance of the figure, we found that (1) "w / o propensity" performs the worst, indicating the importance of propensity scores in unbiased sequential recommendations; (2) "w / o user-oriented IPS" and "w / o item-oriented IPS" perform much better, indicating that the propensity scores estimated from these two perspectives are effective; (3) DEPS with dual propensity scores performs the best, verifying the effectiveness of DEPS by using both perspectives for propensity score estimation.

[0135] In this section, we studied the impact of the sequential IPS estimation method compared with the non-sequential IPS estimation method. In our model, user-oriented and item-oriented IPS are estimated by GRU. We compared it with the non-sequential IPS estimation method (frequency-based propensity score). The user-oriented IPS is calculated as p u,* =m u / max u' m u' , and the item-oriented IPS is calculated as: p *,i =m i / max i' m i' , where m u and m i represent the number of interactions between user u and item i.

[0136] When the unbiased loss function When optimized, we studied their unbiased performance in the second stage of training. Specifically, we presented the NDCG@K and HR@K of several DEPS variants, including using a single propensity score p *,i (denoted as "Item-Pro") to learn the recommendation model, using a single propensity score p u,* (denoted as "User-Pro"), and having a dual propensity score p *,i , p u,* (i.e., replacing the estimated propensity score in DEPS with with p *,i , p u,* denoted respectively as: Item-User-Pro:).

[0137] From the Figure 8 performance shown, we found that (1) "DEPS" was significantly better than "Item-User-Pro", indicating the importance of estimating propensity scores sequentially; (2) "Item-Pro" and "User-Pro" performed worse than "Item-User-Pro", indicating that the propensity scores estimated from two perspectives (users or items) were effective and complementary.

[0138] Although DEPS designed a Transformer-based model for recommendation, it can also be used as a model-agnostic framework by replacing the underlying model with other model sequential recommendation models. In the experiment, we replaced it with GRU4Rec+ and FPMC, implementing two new models, denoted as "DEPS(GRU4Rec+)" or "DEPS(FPMC)". From the Figure 9 reported results, we found that DEPS(GRU4Rec+) and DEPS(FPMC) achieved improvements on the GRU4Rec+ and FPMC base models respectively. The results showed that the propensity scores estimated by DEPS were general. They could be used to improve other sequential recommendation models in a model-agnostic way.

[0139] Finally, according to Theorem 2, the clip value M balanced the unbiasedness and variance in DEPS. In this experiment, we studied how NDCG@K and HR@K changed when the clip value M was set to different values from [0.05, 0.2]. From the Figure 10 curves shown, we found that the performance improved when M ∈ [0.01, 0.05], and then decreased between [0.05, 0.2]. The results verified the theoretical analysis that too small M (e.g., M = 0.01) would lead to large variance estimates, while too large M (e.g., M = 0.2) would lead to large biases. Balancing unbiasedness and variance is important in practical applications.

Claims

1. A method for estimating dual-enhanced propensity scores in sequence recommendation, characterized in that: When a user behavior occurs, the previous recommendation model composed of a Transformer layer and a Prediction layer is used as the architecture of the basic model. The real-valued vector e(u) composed of the context information of the target user-item pair (u, i) collected by the system is used as the input to obtain the final predicted recommendation result for the user behavior. And a network structure for learning the propensity score is designed to weight the user behavior. Based on the weighted data, the architecture of the basic model is learned by combining the training method of the unbiased previous model until the recommendation model obtains an accurate predicted recommendation result for the user behavior; For the network structure for learning the propensity score, a dual GRU neural network is used to estimate the propensity score. When a user behavior occurs, the propensity score is estimated from two perspectives: the user side and the commodity side, which is to weight the user behavior; For the training method of the unbiased previous model, the propensity score is learned synchronously and the learned propensity score is utilized to obtain an accurate previous recommendation model; The method of using the dual GRU neural network to estimate the propensity score is as follows: First, two GRU units are respectively used to estimate the propensity scores from two perspectives. Among them, GRU1 is a GRU unit that processes the sequence from the item perspective, and GRU2 is a GRU unit that processes the sequence from the user perspective; The method for processing the propensity score of sequence estimation from the item perspective is as follows: Given a tuple which represents that at timestamp t, user u accesses the system and interacts with item i, and the sequence of items that user u interacted with before t is and the sequence of users who interacted with item i before time t is where represents the propensity score estimated from the item perspective, that is, it is expressed as the embedding e(i) of the item expression and the output of the last layer in the GRU network, that is, the i l(u,t) -th is used as the input of the GRU network, and thus the estimated propensity score from the item perspective is written as: where y(i l(u , t) ) is the output of the GRU at its last layer, that is, the output corresponding to the GRU of layer l(u,t), where l(u,t) represents the number of items interacted by user u before time t. The GRU scans the items in as follows: at the k-th layer, the k-th item e(i k ) of the item embedding is used as the input embedding and outputs y(i k ), and its representation at each layer is as follows, k = 1, …, l(u,t): y(i k ),z k = GRU1(e(i k ),z k-1 ) where z k and z k-1 are the hidden vectors of the k-th and (k-1)-th steps; where is the propensity score estimated from the item perspective, i.e., it is expressed as the dot product of the embedded e(u) expressed by the user and the output of the i l(u,t) in the last layer of the GRU network, and the sequence is used as the input of the GRU network, and the estimated propensity score from the user perspective is written as: where y(u l(i,t) ) is its last layer, i.e., the output of the GRU at the i l(u,t) -th layer, which corresponds to the output of the GRU at layer l(i,t). The GRU scans the users as follows: at the k-th layer, k = 1, …, l(i,t), where l(i,t) represents the number of users who have interacted with item i before time t. The k-th item e(u k ) is required as the input embedding and outputs y(u k ), which represents the subsequence [u 1 , u 2 , …, u k : y(u k ), z k = GRU2(e(u k ), z k-1 ) where z k and z k-1 are the hidden vectors of the k-th and k-1-th steps.

2. The dual-enhanced propensity score estimation method in sequence recommendation according to claim 1, wherein: The Transformer layer consists of two transformers. One transformer converts the item sequence into a representation vector, and the other transformer also converts the user sequence into another representation vector. Finally, the two vectors are concatenated and input into the MLP predictor to obtain the final user-item preference score; Specifically, the overall item representation of the input tuple (u, i, t) of the Transformer layer is the concatenation of the item id embedding and the user historical sequence embedding: where the operator | represents concatenating two vectors, e(i) is the embedding of an item, is the user sequence related to the target user u at timestamp t, is the vector encoding the sequence, defined as the mean of the vectors output by the transformer: where Mean is the average pooling operation of all input vectors, and Transformer1 is a transformer architecture; Meanwhile, the overall user representation of the input tuple (u, i, t) of the Transformer layer is the concatenation of the user id embedding and the item historical sequence embedding: e(u) is the embedding of the user, is the user sequence related to the target item i at time stamp t, is the vector encoding the sequence, defined as the mean of the vectors output by the transformer: After that, the obtained expression is input into the MLP through the said Prediction layer to obtain the final prediction:

3. The method for estimating the dual enhanced propensity score in sequence recommendation according to claim 2, wherein: The training method of the unbiased pre - order model first has a set of parameters to be learned, denoted as Θ = {θ e , θ p , θ t , θ m}, where the parameter θ e represents the parameter in the embedding that outputs the user and item embeddings of the model, the parameter θ p in GRU1 and GRU2 used to estimate the propensity score, the parameter θ t in the transformer, and the parameter θ m in the MLP are used to make the final recommendation; Design a two-stage learning process to learn the model parameters. In the first stage, the parameters of {θ e , θ p , θ t} are trained in an unsupervised learning manner to achieve a good initialization. In the second stage, all parameters with an unbiased learning objective are learned, and the Adam optimizer is used for optimization; The specific algorithm process is as follows: The dataset to be trained is defined as where N is the number of data entries; the number of learning rounds is defined as n p , n u , n b , and the loss weight is defined as λ p , and Θ is randomly initialized; First, perform the pre-training process in the first stage to optimize the pre-training loss function The number of cycles is n p , in each cycle, first update the parameter θ p , that is, through the optimization function: where ; Update the parameter θ later e , θ t , that is, by means of the optimization function: where is the user sequence related to the target user u at time stamp t, and respectively represent the set of users and items in the system. l(u, t) represents the number of users who have interacted with user u before time t, and l(u, t) represents the number of users who have interacted with user u before time t; Then, the second - stage unbiased learning is carried out, and the number of loop rounds is n u , and the number of nested rounds n b in each loop is used to execute the parameter update θ p By optimizing the above - mentioned function: to obtain a stable propensity score estimation function, where is the user sequence related to the target user u at timestamp t, and respectively represent the sets of users and items in the system. l(u, t) represents the number of users who have interacted with user u before time t, and l(u, t) represents the number of users who have interacted with user u before time t. Then, through the optimization function: update the parameter θ e , θ t , θ m , where The final output parameter Θ = {θ e , θ p , θ t , θ m}.

Citation Information

Patent Citations

  • Click rate prediction method based on heterogeneous behaviors in session

    CN114529077A

  • GCN-GRU-based strip mine truck staying area activity identification method

    CN115062713A