An item recommendation method based on time-aware inverse propensity score
By adopting a time-aware inverse tendency scoring method based on time perception and combining a factor decomposition machine of deep learning, the problem of inaccurate user interest caused by exposure deviation in the traditional recommendation method is solved, and a project recommendation list that is more in line with user expectations is achieved.
Patent Information
- Application Number
- CN202211044162.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-08-30
AI Technical Summary
Due to the exposure deviation problem, traditional recommendation methods cannot accurately grasp the user's real interests, resulting in a huge gap between the recommendation list and user satisfaction.
The project recommendation method based on time-aware inverse tendency scores is adopted. By collecting time-stamped data, extracting historical interaction characteristics, calculating the relative time intervals in the time series, and using a deep learning factor decomposing machine for model training, a project recommendation list with descending click-through rate is generated.
By fully considering the causal relationship between exposure and click between data, the generated project recommendation list is more in line with user expectations, reducing the probability of exposure deviation of the recommendation list.
Smart Images

Figure CN115329202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet, and in particular to a project recommendation method based on time-aware inverse propensity score. Background Art
[0002] Traditional recommendation methods usually build user preference prediction models based on observed user-item interaction data and formulate recommendation lists. However, observed interaction data may lead to exposure bias, which cannot accurately grasp the real interests of users. On the other hand, there is a huge gap between recommendation lists and user satisfaction. Usually, users will click on an item because of an attractive title / cover, but the clicked item does not meet the user's preferences. At present, causal theory and deep learning have performed well in eliminating exposure bias. For example, DICE creates a prediction model by separating user interests and item exposure into two embeddings and combining them together to capture more causal relationships from user interaction history. However, similar methods only use character-level representations of text calculated based on deep learning, and do not reflect the causal relationship between data. FAWMF introduces counterfactual exposure variables to represent the probability of a specific item being exposed to a user and uses it as the confidence weight of personalized data to reduce the bias effect of the recommendation model. However, the debiasing ability of the model directly depends on the accuracy of the confidence, and the latent features of the user's historical interaction records do not take into account the causal effect between exposure and clicks in the user's interaction history. By studying the causal relationship between exposure and clicks in user interaction history under the premise of the existence of correlation in the data, the authenticity and effectiveness of the recommendation list can be improved, and it is of practical significance to make high-quality recommendations. Summary of the invention
[0003] In order to solve the problems existing in the above-mentioned prior art, the present invention provides an item recommendation method based on time-aware inverse propensity score, which is intended to solve the problem that there is a large gap between the recommendation list obtained by the traditional recommendation method and the user's expectation.
[0004] A project recommendation method based on time-aware inverse propensity score, characterized by comprising the following steps:
[0005] Step 1: Collect news recommendation data and short video recommendation data with timestamps, and pre-process the news recommendation data and short video data with timestamps to obtain exposure-click data;
[0006] Step 2: Extract features from the exposure-click data to obtain historical interaction features;
[0007] Step 3: Extract items and timestamps from historical interaction features and calculate relative time intervals in the time series;
[0008] Step 4: Divide the historical interaction records into positive and negative samples based on the time interval and the set threshold;
[0009] Step 5: Perform equal probability sampling on positive and negative samples to obtain an unbiased test set;
[0010] Step 6: Based on the historical interaction characteristics, the inverse propensity score calculation method is used to calculate the user's exposure preference characteristics and the item's exposure tendency characteristics;
[0011] Step 7: Combine the user's historical interaction characteristics and exposure tendency characteristics to train the model using a factorization machine based on deep learning and calculate the click probability of whether item i is clicked by user u after being exposed;
[0012] Step 8: Verify the debiasing effect of the factorization machine based on the unbiased test set;
[0013] Step 9: Generate a list of recommended items in descending order of click-through rate based on the trained factorization machine.
[0014] The present invention calculates the exposure tendency through an inverse tendency score calculation method, and the inverse tendency score calculation method is related to the project exposure characteristics, the project exposure tendency, and the user project exposure characteristic preference, so that the method proposed by the present invention fully considers the causal relationship between exposure and clicks between data, so that the project recommendation list generated by the present invention is more in line with user expectations and the exposure deviation probability of the recommendation list; therefore, the problem that the recommendation list obtained by the traditional recommendation method is significantly different from the user expectation is solved.
[0015] Preferably, step 3 comprises the following steps:
[0016] Step 3.1: Extract the items and timestamps in the historical interaction features to obtain data with time intervals in the historical interaction sequence, and sort the historical interaction features based on the time series to obtain historical interaction feature sequence data;
[0017] Step 3.2: For the historical interaction feature sequence data, calculate the relative time interval between two adjacent historical interaction records according to the time nodes of the historical interaction features; and calculate the earliest time and the latest time of all clicked items of the user according to the user group, and calculate the total time span of the user's interaction records:
[0018] The calculation formula of the relative time interval is:
[0019]
[0020] Among them, T * is the relative time interval, T i is the timestamp of the current item the user is interacting with, T i-1is the timestamp of the user's last interaction item, and T is the total time span of the user interaction record.
[0021] Preferably, step 4 comprises the following steps:
[0022] Step 4.1: Take the interaction records whose relative time interval in the historical interaction features is less than the time length threshold as negative samples;
[0023] Step 4.2: Take the interaction records whose relative time interval in the historical interaction features is greater than the time length threshold as positive samples.
[0024] Preferably, step 6 comprises the following steps:
[0025] Step 6.1: Calculate item clicks based on the time decay function:
[0026]
[0027] Where Y u,i,t To calculate the time decay of item clicks, is the time decay function, which decays exponentially; β is the cooling coefficient, which is a hyperparameter used to control the intensity of time decay; Y u,i The number of item clicks without calculating time decay;
[0028] Step 6.2: Sum the item clicks that take time decay into account and divide it by the maximum exposure of the item in the data to obtain the item exposure tendency and user exposure preference:
[0029]
[0030]
[0031] Where: p i,t is the exposure tendency of item i at time t, Y u,i,t is the number of item clicks with time decay calculated, ∑ is the summation symbol, represents the maximum exposure among all items, τ is the smoothing term of item exposure tendency; p u,t Expose preferences to users; is the cumulative sum of item clicks with time decay; |I u | is the number of items representing user interactions;
[0032] Step 6.3: Based on the item exposure tendency and user exposure preference, the average value of the individual user and the items they interact with is re-summed to obtain the final representation of the user exposure preference:
[0033] θ u,i,t =α-|p u,t -pi,t |;
[0034] Where: θ u,i,t represents the improved propensity score, α is the threshold used to limit the upper limit of exposure tendency, and p u,t is the user’s exposure preference, p i,t is the exposure tendency of the project, and the tendency score θ u,i,t It is expressed as the matching degree between the user's exposure preference and the item's exposure tendency. The more matching the user's exposure preference and the item's exposure degree are, the greater the probability that the item will be clicked.
[0035] Step 6.4: Convert the propensity score θ u,i,t Concatenate with the historical interaction features in the training set, validation set, and test set respectively:
[0036] h′=[u,i,t,Y u,i ,θ u,i,t ];
[0037] Where: θ u,i,t is the improved propensity score, which takes a value between [0,1]; u is the user feature; i is the item feature; t is the timestamp; Y u,i Indicates whether the sample is clicked, 0 means no click, 1 means click; [u, i, t, Y u,i ,θ u,i,t ] is the concatenation of the one-dimensional vector of the project exposure propensity score and the user's historical interaction characteristics, and j′ is the concatenated user-project historical interaction record.
[0038] Preferably, step 7 comprises the following steps:
[0039] Step 7.1: Calculate the absolute value distance between the feature vector of the item exposure tendency and the user's exposure preference feature, and subtract the absolute value of the threshold and the difference as the initial tendency score;
[0040] Step 7.2: Based on the initial propensity score, calculate the prediction score error of the deep learning-based factorization machine;
[0041] Step 7.3: Calculate the prediction score of the deep learning-based factorization machine based on the exposure preference feature and the exposure tendency feature of the item;
[0042] Step 7.4: Calculate the rating of the project being exposed based on the predicted rating error and the predicted rating.
[0043] Preferably, the calculation formula of the prediction score error is as follows:
[0044]
[0045] Where: is the exposure rate of positive samples; is the exposure rate of negative samples, D is the total number of items, θ u,i is the propensity score based on the matching of the item exposure propensity vector and the user's exposure preference, also expressed as the exposure probability of the sample (u,i), ∑ is the sum, It is represented as a positive sample, that is, the item recommended at the current time point not only meets the user's exposure preference but also will be clicked. is the negative sample matched by user-item exposure, indicating that the recommended item will not be clicked at the current time; Loss is the prediction score error, r is the prediction score without the introduction of inverse propensity score in DeepFM, Introducing the inverse propensity score into DeepFM’s prediction score; Regularization term to avoid overfitting.
[0046] Preferably, the calculation formula of the prediction score is as follows:
[0047]
[0048] Where: U u is the user potential feature matrix U∈R n×d , V i is the potential feature matrix V∈R of the project m×d , The predicted score after calculating the gradient according to the loss function and updating the potential feature matrix UV, * is the dot product symbol; T represents the transpose of the vector, U v *V i T Represents the product of matrices, randomly takes a sample from the time series, calculates the gradient according to the loss function, and updates the potential feature matrix.
[0049] The beneficial effects of the present invention include:
[0050] 1. The present invention calculates the exposure tendency through an improved inverse tendency score calculation method, and the inverse tendency score calculation method is related to the project exposure characteristics, the project exposure tendency, and the user project exposure characteristic preference, so that the method proposed by the present invention fully considers the causal relationship between exposure and clicks between data, so that the project recommendation list generated by the present invention is more in line with user expectations and the exposure deviation probability of the recommendation list; therefore, the problem that the recommendation list obtained by the traditional recommendation method is significantly different from the user expectations is solved.
[0051] 2. The present invention can better capture the causal relationship between the user's exposure preference and the project's exposure tendency, adjust the influence of the project exposure characteristics and the user's project exposure preference on the user's score, thereby reducing the distribution drift caused by the exposure bias problem of the user's exposure characteristics and content characteristics not matching, and effectively alleviate the exposure bias problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0053] Figure 2 Schematic diagram of the causal structure of click-exposure.
[0054] Figure 3 This is a schematic diagram of equal probability sampling. DETAILED DESCRIPTION
[0055] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0056] The following is combined with Figure 1 To the attached Figure 3 The present invention is described in detail:
[0057] See also Figure 1 As shown, a project recommendation method based on time-aware inverse propensity score is characterized by comprising the following steps:
[0058] Step 1: Collect news recommendation data and short video recommendation data with timestamps, and pre-process the news recommendation data and short video data with timestamps to obtain exposure-click data;
[0059] The data collected mainly includes news recommendations and product recommendations that are closely related to exposure-clicks, and incomplete data and unnecessary interaction records with less than 5 user-item interactions are removed in the preprocessing stage; secondly, the processing principle for missing values: mainly through regression and other methods, the most likely values are used to replace missing values, so that the relationship between missing values and other values is maximized.
[0060] Step 2: Extract features from the exposure-click data to obtain historical interaction features; historical interaction features include: user ID, item ID, timestamp, whether clicked, rotation rate, etc.
[0061] The feature extraction operation includes: selecting effective features, converting four types of features (u, i, t, y) such as user ID, project ID, timestamp, and whether the feature is clicked, and performing standard scaling on numerical features: normalization, standardization, etc. This reduces feature redundancy caused by high correlation between some features, avoids consuming computing performance, and reduces noise.
[0062] Step 3: Extract items and timestamps from historical interaction features and calculate relative time intervals in the time series;
[0063] The step 3 comprises the following steps:
[0064] Step 3.1: Extract the items and timestamps in the historical interaction features to obtain data with time intervals in the historical interaction sequence, and sort the historical interaction features based on the time series to obtain historical interaction feature sequence data;
[0065] Step 3.2: For the historical interaction feature sequence data, calculate the relative time interval between two adjacent historical interaction records according to the time nodes of the historical interaction features; and calculate the earliest time and the latest time of all clicked items of the user according to the user group, and calculate the total time span of the user's interaction records:
[0066] The calculation formula of the relative time interval is:
[0067]
[0068] Among them, T * is the relative time interval, T i is the timestamp of the current item the user is interacting with, T i-1 is the timestamp of the user's last interaction item, and T is the total time span of the user interaction record.
[0069] Step 4: Divide the historical interaction records into positive and negative samples based on the time interval and the set threshold;
[0070] The step 4 comprises the following steps:
[0071] Step 4.1: Take the interaction records whose relative time interval in the historical interaction features is less than the time length threshold as negative samples;
[0072] Step 4.2: Take the interaction records whose relative time interval in the historical interaction features is greater than the time length threshold as positive samples.
[0073] See also Figure 3, Step 5: Perform equal probability sampling on positive and negative samples to obtain an unbiased test set; adopt the equal probability sampling method to extract user interaction records in positive and negative samples with equal probability and combine them to obtain a theoretically unbiased test data set, in which the probability of users being exposed to different items is consistent.
[0074] Step 6: Based on the historical interaction characteristics, the inverse propensity score calculation method is used to calculate the user's exposure preference characteristics and the item's exposure tendency characteristics;
[0075] The step 6 comprises the following steps:
[0076] Step 6.1: Calculate item clicks based on the time decay function:
[0077]
[0078] Where Y u,i,t To calculate the time decay of item clicks, is the time decay function, which decays exponentially; β is the cooling coefficient, which is a hyperparameter used to control the intensity of time decay; Y u,i The number of item clicks without calculating time decay;
[0079] Step 6.2: Sum the item clicks that take time decay into account and divide it by the maximum exposure of the item in the data to obtain the item exposure tendency and user exposure preference (divided into hot type, diverse type, etc.):
[0080]
[0081] Where: p i,t is the exposure tendency of item i at time t, Y u,i,t is the number of item clicks with time decay calculated, ∑ is the summation symbol, represents the maximum exposure among all items, τ is the smoothing term of item exposure tendency; p u,t Expose preferences to users; is the cumulative sum of item clicks with time decay; |I u | is the number of items representing user interactions;
[0082] Step 6.3: Based on the item exposure tendency and user exposure preference, the average value of the individual user and the items they interact with is re-summed to obtain the final representation of the user exposure preference:
[0083] θ u,i,t =α-|p u,t -p i,t |;
[0084] Where: θ u,i,trepresents the improved propensity score, α is the threshold used to limit the upper limit of exposure tendency, and p u,t is the user’s exposure preference, p i,t is the exposure tendency of the project, and the tendency score θ u,i,t It is expressed as the matching degree between the user's exposure preference and the item's exposure tendency. The more matching the user's exposure preference and the item's exposure degree are, the greater the probability that the item will be clicked.
[0085] Step 6.4: Convert the propensity score θ u,i,t Concatenate with the historical interaction features in the training set, validation set, and test set respectively:
[0086] h′=[u,i,t,Y u,i ,θ u,i,t ];
[0087] Where: θ u,i,t is the improved propensity score, which takes a value between [0,1]; u is the user feature; i is the item feature; t is the timestamp; Y u,i Indicates whether the sample is clicked, 0 means no click, 1 means click; [u, i, t, Y u,i ,θ u,i,t ] is the concatenation of the one-dimensional vector of the project exposure propensity score and the user's historical interaction features, and h′ is the concatenated user-project historical interaction record.
[0088] Step 7: Combine the user's historical interaction characteristics and exposure tendency characteristics to train the model using a factorization machine based on deep learning and calculate the click probability of whether item i is clicked by user u after being exposed;
[0089] The step 7 comprises the following steps:
[0090] Step 7.1: Calculate the absolute value distance between the feature vector of the item exposure tendency and the user's exposure preference feature, and subtract the absolute value of the threshold and the difference as the initial tendency score;
[0091] Step 7.2: Based on the initial propensity score, calculate the prediction score error of the deep learning-based factorization machine;
[0092] The calculation formula of the prediction score error is as follows:
[0093]
[0094] Where: is the exposure rate of positive samples; is the exposure rate of negative samples, D is the total number of items, θ u,iis the propensity score based on the matching of the item exposure propensity vector and the user's exposure preference, also expressed as the exposure probability of the sample (u,i), ∑ is the sum, It is represented as a positive sample, that is, the item recommended at the current time point not only meets the user's exposure preference but also will be clicked. is the negative sample matched by user-item exposure, indicating that the recommended item will not be clicked at the current time; Loss is the prediction score error, r is the prediction score without the introduction of inverse propensity score in DeepFM, Introducing the inverse propensity score into DeepFM’s prediction score; Regularization term to avoid overfitting.
[0095] Step 7.3: Calculate the prediction score of the deep learning-based factorization machine based on the exposure preference feature and the exposure tendency feature of the item;
[0096] The calculation formula of the prediction score is as follows:
[0097]
[0098] Where: U u is the user potential feature matrix U∈R n×d , V i is the potential feature matrix V∈R of the project m×d , The predicted score after calculating the gradient according to the loss function and updating the potential feature matrix UV, * is the dot product symbol; T represents the transpose of the vector, U u *V i T Represents the product of matrices, randomly takes a sample from the time series, calculates the gradient according to the loss function, and updates the potential feature matrix.
[0099] Step 7.4: Calculate the rating of the project being exposed based on the predicted rating error and the predicted rating.
[0100] Step 8: Verify the debiasing effect of the factorization machine based on the unbiased test set: Generate a list of item recommendations in descending order of click-through rate based on the trained deep learning factorization machine network DeepMF.
[0101] According to the predicted scores of the improved inverse propensity score exposure probability calculation method, a list of recommended items is generated in descending order. For each user, a strategy is proposed:
[0102] According to the exposure-click probability causal theory, the top 20 recommended items are re-sorted in descending order during the reasoning process. For each item, the final ranking is calculated based on the final adjusted predicted score.
[0103] Step 9: Generate a list of recommended items in descending order of click-through rate based on the trained factorization machine.
[0104] The present invention calculates the exposure tendency through an inverse tendency score calculation method, and the inverse tendency score calculation method is related to the project exposure characteristics, the project exposure tendency, and the user project exposure characteristic preference, so that the method proposed by the present invention fully considers the causal relationship between exposure and clicks between data, so that the project recommendation list generated by the present invention is more in line with user expectations and the exposure deviation probability of the recommendation list; therefore, the problem that the recommendation list obtained by the traditional recommendation method is significantly different from the user expectation is solved.
[0105] The key to the problem of project exposure and user clicks in the present invention is that the user's click is obtained after exposure, and the user's preference can be reflected through post-click behavior, such as page dwell time, etc. Many studies have used this information as an indicator of user exposure preference. In fact, combined with the causal inference method, the user's exposure preference and the project's exposure tendency can be more clearly reflected. The degree of match between the user's exposure preference and the project's exposure tendency and the user's preference for exposure information will affect whether the user clicks on the project, thereby affecting the click-through rate. Using the causal inference method, the causal relationship between exposure and click can be displayed, and the logic of the recommendation system based on click optimization that unilaterally recommends based on the project's exposure characteristics and the user's preferences for it can be corrected. The project click prediction results can be corrected to alleviate the exposure bias problem in the recommendation process.
[0106] The above-mentioned embodiments only express the specific implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the protection scope of the present application. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the technical solution concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. A project recommendation method based on time-aware inverse propensity score, It is characterized in that The following steps are involved: Step 1: Collect news recommendation data and short video recommendation data with timestamps, and pre-process the news recommendation data and short video data with timestamps to obtain exposure-click data; Step 2: Extract features from the exposure-click data to obtain historical interaction features; Step 3: Extract items and timestamps from historical interaction features and calculate relative time intervals in the time series; Step 4: Divide the historical interaction records into positive and negative samples based on the time interval and the set threshold; Step 5: Perform equal probability sampling on positive and negative samples to obtain an unbiased test set; Step 6: Based on the historical interaction characteristics, the inverse propensity score calculation method is used to calculate the user's exposure preference characteristics and the item's exposure tendency characteristics; The step 6 comprises the following steps: Step 6.1: Calculate item clicks based on the time decay function: Where Y u,i,t To calculate the time decay of item clicks, is the time decay function, which decays exponentially; β is the cooling coefficient, which is a hyperparameter used to control the intensity of time decay; Y u,i The number of item clicks without calculating time decay; Step 6.2: Sum the item clicks that take time decay into account and divide it by the maximum exposure of the item in the data to obtain the item exposure tendency and user exposure preference: Where: p i,t is the exposure tendency of item i at time t, Y u,i,t is the number of item clicks with time decay calculated, ∑ is the summation symbol, represents the maximum exposure among all items, τ is the smoothing term of item exposure tendency; p u,t Expose preferences to users; is the cumulative sum of item clicks with time decay; |I u | is the number of items representing user interactions; Step 6.3: Based on the item exposure tendency and user exposure preference, the average value of the individual user and the items they interact with is re-summed to obtain the final representation of the user exposure preference: i u,i,t =α-|p u,t -p i,t |; Where: θ u,i,t represents the improved propensity score, α is the threshold used to limit the upper limit of exposure tendency, and p u,t is the user’s exposure preference, p i,t is the exposure tendency of the project, and the tendency score θ u,i,t It is expressed as the matching degree between the user's exposure preference and the item's exposure tendency. The more matching the user's exposure preference and the item's exposure degree are, the greater the probability that the item will be clicked. Step 6.4: Convert the propensity score θ u,i,t Concatenate with the historical interaction features in the training set, validation set, and test set respectively: h′=[u,i,t,Y u,i ,θ u,i,t ]; Where: θ u,i,t is the improved propensity score, which takes a value between [0,1]; u is the user feature; i is the item feature; t is the timestamp; Y u,i Indicates whether the sample is clicked, 0 means no click, 1 means click; [u, i, t, Y u,i ,θ u,i,t ] is the concatenation of the one-dimensional vector of the project exposure propensity score and the user's historical interaction characteristics, and h′ is the concatenated user project historical interaction record; Step 7: Combine the user's historical interaction characteristics and exposure tendency characteristics to train the model using a factorization machine based on deep learning and calculate the click probability of whether item i is clicked by user u after being exposed; Step 8: Verify the debiasing effect of the factorization machine based on the unbiased test set; Step 9: Generate a list of recommended items in descending order of click-through rate based on the trained factorization machine.
2. The project recommendation method based on time-aware inverse propensity score according to claim 1, It is characterized in that The step 3 comprises the following steps: Step 3.1: Extract the items and timestamps in the historical interaction features to obtain data with time intervals in the historical interaction sequence, and sort the historical interaction features based on the time series to obtain historical interaction feature sequence data; Step 3.2: For the historical interaction feature sequence data, calculate the relative time interval between two adjacent historical interaction records according to the time nodes of the historical interaction features; and calculate the earliest time and the latest time of all clicked items of the user according to the user group, and calculate the total time span of the user's interaction records: The calculation formula of the relative time interval is: Among them, T * is the relative time interval, T i is the timestamp of the current item the user is interacting with, T i-1 is the timestamp of the user's last interaction item, and T is the total time span of the user interaction record.
3. The project recommendation method based on time-aware inverse propensity score according to claim 1, It is characterized in that The step 4 comprises the following steps: Step 4.1: Take the interaction records whose relative time interval in the historical interaction features is less than the time length threshold as negative samples; Step 4.2: Take the interaction records whose relative time interval in the historical interaction features is greater than the time length threshold as positive samples.
4. The project recommendation method based on time-aware inverse propensity score according to claim 1, It is characterized in that The step 7 comprises the following steps: Step 7.1: Calculate the absolute value distance between the feature vector of the item exposure tendency and the user's exposure preference feature, and subtract the absolute value of the threshold and the difference as the initial tendency score; Step 7.2: Based on the initial propensity score, calculate the prediction score error of the deep learning-based factorization machine; Step 7.3: Calculate the prediction score of the deep learning-based factorization machine based on the exposure preference feature and the exposure tendency feature of the item; Step 7.4: Calculate the rating of the project being exposed based on the predicted rating error and the predicted rating.
5. The project recommendation method based on time-aware inverse propensity score according to claim 4, It is characterized in that The calculation formula of the prediction score error is as follows: Where: is the exposure rate of positive samples; is the exposure rate of negative samples, D is the total number of items, θ u,i is the propensity score based on the matching of the item exposure propensity vector and the user's exposure preference, also expressed as the exposure probability of the sample (u,i), ∑ is the sum, It is represented as a positive sample, that is, the item recommended at the current time point not only meets the user's exposure preference but also will be clicked. is the negative sample matched by user-item exposure, indicating that the recommended item will not be clicked at the current time; Loss is the prediction score error, r is the prediction score without the introduction of inverse propensity score in DeepFM, Introducing the inverse propensity score into DeepFM’s prediction score; Regularization term to avoid overfitting.
6. The project recommendation method based on time-aware inverse propensity score according to claim 4, It is characterized in that The calculation formula of the prediction score is as follows: Where: U u is the user potential feature matrix U∈R n×d , V i is the potential feature matrix V∈R of the project m×d , To calculate the gradient according to the loss function and update the predicted score after the potential feature matrix UV, * is the dot product symbol; T represents the transpose of the vector, Represents the product of matrices, randomly takes a sample from the time series, calculates the gradient according to the loss function, and updates the potential feature matrix.
Citation Information
Patent Citations
Information delivery method and device, and computer readable storage medium
CN111681058A
Financial management recommendation method, device and system and storage medium
CN114297511A