A multi-bias recommendation debiasing method and storage medium based on causal inference

By introducing gated decision network and gated debias converged network in the recommendation system, the most suitable debias algorithm is automatically selected, and the problem of poor debias effect in the multi-deviation recommendation system in the existing technology is solved, and effective modeling and removal of multiple deviations is achieved.

CN119202548BActive Publication Date: 2025-05-13CHINA ORDNANCE SCI INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411716698.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-13
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The existing recommended algorithms lack universality when dealing with multiple deviations, and cannot effectively model and remove various deviations in the data. The fusion of the mean methods of multiple debias algorithms cannot determine the main debias effect.

Method used

A multi-bias recommended debiasing method based on causal inference is adopted, through the gated decision network and the gated debiasing convergence network, the weights of different types of deviations are used for modeling, and the most suitable debiasing algorithm is automatically selected for matching debiasing.

Benefits of technology

Effective modeling and removal of multiple deviations in the data is achieved, avoiding the problem of no modeling bias in meta-learning, and capturing the correlation and differences between multiple debias tasks through a multi-task learning architecture, improving the accuracy of the debias effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202548B_ABST
    Figure CN119202548B_ABST
Patent Text Reader

Abstract

A multi-bias recommendation debiasing method based on causal inference and a storage medium, comprising a basic data definition step, a gated decision network training step, a gated debiasing fusion network training step, and a score calculation and recommendation step; the present invention introduces a gated decision network and uses weights of different types of biases for modeling, thereby avoiding the problem of meta-learning without modeling bias; uses a multi-task learning architecture to fuse multiple debiasing algorithms, and learns weight parameter groups to capture the correlation and distinction between multiple debiasing tasks, thereby avoiding the problem of lack of focus when using the mean method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of recommendation algorithms, and more specifically, to a multi-bias recommendation debiasing method and storage medium based on causal inference, which can integrate multiple debiasing algorithms and automatically use the most suitable debiasing algorithm combination for debiasing. Background Art

[0002] In recent years, with the rise of causal inference, methods such as propensity scores, counterfactual thinking, and removal of confounding factors have been widely studied and applied in the field of recommendation systems. The corresponding application areas of recommendation systems include video recommendation, article recommendation, and e-commerce recommendation. The characteristics of these systems are that the recommendation is based on multiple indicators. If only a single indicator is relied on for recommendation, bias will occur. At present, a large number of debiasing methods have been proposed at home and abroad to reduce the bias or unfairness in recommendation systems, such as selection bias, consistency bias, exposure bias, position bias, popularity bias, and unfairness.

[0003] Existing recommendation algorithms, such as IPS and data imputation methods, focus on one or two specific biases and lack universality. However, the data collected in the real world has various biases due to user personal preferences, location preferences, models, etc., and a single solution to a certain type of bias problem cannot meet the debiasing problem of the recommendation system.

[0004] The existing mainstream multi-bias removal method is AutoDebias and its variants, which is a universal and adaptive debiasing method based on meta-learning. First, a general debiasing framework is established. The framework first uses the re-weighting method to add specific weights to each sample in the training set, and then further uses the data filling method to process the uncovered part of the training set, that is, to construct pseudo-label data. The overall framework is expressed as an averaging method.

[0005]

[0006] In this framework, the problem of finding the optimal debiasing strategy is transformed into the problem of setting appropriate debiasing parameters in the framework. At the same time, due to the huge number of debiasing parameters in the framework, direct optimization will lead to overfitting problems and lack generalization performance. Therefore, the linear mean method is used to fuse the results of different debiasing algorithms to finally achieve adaptive debiasing.

[0007] In summary, the debiasing algorithms in the prior art have the following problems: (1) they do not model the bias in the data, but instead use a meta-learning method to convert the problem of solving the bias into the problem of finding debiasing parameters; (2) they use multiple debiasing algorithms in a progressive mean manner to solve the bias, and it is impossible to know which algorithm has the main debiasing effect.

[0008] Therefore, how to apply multi-bias recommendation, model and classify the deviations of recommended data, and automatically adopt the most appropriate debiasing algorithm for debiasing has become a technical problem that needs to be urgently solved by existing technologies. Summary of the invention

[0009] The purpose of the present invention is to propose a multi-deviation recommendation debiasing method and storage medium based on causal inference, which can make it possible to model and classify the deviations in the data and automatically use the most suitable debiasing algorithm to perform debiasing.

[0010] To achieve this object, the present invention adopts the following technical solutions:

[0011] A multi-bias recommendation debiasing method based on causal inference, comprising the following steps:

[0012] Basic data definition step S110:

[0013] According to the historical recommendation data, the basic data of users' recommendations for different items are collected and sorted, among which the user is defined , user collection ,in ,

[0014] Defining Items , item collection ,in ,

[0015] definition , indicating user forward A historical item sequence consisting of item subsets, where Representing a subset of items, and obtaining the user's ratings of historically recommended items, and establishing a sample data set based on the above historical recommendation data, wherein the sample data includes training data and verification data;

[0016] Gated decision network training step S120:

[0017] Select k independent debiasing algorithms for different biased data, and use the uniformly distributed Xavier algorithm to debias the gated decision network. Initialize and use the selected debiasing algorithm to pass through the gated decision network Calculate the unbiased output for the item and fuse it to get the output result y. Calculate the cross entropy value between y and the expected value in the training data. Use the cross entropy value to optimize the loss function. When the function value keeps decreasing and tends to be stable, the training ends.

[0018] Gated debiasing fusion (ADMF4Rec) network training step S130:

[0019] Select the k debiasing algorithms to process the training data and the validation data respectively, remove the corresponding deviations, and use the trained gated decision network to fuse the results of multiple debiasing algorithms again. Each debiasing task uses a trained gated decision network separately, with different weights in the fusion. , the fusion model is optimized using the squared error as the loss function to obtain the optimized gated debiasing fusion network;

[0020] Score calculation and recommendation step S140:

[0021] Use the trained gated debiasing fusion network to Predict the rating of the item that needs to be interacted with next time.

[0022] Optionally, in the basic data definition step S110,

[0023] The sample data set is used for training and verification. Each sample data set includes multiple samples, and each sample includes a single user The interaction history of the user is as follows: Item access history , the user The ratings of the items in the sequence and the next item that the user is expected to visit.

[0024] Optionally, in the basic data definition step S110,

[0025] The ratings are saved in a matrix format, defining the user's ratings of items. ,in R is a scoring matrix, which is used to select recommended items after calculating the recommendation points in the subsequent step.

[0026] Optionally, the sample data set is divided into a training set, a validation set and a test set, wherein the training set is used to estimate the model and accounts for 50% of the total sample data, the validation set accounts for 20% of the total sample data, and the test set accounts for 30% of the total sample data.

[0027] Optionally, the gated decision network training step S120 is specifically as follows:

[0028] Initialized gated decision network As shown in formula (1),

[0029] (1)

[0030] in, is the model parameter of the gated unit, that is, the parameter that needs to be learned for gated network training. The user's access sequence to items and the ratings of items in the training set are converted into word vectors through Word2Vec and concatenated, as shown in formula (2):

[0031] (2)

[0032] Use the selected debiasing algorithm to process the data in the training set to obtain the user's unbiased output of the item , after the gating network fusion, the output result y is obtained, as shown in formula (3):

[0033] (3)

[0034] The cross entropy is used as the loss function, as shown in formula (4), and the Adam algorithm is used to optimize the loss function, where represents the expected output in the training set, is the regularization factor. When the function value keeps decreasing and approaches stability, the training ends.

[0035] (4).

[0036] Optionally, in the gated decision network training step S120,

[0037] The debiasing algorithm includes selectivity bias, consistency bias, exposure bias, position bias, popularity bias, and unfairness.

[0038] Optionally, the gated debiasing fusion network training step S130 is specifically as follows:

[0039] Select the debiasing algorithm to process the user data in the training set to obtain unbiased output , the results of multiple debiasing algorithms are fused using the trained gated decision network, as shown in formula (5).

[0040]

[0041] (5)

[0042] in, is the gated decision network, Represents the weight parameter group in the gated debiasing fusion network, including ( );

[0043] The gated debiasing fusion network is optimized using the squared error loss function, as shown in formula (6):

[0044] (6)

[0045] in, represents the expected output in the training set, is the regularization factor, and the optimization algorithm is the Adam algorithm. When the loss function converges, that is, the value of formula (6) no longer decreases and tends to be stable, it is considered to have converged, and the weight parameter group to be optimized is determined. , and finally obtain the gated debiasing fusion network.

[0046] Optionally, the score calculation and recommendation step S140 is specifically as follows:

[0047] The result is calculated using formula (5): , through the user Based on the historical ratings of items, find the item with the closest predicted rating and recommend it to the user.

[0048] The present invention further discloses a storage medium for storing computer executable instructions, which, when executed by a processor, execute the above-mentioned multi-bias recommendation debiasing method based on causal inference.

[0049] In summary, the present invention has the following advantages:

[0050] (1) The present invention avoids the problem of no modeling bias in meta-learning by introducing a gated decision network and using the weights of different types of biases for modeling.

[0051] (2) A multi-task learning architecture is used to integrate multiple debiasing algorithms. The weight parameter group is learned to capture the correlation and differences between multiple debiasing tasks, avoiding the problem of lack of focus when using the mean method. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flow chart of a multi-bias recommendation debiasing method based on causal inference according to the present invention;

[0053] Figure 2 is a basic flow chart of training a gated decision network according to the present invention;

[0054] Figure 3 It is a basic flow chart of training the gated debiasing fusion model according to the present invention. DETAILED DESCRIPTION

[0055] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0056] The present invention is to introduce a gated decision network for training using a multi-bias algorithm, and then use a fusion of multiple debiasing algorithms. All debiasing tasks use a separate gated network, and selectively utilize different debiasing models through different final output weights, thereby reducing the mutual influence between multiple debiasing models.

[0057] For details, see Figure 1 A flowchart of a multi-bias recommendation debiasing method based on causal inference according to a specific embodiment of the present invention is shown, comprising the following steps:

[0058] Basic data definition step S110:

[0059] According to the historical recommendation data, the basic data of users' recommendations for different items are collected and sorted, among which the user is defined , user collection ,in , define items , item collection ,in ,definition , indicating user forward A historical item sequence consisting of item subsets, where Represent a subset of items, obtain users' ratings of historically recommended items, and establish a sample data set based on the above historical recommendation data. The sample data includes training data, verification data, and may also include test data.

[0060] Furthermore, the ratings can be saved in a matrix format. Specifically, the user's ratings for the items can be defined. ,in R is a scoring matrix, which is used to select recommended items after calculating the recommendation points in the subsequent step.

[0061] Furthermore, the sample data set is used for training and verification, and each sample data set includes multiple samples, each sample includes a single user The interaction history of the user is as follows: Item access history , the user Ratings for items that appear in a sequence and items that the user is expected to visit in the future.

[0062] The sample data set is divided into a training set, a validation set and a test set, wherein the training set is used to estimate the model and accounts for 50% of the total sample data, the validation set accounts for 20% of the total sample data, and the test set accounts for 30% of the total sample data.

[0063] Through the above basic data, the present invention aims to The interactive history sequence , for this user The user will be recommended items with high scores, that is, the items that are expected to be visited by the user.

[0064] There may be some differences between the items that users are expected to visit in the future and the items that are expected to be visited by users through calculation. By analyzing the above differences, the model can be further trained, judged and studied until it meets the needs and is put into subsequent use.

[0065] The difference between the two items can be analyzed by looking at the historical scores of the items, such as the score matrix R , find the scores of the two items respectively, determine the difference between the two scores, and thus determine the results of the model training and make improvements.

[0066] It is particularly noted that, since the sample data belongs to training, validation and testing data, the items that users are expected to access in the future are constructed to conduct complete training, validation and testing.

[0067] Gated decision network training step S120:

[0068] See also Figure 2 , select k independent debiasing algorithms for different biased data, and use the uniformly distributed Xavier algorithm to gate the decision network Initialize and use the selected debiasing algorithm to pass through the gated decision network The unbiased output for the item is calculated and fused to get the output result y. The cross entropy value between y and the expected value in the training set is calculated, and the cross entropy value is used to optimize the loss function. When the function value continues to decrease and tends to be stable, the training ends.

[0069] Specifically, the gated decision network after initialization As shown in formula (1),

[0070] (1)

[0071] in, is the model parameter of the gated unit, that is, the parameter that needs to be learned for gated network training. The user's access sequence to items and the ratings of items in the training set are converted into word vectors through Word2Vec and concatenated, as shown in formula (2):

[0072] (2)

[0073] Use the selected debiasing algorithm to process the data in the training set to obtain the user's unbiased output of the item , after the gating network fusion, the output result y is obtained, as shown in formula (3):

[0074] (3)

[0075] The cross entropy is used as the loss function, as shown in formula (4), and the Adam algorithm is used to optimize the loss function, where represents the expected output in the training set, is the regularization factor. When the function value keeps decreasing and approaches stability, the training ends.

[0076] (4).

[0077] The debiasing algorithm includes but is not limited to selection bias, consistency bias, exposure bias, position bias, popularity bias and unfairness.

[0078] Gated debiasing fusion (ADMF4Rec) network training step S130:

[0079] See also Figure 3 , shows a gated debiasing fusion network according to a specific embodiment of the present invention,

[0080] Select k debiasing algorithms to process the user data in the training set and the validation set respectively to remove the corresponding biases. Use the trained gated decision network to fuse the results of multiple debiasing algorithms again. Each debiasing task uses a trained gated decision network separately, with different weights in the fusion. , the fusion model is optimized using the squared error as the loss function to obtain the optimized gated debiasing fusion network.

[0081] Specifically, we select the debiasing algorithm to process the user data in the training set and obtain the unbiased output. , the results of multiple debiasing algorithms are fused using the trained gated decision network, as shown in formula (5).

[0082]

[0083] (5)

[0084] in, is the gated decision network, Represents the weight parameter group in the gated debiasing fusion network, which is the parameter that needs to be optimized, including ( ), which has different values ​​for each debiasing algorithm;

[0085] The gated debiasing fusion network is optimized using the squared error loss function, as shown in formula (6):

[0086] (6)

[0087] in, represents the expected output in the training set, is the regularization factor, and the optimization algorithm is the Adam algorithm. When the loss function converges, that is, the value of formula (6) continues to decrease and tends to be stable, it is considered to have converged, and the weight parameter group to be optimized is determined. , and finally obtain the gated debiasing fusion network.

[0088] Score calculation and recommendation step S140:

[0089] Use the trained gated debiasing fusion network to The score of the item that needs to be interacted with next time is predicted. Specifically, the result is calculated using formula (5): , through the user For the historical ratings of items, find the item that is closest to the predicted rating, which is the item recommended to the user. For example, in formula (7), in the predicted rating matrix , find items recommended to the user .

[0090] (7)

[0091] The data to be predicted may be data in a test set or new data to be predicted.

[0092] The present invention further discloses a storage medium for storing computer executable instructions, which, when executed by a processor, execute the above-mentioned multi-bias recommendation debiasing method based on causal inference.

[0093] Example:

[0094] Using the Huawei Causal Inference Competition PCIC Track 2 dataset, the MF_IPS algorithm that removes selectivity bias, the CausE algorithm that removes exposure bias, and the MF_PDA algorithm that removes popularity bias were selected in the experiment. The gated decision network in the present invention was replaced with the AutoDebias model.

[0095] As shown in Table 1, the performance of the solution of the present invention is improved compared with the AutoDebias model, and the core evaluation index AUC performance is improved by 0.45%, which proves that the debiasing effect of the present invention is better than that of the AutoDebias model.

[0096]

[0097] Table 1

[0098] In summary, the present invention has the following advantages:

[0099] (1) The present invention avoids the problem of no modeling bias in meta-learning by introducing a gated decision network and using the weights of different types of biases for modeling.

[0100] (2) A multi-task learning architecture is used to integrate multiple debiasing algorithms. The weight parameter group is learned to capture the correlation and differences between multiple debiasing tasks, avoiding the problem of lack of focus when using the mean method.

[0101] Obviously, those skilled in the art should understand that the above-mentioned units or steps of the present invention can be implemented by a general-purpose computing device, they can be concentrated on a single computing device, and optionally, they can be implemented by a program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0102] The above content is a further detailed description of the present invention in combination with a specific preferred embodiment. It cannot be determined that the specific embodiments of the present invention are limited to this. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention determined by the submitted claims.

Claims

1. A multi-bias recommendation debiasing method based on causal inference, characterized in that: The steps include: Basic data definition step S110: According to the historical recommendation data, the basic data of users' recommendations for different items are collected and sorted, among which the user is defined , user collection ,in , Defining Items , item collection ,in , definition , indicating user forward A historical item sequence consisting of item subsets, where Representing a subset of items, and obtaining the user's ratings of historically recommended items, and establishing a sample data set based on the above historical recommendation data, wherein the sample data includes training data and verification data; Gated decision network training step S120: Select k independent debiasing algorithms for different biased data, and use the uniformly distributed Xavier algorithm to debias the gated decision network. Initialize and use the selected debiasing algorithm to pass through the gated decision network Calculate the unbiased output for the item and fuse it to get the output result y. Calculate the cross entropy value between y and the expected value in the training data. Use the cross entropy value to optimize the loss function. When the function value keeps decreasing and tends to be stable, the training ends. Gated debiasing fusion network training step S130: Select the k debiasing algorithms to process the training data and the validation data respectively, remove the corresponding deviations, and use the trained gated decision network to fuse the results of multiple debiasing algorithms again. Each debiasing task uses a trained gated decision network separately, with different weights in the fusion. , the fusion model is optimized using the squared error as the loss function to obtain the optimized gated debiasing fusion network; Score calculation and recommendation step S140: Use the trained gated debiasing fusion network to Predict the rating of the item that needs to be interacted with next time; In the basic data definition step S110, The sample data set is used for training and verification. Each sample data set includes multiple samples, and each sample includes a single user The interaction history of the user is as follows: Item access history , the user The ratings of the items in the sequence and the next item that the user is expected to visit; The ratings are saved in a matrix format, defining the user's ratings of items. ,in R is a scoring matrix, which is used to select recommended items after obtaining and calculating recommendation scores in a subsequent step; Step S120 is specifically as follows: Initialized gated decision network As shown in formula (1), (1) in, is the model parameter of the gated unit, that is, the parameter that needs to be learned for gated network training. The user's access sequence to items and the ratings of items in the training set are converted into word vectors through Word2Vec and concatenated, as shown in formula (2): (2) Use the selected debiasing algorithm to process the data in the training set to obtain the user's unbiased output of the item , after the gating network fusion, the output result y is obtained, as shown in formula (3): (3) The cross entropy is used as the loss function, as shown in formula (4), and the Adam algorithm is used to optimize the loss function, where represents the expected output in the training set, is the regularization factor. When the function value keeps decreasing and approaches stability, the training ends. (4); Step S130 is specifically as follows: Select the debiasing algorithm to process the user data in the training set to obtain unbiased output , the results of multiple debiasing algorithms are fused using the trained gated decision network, as shown in formula (5). (5) in, is the gated decision network, Represents the weight parameter group in the gated debiasing fusion network, including ( ); The gated debiasing fusion network is optimized using the squared error loss function, as shown in formula (6): (6) in, represents the expected output in the training set, is the regularization factor, and the optimization algorithm is the Adam algorithm. When the loss function converges, that is, the value of formula (6) continues to decrease and tends to be stable, it is considered to have converged, and the weight parameter group to be optimized is determined. , and finally obtain the gated debiasing fusion network.

2. The multi-bias recommendation debiasing method according to claim 1, characterized in that: The sample data set is divided into a training set, a validation set and a test set, wherein the training set is used to estimate the model and accounts for 50% of the total sample data, the validation set accounts for 20% of the total sample data, and the test set accounts for 30% of the total sample data.

3. The multi-bias recommendation debiasing method according to claim 2, characterized in that: In the gated decision network training step S120, The debiasing algorithm includes selectivity bias, consistency bias, exposure bias, position bias, popularity bias, and unfairness.

4. The multi-bias recommendation debiasing method according to claim 3, characterized in that: The score calculation and recommendation step S140 is specifically as follows: The result is calculated using formula (5): , through the user Based on the historical ratings of items, find the item with the closest predicted rating and recommend it to the user.

5. A storage medium for storing computer executable instructions, characterized in that: When the computer executable instructions are executed by a processor, the multi-bias recommendation debiasing method based on causal inference is performed as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Dual enhancement tendency score estimation method in sequence recommendation

    CN115599972A

  • Recommendation system popularity depolarization method and system, and storage medium

    CN116664226A