A Knowledge Distillation Debiased Recommendation Method and System Based on Causal Theory
Through the knowledge distillation and de-bias recommendation method based on causal theory, unbiased soft labels are generated using display and implicit feedback teacher models, the problem of selection bias in the recommendation system is solved and the accuracy and robustness of the recommendation system is improved.
Patent Information
- Application Number
- CN202510292925.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing recommendation system relies on biased data when generating soft labels, which leads to selection bias problems, affecting recommendation performance. The traditional knowledge distillation method fails to effectively solve the causal path of unobserved data, resulting in further intensification of deviations.
Using the knowledge distillation and debias recommendation method based on causal theory, we generate unbiased soft labels by constructing two teacher models and one student model, and use display feedback and implicit feedback teacher models to process the causal paths of observed and unbiased data respectively, generate unbiased soft labels, and perform knowledge distillation to guide students' model learning.
It significantly improves the accuracy and robustness of the recommendation system. By supplementing the causal paths of unobserved data, unbiased soft labels are generated, select bias is reduced, and the performance of the recommendation system is improved.
Smart Images

Figure CN119808836B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge distillation, and in particular to a knowledge distillation debiasing recommendation method and system based on causal theory. Background Art
[0002] In a recommendation system, user preferences are usually learned through the interaction feedback between users and recommended items. However, there is an important problem in such a preference learning process, that is, the selection bias in the user interaction process. Currently, most recommendation systems are built only based on biased data, which cannot provide completely accurate recommendations.
[0003] In recent years, to solve the selection bias problem in recommendation systems, researchers have proposed various methods. Among these methods, the knowledge distillation-based technology has shown relatively advanced performance, demonstrating its great potential in alleviating the bias problem in recommendation systems. Traditional knowledge distillation methods first pre-train a large teacher model, and then use the soft labels distilled from this teacher model to supervise the learning of a small student model. Since these soft labels encode the knowledge learned by the large teacher model, the student model can obtain additional user preference information from them and perform better than learning directly from the original training data. However, there are still some problems. The data collected by the recommendation system usually has serious selection bias, and the learning process of the large teacher model mainly depends on these biased user feedback data. Therefore, the teacher model will inevitably inherit this bias and pass it on to the soft labels. When the student model uses these biased soft labels for learning, it not only cannot correct its original bias, but may further exacerbate this bias, thus affecting its performance. The main problem lies in that the teacher model generates soft labels only from the causal path of the observed data, while ignoring the soft labels on the causal path of the unobserved data. This approach results in the generated soft labels only reflecting the user preferences on some specific causal paths, thus carrying selection bias. Therefore, how to make the large teacher model generate unbiased soft labels to assist the student model in learning has become a key issue in reducing the selection bias in the knowledge distillation process.
[0004] Based on the above problems, there is an urgent need to propose a new knowledge distillation debiasing method, namely Debiasing-Knowledge Distillation (Debiasing-KD), to provide more unbiased soft labels, thereby guiding the student model to perform debiasing learning. Summary of the Invention
[0005] To solve the selection bias problem caused by the lack of historical interaction data of users in a knowledge distillation recommendation system, the present invention provides a knowledge distillation debiasing recommendation method and system based on causal theory.
[0006] In the first aspect, a knowledge distillation debiasing recommendation method based on causal theory provided by the present invention adopts the following technical solutions:
[0007] A knowledge distillation debiasing recommendation method based on causal theory includes:
[0008] Obtain a user rating dataset;
[0009] Construct a knowledge distillation debiasing recommendation model. Among them, construct two teacher models and one student model, generate the causal path of the observed data according to the user rating dataset; supplement the causal path of the unobserved data based on the causal effect of quantifying the selection bias; use the two teacher models to make predictions on the causal path of the observed data and the causal path of the unobserved data respectively to generate unbiased soft labels;
[0010] Perform knowledge distillation on the student model based on the unbiased soft labels to obtain a knowledge distillation debiasing recommendation model;
[0011] Use the knowledge distillation debiasing recommendation model to make user predictions.
[0012] Further, the obtaining of the user's observed rating data includes obtaining the observed rating data of the user for the items. Among them, let be a set of M users, be a set of N items; represent as the true feedback label matrix collected from the user interaction history, and each item in it represents the true rating of user for item .
[0013] Further, the generating of the causal path of the observed data according to the user rating dataset includes inputting the user's observed rating data into the explicit feedback teacher model, and the explicit feedback teacher model makes predictions according to the observed data to generate soft labels, and the path of generating the soft labels is used as the causal path of the observed data.
[0014] Further, the supplementing of the causal path of the unobserved data based on the causal effect of quantifying the selection bias includes based on the click data in the case of the biased causal effect formula and the unobserved data in the case of the causal effect formula, and obtaining the unbiased causal effect formula by combining the biased causal effect formula and the causal effect formula, which is expressed as:
[0015]
[0016] Among them, user ; item ; unbiased soft label .
[0017] Further, the method of using Debiasing-KD to perform debiased knowledge distillation on the causal path of observed data and the causal path of unobserved data respectively to generate unbiased soft labels includes performing debiased knowledge distillation based on the Debiasing-KD method, wherein the causal path obtained by using the explicit feedback teacher model to act on the observed user feedback data; the causal path obtained by using the implicit feedback teacher model to act on the unobserved user feedback data, and finally generating unbiased soft labels.
[0018] Further, the method of performing knowledge distillation on the student model based on the unbiased soft labels to obtain a knowledge distillation debiased recommendation model includes using the student model to learn with the unbiased soft labels as the fitting target, based on the knowledge learned by the teacher model encoded by the unbiased soft labels, and based on the loss function, enabling the student model to learn the knowledge of the teacher model in the unbiased soft labels during the learning process. The training loss function of the student model is expressed as:
[0019]
[0020] where is the fitting loss function between the predicted value of the student model and the unbiased soft label.
[0021] In a second aspect, a knowledge distillation debiased recommendation system based on causal theory includes:
[0022] A data acquisition module configured to acquire a user rating data set;
[0023] A model construction module configured to construct a knowledge distillation debiased recommendation model, wherein two teacher models and one student model are constructed, the causal path of the observed data is generated according to the user rating data set; the causal path of the unobserved data is supplemented based on the causal effect of the quantified selection bias; the two teacher models are used to make predictions on the causal path of the observed data and the causal path of the unobserved data respectively to generate unbiased soft labels;
[0024] A knowledge distillation module configured to perform knowledge distillation on the student model based on the unbiased soft labels to obtain a knowledge distillation debiased recommendation model;
[0025] A prediction module configured to perform user prediction using the knowledge distillation debiased recommendation model.
[0026] In a third aspect, the present invention provides a computer-readable storage medium storing multiple instructions, and the instructions are adapted to be loaded and executed by a processor of a terminal device to perform the method of a knowledge distillation debiased recommendation based on causal theory.
[0027] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the knowledge distillation debiasing recommendation method based on causal theory.
[0028] In summary, the present invention has the following beneficial technical effects:
[0029] From the perspective of causal inference, the present invention analyzes the causal reasons for selection bias in the knowledge distillation process, and solves the selection bias problem by supplementing the causal paths of unobserved data; at the same time, it quantifies the causal effect of selection bias, gives the causal effect formulas on two paths respectively, and further gives the unbiased causal effect formula on the comprehensive path.
[0030] The present invention innovatively introduces two teacher models, generates soft labels from the causal paths of observed data and unobserved data respectively, and combines these two soft labels to generate unbiased soft labels, so as to guide the student model to perform debiasing learning.
[0031] Through experimental analysis, it demonstrates the significant performance improvement of the implicit feedback teacher model, fully illustrating the innovation of the implicit feedback teacher model (IFTM) in the Debiasing-KD debiasing method. By supplementing the implicit feedback causal path in the knowledge distillation causal graph, this method can better fit user preferences and generate more unbiased distilled soft labels. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is the causal graph of the selection bias problem in the knowledge distillation of Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The present invention will be further described in detail below with reference to the accompanying drawings.
[0034] Embodiment 1
[0035] Refer to Figure 1 , a knowledge distillation debiasing recommendation method based on causal theory in this embodiment includes:
[0036] Obtain the user rating dataset;
[0037] Construct a knowledge distillation debiasing recommendation model, wherein two teacher models and one student model are constructed, generate the causal path of observed data according to the user rating dataset; supplement the causal path of unobserved data based on the quantified causal effect of selection bias; use the two teacher models to perform predictions on the causal path of observed data and the causal path of unobserved data respectively to generate unbiased soft labels;
[0038] Perform knowledge distillation on the student model based on unbiased soft labels to obtain a knowledge distillation debiased recommendation model;
[0039] Use the knowledge distillation debiased recommendation model to perform user prediction.
[0040] Specifically:
[0041] S1. Obtain user observed rating data,
[0042] Obtained from the existing public datasets Coat Shopping, Yahoo, KuaiRec, and its data form is the rating of items by users. For example, (1, 2, 4) means that user 1 gives a rating of 4 to item 2.
[0043] Among them, let be the set of M users, be the set of N items. Denote as the true feedback label matrix collected from the user interaction history, and each item in it represents the true rating of user on item . In the recommendation system, since users usually only give feedback on some items, most are missing, which leads to the selection bias problem. represents the soft label output by the large teacher model, containing the knowledge learned by the large teacher model. represents whether the user feedback is observed. If it can be observed, , otherwise .
[0044] S2. Generate the causal path of the observed data according to the user rating dataset, input the observed data of the user into the display feedback teacher model, and the display feedback teacher model makes predictions based on the observed data to generate soft labels. This path of generating soft labels is called the causal path of the observed data. (Observed data -> Display feedback teacher model -> Soft label)
[0045] For example, Figure 1 in (a) shows the causal relationship in the knowledge distillation process of the traditional recommendation system, which contains two causal paths affecting the generation of soft labels. represents the causal process of generating unbiased soft labels according to the characteristics of user and item in the ideal situation where all user feedback can be observed, and guiding the student model to learn. However, in the actual knowledge distillation process, since users only give feedback on some items, the causal path of generating soft labels changes from the ideal state to , that is, only based on the observed user feedback data to generate soft labels while ignoring the unobserved user feedback data , which leads to the problem of selection bias in soft labels.
[0046] According to causal relationship analysis, due to the missing causal path corresponding to the unobserved user feedback data a biased soft label is generated, and the selection bias is further transmitted to the student model during the distillation process. Therefore, the key to solving the selection bias problem lies in restoring the causal path for generating soft labels based on unobserved user feedback . Specifically, as shown in (b) of Figure 1 , in addition to the existing path , we also need to supplement another path . This can ensure that all possible user-item interaction situations are comprehensively considered, thereby reducing the impact of selection bias.
[0047] S3. Supplement the causal path of unobserved data based on the causal effect of quantifying selection bias,
[0048] where supplementing the causal path is crucial for eliminating selection bias in knowledge distillation. In this section, we use the language of causal inference to infer the formula for the unbiased causal effect to be estimated in knowledge distillation. First, the formula for the biased causal effect in the case of only click data is given, and then the formula for the causal effect in the case of unobserved data that needs to be supplemented is given. Finally, the formulas for the causal effects in these two cases are combined to derive the formula for the unbiased causal effect.
[0049] Let denote the random variable under the conditional distribution , where the variable is intervened to a specific value , and the causal effect of the variable on can be expressed as the change of value as the value changes. Therefore, for user , along the path , the formula for the causal effect of on can be expressed as:
[0050]
[0051] where Represents the baseline case for comparison. The above formula can be expressed as under conditions and the baseline case of the difference.
[0052] Similarly, for the user , along the path when for the causal effect formula can be expressed as:
[0053]
[0054] The above formula can be expressed as under conditions and the baseline case of the difference.
[0055] Traditional knowledge distillation has only one causal path. Therefore, eliminating selection bias can be achieved by increasing the causal path and the total causal effect can be expressed as:
[0056]
[0057] This formula represents the unbiased causal effect when combining the two paths and can be used to generate unbiased soft labels to guide the learning of the student model.
[0058] S4. Use two teacher models to make predictions on the causal paths of observed data and unobserved data respectively to generate unbiased soft labels.
[0059] Among them, to achieve debiased knowledge distillation, this paper proposes the Debiasing-Knowledge Distillation (Debiasing-KD) method. Specifically, this paper introduces two teacher models: the Explicit Feedback Teacher Model (EFTM) and the Implicit Feedback Teacher Model (IFTM), which act on the causal paths and respectively to generate comprehensive unbiased soft labels.
[0060] The Explicit Feedback Teacher Model directly generates soft labels on the observed user feedback , which is easy to implement. Its learning process can be expressed as:
[0061] (1)
[0062] Among them, Soft labels learned by the teacher model.
[0063] However, for the implicit feedback teacher model, it is unrealistic to directly fill in the true labels of the unobserved data . Therefore, in this paper, the click-through conversion rate of the items unobserved by the user is predicted to generate pseudo-labels to replace the true user feedback. During the learning process of the model, the loss function is used to measure the difference between the pseudo-labels and the click-through conversion rate predicted by the model , so as to improve the accuracy of the pseudo-labels. Its form can be expressed as:
[0064]
[0065] In the above formula, although ideally the pseudo-labels should be compared with the true user feedback to reduce the loss, due to the lack of the true labels of these unobserved data, this paper proposes a new idea, that is, using the click-through conversion rate predicted by the explicit feedback teacher model to compare with the pseudo-labels to narrow the gap between the two. The reason for doing this is that is obtained through comparison training with the true labels , and as the explicit feedback teacher model learns, will continuously approach the true labels and tend to be accurate. Therefore, by making compare with and narrow the gap during the learning process, will also gradually tend to be accurate.
[0066] Finally, the soft labels generated by the explicit feedback teacher model and the soft labels generated by the implicit feedback teacher model are combined to obtain unbiased soft labels to guide the learning of the student model.
[0067]
[0068] Among them, is a weight parameter used to adjust the contribution ratio of the explicit feedback soft labels and the implicit feedback soft labels in the final unbiased soft labels.
[0069] As a further implementation method, the unbiasedness analysis of the Debiasing-Knowledge Distillation method
[0070] In this section, we will analyze the unbiasedness of the Debiasing-KD method. By introducing the explicit feedback teacher model and the implicit feedback teacher model, the Debiasing-KD method aims to generate comprehensive unbiased soft labels , so as to guide the student model to learn and reduce the impact of selection bias. To verify the unbiasedness of this method, we will conduct a detailed discussion from a theoretical aspect. The following Lemma 1 is proposed in this paper:
[0071] Lemma 1. For each user , this paper supplements the causal path of the unobserved user data , thus generating unbiased soft labels that integrate the two paths . Therefore, knowledge distillation based on is approximately equal to knowledge distillation based on the total causal effect .
[0072] For the above Lemma 1, the proof is as follows:
[0073] Proof. Consider any two items and , and assume that for user , the probabilities of giving feedback or not giving feedback to these two items are the same, that is , so we have:
[0074]
[0075]
[0076]
[0077] Because is the comprehensive unbiased causal effect on path and path , the distilled soft label generated by the Debiasing-KD method is also unbiased. So far, Lemma 1 is proved, that is, the unbiasedness of the Debiasing-KD method is proved.
[0078] S5. Knowledge distillation is performed on the student model based on the unbiased soft labels to obtain a knowledge distillation debiased recommendation model.
[0079] Among them, the unbiased soft labels are obtained by weighted combination of the soft labels output by the two teacher models; and the unbiased soft labels are passed to the student model;
[0080] During the learning process, the student model uses unbiased soft labels as the fitting target. Since the unbiased soft labels encode the knowledge learned by the large teacher model, the student model can learn the rich knowledge in the unbiased soft labels during the learning process. The training loss function of the student model is as follows:
[0081]
[0082] where is the fitting loss function between the prediction value of the student model and the unbiased soft label.
[0083] In this embodiment, the Debiasing-KD method can theoretically be integrated into various debiasing models. In this section, we apply the Debiasing-KD method to the double-robust (DR) debiasing model for an instantiation study.
[0084] Traditional DR models mainly rely on the observed user feedback data during the learning process. Although some existing methods balance the data distribution through weighted loss functions to further reduce bias, traditional DR models have fewer parameters, weaker learning ability, and limited user feedback data volume, resulting in the DR model being difficult to achieve an ideal debiasing effect.
[0085] Therefore, this paper applies the Debiasing-KD method to the DR debiasing model to further reduce bias. Specifically, first, two large teacher models with more parameters and stronger learning ability are introduced, namely the explicit feedback teacher model and the implicit feedback teacher model, and the methods in Section 4.3 are used to pre-train these two large teacher models. Then, unbiased soft labels are generated . Finally, knowledge distillation is performed based on to train the DR model. The training loss of the DR model can be expressed as:
[0086]
[0087] where is the training loss between the prediction value of the DR model and the true label, is the distillation loss. β is a hyperparameter used to balance the contributions of the two losses.
[0088] Experimental verification:
[0089] Yahoo: This dataset is a dataset of user ratings for songs collected in the music recommendation scenario, including a biased subset and an unbiased subset. The biased subset contains 300,000 rating records of 1,000 songs by 15,400 users, which are collected from the users' normal historical interactions. The unbiased subset is collected from the interactions of the first 5,400 users among these 15,400 users. Each user is required to randomly select 10 songs from 1,000 songs for rating, and this random strategy ensures the unbiasedness of this subset. In addition, in this paper, the ratings are binarized with a threshold of 3, that is, ratings less than 3 are binarized to 0, otherwise 1.
[0090] Coat Shopping: This dataset is a user rating dataset collected from the Amazon platform, including a biased subset and an unbiased subset. The biased subset contains rating records of 300 coats by 290 users, which are collected from the users' normal historical interactions. The unbiased subset is obtained by randomly selecting 16 coats from these 290 users for rating. The way of binarizing the ratings of the dataset is the same as that of the Yahoo dataset.
[0091] KuaiRec: This dataset is derived from the recommendation logs of the Kuaishou platform, recording 4,676,570 viewing records of 3,237 videos by 1,411 users. In this paper, the dataset is binarized with a threshold of 2, that is, those less than 2 are set to 0, otherwise 1.
[0092] Four evaluation metrics are adopted to measure the debiasing performance: Mean Squared Error (MSE), Area Under the Curve (AUC), Recall@N, and Normalized Discounted Cumulative Gain (NDCG@N). On the Yahoo dataset and the Coat Shopping dataset, we set N to 5 and 10; while on the KuaiRec dataset, we set N to 20 and 50. All experiments are implemented on the PyTorch framework.
[0093] In the experiment, we selected the Graph Convolutional Network (GCN) and Neural Collaborative Filtering (NCF) models with relatively rich parameters for the teacher model to fully utilize their powerful representation capabilities and prediction accuracy. For the student model, a relatively simple Matrix Factorization (MF) model was used as the backbone model, aiming to achieve efficient knowledge transfer through knowledge distillation. This experiment followed previous research, and all models were trained using the Adam optimizer. The learning rate was adjusted in the range of {0.005, 0.01, 0.05, 0.1}, the weight decay was adjusted in the range of {1e-5, 5e-5, 1e-4, 5e-4, 1e-3, 5e-3, 1e-2}, the hyperparameter α was fixed at 0.5, the hyperparameter β was in the range of {0.1, 0.2, 0.3, …, 1.0}, the batch size on the Coat Shopping dataset was {128, 256, 512, 1024, 2048}, and the batch size on the Yahoo and KuaiRec datasets was {1024, 2048, 4096, 8192, 16384}.
[0094] In this section, we verified the effectiveness of the Debiasing-KD debiasing method on 8 baseline models through comparative experiments.
[0095] Comparative experiment analysis:
[0096] Tables 1 and 2 show the comparison test results with other baseline models on three datasets. The model applying the Debiasing-KD debiasing method outperformed the baseline models in all cases. Specifically, compared with the latest baseline model BEDR-JL, on the Yahoo dataset, after applying the Debiasing-KD debiasing method, the MSE, AUC, Recall@5, NDCG@5, and NDCG@10 increased by 1.4%, 0.4%, 0.8%, 0.83%, and 0.44% respectively; on the Coat Shopping dataset, they increased by 0.9%, 1.2%, 3.6%, 3.4%, and 1.8% respectively; on the KuaiRec dataset, the MSE, AUC, NDCG@20, NDCG@50, Recall@20, and Recall@50 increased by 3.17%, 2.6%, 1.48%, 1.18%, 4.08%, and 2.19% respectively.
[0097] This indicates that the Debiasing-KD method effectively solves the selection bias problem in the knowledge distillation process by supplementing the causal path of unobserved feedback. This method can learn unbiased soft labels through the teacher model and effectively guide the student model to reduce bias, thus significantly improving the accuracy and robustness of the recommendation system.
[0098] Table 1 Comparative experimental results on Yahoo and Coat Shopping datasets
[0099]
[0100] Table 2 Comparative experimental results on KuaiRec dataset
[0101]
[0102] Example 2
[0103] This example provides a causal theory-based knowledge distillation debiasing recommendation system, including:
[0104] A data acquisition module, configured to:
[0105] A computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the causal theory-based knowledge distillation debiasing recommendation method.
[0106] A terminal device, including a processor and a computer-readable storage medium, the processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the causal theory-based knowledge distillation debiasing recommendation method.
[0107] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A knowledge distillation debiasing recommendation method based on causal theory, characterized in that including: Obtain a user rating dataset; Construct a knowledge distillation debiased recommendation model, in which, construct two teacher models and one student model, generate the causal path of the observed data according to the user rating dataset; supplement the causal path of the unobserved data based on the causal effect of quantifying the selection bias; use the two teacher models to make predictions on the causal path of the observed data and the causal path of the unobserved data respectively to generate unbiased soft labels; Perform knowledge distillation on the student model based on the unbiased soft labels to obtain a knowledge distillation debiased recommendation model; Use the knowledge distillation debiased recommendation model to make user predictions; The obtaining of user observed rating data includes obtaining the observed rating data of users for products. Among them, let be a set of M users, and be a set of N items; is expressed as the true feedback label matrix collected from the user interaction history, and each item in it represents that user gives a true rating to item The generating the causal path of the observed data according to the user rating dataset includes inputting the user observed rating data into the display feedback teacher model, and the display feedback teacher model makes predictions according to the observed data to generate soft labels, and the path of generating the soft labels is used as the causal path of the observed data; The causal path for supplementing unobserved data with causal effects based on quantized selection bias includes solving click data The biased causal effect formula and unobserved data in the case of The causal effect formula in the case of, where Denote the random variable under the conditional distribution Among them, the variable Is intervened to a specific value The variable To The causal effect of is expressed as as the Value changes, The change of the value. Therefore, for the user Along the path When To The biased causal effect formula of is expressed as: , Among them represents the baseline case for comparison, and this formula represents the difference under the condition of and the baseline case ; For the user when following the path the causal effect formula for the unbiased soft label is expressed as the difference between the condition and the baseline case and is specifically expressed as: , The supplementing the causal path of the unobserved data based on the causal effect of quantifying the selection bias further includes obtaining an unbiased causal effect formula by combining the biased causal effect formula and the causal effect formula, and synthesizing the unbiased causal effects of the two paths through the unbiased causal effect formula for generating unbiased soft labels, and the unbiased causal effect formula is expressed as: , wherein, the user ; the article ; represents a reference case; The above method of using two teacher models to make predictions on the causal paths of observed data and unobserved data respectively to generate unbiased soft labels includes debiased knowledge distillation based on the Debiasing-KD method. Specifically, the causal path obtained by using the explicit feedback teacher model to act on the observed user feedback data; the causal path obtained by using the implicit feedback teacher model to act on the unobserved user feedback data, and finally generate unbiased soft labels. Among them, two teacher models are introduced: the explicit feedback teacher model and the implicit feedback teacher model, which act on the causal paths respectively and , thereby generating comprehensive unbiased soft labels, where The display feedback teacher model directly generates soft labels on the observed user feedback and its learning process is expressed as: (1) Among them, is the soft label learned by the teacher model; For the implicit feedback teacher model, the click-through conversion rate of unobserved items by users is estimated to generate pseudo-labels to replace the real user feedback; during the learning process of the model, the loss function is used to measure the pseudo-labels and the click-through conversion rate predicted by the model The difference between them is used to improve the accuracy of the pseudo-labels, and its form is expressed as: , Finally, the soft labels generated by the display feedback teacher model and the soft labels generated by the implicit feedback teacher model are combined to obtain unbiased soft labels to guide the learning of the student model; , Among them, is a weight parameter used to adjust the contribution ratio of the display feedback soft label and the implicit feedback soft label in the final unbiased soft label; The performing knowledge distillation on the student model based on the unbiased soft labels to obtain a knowledge distillation debiased recommendation model includes using the student model to learn with the unbiased soft labels as the fitting target, based on the knowledge learned by the teacher model encoded by the unbiased soft labels, and making the student model learn the knowledge of the teacher model in the unbiased soft labels during the learning process based on the loss function, and the training loss function of the student model is expressed as: , where is the fitting loss function between the predicted value of the student model and the unbiased soft label; is the training loss between the predicted value of the DR model and the true label, is a hyperparameter used to balance the contributions between the two losses.
2. A knowledge distillation debiased recommendation system based on causal theory, which executes a knowledge distillation debiased recommendation method based on causal theory as described in claim 1, characterized in that, including: A data acquisition module configured to obtain a user rating dataset; A model construction module configured to construct a knowledge distillation debiased recommendation model, in which, construct two teacher models and one student model, generate the causal path of the observed data according to the user rating dataset; supplement the causal path of the unobserved data based on the causal effect of quantifying the selection bias; use the two teacher models to make predictions on the causal path of the observed data and the causal path of the unobserved data respectively to generate unbiased soft labels; A knowledge distillation module configured to perform knowledge distillation on the student model based on the unbiased soft labels to obtain a knowledge distillation debiased recommendation model; A prediction module configured to use the knowledge distillation debiased recommendation model to make user predictions.
3. A computer-readable storage medium storing multiple instructions, characterized in that, The instructions are adapted to be loaded and executed by a processor of a terminal device to perform the method according to claim 1.
4. A terminal device, comprising a processor and a computer-readable storage medium, the processor being configured to implement each instruction; the computer-readable storage medium being configured to store a plurality of instructions, characterized in that, The instructions are adapted to be loaded and executed by a processor to perform the method according to claim 1.
Citation Information
Patent Citations
Method for relieving backdoor attack by using label-free data
CN115879538A
Neural network distillation method and device
CN116249991A