Personalized news recommendation method based on causal intervention

Through causal intervention technology, the impact of popularity between news and users is decoupled, and the problem of popularity confusion in the news recommendation system is solved, more accurate and personalized news recommendation is achieved, and AUC and MRR indicators are improved.

CN120492716APending Publication Date: 2025-08-15TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510493139.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the existing news recommendation system, the confusing effect of popularity characteristics in the process of user personalized modeling leads to a decrease in recommendation accuracy, and the existing debiasing methods have failed to effectively alleviate the separation of news and user representation.

Method used

Through causal intervention technology, the influence of popularity of news and users is decoupled, and news representation learning, user representation learning and popularity decoupling training are used to design multi-task learning loss functions and causal frameworks to remove the confusing influence of popularity on news and user modeling, and reorder it through causal reasoning.

Benefits of technology

It effectively alleviates the data deviation problem caused by popularity, improves the accuracy and personalization of news recommendations, improves the AUC and MRR indicators, and recommends news that is more in line with user interests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492716A_ABST
    Figure CN120492716A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of news recommendation, and provides a personalized news recommendation method based on causal intervention. According to the method, firstly, features of news are constructed through news titles, categories and news click behaviors; secondly, combining popularity features with a candidate perception recommendation model, and learning a news recommendation model containing popularity deviation; then, a causal intervention technology is introduced, false correlation between news features and popularity is cut off through a do-dry budget, and interference of popularity deviation on user interest modeling is effectively restrained; and finally, an inference mechanism based on intervention is designed, so that the adverse effect of news popularity deviation on user modeling is weakened. According to the method, the correlation of the user news is reordered by utilizing a causal intervention technology, so that the accuracy of news recommendation is improved, and technical support is provided for a personalized news recommendation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of news recommendation, and in particular relates to a personalized news recommendation method based on causal intervention. Background Art

[0002] The rise of digital news platforms has greatly expanded the audience engaging with online news. To alleviate the problem of information overload, news recommendation systems have emerged. These technologies, which help recommend news that matches each user's unique preferences, are attracting increasing attention.

[0003] Due to the dynamic and interactive nature of news platforms, data bias is particularly prominent in the news recommendation field. The popularity distribution of news content often exhibits a long-tail characteristic, meaning a small number of trending news stories capture the majority of user attention, while a large number of niche or emerging topics receive little attention. If left unaddressed, this uneven distribution can lead recommendation systems into a "popularity trap." From a news perspective, excessive exposure to trending news can hinder the cold start of less popular news, hindering its effective exposure and potentially leading to a "Matthew effect." From a user's perspective, the system may favor highly popular news over less popular content that may align with a user's potential interests, thus impacting recommendation effectiveness.

[0004] However, existing methods have the following two problems:

[0005] 1. Most methods simply incorporate popularity as an additional feature into the news and user personalization modeling process without considering the confounding effect of popularity. This confounding effect can lead to data bias in the recommendation system, resulting in reduced recommendation accuracy.

[0006] 2. Popularity debiasing methods in other recommendation domains focus on mitigating the impact of popularity on news exposure. This approach, which reweights news popularity, ignores the fact that popularity can also confound news representations and user personalized modeling during the recommendation process, leading to a separation between news and user representations based on popularity.

[0007] In order to solve the above problems, the present invention proposes a user modeling debiasing method based on causal intervention to alleviate the bias impact of popularity in the recommendation system on both news and user modeling. Summary of the Invention

[0008] To address the shortcomings of current solutions, this paper proposes a news recommendation method that mitigates the influence of popularity on user modeling. To mitigate the confounding effect of popularity on personalized user modeling, this paper decouples news, users, and popularity through popularity perception. Furthermore, through causal intervention, it cuts off the backdoor path of popularity to news and users, resulting in a recommendation model that eliminates popularity bias.

[0009] The present invention provides a personalized news recommendation method based on causal intervention, comprising the following steps:

[0010] S1 News Representation Learning: For the title features and category features of news, the classic pre-trained word embedding model BERT and the vector production model GloVe are used to learn the intrinsic semantic information of the title and category, and obtain the semantic representation of the title. t and semantic representation of categories c , the overall characterization of the last news news is the aggregation of these two representations;

[0011] In addition to the features of the text content, the present invention also calculates the number of clicks on the news as the popularity of the news on a daily basis, and maps it into a dense vector e after bucketing through a Dense network. pop ;

[0012] S2 User Representation Learning: When modeling user interests, the representation of user interests should have different tendencies depending on the candidate news. Therefore, this method uses a candidate-aware user encoder to learn personalized user representations based on candidate news from the user's historical interaction sequences.

[0013] First, the news in the candidate user’s historical interaction sequence is converted into a vector sequence through the S1 news representation learning method. Then, a candidate-perceived interactive attention network is used to learn the correlation between historical news representations and candidate news. Finally, the correlation is converted into weights through the softmax function, and the news representations in the historical sequence are weighted and summed according to the weights to obtain the candidate-perceived user-personalized representation e user .

[0014] S3 Popularity Decoupling Training: This method decouples the confusing effects of news popularity on user modeling, calculates the relevance between users and news through a click prediction module, and trains the model to recommend highly relevant news.

[0015] The detailed steps for S3 are as follows:

[0016] S3-1 decouples the confusing influence of user and news popularity by modeling the impact of popularity on news and user personalization. The confusing influence of candidate news is reflected as the exposure of news, which is decoupled into the popularity vector of news e n-pop The influence of popularity on users is learned from the popularity representation of users' historical news through the recurrent neural network GRU.

[0017] S3-2 Click Prediction Module, after decoupling the influence of popularity on news and user confusion, designs an aggregation module to aggregate the influence of popularity with news representation and user representation to obtain news representation and user representation with popularity confusion, where the aggregation method is Hadamard product. The estimated probability of user click is s u,n It is obtained by the inner product of the two representations.

[0018] S3-3 Design a multi-task learning loss function to guide model training. Using contrastive learning loss L nce Optimize the model parameters so that the model can learn the correlation between users and positive sample news. At the same time, introduce the regression loss mean square error L mse Learn the click-through rate behavior of users. Using the multi-task learning idea, the overall loss function is L = λ·L nce +(1-λ)·L mse .

[0019] S4 Causal Framework Construction: Causal graphs play an important role in removing the causal influence of popularity in news recommendation frameworks. They provide a detailed representation of the various factors in the news recommendation task and the causal links between them. In a causal graph, nodes represent key variables, and sidebands represent the correlations between variables.

[0020] The main variables in news recommendation include: ① Node N represents candidate news, that is, the news presented to the user; ② Node U represents the platform user, including the news that the user has previously clicked; ③ Node I represents the degree of user interest matching; ④ Node P represents the popularity of the news; ⑤ Node M represents the user click probability.

[0021] The relationship between these variables is as follows: ① Popularity P affects whether candidate news N is displayed; ② Interest matching I is affected by the user's historical clicked news U and the specific candidate news N. ③ The user's click probability M is determined by the user's interest matching degree I and the popularity P of the news.

[0022] In this causal graph, popularity becomes a backdoor path between news and users, introducing data bias into news and user modeling. To mitigate the confounding effect of popularity P on news N and the resulting user interest match I, this paper decouples the influence of popularity and, through causal intervention analysis in the inference phase, calculates the actual click probability of users after removing the confounding effect of popularity P on news.

[0023] S5 Popularity Debiasing Inference: In order to not affect the direct effect of popularity on end-user clicks and remove the confusing influence of popularity on the news and user modeling process, the present invention calculates the similarity between user and news representations without popularity influence in the inference stage and reweights them according to the popularity of the news, which can effectively remove the confusing influence brought by popularity.

[0024] Finally, the present invention calculates the estimated score of user click probability, which can remove the confusing influence of popularity on user personalized modeling and news, and provide a more personalized news recommendation service.

[0025] Beneficial effects

[0026] 1. The causal debiasing recommendation model of the news recommendation system of the present invention can effectively alleviate the data bias problem caused by popularity.

[0027] 2. The present invention simulates the causal relationship between user click history, candidate news, news popularity and user-news interaction.

[0028] By implementing causal intervention techniques, the possible adverse effects of popularity are effectively separated from user interest modeling.

[0029] 3. In order to enhance the personalized recommendation capability of the model, an inference method based on causal intervention is designed to re-rank the estimated intervention results and more unbiasedly estimate the estimated probability of users clicking on news.

[0030] 4. The present invention effectively improves the accuracy of news recommendations. At the same time, by decoupling the influence of popularity, it can separate the confusing influence of popularity at the news and user levels, and recommend news that better meets the personalized needs of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is the model structure of the present invention. DETAILED DESCRIPTION

[0032] The present invention will be further described below with reference to the accompanying drawings.

[0033] S1 News Representation Learning: This paper uses the classic pre-trained word embedding model BERT and the vector production model GloVe to learn the intrinsic semantic information of the title and category for the news title features and category features, and obtains the semantic representation of the title. t and semantic representation of categories c .

[0034] like Figure 1 As shown, the text features of the news include title t = [w1, w2, ..., w l ] and category c, where w i is a specific word. Input it into the BERT and GloVe models to obtain the representation of the title and category:

[0035] e t =BERT([CLS]w1,w2,...,w l )

[0036] e c =GloVe(c)

[0037] Finally, the overall representation of news is the aggregation of these two representations, namely:

[0038] e news =[e t :e c ]

[0039] In addition to the features of the text content, the present invention also calculates the popularity of the news on a daily basis and divides it into buckets of equal frequency from small to large according to the popularity, and divides it into p buckets [bucket1, bucket2, ..., bucket p After discretizing the popularity, it is mapped into a dense vector e through the Dense network. pop ,Right now:

[0040] e pop =Dense(bucket i )

[0041] S2 User Representation Learning: When modeling user interests, the representation of user interests should vary depending on the candidate news. Therefore, this method uses a candidate-aware user encoder to learn personalized user representations based on candidate news from the user's historical interaction sequences.

[0042] Figure 1 As shown, first, the news in the candidate user’s historical interaction sequence is converted into a vector sequence through the S1 news representation learning method. Then, a candidate-aware interactive attention network is used to learn the correlation between historical news representations and candidate news. The correlation function φ(·,·) is calculated as follows:

[0043] φ(e news ,e his )=GELU(e news ·e his )

[0044] GELU(·) is the activation function. Finally, the correlation is converted into weights through the softmax function, and the news representations in the historical sequence are weighted and summed according to the weights to obtain the user personalized representation of the candidate perception e user ,Right now:

[0045]

[0046] S3 Popularity Decoupling Training: This method decouples the confusing effects of news popularity on user modeling, calculates the relevance between users and news through a click prediction module, and trains the model to recommend highly relevant news.

[0047] The detailed steps for S3 are as follows:

[0048] S3-1 decouples the confusing influence of user and news popularity by modeling the impact of popularity on news and user personalization. The confusing influence of candidate news is reflected as the exposure of news, which is decoupled into the popularity vector of news e n-pop The influence of popularity on users is learned from the popularity representation of historical news through the recurrent neural network GRU, that is,

[0049] S3-2 Click Prediction Module, after decoupling the influence of popularity on news and user confusion, designed an aggregation module to aggregate the influence of popularity with news representation and user representation to obtain news representation and user representation with popularity confusion, where the aggregation method is Hadamard product. The estimated probability of user click is s u,n Then we can get it by taking the inner product of the two representations:

[0050] s u,n =(e n-pop ⊙e news )·(e u-pop ⊙e user )

[0051] Where ⊙ represents the vector Hadamard product operation and · represents the vector inner product operation.

[0052] S3-3 designed a multi-task learning loss function to guide the training of the model. Using contrastive learning loss L nceOptimize model parameters so that it learns the correlation between users and positive sample news.

[0053]

[0054] in, and Represents the estimated click probability of positive samples and negative samples, D - Represents the negative sample set. At the same time, the regression loss mean square error L is introduced mse To learn users’ click-through rate behavior from popularity vector.

[0055] L mse =(e u-pop W T +b-ctr u ) 2

[0056] W and b are both trainable parameters, and ctr u is the user’s click rate. Using the multi-task learning idea, the overall loss function is L = α·L nce +(1-α)·L mse .

[0057] S4 builds a causal framework: Causal graphs play a crucial role in removing the causal influence of popularity in news recommendation frameworks. They provide a detailed representation of the various factors in the news recommendation task and the causal links between them. In a causal graph, nodes represent key variables, while edges represent the correlations between variables.

[0058] The main variables in news recommendation include: ① Node N represents candidate news, that is, the news presented to the user; ② Node U represents the platform user, including the news that the user has previously clicked; ③ Node I represents the degree of user interest matching; ④ Node P represents the popularity of the news; ⑤ Node M represents the user click probability.

[0059] The relationship between these variables is as follows: ① Popularity P affects whether candidate news N is displayed; ② Interest matching I is affected by the user's historical clicked news U and the specific candidate news N. ③ The user's click probability M is determined by the user's interest matching degree I and the popularity P of the news.

[0060] In this causal graph, popularity becomes a backdoor path between news and users, introducing data bias into news and user modeling. To mitigate the confounding effect of popularity P on news N and the resulting user interest match I, this paper decouples the influence of popularity and, through causal intervention analysis in the inference phase, calculates the actual click probability of users after removing the confounding effect of popularity P on news.

[0061] S5 Popularity Debiasing Inference: Considering the impact of popularity on the final click, a reweighting approach is used to correct the final result. By decoupling the confounding effect of popularity and calculating correlation using personalized news and user representations, an unbiased estimated score is obtained. The final estimated score can be calculated using the following formula:

[0062]

[0063] Where p is any bucket label, and is converted into a reweighted weight through the weight function f(p) and participates in the model calculation. The weight function is designed as:

[0064]

[0065] Where σ is a model hyperparameter, and AC(p) is the average number of clicks for buckets with bucket label p. After obtaining unbiased user click prediction scores, the recommendation model ranks candidate news from highest to lowest according to the predicted scores and displays the top-k news items to the user. To evaluate the effectiveness of the model, the AUC and MRR metrics are used. AUC represents the probability that the predicted score of a positive sample is greater than that of a negative sample. A higher AUC indicates better recommendation efficiency. MRR evaluates the reciprocal of the ranking of the first positive sample in the dataset; a higher MRR indicates a higher ranking of the positive sample. Through the dual evaluation of AUC accuracy and MRR, the accuracy of the model in recommending news of interest to users can be comprehensively measured.

[0066] Finally, the present invention integrates the textual semantic features of news and users with the dynamic popularity features, designs a popularity confusion decoupling module, and effectively separates the influence of popularity on the representation of users and news. Through causal intervention reasoning, on the widely used news dataset MIND, compared with the baseline model UNBERT, the AUC and MRR are improved by 1.36% and 2.06% respectively, significantly improving the accuracy of recommending personalized news to users.

Claims

1. A personalized news recommendation method based on causal intervention, characterized in that: The steps include: S1 News Representation Learning: For the title features and category features of news, the classic pre-trained word embedding model BERT and the vector production model GloVe are used to learn the intrinsic semantic information of the title and category, and obtain the semantic representation of the title. t and semantic representation of categories c , the overall characterization of the last news news is the aggregation of these two representations; In addition to the features of the text content, the popularity of the news is calculated by counting the number of clicks on the news on a daily basis, and after bucketing, it is mapped into a dense vector e through the Dense network. pop ; S2 User Representation Learning: Through the candidate-aware user encoder, we learn personalized user representations based on candidate news from the user’s historical interaction sequence: S3 popularity decoupling training: This decouples the confounding effects of news popularity on user modeling. The click prediction module calculates the relevance between users and news, and trains the model to recommend highly relevant news. S4 Causal framework construction: In the causal graph, nodes represent key variables, and sidebands represent the correlation between variables; The main variables in news recommendation include: ① Node N represents candidate news, that is, the news presented to the user; ② Node U represents the platform user, including the news that the user has previously clicked; ③ Node I represents the degree of user interest matching; ④ Node P represents the popularity of the news; ⑤ Node M represents the user click probability; The following relationships exist between these variables: ① Popularity P affects whether candidate news N is displayed; ② Interest matching I is affected by the user's historical clicked news U and the specific candidate news N; ③ The user's click probability M is determined by the user's interest matching degree I and the popularity P of the news; The data bias problem is introduced into news and user modeling. By decoupling the influence of popularity and performing causal intervention analysis in the inference phase, the actual click probability of users is calculated after removing the confounding effect of popularity P on news. S5 Popularity Debiasing Inference: In the inference stage, the similarity between user and news representations without popularity influence is calculated, and then reweighted according to the popularity of the news to remove the confusing effect caused by popularity.

2. The method according to claim 1, characterized in that The S2 is specifically as follows: First, the news in the candidate user’s historical interaction sequence is converted into a vector sequence through the S1 news representation learning method. Then, the correlation between historical news representation and candidate news is learned through the candidate-aware interactive attention network; Finally, the correlation is converted into weights through the softmax function, and the news representations in the historical sequence are weighted and summed according to the weights to obtain the user personalized representation of the candidate perception e user .

3. The method according to claim 1, characterized in that The S3 is as follows: S3-1 decouples the confusion between user and news popularity by modeling the impact of popularity on news and user personalization. The confusion of candidate news is reflected as the exposure of news, which is decoupled into the popularity vector e of news. n -pop The influence of popularity on users is learned from the popularity representation of users’ historical news through the recurrent neural network GRU; S3-2 Click Prediction Module, after decoupling the influence of popularity on news and user confusion, designs an aggregation module to aggregate the influence of popularity with news representation and user representation to obtain news representation and user representation with popularity confusion. The aggregation method is Hadamard product, and the estimated probability of user click is s u,n Then it is obtained by the inner product of the two representations; S3-3 Design a multi-task learning loss function to guide model training: Using contrastive learning loss L nce Optimize the model parameters so that the model can learn the correlation between users and positive sample news. At the same time, introduce the regression loss mean square error L mse Learn users’ click-through rate behavior; Using the multi-task learning idea, the overall loss function is L = λ·L nce +(1-λ)·L mse .

4. The method according to claim 1, wherein Calculate the estimated user click probability score to remove the confounding effect of popularity on user personalized modeling and news.