A Long-Term and Short-Term Recommendation Method Incorporating Hierarchical Structure
By using hierarchical structure and GRU network in news recommendations, we have solved the problem that existing recommendation methods cannot distinguish between users' long and short-term preferences, and achieved more accurate user feature construction and personalized recommendation effects.
Patent Information
- Application Number
- CN202210390624.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-14
AI Technical Summary
The existing recommendation method only learns the single representation of the user, and cannot effectively distinguish between the long-term and short-term preferences of users, and has shortcomings in the diversification of user interests and multi-grain mining.
A long-term and short-term recommendation method integrated into a hierarchical structure is adopted to extract news representations through a news encoder, a three-level hierarchical structure extracts users' long-term interest representations, and a gated recurrent unit (GRU) extracts users' short-term interest representations, and calculates matching probability through vector internal product for recommendation.
Through experimental verification, the recommendation effect is better than that of traditional models, and indicators such as AUC, MRR, nDCG@5 and nDCG@10 have been improved, which can more accurately construct user characteristics and provide personalized recommendations.
Smart Images

Figure CN114722287B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a recommendation method based on user interests, and more specifically to a long-term and short-term recommendation method incorporating a hierarchical structure. Background Art
[0002] News recommendation is an important task in the field of natural language processing and has received increasing attention from scholars in recent years. For news recommendation, it is crucial to learn accurate user and news representations. Early news recommendation methods often relied on the association and semantic similarity between news, but these methods usually had difficulty effectively modeling users' reading preferences for personalized recommendations. The collaborative filtering algorithm is one of the earliest studied recommendation techniques, which has greatly promoted the development of personalized recommendations and is often used in news recommendation. However, the news recommendation method based on collaborative filtering has serious cold start problems, so many scholars have turned their attention to content-based recommendations. For example, Phelan et al. proposed to combine users' behaviors on Twitter with news browsing records for user modeling; Liu et al. proposed to represent news and users with news categories and user interest features generated by a Bayesian model respectively for news recommendation. However, in these traditional methods, constructing user and news representations usually relies on manually designed features and requires a large amount of domain knowledge and time.
[0003] In recent years, with the successful application of deep learning technology in fields such as image recognition and text classification, the research on combining it with recommendation technology has also received extensive attention from scholars. For example, Lian et al. proposed a news recommendation method based on a DeepFusionModel (DFM), which represents news and users by constructing features. Their method includes two core modules, one module is used to model different interaction information between features, and the other module is used to assign different weights to features in different channels and achieved good results on the Bing dataset. However, this method only uses coarse-grained information such as title length and entity names to model news representation and does not mine finer-grained semantic information. Wang et al. proposed to fuse a knowledge graph and a convolutional neural network, then learn news representation from the title, calculate the similarity between candidate news and historical articles browsed by users, and use the result as an attention weight to perform weighted summation on the news representations in the user's browsing history to obtain the user representation. Wu et al. proposed a news recommendation method of a personalized attention network, which uses user ID embedding to generate an attention query vector. However, the above two methods only learn a single representation of the user and cannot distinguish the long-term and short-term preferences of the user, which is far from sufficient for accurately learning the user representation. Summary of the Invention
[0004] In view of the deficiencies of the existing technologies, the present invention proposes a long-term and short-term recommendation method incorporating a hierarchical structure, aiming to solve the problems that existing recommendation methods only learn a single representation of users and lack in diverse user interests and multi-granularity mining. The technical solution of the present invention mainly includes the following steps:
[0005] 1. Extract news features: Use a news encoder to learn news titles, topics, and sub-topics, and then extract news representations; 2. Extract long-term behavioral features of users: Adopt a three-level hierarchical structure to obtain long-term interest representations of users. The bottom layer is used to obtain interest representations at the sub-topic level, the middle layer is used to obtain interest representations at the topic level, and the top layer is used to obtain long-term interest representations at the user level; 3. Extract short-term behavioral features of users: Use a Gated Recurrent Unit (GRU) to obtain short-term interest representations of users, and initialize the GRU with the long-term interest representations. The obtained short-term interest representations are the final user representations; 4. Calculate the matching probability and complete the recommendation: Match the final user representations with candidate news representations to obtain a recommendation list and complete the recommendation.
[0006] The effects of the present invention are as follows: The method of the present invention is experimentally verified by applying it to the MIND dataset. The AUC, MRR, nDCG@5, and nDCG@10 values of the optimal experimental results are 60.84%, 29.44%, 31.45%, and 39.58% respectively, and the recommendation effect is better than that of traditional models. Description of the Drawings
[0007] Figure 1 Model Structure Diagram
[0008] Figure 2 Three-Level Hierarchical Structure Diagram Detailed Embodiments
[0009] To solve the problems that existing recommendation methods usually only learn a single representation of users and lack in diverse user interests and multi-granularity mining, the solution provided by the present invention is: a long-term and short-term recommendation method incorporating a hierarchical structure. First, use a news encoder to extract news representations, and then use a three-level hierarchical structure to extract long-term interest representations of users. This structure from the bottom layer to the top layer is: sub-topic level interest representation layer, topic level interest representation layer, user level long-term interest representation layer. Then use GRU to obtain short-term preferences from the user's recent browsing history, and at the same time use the long-term interest representations of users to initialize the GRU, strengthen the influence of long-term preferences, and obtain the final user representations. Finally, according to the user representations and candidate news representations, obtain the probability scores of the news that the user may click through the vector inner product method, and then obtain the recommendation list. The structure diagram of this method is as Figure 1 shown:
[0010] Figure 1Model structure diagram
[0011] (1) Extract news features
[0012] First, the news title is converted into a vector sequence through word embedding, and then this vector sequence is input into a convolutional neural network to capture local context information and learn the word representation of the context. Then, a word-level attention mechanism is used to select important words in the title to obtain the title representation. Finally, the title representation, the topic and subtopic representations are concatenated to obtain the final representation of the news. The concatenation formula is as follows:
[0013] n = concat(n t , n v , n sv )
[0014] where n is the final representation of the news, n t is the news title representation, n v is the news topic representation, obtained through the embedding of topic words, and n sv is the subtopic representation, obtained through the embedding of subtopic words.
[0015] (2) Extract the long-term behavior features of users
[0016] The three-level hierarchical structure diagram for extracting the long-term behavior features of users is as Figure 2 shown.
[0017] Figure 2 Three-level hierarchical structure diagram
[0018] The subtopic-level interest representation layer is used to obtain fine-grained user interests and consists of multiple subtopic-level interest representations. It learns through the user's browsing history in subtopic-class news (such as all the sub-news of basketball under the sports topic). A subtopic-level attention network is used to obtain the important news vector representation c ij , and then the vector representation s ij of the subtopic word is obtained through word embedding. Finally, these two vector representations are fused to obtain the subtopic-level interest representation. The formula is as follows:
[0019]
[0020]
[0021] where α k is the attention weight of the k-th clicked news in N ij , is the vector representation of the k-th clicked news in N ij , N ij represents the set of all news corresponding to the j-th subtopic of the i-th clicked topic, For the final sub-topic level interest representation.
[0022] The topic-level interest representation layer is used to obtain the coarse-grained user interest, which is composed of multiple topic-level interest representations and is learned from the sub-topic-level interest representations. A topic-level attention network is used to obtain the important sub-topic-level user interest vector representation z i , and then the vector representation t of the topic word is obtained by word embedding i . Finally, these two vector representations are fused to obtain the topic-level interest representation, and the formula is as follows:
[0023]
[0024]
[0025] where β j represents the attention weight of the sub-topic level interest representation , and is the topic-level interest representation.
[0026] The user-level long-term interest representation is learned from the topic-level interest representation. Similar to the previous two representations, the user-level attention network is used here to select the important topic vector representation, which is the final long-term interest representation. The formula is as follows:
[0027]
[0028] u l is the user-level long-term interest representation, and γ i is the attention weight of the topic-level interest representation .
[0029] (3) Extract the short-term behavior features of the user
[0030] Learn the short-term interest representation of the user from the news that the user has recently browsed. Input the vector representation of the news sequence into the GRU to capture the sequential news reading pattern, that is, arrange the news that the user has recently browsed in ascending order of time stamps. At the same time, use the user's long-term interest representation to initialize the GRU to strengthen the influence of long-term preferences, and the obtained short-term interest representation is the user's final representation.
[0031] (4) Calculate the matching probability and complete the recommendation
[0032] The probability score of the user clicking on the candidate news is obtained by taking the vector inner product of the user's final representation and the candidate news representation, and then the recommendation list is obtained.
[0033] Example 1: News recommendation
[0034] News recommendation is a method of recommending news to users according to their preferences. The similarity between the user representation and the candidate news representation can be calculated to obtain a recommendation list. The evaluation metrics for the recommendation results are AUC, MRR, nDCG@5, and nDCG@10. The formula for AUC is:
[0035]
[0036] where rank is the rank of the sample prediction value, M and N are the number of positive samples and negative samples respectively, and p is the number of samples.
[0037] The formula for MRR is:
[0038]
[0039] where Q is the number of queries. If the first positive news sample is ranked at the rank-th position, the score of MRR is The formula for nDCG@K is:
[0040]
[0041]
[0042]
[0043] where rel i represents the true relevance score of the i-th result, IDCG is the ideal DCG, and |REL| represents the number of elements in the set composed of the top K results sorted by true relevance from large to small.
[0044] Table 1: News Recommendation Task
[0045]
[0046]
[0047] It can be observed from Table 1 that all indicators of the present invention are improved compared with other models. The reasons are as follows: (1) The present invention separately learns the long-term interest and short-term interest of users. Compared with the single learning of user representation in the baseline method (such as Wide&Deep, DeepFM, DFM), the present invention can construct user features more accurately. (2) When extracting the long-term interest preference of users, the present invention uses a hierarchical structure to represent, fully learning the diverse and multi-granularity interest features of users; when extracting the short-term interest preference of users, the time sequence is taken into account. Compared with DeepFM that does not consider it, the recommendation effect is significantly improved. (3) Different from other methods (such as CNN, DKN) that only learn the title features, the present invention combines the news title, theme and sub-theme to learn the news features.
[0048] In summary, the present invention proposes a long-term and short-term recommendation method incorporating a hierarchical structure, which can separately learn the long-term interest and short-term interest of users. The long-term interest representation is obtained using a three-level hierarchical structure. The bottom layer is used to obtain the interest representation at the sub-theme level, the middle layer is used to obtain the interest representation at the theme level, and the top layer is used to obtain the long-term interest representation at the user level. The short-term interest is learned from the user's recent browsing history using a GRU network, and the GRU is initialized with the long-term interest representation of the user to strengthen the influence of long-term preferences and optimize the model training effect. Finally, the effectiveness of the present method is verified on the MIND dataset, and the recommendation effect of the present invention is better than that of other models.
[0049] The above examples are only illustrative of the present invention and do not constitute a limitation on the protection scope of the present invention. Any design identical or similar to the present invention belongs to the protection scope of the present invention.
Claims
1. A long - term and short - term recommendation method integrating a hierarchical structure, characterized in that, Including the following steps: (1) Extract news features: First, convert the news title into a vector sequence through word embedding, then input the vector sequence into a convolutional neural network to capture local context information and learn the word representation of the context. Then, use the word-level attention mechanism to select important words in the title to obtain the title representation. Finally, concatenate the title representation, the topic and subtopic representations to obtain the final representation of the news. The concatenation formula is as follows: n = concat(n t , n v , n sv ) Among them, n is the final representation of the news, n t is the representation of the news title, n v is the representation of the news theme, obtained through the embedding of the theme words, n sv is the representation of the sub-theme, obtained through the embedding of the sub-theme words; (2) Extract the long-term behavior features of users: Adopt a three-level hierarchical structure to obtain the long-term interest representation of users. The bottom layer is used to obtain the interest representation at the subtopic level, the middle layer is used to obtain the interest representation at the topic level, and the top layer is used to obtain the long-term interest representation at the user level; The interest representation layer at the sub-topic level is used to obtain fine-grained user interests. It consists of multiple sub-topic level interest representations. It learns through the user's browsing history in sub-topic category news and uses a sub-topic level attention network to obtain an important news vector representation c ij , and then obtains the vector representation s of the sub-topic word by using word embedding ij , and finally fuses these two vector representations to obtain the interest representation at the sub-topic level. The formula is as follows: Among them, α k is the attention weight of the k-th clicked news in N ij , is the vector representation of the k-th clicked news in N ij , N ij represents the set of all news corresponding to the j-th sub-topic of the i-th clicked topic, is the final sub-topic level interest representation; The topic-level interest representation layer is used to obtain coarse-grained user interests. It consists of multiple topic-level interest representations, which are learned from sub-topic-level interest representations. A topic-level attention network is adopted to obtain an important sub-topic-level user interest vector representation z i , and then the vector representation t of the topic word is obtained by word embedding i , and finally, these two vector representations are fused to obtain the topic-level interest representation. The formula is as follows: Among them, β j represents the attention weight of the sub-topic level interest representation ; and is the topic level interest representation. The long-term interest representation at the user level is learned from the interest representation at the topic level. Similar to the previous two representations, here use the attention network at the user level to select the important topic vector representation, which is the final long-term interest representation. The formula is as follows: u l is the long-term interest representation at the user level, γ i is the topic-level interest representation is the attention weight; (3) Extract the short-term behavior features of users: Use the gated recurrent unit to obtain the short-term interest representation of users, and initialize the GRU with the long-term interest representation. The obtained short-term interest representation is the final representation of the user; (4) Recommendation: Match the final representation of the user and the candidate news representation to obtain a recommendation list and complete the recommendation.
Citation Information
Patent Citations
Content recommendation method and device, equipment and storage medium
CN111552888A
Session recommendation method based on modeling of long-term interests and short-term interests of users
CN113722599A