Intelligent News Recommendation Method and System Based on User Intention Features
By building an intelligent news recommendation model based on user intention characteristics, using neural networks and deep learning methods, the problem of uncatched user intention characteristics in the prior art is solved, and more accurate news recommendation is achieved.
Patent Information
- Application Number
- CN202310412911.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-04-13
AI Technical Summary
Existing news recommendation methods cannot accurately capture the user's high-level intention characteristics, resulting in inaccurate recommendation results.
A smart news recommendation model based on user intention characteristics is constructed, and a news encoder, TF-IDF algorithm, graph neural network and intent decoder are used to extract and represent the user's high-order intent characteristics through neural networks and deep learning methods.
It improves the accuracy of news recommendations, can more comprehensively model user characteristics, and achieve accurate news recommendations.
Smart Images

Figure CN116431919B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and natural language processing, and particularly relates to an intelligent news recommendation method and system based on user intention features. Background Art
[0002] With the popularization and rapid development of recommendation system technology, more and more users like to read news through online news platforms such as NetEase, Sina, and Toutiao. Although these platforms attempt to provide personalized news recommendation services for users, they are still inevitably troubled by the problem of inaccurate recommendation results. If the news recommendation system cannot accurately recommend appropriate content to users, it will affect the user's reading experience. The key to solving the above problems is to accurately model user intentions, that is, to accurately capture user intention features.
[0003] In reality, users' behaviors are often influenced by their intentions. Relative to explicit and specific user interests, user intentions are implicit user features and are high-level signals indicating the direction of user behavior. In most cases, users tend to purposefully read some news related to a theme or event, such as economic news, which indicates their reading intentions. Then, for the same purpose (such as reading economic news), different users may have different preferences for specific news, and thus may read different economic news. However, existing news recommendation methods only consider low-order interest features of users and ignore high-order intention features of users. Existing news recommendation methods usually establish a user interest representation based on the news text content in the user's browsing records, and then perform a recommendation task based on the user's interest features. Although these methods have improved the accuracy of recommendation results to a certain extent in the news recommendation task, they all ignore the high-order intention features of users; only modeling low-order user interest features often cannot accurately capture the true intention behaviors of users, which will inevitably affect the accuracy of news recommendation results. Summary of the Invention
[0004] The technical task of the present invention is to provide an intelligent news recommendation method and system based on user intention features to solve the problem of inaccurate recommendation results in the news recommendation system.
[0005] The technical task of the present invention is achieved in the following manner. An intelligent news recommendation method based on user intention features, the method includes the following steps:
[0006] S1. Construct a training data set for the news recommendation model: First, download the publicly available news data set on the network, then preprocess the data set, and finally construct positive example data and negative example data, and combine them to generate the final training data set;
[0007] S2. Build a news recommendation model based on user intention features: Use neural network and deep learning methods to build a news recommendation model;
[0008] S3. Train the model: Train the news recommendation model built in step S2 on the training dataset obtained in step S1.
[0009] An intelligent news recommendation system based on user intention features, the system includes:
[0010] A training dataset generation unit, first obtains the browsing record information of users on online news websites, and then performs preprocessing operations on it to obtain user browsing records and their news text content that meet the training requirements;
[0011] A news recommendation model construction unit based on user intention features, used to load the training dataset, construct a news encoding module, construct a TF-IDF algorithm module, construct an intention encoding module, construct a graph neural network module, construct an intention decoding module, and construct a click-through rate prediction module;
[0012] A model training unit, used to construct the loss function required during model training and complete the optimization training of the model.
[0013] A storage medium, in which multiple instructions are stored, and the instructions are loaded by a processor to execute the steps of the above-mentioned intelligent news recommendation method based on user intention features.
[0014] An electronic device, the electronic device includes: the above storage medium; and a processor, used to execute the instructions in the storage medium.
[0015] Preferably, the construction process of the news encoder is as follows:
[0016] Construct a news encoder, including four modules: a title learning module, an abstract learning module, a category learning module, and an attention module, specifically as follows:
[0017] Construct a title learning module: Build a word mapping table for each word in the dataset, and map each word in the table to a unique digital identifier. The mapping rule is: starting from the number 1, and then sorting them in ascending order in the order in which each word is entered into the word mapping table, so as to form a word mapping conversion table; Use the Glove pre-trained language model to obtain the word vector representation of each word; In the word embedding layer, convert each news title T = [w1, w2,..., w N into a vector representation, denoted as E = [e1, e2,..., e N , where N represents the length of a news title, and e N represents the vector representation of each word, and w represents a word in the news title.
[0018] For \(E = [e_1, e_2, \cdots, e\) N , use a convolutional neural network (CNN) for feature extraction to obtain the context feature vector \([c_1, c_2, \cdots, c\) N , and the formula is as follows:
[0019] c i = ReLU(Q w = e (i-k):(i+k) + b w );
[0020] Among them, \(i\) represents the relative position of the corresponding character vector in the news title, \(k\) represents the difference in the relative position from \(i\), \(Q\) w represents the convolution kernel of the CNN filter, \(b\) w represents the bias, ReLU is an activation function, and the operator \(\times\) is matrix multiplication.
[0021] For the context feature vector \([c_1, c_2, \cdots, c\) N , use the attention mechanism to further extract the key features to obtain the final news title vector \(r\) t , and the formula is as follows:
[0022] a i = q T tanh(V \times c i + v);
[0023]
[0024]
[0025] Among them, \(q\) is the attention query vector randomly initialized according to the dimension of the context feature vector, \(V\) and \(v\) are parameters learned during the training process, tanh is an activation function, the operator \(\times\) is matrix multiplication, exp is the logarithmic function operation, \(a\) i is the attention score of the \(i\)-th word, \(\alpha\) i is the attention weight of the \(i\)-th word, and \(N\) is the length of the context feature vector \([c_1, c_2, \cdots, c\) N .[[]]
[0026] Construct a summary learning module: Use the news summary as the input, and the specific implementation method is the same as the title learning module, and then obtain the summary vector representation \(r\) a .
[0027] Construct a category learning module: In the word embedding layer, map the main category label and the sub-category label to the low-dimensional space vector respectively by the word vector method to obtain the word vector representation \(e\) c and \(e\) sc, and then use the ReLU activation function to generate the final vector r of the class label c and r sc , the formula is as follows:
[0028] r c = ReLU(V c × e c + v c );
[0029] r sc = ReLU(V sc × e sc + v sc );
[0030] Among them, ReLU is an activation function, V c , V sc , v sc and v c are parameters learned from the training process, and the operator × is matrix multiplication.
[0031] Construct the attention module: For the vectors r t 、r a 、r c and r sc of the title, abstract, main class label, and sub-class label, use the activation function tanh to calculate their respective attention scores, that is, a t 、a a 、a c 、a sc , and then further obtain their respective attention weights through the attention mechanism. The formula is as follows:
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040] Among them, V t 、V a 、v c 、V sc 、vt , v a , v c , v sc For calculating the title attention score a t , the abstract attention score a a , the main category label attention score a c and the sub-category label attention score a sc of randomly initialized learnable parameters, is the attention query vector generated by the title vector r t ; is the attention query vector generated by the abstract vector r a ; is the attention query vector generated by the main category label vector r c ; is the attention query vector generated by the sub-category label vector r sc ; tanh is an activation function, the operator × is matrix multiplication, exp is the logarithmic function operation, α t is the attention weight of the title, α a is the attention weight of the abstract, α c is the attention weight of the main category label, α sc is the attention weight of the sub-category label.
[0041] The final news vector r is determined by the title vector r t , the abstract vector r a , the main category label vector r c and the sub-category label vector r sc and their respective attention weights, and the formula is as follows:
[0042] r = [α t r t ; α a r a ; α c r c ; α sc r sc ;
[0043] where the symbol ; represents the concatenation operation.
[0044] Preferably, the process of constructing the news recommendation model based on user intention features is specifically as follows:
[0045] Construct a TF-IDF algorithm module: First, input a user browsing record C u = {v1,..., v i ,..., v t-1} into this module, where v iRepresent each user browsing record; then, use the TF-IDF algorithm to extract keywords from the user browsing records; finally, map these keywords to a keyword vector matrix K through the word embedding layer i 。
[0046] Construct an intent encoder, which aims to infer the user's intent from the user browsing records, as follows:
[0047] Construct a convolutional neural network:
[0048] Take the keyword vectors {K1,...,K i ,...,K t-1} of the user's historical news sequence as the input of the convolutional neural network, and use the convolutional neural network to encode these vectors, which is expressed by the formula as follows:
[0049] c i = ReLU(W′ * K (i-f):(i+f) + b′);
[0050] where, K (i-f):(i+f) represents the keyword vector connected from position (i - f) to (i + f), W′ represents the learnable parameter of the CNN filter, b′ is the bias, * represents the convolution operation, ReLU represents the ReLU activation function, and c i represents the convolved keyword vector.
[0051] Construct an information mapping module:
[0052] To infer the user's intent from the keyword vectors of the historical news sequence, first extract the keywords of all news from the news recommendation dataset using the TF-IDF method, then generate a global keyword vector matrix H using the word embedding layer, then filter the possible keywords through a learnable mapping matrix M, and then obtain the possible keyword vector matrix by calculating the distribution probability of the possible keywords in the global keyword vector matrix H. The specific process formula is expressed as follows:
[0053] C = [c1; c2;...; c t-1 ;
[0054] W p = softmax(HMC);
[0055] C p = W p H;
[0056] where, softmax represents the softmax normalization function, C represents the aggregation matrix after connecting the convolved keyword vectors, and W p represents the learnable weight matrix. C pRepresents a possible keyword vector matrix, which contains all the initial intent feature vectors.
[0057] Construct a graph neural network: Once this module obtains the possible keyword vector matrix C p , it inputs this matrix into the graph neural network. Specifically, the operation process of the l-th layer of the graph neural network is as follows:
[0058]
[0059] where σ represents the activation function, H l is the node representation of the l-th layer of the graph neural network, W l represents the learnable parameter of the l-th layer of the graph neural network, D is the degree matrix. A = A + I, where A is the adjacency matrix and I is the identity matrix. Specifically, the input of the first layer is C p , then its output is H 0 = C p . After passing through n layers of the graph neural network, the user intent at time t can be represented as C t = H n , where C t contains the updated intent feature vectors.
[0060] Construct an intent decoder: After the user intent C t is generated by the graph neural network, the intent decoder uses the attention mechanism as the decoder to generate the final intent feature vector, and the formula is as follows:
[0061]
[0062]
[0063] where, represents the attention weight, represents the j-th row vector of the C t matrix, R is the total number of rows of the matrix, exp is the logarithmic function operation, represents the activation function. u o is the final intent feature vector.
[0064] More preferably, the construction process of the click-through rate predictor is specifically as follows:
[0065] Construct an attention network based on candidate news: It is designed to integrate the features of candidate news into the final intent feature vector to generate the final user vector. The formula is as follows:
[0066] α = Att(W Q d, W K u o );
[0067]
[0068] Among them, W Q and W K are learnable parameters, d is the news vector representation of the candidate news generated by the news encoder, U is the length of a user browsing record, u is the final user vector, Att represents the attention mechanism function, and α is the attention weight.
[0069] Construct a prediction module: It takes the vector representation d of the candidate news and the final user vector u as inputs. This module uses the dot product operation to predict the click-through rate of the candidate news. The formula is as follows:
[0070]
[0071] Among them, represents the click-through rate of the candidate news.
[0072] When the method model has not been fully trained, it needs to be trained on the training dataset to optimize the model parameters; when the model training is completed, the click-through rate predictor can predict the recommendation score of each candidate news, and according to the score, recommend appropriate news to the user.
[0073] More preferably, the construction process of the training dataset is specifically as follows:
[0074] Construct a news dataset or select a publicly available news dataset;
[0075] Preprocess the news dataset: Preprocess each news text in the news dataset, remove the stop words and special characters in the news dataset; extract the title, category, sub-category and abstract information of each news text respectively;
[0076] Construct training positive examples: Use the historical news sequence and the news numbers with label 1 in the interaction behavior sequence in the user browsing record, that is, the numbers of the news clicked by the user, to construct training positive examples;
[0077] Construct training negative examples: Use the historical news sequence and the news numbers with label 0 in the interaction behavior sequence in the user browsing record, that is, the numbers of the news not clicked by the user, to construct training negative examples;
[0078] Construct a training dataset: Combine all the positive example data and negative example data and shuffle their order to construct the final training dataset;
[0079] After the news recommendation model is constructed, it is trained and optimized through the training dataset as follows:
[0080] Construct the loss function: Using the negative sampling technique, the news that a user has clicked is defined as a positive example, and the news that the user has not clicked is defined as a negative example, and calculate the click prediction value p of the positive example i . The formula is as follows:
[0081]
[0082] where, is the click-through rate of the j-th negative example relative to the i-th positive example in the same click sequence, is the i-th positive example, and G is the number of negative examples.
[0083] The loss function of news recommendation is the negative log-likelihood function of all positive examples, and the formula is as follows:
[0084]
[0085] where, is the set of positive examples.
[0086] Optimize the training model: Select the Adam optimization function as the optimization function of this model. Among them, the learning rate is set to 0.001, the smoothing constants are set to (0.9, 0.999), eps is set to 1e-8, and the L2 penalty value is set to 0.
[0087] An intelligent news recommendation system based on user intention features, the system includes,[[]]
[0088] The training dataset generation unit first obtains the browsing record information of users on the online news website, and then performs preprocessing operations on it to obtain the user browsing records and their news text content that meet the training requirements; the training dataset generation unit includes,[[]]
[0089] The original data acquisition unit is responsible for downloading the news website dataset that has been publicly available on the network and using it as the original data for constructing the training dataset;
[0090] The original data preprocessing unit is responsible for preprocessing each news text in the news dataset, removing the stop words and special characters in the news dataset; respectively extracting the key information of each news text, such as the title, category, and summary; thereby constructing the training dataset;
[0091] The news recommendation model construction unit based on user intention features is used to load the training dataset, construct the news encoding module, construct the TF-IDF algorithm module, construct the intention encoding module, construct the graph neural network module, construct the intention decoding module, and construct the click-through rate prediction module. The news recommendation model construction unit based on user intention features includes,[[]]
[0092] Training dataset loading unit, responsible for loading the training dataset;
[0093] News encoding module construction unit, responsible for training news vectors based on the Glove word vector model in the training dataset and defining all news vector representations; First, use a convolutional neural network and an attention mechanism to encode the news title and abstract respectively to obtain news title and abstract vectors; At the same time, use a fully connected layer to encode the main news category and sub-category respectively to obtain main news category and sub-category vectors; Then, connect the news title, abstract, main category and sub-category vectors and input them into the attention mechanism to obtain the final news vector;
[0094] TF-IDF algorithm module construction unit, responsible for using the TF-IDF algorithm to extract news keywords from the user browsing record, and then using the word embedding method to map each keyword to the same vector space, so as to obtain the keyword vector of the news content.
[0095] Intention encoding module construction unit, responsible for using CNN to extract the main features of the keyword vector and generating a keyword vector matrix through an aggregation operation, then filtering possible keywords through a learnable mapping matrix M, and then calculating the distribution probability of possible keywords in the keyword vector matrix to obtain a possible keyword vector matrix.
[0096] Graph neural network module construction unit, responsible for using the graph neural network to propagate and aggregate the keyword vector features, so as to obtain the keyword vector after information aggregation.
[0097] Intention decoding module construction unit, responsible for using the attention network to decode the user intention features in the keyword vector, so as to obtain the intention feature vector.
[0098] Click-through rate prediction module construction unit, responsible for using the attention network based on the candidate news to fuse the intention feature vector and the candidate news feature vector to obtain the user vector, then using the user vector and the candidate news feature vector as inputs, and then generating the score of each candidate news, that is, the click-through rate, through the vector inner product operation, and then sorting all candidate news from high to low according to the click-through rate, and recommending the top-K news to the user;
[0099] Model training unit, used to construct the loss function required in the model training process and complete the optimization training of the model; The model training unit includes,
[0100] Loss function construction unit, responsible for calculating the error between the predicted candidate news and the real target news;
[0101] Model optimization unit, responsible for training and adjusting the parameters in the model training to reduce the prediction error.
[0102] A storage medium stores multiple instructions, characterized in that the instructions are loaded by a processor to execute the steps of the above-mentioned intelligent news recommendation method based on user intention features.
[0103] An electronic device, characterized in that the electronic device includes:
[0104] The above-mentioned storage medium; and
[0105] A processor for executing the instructions in the storage medium.
[0106] The intelligent news recommendation method and system based on user intention features of the present invention have the following advantages:
[0107] (1) The present invention proposes an intelligent news recommendation method based on user intention features, mines high-order intention features of users, and can more comprehensively model the feature representation of users compared with the representation of low-order interest features of users in existing methods, thereby improving the accuracy of news recommendation.
[0108] (2) The present invention first extracts keyword features from news content in user browsing records through the TF-IDF method, and then uses an intention encoder, a graph neural network, and an intention decoder to establish a user intention feature representation, thereby obtaining a high-order user intention representation.
[0109] (3) Through the click-through rate predictor module of the present invention, the prediction scores of candidate news sequences can be accurately output according to accurate news representations and user representations.
[0110] (4) The present invention can use deep learning technology and information of knowledge graphs to abstractly model users, thereby realizing accurate news recommendation.
[0111] (5) The present invention can define and implement a complete news recommendation system model to directly provide personalized news recommendation services for users. Description of the Drawings
[0112] The present invention will be further described below with reference to the accompanying drawings.
[0113] Figure 1 It is a flowchart of an intelligent news recommendation method based on user intention features
[0114] Figure 2 It is a flowchart of constructing a training data set for a news recommendation model
[0115] Figure 3 It is a flowchart of constructing a news recommendation model based on user intention features
[0116] Figure 4 It is a flowchart of training a news recommendation model based on user intention features
[0117] Figure 5 Schematic diagram of a news recommendation model based on user intention features
[0118] Figure 6 Schematic diagram of a news encoder Detailed implementation manners
[0119] The intelligent news recommendation method and system based on user intention features of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0120] Embodiment 1:
[0121] The overall model framework of the present invention is as shown in Figure 5 It can be seen from that Figure 5 the main framework structure of the present invention includes a news encoder, a term frequency-inverse document frequency TF-IDF algorithm module, an intention encoder, a graph neural network, an intention decoder, and a click-through rate predictor module. Among them, the news encoder is responsible for extracting the main features of the candidate news content from the candidate news vectors by using a convolutional neural network and an attention network, and generating candidate news feature vectors; the TF-IDF algorithm module is responsible for extracting the keywords of the news content from the user browsing records, and then generating keyword vectors by using a word embedding layer; the intention encoder is responsible for encoding the keyword vectors by using a convolutional neural network to obtain initial intention feature vectors; the graph neural network is responsible for propagating the input initial intention feature vectors by using a graph convolution method, and obtaining updated intention feature vectors through information aggregation; the intention decoder is responsible for decoding the updated intention feature vectors obtained from the graph neural network by using an attention network, and obtaining the final intention feature vectors; the click-through rate predictor module first fuses the final intention feature vectors and the candidate news feature vectors by using an attention network based on the candidate news to obtain user vectors, then takes the user vectors and the candidate news feature vectors as inputs, and finally generates the scores of each candidate news, that is, the click-through rates, through vector dot product operations, and sorts all candidate news from high to low according to the click-through rates, so as to recommend the top-K news to the user. The above is a brief introduction to the structure of this model invention.
[0122] Embodiment 2:
[0123] As shown in the appendix Figure 1 The intelligent news recommendation method based on user intention features of the present invention is as follows:
[0124] S1. Construct the training dataset for the news recommendation model: The news dataset consists of two parts of data files: user browsing records and news text content. Among them, the user browsing records include user ID, time, historical news sequence, and interaction behavior sequence. The news text content includes news ID, category, sub-category, title, abstract, and entity. Select the historical news sequence and interaction behavior sequence in the user browsing records to construct the user behavior data of the training dataset, and select the title, category, sub-category, and abstract of the news text content to construct the news text data of the training dataset. Among them, the user behavior data will be used for the extraction of user intention features, and the news text content data will be used for the extraction of news features. The method for constructing the training dataset is as follows:
[0125] S101. Construct a news dataset or select a publicly available news dataset.
[0126] Example: Download the publicly available MIND news dataset on the Internet from Microsoft and use it as the original data for news recommendation. MIND is currently the largest English news recommendation system dataset, containing 1,000,000 users in 200,000 categories and 161,013 news, divided into training set, validation set, and test set. The MIND dataset also provides detailed information on the news text content. Each news has a news ID, link, title, abstract, category, and entity:
[0127]
[0128] In addition, the MIND dataset also provides user browsing records, and each record contains user ID, time, historical news sequence, and interaction behavior sequence:
[0129]
[0130]
[0131] Among them, the user ID represents the unique ID of each user on the news platform; the time represents the start time when the user clicks and browses a series of news; the historical news sequence represents a sequence of news IDs browsed by the user; the interaction behavior sequence represents the actual interaction behavior of the user on a series of news recommended by the system, where 1 represents click and 0 represents non-click.
[0132] S102. Preprocess the news dataset: Preprocess each news text in the news dataset, remove the stop words and special characters in the news dataset, and extract the title, category, sub-category, and abstract information of each news text respectively.
[0133] S103. Construct training positive examples: Use the historical news sequence in the user browsing record and the news numbers with label 1 in the interaction behavior sequence, that is, the numbers of the news clicked by the user, to construct training positive examples.
[0134] Example: For the news examples shown in step S101, the formalized positive example data is: (N29038, N15201, N8018, N32012, N30859, N26552, N25930). The last number is the number of the news clicked by the user.
[0135] S104. Construct training negative examples: Use the historical news sequence in the user browsing record and the news numbers with label 0 in the interaction behavior sequence, that is, the numbers of the news not clicked by the user, to construct training negative examples.
[0136] Example: For the news examples shown in step S101, the formalized negative example data is: (N29038, N15201, N8018, N32012, N30859, N26552, N17825). The last number is the number of the news not clicked by the user.
[0137] S105. Construct a training data set: Combine all the positive example data and negative example data obtained after the operations in steps S103 and S104, and shuffle their order to construct the final training data set.
[0138] S2. Construct a news recommendation model based on user intention features: As shown in the appendix Figure 3 This news recommendation model includes a news encoder, a term frequency-inverse document frequency (TF-IDF) algorithm module, an intention encoder, a graph neural network, an intention decoder, and a click-through rate predictor module. Specifically as follows:
[0139] S201. Construct a news encoder, as shown in the appendix Figure 6 This includes four modules: a title learning module, an abstract learning module, a category learning module, and an attention module. Specifically as follows:
[0140] S20101. Construct a title learning module, specifically as follows:
[0141] S2010101. For each word in the data set, construct a word mapping table, and map each word in the table to a unique digital identifier. The mapping rule is: starting from the number 1, and then increasing sequentially according to the order in which each word is entered into the word mapping table, so as to form a word mapping conversion table; use the Glove pre-trained language model to obtain the word vector representation of each word; in the word embedding layer, for each news title T = [w1, w2,..., w NConvert it into a vector representation, denoted as E = [e1, e2,..., e N , where N represents the length of a news title, and e N represents the vector representation of each word, and w represents a word in the news title.
[0142] For example, with the help of pre-trained word vectors Glove, each news title T = [w1, w2,..., w N can be converted into a word vector representation, denoted as E = [e1, e2,..., e N .
[0143] S2010102. Use a convolutional neural network CNN to extract features from E = [e1, e2,..., e N to obtain context feature vectors [c1, c2,..., c N . The formula is as follows:
[0144] c i = ReLU(Q w × e (i-k):(i+k) + b w );
[0145] Among them, i represents the relative position of the corresponding character vector in the news title, k represents the difference in the relative position from i, Q w represents the convolution kernel of the CNN filter, b w represents the bias, ReLU is an activation function, and the operator × is matrix multiplication.
[0146] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0147]
[0148]
[0149] Among them, nn.Conv2d and F.dropout are built-in convolutional neural network methods and methods to prevent overfitting in pytorch. title_vector is the title vector after being processed by pre-trained word vectors, convoluted_title_vector is the context feature vector after the title vector is processed by the convolutional neural network, and activated_title_vector is the context feature vector after being processed by the activation function ReLU function.
[0150] S2010103. For the context feature vectors [c1, c2,..., c N, the attention mechanism is used to further extract key features to obtain the final news title vector r t , and the formula is as follows:
[0151] a i = q T tanh(V × c i + v);
[0152]
[0153]
[0154] Among them, q is the attention query vector obtained by random initialization according to the dimension of the context feature vector, V and v are parameters learned from the training process, tanh is an activation function, the operator × is matrix multiplication, exp is logarithmic function operation, a i is the attention score of the i-th word, α i is the attention weight of the i-th word, and N is the length of the context feature vector [c1, c2,..., c N .
[0155] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0156] self.title_attention = AdditiveAttention(config.query_vector_dim, config.num_filters)
[0157] weighted_title_vector = self.title_attention(activated_title_vector.transpose(1, 2))
[0158] Among them, self.title_attention, that is, AdditiveAttention, is a method customized by the present invention according to the principle of the attention mechanism, weighted_title_vector is the attention weight of the title vector, and config.query_vector_dim and config.num_filters are custom vector dimension parameters.
[0159] S20102. Construct a summary learning module: The specific steps are the same as those in S20101 for constructing a title learning module to obtain a summary vector r a .
[0160] S20103. Construct a category learning module:
[0161] In the word embedding layer, the main category label and the sub-category label are respectively mapped to low-dimensional space vectors through the word vector method to obtain the word vector representation e of each category label c and e sc , and then the activation function ReLU is used to generate the final vectors r c and r sc , and the formula is as follows:
[0162] r c = ReLU(V c × e c + v c );
[0163] r sc = ReLU(V sc × e sc + v sc );
[0164] Among them, ReLU is an activation function, V c , V sc , v sc and v c are parameters learned from the training process, and the operator × is matrix multiplication.
[0165] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0166]
[0167]
[0168] Among them, nn.Embedding, nn.Linear, and F.relu are the built-in word vector embedding method, connection layer method, and activation function in pytorch respectively. config.category_embedding_dim and config.num_filters are custom vector dimension parameters, and activated_category_vector and activated_subcategory_vector are the final generated main category label vector r c and sub-category label vector r sc .
[0169] S20104. Construct the attention module: For the vectors r t , r a , r c and r sc of the title, abstract, main category label, and sub-category label, use the activation function tanh to calculate their respective attention scores, that is, at 、a a 、a c 、a sc , and then, through the attention mechanism, the respective attention weights are obtained as follows:
[0170]
[0171]
[0172]
[0173]
[0174]
[0175]
[0176]
[0177]
[0178] Among them, V t 、V a 、V c 、V sc 、v t 、v a 、v c 、v sc are learnable parameters initialized randomly for calculating the title attention score a t 、abstract attention score a a 、main category label attention score a c and sub-category label attention score a sc ; is the attention query vector generated by the title vector r t ; is the attention query vector generated by the abstract vector r a ; is the attention query vector generated by the main category label vector r c ; is the attention query vector generated by the sub-category label vector r sc ; tanh is an activation function, the operator × is matrix multiplication, exp is logarithmic function operation, α t is the attention weight of the title, α a is the attention weight of the abstract, α c is the attention weight of the main category label, α sc is the attention weight of the sub-category label.
[0179] The final news vector r is determined by the title vector r t , the abstract vector r a , the main category label vector r c and the subcategory label vector r sc as well as their respective attention weights, and the formula is as follows:
[0180] r = [α t r t ; α a r a ; α c r c ; α sc r sc ;
[0181] Among them, the symbol ; represents the concatenation operation.
[0182] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0183]
[0184] Among them, self.final_attention is a method customized by the present invention according to the principle of the attention mechanism; weighted_title_vector, weighted_abstract_vector, activated_category_vector, and activated_subcategory_vector are the vectors r t , r a , r c and r sc of the title, abstract, main category label, and subcategory label respectively; news_vector is the final news vector r.
[0185] S202. Construct a TF-IDF algorithm module, specifically as follows: First, input a user's browsing record C u = {v1,..., v i ,..., v t-1} into this module, where v i represents each user browsing record; then, use the TF-IDF algorithm to extract keywords from the user browsing record; finally, map these keywords to a keyword vector matrix K i .
[0186] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0187] tfidf_model = TfidfVectorizer().fit(df)
[0188] sparse_result = tfidf_model.transform(df)
[0189] Among them, df is the input data, TfidfVectorizer() is the vectorization method of the TF-IDF algorithm, and tfidf_model.transform() is the method for sparse matrix conversion.
[0190] S203. Construct an intent encoder, which aims to infer the user's intent from the user's browsing history, as follows:
[0191] S20301. Construct a convolutional neural network:
[0192] Take the keyword vectors {K1,...,K i ,...,K t-1} of the user's historical news sequence as the input of the convolutional neural network, and use the convolutional neural network to encode these vectors. The formula is as follows:
[0193] c i = ReLU(W' * K (i-f):(i+f) + b');
[0194] Among them, K (i-f):(i+f) represents the keyword vector connected from position (i - f) to (i + f), W' represents the learnable parameter of the CNN filter, b' is the bias, * represents the convolution operation, ReLU represents the ReLU activation function, and c i represents the convolved keyword vector.
[0195] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0196] self.conv = Conv1D(config.cnn_method, self.concept_num_per_news,
[0197] self.num_concepts, config.cnn_window_size)
[0198] c = self.dropout_(self.conv(clicked_concept_emebedding))
[0199] Among them, Conv1D is the convolutional neural network method of the pytorch toolkit
[0200] S20302. Construct an information mapping module:
[0201] To infer the user's intention from the keyword vectors of the historical news sequence, first extract the keywords of all news from the news recommendation dataset using the TF-IDF method, then generate a global keyword vector matrix H using the word embedding layer, then filter the possible keywords through a learnable mapping matrix M, and then obtain the possible keyword vector matrix by calculating the distribution probability of the possible keywords in the global keyword vector matrix H. The specific process formula is as follows:
[0202] C = [c1; c2;...; c t-1 ;
[0203] W p = softmax(HMC);
[0204] C p = W p H;
[0205] Among them, softmax represents the softmax normalization function, C represents the aggregation matrix after the convolutionized keyword vectors are joined, and W p represents the learnable weight matrix. C p represents the possible keyword vector matrix, which contains all the initial intention feature vectors.
[0206] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0207] temp = torch.matmul(c.reshape(-1, self.word_embedding_dim), self.transform_matrix)
[0208] t = torch.matmul(temp, self.pretrained_concept_embedding.transpose(0, 1))
[0209] concept_weight = F.softmax(t, dim = 1)
[0210] personalized_concept_vector = torch.matmul(concept_weight, self.pretrained_concept_embedding).reshape(batch_size, -1, self.word_embedding_dim)
[0211] Among them, torch.matmul is matrix multiplication, and F.softmax is the softmax normalization function.
[0212] S204. Construct a graph neural network:
[0213] Once this module obtains the possible keyword vector matrix C p , this matrix will be input into the graph neural network. Specifically, the operation process of the l-th layer of the graph neural network is expressed as follows:
[0214]
[0215] Among them, σ represents the activation function, H l is the node representation of the l-th layer of the graph neural network, W l represents the learnable parameter of the l-th layer of the graph neural network, and D is the degree matrix. A = A + I, where A is the adjacency matrix and I is the identity matrix. Specifically, the input of the first layer is C p , then its output is H 0 = C p . After passing through n layers of the graph neural network, the user intention at time t can be represented as C t = H n , where C t contains the updated intention feature vector.
[0216] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0217]
[0218] Among them, GCN is the graph neural network method of the pytorch_geometric toolkit, in_dim is the size of the input vector, out_dim is the size of the output vector, hidden_dim is the size of the hidden layer vector, and num_layers is the number of layers of the graph neural network.
[0219] S205. Construct an intention decoder:
[0220] In the user intention C tAfter being generated by the graph neural network, the intent decoder uses the attention mechanism as the decoder to generate the final intent feature vector, as shown in the following formula:
[0221]
[0222]
[0223] Among them, represents the attention weight, represents the j-th row vector of the C t matrix, and R is the total number of rows of the matrix. U o is the final intent feature vector, exp is the logarithmic function operation, represents the activation function.
[0224] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0225] self.attention_decoder = Attention(config.word_embedding_dim,config.attention_dim)
[0226] user_vector = self.attention_decoder(gcn_feature.reshape(batch_news_num,-1,self.word_embedding_dim)).reshape(batch_size,-1,self.word_embedding_dim)
[0227] Among them, Attention is a custom attention mechanism method.
[0228] S206. Build a click-through rate predictor, which mainly includes an attention network based on candidate news and a prediction module. Specifically as follows:
[0229] S20601. Build an attention network based on candidate news, which is designed to integrate the features of candidate news into the final intent feature vector to generate the final user vector. The formula is as follows:
[0230] α = att(W Q d, W K u o );
[0231]
[0232] Among them, W Q 、WK is a learnable parameter, d is the news vector representation of the candidate news generated by the news encoder, U is the length of a user browsing record, u is the final user vector, Att represents the attention mechanism function, and α represents the attention weight.
[0233] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0234] self.FusionAttention = ScaledDotProduct_CandidateAttention(self.news_embedding_dim, self.news_embedding_dim, self.attention_dim)
[0235] inter_cluster_feature = self.FusionAttention(final_user_representation.view([batch_news_num, 1, self.news_embedding_dim]), candidate_news_representation.view([batch_news_num, self.news_embedding_dim])).view([batch_size, news_num, self.news_embedding_dim])
[0236] Among them, ScaledDotProduct_CandidateAttention is a custom dot product attention method.
[0237] S20602. Build a prediction module, which takes the vector representation d of the candidate news and the final user vector u as inputs. This module uses the dot product operation to predict the click-through rate of the candidate news. The formula is as follows:
[0238]
[0239] Among them, represents the click-through rate of the candidate news.
[0240] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0241] probability = torch.bmm(user_vector.unsqueeze(dim = 1),
[0242] (candidate_news_vector.unsqueeze(dim=2)).flatten()
[0243] Among them, torch.bmm is the vector inner product operation, user_vector is the Aspect-level user vector representation u, and candidate_news_vector is the Aspect-level news vector representation n.
[0244] S3. Train the model: As shown in the appendix Figure 4 as follows:
[0245] S301. Construct the loss function: Adopt the negative sampling technique. Define the news that a user has clicked as the positive example, and the news that has not been clicked as the negative example, and calculate the click prediction value pi of the positive example. The formula is as follows:
[0246]
[0247] Among them, is the click-through rate of the j-th negative example relative to the i-th positive example in the same click sequence, is the i-th positive example, and G is the number of negative examples.
[0248] The loss function of news recommendation is the negative log-likelihood function of all positive examples. The formula is as follows:
[0249]
[0250] Among them, is the set of positive examples.
[0251] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0252] loss = torch.stack([x[0] for x in -F.log_softmax(y_pred, dim=1)]).mean()
[0253] Among them, F.log_softmax is the built-in log_softmax loss function in pytorch, and y_pred is the click prediction value p.
[0254] S302. Optimize the model: Select the Adam optimization function as the optimization function of this model. Among them, the learning rate is set to 0.0001, the smoothing constants are set to (0.9, 0.999), eps is set to 1e-8, and the L2 penalty value is set to 0.
[0255] In the experiment, the present invention selects the area under the ROC curve (AUC), mean reciprocal rank (MRR), and normalized discounted cumulative gain (nDCG) as evaluation indicators.
[0256] For example, the optimization function described above is expressed in code in PyTorch as follows:
[0257] optimizer = torch.optim.Adam(model.parameters(), lr = learning_rate)
[0258] Among them, torch.optim.Adam is the Adam optimization function embedded in PyTorch, model.parameters() is the set of parameters for model training, and learning_rate is the learning rate.
[0259] The model of the present invention has achieved better results than the current model on the MIND public dataset. The comparison of the experimental results is shown in the following table:
[0260]
[0261] The model of the present invention is compared with the existing models, and it can be seen that the method of the present invention has the best performance among other methods. Among them, NRMS comes from the literature "Neural news recommendation with multi-head self-attention", NPA comes from the literature "NPA: neural news recommendation with personalized attention", and NNR comes from the literature "Neural News Recommendation with Collaborative News Encoding and Structural User Encoding".
[0262] Example 3:
[0263] Based on Example 2, an intelligent news recommendation system based on user intention features is constructed. The system includes:
[0264] A training dataset generation unit, which first obtains the browsing record information of users on online news websites, and then performs preprocessing operations on it to obtain user browsing records and their news text content that meet the training requirements; the training dataset generation unit includes,
[0265] A raw data acquisition unit, which is responsible for downloading the publicly available news website dataset on the network and using it as the raw data for constructing the training dataset;
[0266] The original data preprocessing unit is responsible for preprocessing each news text in the news dataset, removing stop words and special characters in the news dataset; extracting key information of each news text respectively, such as title, category, and abstract; thus constructing a training dataset;
[0267] The news recommendation model construction unit based on user intention features is used to load the training dataset, construct a news encoding module, construct a TF-IDF algorithm module, construct an intention encoding module, construct a graph neural network module, construct an intention decoding module, and construct a click-through rate prediction module. The news recommendation model construction unit based on user intention features includes,
[0268] The training dataset loading unit is responsible for loading the training dataset;
[0269] The news encoding module construction unit is responsible for training news vectors based on the Glove word vector model in the training dataset and defining all news vector representations; first, using a convolutional neural network and an attention mechanism to encode the news title and abstract respectively to obtain news title and abstract vectors; at the same time, using a fully connected layer to encode the main category and sub-category of the news respectively to obtain main category and sub-category vectors of the news; then, concatenating the news title, abstract, main category, and sub-category vectors and inputting them into the attention mechanism to obtain the final news vector;
[0270] The TF-IDF algorithm module construction unit is responsible for using the TF-IDF algorithm to extract news keywords in the user browsing record, and then using the word embedding method to map each keyword to the same vector space, so as to obtain keyword vectors of the news content.
[0271] The intention encoding module construction unit is responsible for using CNN to extract the main features of the keyword vectors and generating a keyword vector matrix through an aggregation operation, then filtering possible keywords through a learnable mapping matrix M, and then obtaining a possible keyword vector matrix by calculating the distribution probability of possible keywords in the keyword vector matrix.
[0272] The graph neural network module construction unit is responsible for using the graph neural network to propagate and aggregate keyword vector features, so as to obtain keyword vectors after information aggregation.
[0273] The intention decoding module construction unit is responsible for using the attention network to decode the user intention features in the keyword vectors, so as to obtain intention feature vectors.
[0274] The click-through rate prediction module construction unit is responsible for using the attention network based on candidate news to fuse the intent feature vector and the candidate news feature vector to obtain the user vector. Then, taking the user vector and the candidate news feature vector as inputs, it generates the score of each candidate news, i.e., the click-through rate, through vector inner product operation. Then, all candidate news are sorted from high to low according to the click-through rate, and the top-K news are recommended to the user;
[0275] The model training unit is used to construct the loss function required in the model training process and complete the optimization training of the model. The model training unit includes,
[0276] The loss function construction unit is responsible for calculating the error between the predicted candidate news and the real target news;
[0277] The model optimization unit is responsible for training and adjusting the parameters in the model training to reduce the prediction error.
[0278] Embodiment 4:
[0279] Based on the storage medium of Embodiment 2, which stores multiple instructions that are loaded and executed by a processor to perform the steps of the intelligent news recommendation method based on user intent features in Embodiment 2.
[0280] Embodiment 5:
[0281] Based on the electronic device of Embodiment 4, the electronic device includes: the storage medium of Embodiment 4; and a processor for executing the instructions in the storage medium of Embodiment 4.
[0282] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent news recommendation system based on user intention features, characterized in that The system includes: A training dataset construction unit that first obtains the browsing record information of users on online news websites and then performs preprocessing operations on it to obtain user browsing records and their news text content that meet the training requirements. The training dataset construction unit includes: An original data acquisition unit responsible for downloading the publicly available news website dataset on the network and using it as the original data for constructing the training dataset; An original data preprocessing unit responsible for preprocessing each news text in the news dataset, removing stop words and special characters in the news dataset; extracting the key information of each news text, such as title, category, and summary; thus constructing the training dataset; A news recommendation model construction unit based on user intention features, which is used to load the training dataset, construct a news encoding module, construct a TF-IDF algorithm module, construct an intention encoding module, construct a graph neural network module, construct an intention decoding module, and construct a click-through rate prediction module. The news recommendation model construction unit based on user intention features includes: A training dataset loading unit responsible for loading the training dataset; A news encoding module construction unit responsible for training news vectors based on the Glove word vector model in the training dataset and defining all news vector representations; first using a convolutional neural network and an attention mechanism to encode the news title and summary respectively to obtain news title and summary vectors; at the same time using a fully connected layer to encode the main news category and subcategory respectively to obtain main news category and subcategory vectors; then concatenating the news title, summary, main category, and subcategory vectors and inputting them into the attention mechanism to obtain the final news vector; A TF-IDF algorithm module construction unit responsible for using the TF-IDF algorithm to extract news keywords in the user browsing records, and then using the word embedding method to map each keyword to the same vector space to obtain the keyword vectors of the news content; An intention encoding module construction unit responsible for using CNN to extract the main features of the keyword vectors and generating a keyword vector matrix through an aggregation operation, then filtering possible keywords through a learnable mapping matrix M, and then obtaining a possible keyword vector matrix by calculating the distribution probability of possible keywords in the keyword vector matrix; A graph neural network module construction unit responsible for using the graph neural network to propagate and aggregate the keyword vector features to obtain the keyword vectors after information aggregation; An intention decoding module construction unit responsible for using the attention network to decode the user intention features in the keyword vectors to obtain intention feature vectors; A click-through rate prediction module construction unit responsible for using the attention network based on candidate news to fuse the intention feature vectors and candidate news feature vectors to obtain user vectors, then using the user vectors and candidate news feature vectors as inputs, and then generating the score of each candidate news, that is, the click-through rate, through vector inner product operation, and then sorting all candidate news from high to low according to the click-through rate, and recommending the top-K news to the user; A model training unit, which is used to construct a loss function required in the model training process and complete the optimization training of the model; the model training unit includes: A loss function construction unit, which is responsible for calculating the error between the predicted candidate news and the true target news; A model optimization unit, which is responsible for training and adjusting the parameters in the model training to reduce the prediction error.
2. An intelligent news recommendation method based on user intention features, characterized in that, This method is an implementation method of an intelligent news recommendation system based on user intention features described in claim 1. This method constructs and trains a news recommendation model composed of a news encoder, a term frequency-inverse document frequency (TF-IDF) algorithm module, an intention encoder, a graph neural network, an intention decoder, and a click-through rate predictor module, sorts all candidate news in descending order according to the click-through rate, and recommends the top-K news to the user; specifically as follows: Construct a news encoder, which takes the title, abstract, main category, and subcategory information of the news as input, and learns news feature vectors from the above four types of information respectively; Construct a news recommendation model based on user intention features, which takes the news feature vectors generated by the news encoder as input, and uses the TF-IDF algorithm, convolutional neural network, attention network, and graph convolutional neural network to obtain the user's intention feature vectors; Construct a click-through rate predictor module. First, use the attention network based on the candidate news to fuse the intention feature vectors and candidate news feature vectors to obtain a user vector, then use the user vector and candidate news feature vectors as input, and finally generate the score of each candidate news, that is, the click-through rate, through vector dot product operation. Sort all candidate news in descending order according to the click-through rate, and recommend the top-K news to the user.
3. The intelligent news recommendation method based on user intention features according to claim 2, wherein The construction process of the news encoder is specifically as follows: Construct a title learning module: construct a word mapping table for each word in the dataset, and map each word in the table to a unique digital identifier. The mapping rule is: starting from the number 1, and then increasing sequentially according to the order in which each word is entered into the word mapping table, so as to form a word mapping conversion table; Use the Glove pre-trained language model to obtain the word vector representation of each word; At the word embedding layer, each news title \(T = [w_1, w_2,\cdots, w N \) is converted into a vector representation, denoted as \(E = [e_1, e_2,\cdots, e N \), where \(N\) represents the length of a news title, and \(e N \) represents the vector representation of each word, and \(w\) represents a word in the news title; For E = [e1, e2, …, e N , the convolutional neural network CNN is used for feature extraction to obtain the context feature vector [c1, c2, …, c N , and the formula is as follows: c i = ReLU(Q w × e (i-k):(i+k) + b w ); Among them, i represents the relative position of the corresponding character vector in the news title, k represents the difference in the relative position from i, Q w represents the convolution kernel of the CNN filter, b w represents the bias, ReLU is an activation function, and the operator × represents matrix multiplication; For the context feature vector [c1, c2, …, c N , the attention mechanism is used to further extract key features to obtain the final news title vector r t , and the formula is as follows: Among them, q is the attention query vector randomly initialized according to the dimension of the context feature vector, V and v are parameters learned during the training process, tanh is an activation function, the operator × represents matrix multiplication, exp represents logarithmic function operation, a i is the attention score of the i-th word, α i is the attention weight of the i-th word, and N is the length of the context feature vector [c1, c2,..., c N ; Construct a summary learning module: Use the news summary as the input. The specific implementation method is the same as that of the title learning module, and then obtain the summary vector representation r a ; Construct a category learning module: in the word embedding layer, the main category label and the sub-category label are respectively mapped to low-dimensional space vectors by the word vector method to obtain the word vector representations e c and e sc of each category label, and then the activation function ReLU is used to generate the final vectors r c and r sc as follows: r c = ReLU(V c × e c + v c ); r sc = ReLU(V sc × e sc + v sc )); Among them, ReLU is an activation function, V c 、V sc 、v sc and v c are parameters learned from the training process, and the operator × represents matrix multiplication; Construct an attention module: vectors r for the title, abstract, main category label, and sub-category label t , r a , r c and r sc , and use the activation function tanh to calculate their respective attention scores, namely a t , a a , a c , a sc , and then further obtain their respective attention weights through the attention mechanism. The formula is as follows: Among them, V t 、V a 、V c 、V sc 、v t 、v a 、v c 、v sc are randomly initialized learnable parameters for calculating the title attention score a t 、abstract attention score a a 、main category label attention score a c and sub-category label attention score a sc ; is the attention query vector generated by the title vector r t ; is the attention query vector generated by the abstract vector r a ; is the attention query vector generated by the main category label vector r c ; is the attention query vector generated by the sub-category label vector r sc ; tanh is an activation function, the operator × is matrix multiplication, exp is logarithmic function operation, α t is the attention weight of the title, α a is the attention weight of the abstract, α c is the attention weight of the main category label, α sc is the attention weight of the sub-category label; The final news vector r is determined by the title vector r t , the abstract vector r a , the main category label vector r c and the sub-category label vector r sc as well as their respective attention weights, and the formula is as follows: r=[α t r t ;α a r a ;α c r c ;α sc r sc ]; Among them, the symbol ; represents the concatenation operation.
4. The intelligent news recommendation method based on user intention features according to claim 2, characterized in that, The construction process of the news recommendation model based on user intention features is specifically as follows: Build the TF-IDF algorithm module: First, input a user's browsing record C u ={v1,...,v i ,...,v t-1} into this module, where v i represents each user browsing record; then, use the TF-IDF algorithm to extract keywords from the user browsing records; finally, map these keywords to a keyword vector matrix K i ; Construct an intention encoder, and the intention encoder aims to infer the user's intention from the user's browsing records, specifically as follows: Construct a convolutional neural network: Use the keyword vectors {K1,...,K i ,...,K t-1} of the user's historical news sequence as the input of the convolutional neural network, and use the convolutional neural network to encode these vectors. The formula is as follows: c i = ReLU(W' * K (i-f):(i+f) + b'); Among them, K (i-f):(i+f) represents the keyword vector connected from position (i - f) to (i + f), W′ represents the learnable parameter of the CNN filter, b′ is the bias, * represents the convolution operation, ReLU represents the ReLU activation function, and c i represents the convoluted keyword vector; Construct an information mapping module: In order to infer the user's intention from the keyword vectors of the historical news sequence, first extract the keywords of all news from the news recommendation dataset using the TF-IDF method, then use the word embedding layer to generate a global keyword vector matrix H, and then filter the possible keywords through a learnable mapping matrix M. Then, the possible keyword vector matrix can be obtained by calculating the distribution probability of the possible keywords in the global keyword vector matrix H; the specific process formula is as follows: C = [c1; c2;...; c t-1 ; W p = softmax(HMC); C p = W p H; Among them, softmax represents the softmax normalization function, C represents the aggregation matrix after the connection of convolutional keyword vectors, and W p represents the learnable weight matrix; C p represents the possible keyword vector matrix, which contains all the initial intent feature vectors; Constructing a Graph Neural Network: Once the possible keyword vector matrix C is obtained p , the matrix is input into the graph neural network; specifically, the operation process of the l-th layer of the graph neural network is expressed as follows: where σ represents the activation function, and H l is the node representation of the l-th layer graph neural network, and W l represents the learnable parameter of the l-th layer graph neural network, D is the degree matrix; A = A + I, where A is the adjacency matrix and I is the identity matrix; specifically, the input of the first layer is C p , then its output is H 0 = C p ; after passing through n layers of graph neural networks, the user intention at time t can be represented as C t = H n , where C t contains the updated intention feature vector. Constructing the intent decoder: After the user intent C t is generated by the graph neural network, the intent decoder uses the attention mechanism to generate the final intent feature vector, as shown in the following formula: Among them, represents the attention weight, represents the j-th row vector of the C t matrix, R is the total number of rows of the matrix; u o is the final intention feature vector, exp is the logarithmic function operation, represents the activation function.
5. The intelligent news recommendation method based on user intention features according to claim 2, wherein The construction process of the click-through rate predictor is specifically as follows: Construct an attention network based on candidate news: It is designed to integrate the features of candidate news into the final intent feature vector to generate the final user vector; the formula is as follows: α = Att(W Q d, W K u o ); Among them, W Q and W K are learnable parameters, d is the news vector representation of the candidate news generated by the news encoder, U is the length of a user browsing record, u is the final user vector, Att represents the attention mechanism function, and α represents the attention weight; Construct a prediction module: It takes the vector representation d of the candidate news and the final user vector u as inputs, and uses the dot product operation to predict the click-through rate of the candidate news. The formula is as follows: Among them, represents the click-through rate of the candidate news; When the model of this method has not been fully trained, it needs to be trained on the training dataset to optimize the model parameters; when the model training is completed, the click-through rate predictor can predict the recommendation score of each candidate news, and according to the score, recommend appropriate news to the user.
6. The intelligent news recommendation method based on user intention features according to claim 2, wherein The construction process of the training dataset is specifically as follows: Construct a news dataset or select a publicly available news dataset; Preprocess the news dataset: Preprocess each news text in the news dataset to remove stop words and special characters in the news dataset; Extract the title, category, subcategory, and abstract information of each news text respectively; Construct training positive examples: Use the news numbers labeled 1 in the historical news sequence and interaction behavior sequence in the user browsing record, that is, the numbers of the news clicked by the user, to construct training positive examples; Construct training negative examples: Use the news numbers labeled 0 in the historical news sequence and interaction behavior sequence in the user browsing record, that is, the numbers of the news not clicked by the user, to construct training negative examples; Construct a training dataset: Combine all the positive example data and negative example data and shuffle their order to construct the final training dataset.
7. The news recommendation method based on user intention features according to claim 2, wherein After the news recommendation model based on user intent features is constructed, it is trained and optimized through the training dataset, specifically as follows: Construct the loss function: Using the negative sampling technique, the news that a user has clicked is defined as the positive example, and the news that has not been clicked is defined as the negative example, and calculate the click prediction value p of the positive example i ; The formula is as follows: wherein, is the click-through rate of the j-th negative example relative to the i-th positive example in the same click sequence, is the i-th positive example, and G is the number of negative examples; The loss function of news recommendation is the negative log-likelihood function of all positive examples. The formula is as follows: Among them, is the set of positive examples; Optimize the training model: Select the Adam optimization function as the optimization function of this model. Among them, the learning rate is set to 0.001, the smoothing constants are set to (0.9, 0.999), eps is set to 1e-8, and the L2 penalty value is set to 0.
8. A storage medium storing a plurality of instructions, characterized in that, The instruction is loaded by the processor and executes the steps of the news recommendation method based on user intent features described in any one of claims 2-7.
9. An electronic device, characterized in that, The electronic device includes: The storage medium described in claim 8; and a processor for executing the instructions in the storage medium.
Citation Information
Patent Citations
Session recommendation method and system based on graph neural network and comment similarity
CN113610610A
Generating recommendation information
US20210027018A1