Intelligent News Recommendation Method and System Based on Explicit and Implicit Interest Features
By building an intelligent news recommendation model based on explicit and implicit interest characteristics, using neural networks and deep learning methods, the problem of inaccurate recommendation results in the existing technology is solved, and comprehensive modeling and accurate recommendation of user interest characteristics is achieved.
Patent Information
- Application Number
- CN202310412932.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-04-13
AI Technical Summary
Existing news recommendation methods cannot accurately identify users' explicit and implicit interest characteristics, resulting in inaccurate recommendation results.
Build an intelligent news recommendation model based on explicit and implicit interest characteristics. Through neural networks and deep learning methods, use news encoder, explicit interest encoder, TF-IDF algorithm module, implicit interest encoder, graph neural network and click-through rate predictor, and combine user browsing records and news text content to generate user feature vectors and predict click-through rate.
It improves the accuracy of news recommendations, can fully model user interest characteristics, and improves the effectiveness of recommendations.
Smart Images

Figure CN116340641B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of recommendation systems and natural language processing, and particularly relates to an intelligent news recommendation method and system based on explicit and implicit interest features. Background Art
[0002] In reality, users' behaviors are often influenced by their explicit and implicit interests. In most cases, users tend to purposefully read some news related to a topic or event, such as economic news, which indicates their explicit interest in reading. Then, for the same purpose (e.g., reading economic news), different users may have different preferences for specific news, and thus may read different economic news. However, existing news recommendation methods usually establish a single interest representation of users based on the news text content in their browsing records, and then perform recommendation tasks based on the single interest feature. Although these methods have achieved certain results in news recommendation tasks, they all ignore the explicit and implicit interests of users; simply modeling a single general user interest feature often fails to accurately capture the comprehensive interest features of users, which inevitably affects the accuracy of news recommendation results. Summary of the Invention
[0003] The technical task of the present invention is to provide an intelligent news recommendation method and system based on explicit and implicit interest features to solve the problem of inaccurate recommendation results in news recommendation systems.
[0004] The technical task of the present invention is realized in the following way. An intelligent news recommendation method based on explicit and implicit interest features, the method includes the following steps:
[0005] S1. Construct a training data set for the news recommendation model: First, download the publicly available news data set on the network, then preprocess the data set, and finally construct positive example data and negative example data, and combine them to generate the final training data set;
[0006] S2. Construct a news recommendation model based on explicit and implicit interest features: Use neural networks and deep learning methods to construct a news recommendation model;
[0007] S3. Train the model: Train the news recommendation model constructed in step S2 in the training data set obtained in step S1.
[0008] An intelligent news recommendation system based on explicit and implicit interest features, the system includes:
[0009] A training data set generation unit, which first obtains the browsing record information of users on online news websites, and then performs preprocessing operations on it to obtain the user browsing records that meet the training requirements and their news text content;
[0010] A news recommendation model construction unit based on explicit and implicit interest features, which is used to load a training data set, construct a news encoding module, construct an explicit interest encoding module, construct a TF-IDF algorithm module, construct an implicit interest encoding module, construct a graph neural network module, construct an implicit interest decoding module, and construct a click-through rate predictor module;
[0011] A model training unit, which is used to construct a loss function required during model training and complete the optimization training of the model.
[0012] A storage medium, in which multiple instructions are stored, and the instructions are loaded by a processor to execute the steps of the above-mentioned intelligent news recommendation method based on user implicit interest features.
[0013] An electronic device, the electronic device includes: the above storage medium; and a processor for executing the instructions in the storage medium.
[0014] Preferably, the construction process of the news encoder is as follows:
[0015] First, a word mapping table is constructed for each word in the data set, and each word in the table is mapped to a unique digital identifier. The mapping rule is: starting from the number 1, and then sorted in ascending order in sequence according to the order in which each word is entered into the word mapping table, so as to form a word mapping conversion table; use the Glove pre-trained language model to obtain the word vector representation of each word; in the word embedding layer, each news title T = [w1, w2,..., w N is converted into a vector representation, denoted as x = [x1, x2,..., x N , where N represents the length of a news title, and x N represents the vector representation of each word, and w represents a word in the news title.
[0016] Then, using the news title vector x as the input, elements in the input are randomly set to zero with a certain probability to obtain a noisy vector Then the noisy vector is input into the fully connected layer to obtain the hidden layer representation h, and the formula is as follows:
[0017]
[0018]
[0019] Among them, represents the noisy vector, x represents the news title vector, q(x) represents the random zeroing process, f(·) represents the sigmoid activation function, and U and u are parameters learned from the training process.
[0020] Finally, using the hidden layer representation h as the input, the news feature vector r is reconstructed through a fully connected layer. The formula is as follows:
[0021] r = f(U'h + u')
[0022] where r is the news feature vector, f(·) represents the sigmoid activation function, and U' and u' are parameters learned during the training process.
[0023] More preferably, the construction process of the news recommendation model based on explicit and implicit interest features is specifically as follows:
[0024] Construct an explicit interest encoder: To generate the explicit interest features of users, the explicit interest encoder uses the Fastformer method to process the user browsing records and outputs the explicit interest feature vector; specifically as follows:
[0025] First, Fastformer converts the input news feature vector into three vector representations of query, key, and value through three linear layers with non-shared parameters. The formula is as follows:
[0026] q i = W q r i
[0027] k i = W k r i
[0028] v i = W v r i
[0029] where W q , W k and W v are all learnable parameters, r i represents the i-th news feature vector, q i represents the query vector of the i-th news, k i represents the key vector of the i-th news, and v i represents the value vector of the i-th news;
[0030] Then, the additive attention mechanism is used to aggregate and compress the query vectors. The formula is expressed as follows:
[0031] q = Att(q1, q2,..., q N )
[0032] where q iDenote the query vector of the $i$-th news as $q_i$, $q$ represents the query vector aggregated with context information, and Att represents the additive attention mechanism;
[0033] After that, the interaction information between the key vector and the query vector is calculated using the additive attention mechanism and element-wise multiplication operation. The formula is as follows:
[0034] $k = Att(q\odot k_1, q\odot k_2, \cdots, q\odot k$ i , \cdots, q\odot k$ N )
[0035] where $k$ i represents the key vector of the $i$-th news, $\hat{k}$ represents the key vector aggregated with context information, $\odot$ represents element-wise multiplication, and Att represents the additive attention mechanism;
[0036] Then, the key vector and the value vector are processed through dot product operation and a linear layer to obtain the news feature vector of a single attention head. The formula is expressed as follows:
[0037]
[0038] where $W$ o is a learnable parameter, $\odot$ represents element-wise multiplication, $v$ i represents the value vector of the $i$-th news, is the $i$-th news feature vector output by a single attention head;
[0039] Finally, based on the outputs of $M$ attention heads and combined with the user browsing record, an explicit interest feature vector is established. The formula is expressed as follows:
[0040]
[0041] $u$ p $= [d_1; d_2; \cdots; d$ k ; \cdots; d$ N
[0042] where $[;]$ represents the concatenation operation, is the $k$-th news feature vector output by the $n$-th attention head, $M$ is the number of attention heads, $N$ is the length of the user browsing record, $d$ k is the $k$-th news feature vector obtained by aggregating and concatenating $M$ attention heads, and $u$ p is the explicit interest feature vector.
[0043] Construct the TF-IDF algorithm module: First, a user browsing record $C$ u $= \{v_1, \cdots, v$ i , \cdots, v$t-1} are input into this module, where v represents each user browsing record; then, the TF-IDF algorithm is used to extract keywords from the user browsing records; finally, the keywords are mapped to a keyword vector matrix K through a word embedding layer, where this matrix contains the keyword vectors of this segment of user browsing records.
[0044] Construct an implicit interest encoder, which aims to infer the implicit interests of users from user browsing records, as follows:
[0045] Construct a multi-layer perceptron:
[0046] Take the keyword vector matrix K of the user browsing records as the input, and use the multi-layer perceptron to encode these vectors. The formula is as follows:
[0047] C = MLP(W′K + b′)
[0048] Where K is the keyword vector matrix, W′ represents the learnable parameters of the multi-layer perceptron, b′ is the bias, C represents the keyword vector output after being processed by the multi-layer perceptron, and MLP is the multi-layer perceptron.
[0049] Construct an interest inference module:
[0050] To infer implicit interests from the keyword vectors of the historical news sequence, first extract the keywords of all news from the news recommendation dataset using the TF-IDF method, map them to a global keyword vector matrix H through a word embedding layer, then filter the possible keywords through a learnable mapping matrix M, and then obtain the possible keyword vector matrix, that is, the initial implicit interest feature vector, by calculating the distribution probability of the possible keywords in the global keyword vector matrix H. The specific process formula is as follows:
[0051] W p = softmax(HMC);
[0052] C p = W p H;
[0053] Where softmax represents the softmax normalization function, and W p represents the learnable weight matrix. C p represents the possible keyword vector matrix, which contains all the initial implicit interest feature vectors.
[0054] Construct a graph neural network: Use the initial implicit interest feature vector C p as the input, and obtain the updated implicit interest feature vector through the graph neural network. Specifically, the operation process of the l-th layer graph neural network is as follows:
[0055]
[0056] Among them, σ represents the activation function; H l is the node representation of the l-th layer of the graph neural network, and W l represents the learnable parameter of the l-th layer of the graph neural network, D is the degree matrix; A = A + I, where A is the adjacency matrix and I is the identity matrix; specifically, the input of the first layer is C p , then its output is H 0 = C p ; After passing through the graph neural network with n layers, the implicitly interested feature vector updated at time t can be expressed as C t = H n .
[0057] Construct an implicit interest decoder: Using the updated implicitly interested feature vector C t as the input, a multi-layer perceptron is used as the decoder to generate the final implicitly interested feature vector. The formula is as follows:
[0058] u o = MLP(WC t + b)
[0059] Among them, C t is the updated implicitly interested feature vector, W is the learnable parameter of the multi-layer perceptron, b is the bias, and u o is the final implicitly interested feature vector, and MLP is the multi-layer perceptron.
[0060] More preferably, the construction process of the click-through rate predictor is specifically as follows:
[0061] Construct a gating network: It is designed to select important feature information and aggregate the explicit interest feature vector and the final implicit interest feature vector; using the explicit interest feature vector u p generated by the explicit interest encoder and the final implicit interest feature vector u o generated by the implicit interest decoder as the input, a user feature vector u g is generated through the gating network; The formula is expressed as follows:
[0062] g = ReLU(W g [u o ; u p + b g )
[0063] u g = g ⊙ tanh(Vu o + v)+(1 - g) ⊙ u p
[0064] Among them, W g, W b , V, and v represent learnable parameters, b g represents the bias, the symbol ; represents the connection operation, u p is the explicit interest feature vector, u o is the final implicit interest feature vector, ReLU and tanh are activation functions, u g is the user feature vector, and g is the gating network.
[0065] Construct an attention network based on candidate news, which is designed to integrate the features of candidate news into the user feature vector to generate the final user feature vector; the formula is as follows:
[0066] α = Att(W Q n, W K u g );
[0067]
[0068] where, W Q , W K are learnable parameters, n is the news feature vector of candidate news generated by the news encoder, u g is the user feature vector, L is the length of a user browsing record, u is the final user feature vector, Att represents the attention mechanism function, and α is the attention weight.
[0069] Construct a prediction module: It takes the news feature vector n of candidate news generated by the news encoder and the final user feature vector u as inputs, and predicts the click-through rate of candidate news through dot product operation. The formula is as follows:
[0070]
[0071] where, represents the click-through rate of candidate news.
[0072] When the model of this method has not been fully trained, it needs to be trained on the training dataset to optimize the model parameters; when the model training is completed, the click-through rate predictor can predict the recommendation score of each candidate news, and recommend appropriate news to users according to the score.
[0073] More preferably, the construction process of the training dataset is specifically as follows:
[0074] Construct a news dataset or select a publicly available news dataset;
[0075] Preprocess the news dataset: Preprocess each news text in the news dataset, removing stop words and special characters from the news dataset; extract the title, category, sub-category, and summary information of each news text respectively;
[0076] Construct training positive examples: Use the historical news sequence and the news numbers with label 1 in the interaction behavior sequence in the user browsing record, that is, the numbers of the news clicked by the user, to construct training positive examples;
[0077] Construct training negative examples: Use the historical news sequence and the news numbers with label 0 in the interaction behavior sequence in the user browsing record, that is, the numbers of the news not clicked by the user, to construct training negative examples;
[0078] Construct the training dataset: Combine all the positive example data and negative example data and shuffle their order to construct the final training dataset;
[0079] After the news recommendation model is constructed, it is trained and optimized through the training dataset, specifically as follows:
[0080] Construct the loss function: Adopt the negative sampling technique, define the news clicked by a user as a positive example, and the news not clicked as a negative example, and calculate the click prediction value p of the positive example i . The formula is as follows:
[0081]
[0082] Among them, is the click-through rate of the j-th negative example relative to the i-th positive example in the same click sequence, is the i-th positive example, and G is the number of negative examples.
[0083] The loss function of news recommendation is the negative log-likelihood function of all positive examples, and the formula is as follows:
[0084]
[0085] Among them, is the set of positive examples.
[0086] Optimize the training model: Select the Adam optimization function as the optimization function of this model. Among them, the learning rate is set to 0.001, the smoothing constants are set to (0.9, 0.999), eps is set to 1e-8, and the L2 penalty value is set to 0.
[0087] An intelligent news recommendation system based on explicit and implicit interest features, the system includes,
[0088] The training dataset generation unit first obtains the browsing record information of users on online news websites and then performs preprocessing operations on it to obtain user browsing records and their news text content that meet the training requirements. The training dataset generation unit includes:
[0089] The original data acquisition unit is responsible for downloading the publicly available news website dataset on the network and using it as the original data for constructing the training dataset.
[0090] The original data preprocessing unit is responsible for preprocessing each news text in the news dataset, removing stop words and special characters in the news dataset; extracting the key information of each news text, such as the title; thereby constructing the training dataset.
[0091] The news recommendation model construction unit based on explicit and implicit interest features is used to load the training dataset, construct a news encoding module, construct an explicit interest encoding module, construct a TF-IDF algorithm module, construct an implicit interest encoding module, construct a graph neural network module, construct an implicit interest decoding module, and construct a click-through rate predictor module. The news recommendation model construction unit based on explicit and implicit interest features includes:
[0092] The training dataset loading unit is responsible for loading the training dataset.
[0093] The news encoding module construction unit is responsible for training news feature vectors based on the Glove word vector model in the training dataset and defining all news feature vectors; first encoding the news title vector using a fully connected layer to obtain a hidden layer representation, and finally decoding the hidden layer representation using a fully connected layer to reconstruct the news feature vectors.
[0094] The explicit interest encoding module construction unit is responsible for constructing explicit interest feature vectors according to user browsing records; among them, the news feature vectors of user browsing records are obtained by the news encoding module construction unit, and the explicit interest feature vectors are obtained using the Fastformer method.
[0095] The TF-IDF algorithm module construction unit is responsible for using the TF-IDF algorithm to extract news keywords in user browsing records, and then using the word embedding method to map each keyword to the same vector space to obtain keyword vectors of news content.
[0096] The implicit interest encoding module construction unit is responsible for using a multi-layer perceptron to extract the main features of keyword vectors and generating a keyword vector matrix through an aggregation operation, then filtering possible keywords through a learnable mapping matrix M, and then obtaining a possible keyword vector matrix by calculating the distribution probability of possible keywords in the keyword vector matrix. This matrix contains the initial implicit interest feature vectors.
[0097] The graph neural network module construction unit is responsible for using the graph neural network to propagate and aggregate the initial implicit interest feature vector, so as to obtain an updated implicit interest feature vector.
[0098] The implicit interest decoding module construction unit is responsible for using a multi-layer perceptron to decode the updated implicit interest feature vector, so as to obtain the final implicit interest feature vector.
[0099] The click-through rate predictor module construction unit first uses a gated network to select important feature information and aggregates the explicit interest feature vector and the final implicit interest feature vector to obtain a user feature vector, then fuses the user feature vector and the news feature vector of the candidate news based on the attention network of the candidate news to obtain the final user feature vector, and finally uses the final user feature vector and the news feature vector of the candidate news as inputs, and generates the score of each candidate news, that is, the click-through rate, through a dot product operation, sorts all candidate news in descending order according to the click-through rate, and recommends the top-K news to the user.
[0100] The model training unit is used to construct the loss function required in the model training process and complete the optimization training of the model; the model training unit includes,
[0101] The loss function construction unit is responsible for calculating the error between the predicted candidate news and the real target news;
[0102] The model optimization unit is responsible for training and adjusting the parameters in the model training to reduce the prediction error.
[0103] A storage medium stores multiple instructions, and is characterized in that the instructions are loaded by a processor to execute the steps of the above-mentioned intelligent news recommendation method based on explicit and implicit interest features.
[0104] An electronic device is characterized in that the electronic device includes:
[0105] The above-mentioned storage medium; and
[0106] A processor for executing the instructions in the storage medium.
[0107] The intelligent news recommendation method and system based on explicit and implicit interest features of the present invention have the following advantages:
[0108] (1) The present invention proposes an intelligent news recommendation method based on explicit and implicit interest features, mines the explicit and implicit interest features of users, can model the feature representation of users more comprehensively, and further improves the accuracy of news recommendation.
[0109] (2) First, the present invention extracts keyword features from the news content in the user browsing record through the TF-IDF method, and then uses a multi-layer perceptron and a graph neural network to encode and decode the user browsing record, so as to accurately model the implicit interest representation.
[0110] (3) The present invention constructs an explicit interest feature vector based on the user browsing record. First, it obtains the news feature vector of the user browsing record through the news encoding module, and then uses the Fastformer method to obtain the explicit interest feature vector, which can accurately model the explicit interest feature, thereby improving the accuracy of news recommendation.
[0111] (4) The present invention fuses the explicit interest feature representation and the implicit interest feature representation through a gating network, and then based on the attention of the candidate news, it fuses the user feature vector and the news feature vector of the candidate news, thereby improving the accuracy of recommendation.
[0112] (5) Through the click-through rate predictor module of the present invention, the prediction scores of the candidate news sequence can be accurately output according to the accurate news representation and user representation. Brief Description of the Drawings
[0113] The present invention will be further described below with reference to the accompanying drawings.
[0114] Figure 1 It is a flowchart of an intelligent news recommendation method based on explicit and implicit interest features
[0115] Figure 2 It is a flowchart of constructing a training data set for a news recommendation model
[0116] Figure 3 It is a flowchart of constructing a news recommendation model based on explicit and implicit interest features
[0117] Figure 4 It is a flowchart of training a news recommendation model based on explicit and implicit interest features
[0118] Figure 5 It is a schematic diagram of a news recommendation model based on explicit and implicit interest features
[0119] Figure 6 It is a schematic diagram of a news encoder Detailed Embodiment
[0120] The intelligent news recommendation method and system based on explicit and implicit interest features of the present invention will be described in detail below with reference to the accompanying drawings of the specification and specific embodiments.
[0121] Embodiment 1:
[0122] The overall model framework of the present invention is as Figure 5 shown. It consists ofFigure 5 It can be seen that the main framework structure of the present invention includes a news encoder, an explicit interest encoder, a term frequency-inverse document frequency (TF-IDF) algorithm module, an implicit interest encoder, a graph neural network, an implicit interest decoder, and a click-through rate predictor module. Among them, the news encoder is responsible for extracting the main features of the news title from the news title vector using a fully connected layer and generating a news feature vector; the explicit interest encoder is responsible for aggregating and encoding the news feature vectors of the user's browsing records using the Fastformer method and generating an explicit interest feature vector; the TF-IDF algorithm module is responsible for extracting the keywords of the news content from the user's browsing records and then generating keyword vectors using a word embedding layer; the implicit interest encoder is responsible for encoding an initial implicit interest feature vector from the keyword vectors using a multi-layer perceptron; the graph neural network is responsible for propagating the input initial implicit interest feature vector using a graph convolutional neural network and obtaining an updated implicit interest feature vector through information aggregation; the implicit interest decoder is responsible for decoding the updated implicit interest feature vector obtained from the graph neural network using a multi-layer perceptron and obtaining the final implicit interest feature vector; the click-through rate predictor module first uses a gated network to select important feature information and aggregates the explicit interest feature vector and the final implicit interest feature vector to obtain a user feature vector, then fuses the user feature vector and the news feature vector of the candidate news based on the attention network of the candidate news to obtain the final user feature vector, and finally uses the final user feature vector and the news feature vector of the candidate news as inputs to generate the score of each candidate news, that is, the click-through rate, through dot product operation, sorts all candidate news in descending order according to the click-through rate, and recommends the top-K news to the user; the above is the structural introduction of the present model invention.
[0123] Example 2:
[0124] As shown in the appendix Figure 1 The intelligent news recommendation method based on explicit and implicit interest features of the present invention is as follows:
[0125] S1. Construct a training data set for the news recommendation model: The news data set includes two parts of data files: user browsing records and news text content; among them, the user browsing records include user ID, time, historical news sequence, and interaction behavior sequence; the news text content includes news ID, category, sub-category, title, abstract, and entity; select the historical news sequence and interaction behavior sequence in the user browsing records to construct the user behavior data of the training data set, and select the title, category, sub-category, and abstract of the news text content to construct the news text data of the training data set; among them, the user behavior data will be used for the extraction of explicit and implicit interest features, and the news text content data will be used for the extraction of news features; the method for constructing the training data set is as follows:
[0126] S101. Construct a news dataset or select a publicly available news dataset.
[0127] For example: Download the MIND news dataset publicly available on the Internet by Microsoft and use it as the original data for news recommendation. MIND is currently the largest English news recommendation system dataset, containing 1,000,000 users in 200,000 categories and 161,013 news, divided into a training set, a validation set, and a test set. The MIND dataset also provides detailed information on the news text content. Each news has a news number, a link, a title, an abstract, a category, and an entity:
[0128]
[0129] In addition, the MIND dataset also provides user browsing records, and each record contains a user number, a time, a historical news sequence, and an interaction behavior sequence:
[0130]
[0131]
[0132] Among them, the user number represents the unique number of each user on the news platform; the time represents the start time when the user clicks to browse a series of news; the historical news sequence represents the sequence of a series of news numbers browsed by the user; the interaction behavior sequence represents the actual interaction behavior of the user on a series of news recommended by the system, 1 means click, and 0 means no click.
[0133] S102. Preprocess the news dataset: Preprocess each news text in the news dataset, remove stop words and special characters in the news dataset; extract the title, category, subcategory, and abstract information of each news text respectively.
[0134] S103. Construct training positive examples: Use the historical news sequence in the user browsing record and the news numbers with the label of 1 in the interaction behavior sequence, that is, the news numbers of the news clicked by the user, to construct training positive examples.
[0135] For example: For the news example shown in step S101, the positive example data is formalized as: (N29038,N15201,N8018,N32012,N30859,N26552,N25930). The last number is the news number clicked by the user.
[0136] S104. Construct training negative examples: Use the historical news sequence in the user browsing record and the news numbers with the label of 0 in the interaction behavior sequence, that is, the news numbers of the news not clicked by the user, to construct training negative examples.
[0137] Example: For the news example shown in step S101, the negative example data constructed is formalized as: (N29038, N15201, N8018, N32012, N30859, N26552, N17825). The last number among them is the number of the news that has not been clicked by the user.
[0138] S105. Construct a training data set: Combine all the positive example data and negative example data obtained after the operations in step S103 and step S104, and shuffle their order to construct the final training data set.
[0139] S2. Construct a news recommendation model based on explicit and implicit interest features: As shown in the appendix Figure 3 This news recommendation model includes a news encoder, an explicit interest encoder, a term frequency-inverse document frequency (TF-IDF) algorithm module, an implicit interest encoder, a graph neural network, an implicit interest decoder, and a click-through rate predictor module. Specifically as follows:
[0140] S201. Construct a news encoder. As shown in the appendix Figure 6 Using the title information of the news as input, learn the news feature vector from the above information. Specifically as follows:
[0141] First, construct a word mapping table for each word in the data set, and map each word in the table to a unique digital identifier. The mapping rule is: starting from the number 1, and then sorting them in ascending order according to the order in which each word is entered into the word mapping table, so as to form a word mapping conversion table; use the Glove pre-trained language model to obtain the word vector representation of each word; in the word embedding layer, convert each news title T = [w1, w2,..., w N into a vector representation, denoted as x = [x1, x2,..., x N , where N represents the length of a news title, and x N represents the vector representation of each word, and w represents a word in the news title.
[0142] Then, using the news title vector x as input, randomly set the elements in the input to zero with a certain probability to obtain a noisy vector Then the noisy vector is input into the fully connected layer to obtain the hidden layer representation h. The formula is as follows:
[0143]
[0144]
[0145] Among them, The vector with noise is denoted as, the news title vector is denoted as x, the random zeroing process is denoted as q(x), the sigmoid activation function is denoted as f(·), and U and u are parameters learned from the training process.
[0146] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0147] self.f1 = nn.Linear(in_features = config.word_embedding_dim, out_features = config.hidden_dim, bias = True)
[0148] corrupted_word_embedding = self.dropout_(word_embedding)
[0149] h = torch.sigmoid(self.f1(corrupted_word_embedding))
[0150] Among them, nn.Linear and torch.sigmoid are the built-in linear layer method and activation function in PyTorch respectively. corrupted_word_embedding is the vector with noise. h is the hidden layer representation.
[0151] Finally, using the hidden layer representation h as the input, the news feature vector r is reconstructed through the fully connected layer. The formula is as follows:
[0152] r = f(U'h + u')
[0153] Among them, r is the news feature vector, f(·) represents the sigmoid activation function, and U' and u' are parameters learned from the training process.
[0154] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0155] self.f2 = nn.Linear(in_features = config.hidden_dim, out_features = config.word_embedding_dim, bias = True)
[0156] news_representation = torch.sigmoid(self.f2(h))
[0157] Among them, torch.sigmoid and nn.Linear are the built-in activation function and connection layer method in PyTorch respectively. news_representation is the news feature vector.
[0158] S202. Construct an explicit interest encoder. To generate the explicit interest features of the user, the explicit interest encoder processes the user's browsing records using the Fastformer method and outputs the explicit interest feature vector. Specifically as follows:
[0159] First, Fastformer converts the input news feature vector into three vector representations of query, key, and value through three linear layers with non-shared parameters. The formula is as follows:
[0160] q i = W q r i
[0161] k i = W k r i ;
[0162] v i = W v r i
[0163] Among them, W q 、W k and W v are all learnable parameters, r i represents the i-th news feature vector, q i represents the query vector of the i-th news, k i represents the key vector of the i-th news, and v i represents the value vector of the i-th news;
[0164] Then, use the additive attention mechanism to aggregate and compress the query vectors. The formula is expressed as follows:
[0165] q = Att(q1, q2,..., q N )
[0166] Among them, q i represents the query vector of the i-th news, q represents the query vector aggregated with context information, and Att represents the additive attention mechanism;
[0167] After that, use the additive attention mechanism and bitwise multiplication operation to calculate the interaction information between the key vector and the query vector. The formula is as follows:
[0168] k = Att(q⊙k1, q⊙k2,..., q⊙k i ,..., q⊙k N )
[0169] where k i represents the key vector of the i-th news, k represents the key vector aggregated with context information, ⊙ represents element-wise multiplication, and Att represents the additive attention mechanism;
[0170] Then, through dot product operation and linear layer processing of the key vector and value vector, the news feature vector of a single attention head is obtained. The formula is as follows:
[0171]
[0172] where W o is a learnable parameter, ⊙ represents element-wise multiplication, v i represents the value vector of the i-th news, is the i-th news feature vector output by a single attention head;
[0173] Finally, based on the outputs of M attention heads and combined with the user browsing history, an explicit interest feature vector is established. The formula is as follows:
[0174]
[0175] u p = [d1;d2;...;d k ;...;d N
[0176] where [;] represents the concatenation operation, is the k-th news feature vector output by the n-th attention head, M is the number of attention heads, N is the length of the user browsing history, d k is the k-th news feature vector obtained by aggregating and concatenating M attention heads, and u p is the explicit interest feature vector.
[0177] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0178] self.fastformer = FastformerEncoder(config)
[0179] h = self.fastformer(history_embedding.view(batch_size, -1, self.news_embedding_dim))
[0180] Among them, FastformerEncoder(config) is a custom Transformer method.
[0181] S203. Build a TF-IDF algorithm module: First, input a user's browsing record C u ={v1,...,v i ,...,v t-1} into this module, where v represents each user's browsing record; then, use the TF-IDF algorithm to extract keywords from the user's browsing record; finally, map the keywords to a keyword vector matrix K through a word embedding layer, where this matrix contains the keyword vectors of this user's browsing record.
[0182] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0183] tfidf_model = TfidfVectorizer().fit(df)
[0184] sparse_result = tfidf_model.transform(df)
[0185] Among them, df is the input data, TfidfVectorizer() is the vectorization method of the TF-IDF algorithm, and tfidf_model.transform() is the method for sparse matrix conversion.
[0186] S204. Build an implicit interest encoder, which aims to infer the user's implicit interest from the user's browsing record, specifically as follows:
[0187] S20401. Build a multi-layer perceptron:
[0188] Take the keyword vector matrix K of the user's browsing record as the input, and use the multi-layer perceptron to encode these vectors. The formula is expressed as follows:
[0189] C = MLP(W′K + b′)
[0190] Among them, K is the keyword vector matrix, W′ represents the learnable parameters of the multi-layer perceptron, b′ is the bias, C represents the keyword vector output after being processed by the multi-layer perceptron, and MLP is the multi-layer perceptron.
[0191] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0192]
[0193] Among them, nn.Sequential, nn.Linear, and nn.ReLU are respectively the methods for building neural network modules, connection layers, and activation functions built into pytorch.
[0194] S20402. Construct an interest inference module:
[0195] In order to infer implicit interests from the keyword vectors of the historical news sequence, first extract the keywords of all news from the news recommendation dataset using the TF-IDF method, map them to a global keyword vector matrix H through a word embedding layer, then filter the possible keywords through a learnable mapping matrix M, and then obtain the possible keyword vector matrix, that is, the initial implicit interest feature vector, by calculating the distribution probability of the possible keywords in the global keyword vector matrix H; the specific process formula is as follows:
[0196] W p = softmax(HMC);
[0197] C p = W p H;
[0198] Among them, softmax represents the softmax normalization function, and W p represents the learnable weight matrix. C p represents the possible keyword vector matrix, which contains all the initial implicit interest feature vectors.
[0199] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0200] temp = torch.matmul(c.reshape(-1, self.word_embedding_dim), self.transform_matrix)
[0201] t = torch.matmul(temp, self.pretrained_concept_embedding.transpose(0, 1)) concept_weight = F.softmax(t, dim = 1)
[0202] personalized_concept_vector = torch.matmul(concept_weight, self.pretrained_concept_embedding).reshape(batch_size, -1, self.word_embedding_dim)
[0203] Among them, torch.matmul is matrix multiplication, and F.softmax is the softmax normalization function.
[0204] S205. Construct a graph neural network:
[0205] Using the initial implicit interest feature vector C p as the input, an updated implicit interest feature vector is obtained through the graph neural network; specifically, the operation process of the l-th layer of the graph neural network is expressed as follows:
[0206]
[0207] Among them, σ represents the activation function; H l is the node representation of the l-th layer of the graph neural network, W l represents the learnable parameter of the l-th layer of the graph neural network, D is the degree matrix; A = A + I, where A is the adjacency matrix and I is the identity matrix; specifically, the input of the first layer is C p , then its output is H 0 = C p ; After passing through n layers of the graph neural network, the updated implicit interest feature vector at time t can be expressed as C t = H n .
[0208] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0209]
[0210] Among them, GCN is the graph neural network method of the pytorch_geometric toolkit, in_dim is the size of the input vector, out_dim is the size of the output vector, hidden_dim is the size of the hidden layer vector, and num_layers is the number of layers of the graph neural network.
[0211] S206. Construct an implicit interest decoder:
[0212] Using the updated implicit interest feature vector C t as the input, a multi-layer perceptron is used as the decoder to generate the final implicit interest feature vector, and the formula is as follows:
[0213] u o = MLP(W C t + b)
[0214] where C t is the updated implicit interest feature vector, W are the learnable parameters of the multi-layer perceptron, b is the bias, and u o is the final implicit interest feature vector, and MLP is the multi-layer perceptron.
[0215] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0216]
[0217]
[0218] where nn.Sequential, nn.Linear, and nn.relu are the built-in methods in PyTorch for building neural network modules, connecting layers, and activation functions respectively, and user_vector is the final implicit interest feature vector.
[0219] S207. Construct a click-through rate predictor, which mainly includes an attention network based on candidate news and a prediction module. Specifically as follows:
[0220] S20701. Construct a gating network, which is designed to select important feature information and aggregate the explicit interest feature vector and the final implicit interest feature vector; using the explicit interest feature vector u p generated by the explicit interest encoder and the final implicit interest feature vector u o generated by the implicit interest decoder as inputs, and generate a user feature vector u g through the gating network; The formula is expressed as follows:
[0221] g = ReLU(W g [u o ; u p + b g )
[0222] u g = g ⊙ tanh(V u o + v) + (1 - g) ⊙ u p
[0223] where W g and W b , V, and v represent learnable parameters, b g represents the bias, the symbol ⊙ represents the concatenation operation, u p is the explicit interest feature vector, uo is the final implicit interest feature vector, ReLU and tanh are activation functions, and u g is the user feature vector, and g is the gating network.
[0224] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0225]
[0226]
[0227] Among them, nn.Sequential is a built-in model construction method in pytorch, nn.Linear is a linear layer method, nn.Sigmoid is an activation function, and torch.cat is a vector concatenation method.
[0228] S20702. Construct an attention network based on candidate news, which is designed to integrate the features of candidate news into the user feature vector to generate the final user feature vector. The formula is as follows:
[0229] α = Att(W Q n, W K u g )
[0230]
[0231] Among them, W Q , W K are learnable parameters, n is the news feature vector of candidate news generated by the news encoder, u g is the user feature vector, L is the length of a user browsing record, u is the final user feature vector, Att represents the attention mechanism function, and α is the attention weight.
[0232] For example: In the pytorch machine learning framework, the code implementation for the above description is as follows:
[0233]
[0234] Among them, ScaledDotProduct_CandidateAttention is a custom dot product attention method.
[0235] S20703. Construct a prediction module, which takes the news feature vector n of candidate news generated by the news encoder and the final user feature vector u as inputs, and predicts the click-through rate of candidate news through dot product operation. The formula is as follows:
[0236]
[0237] Among them, represents the click-through rate of candidate news.
[0238] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0239] probability = torch.bmm(
[0240] user_vector.unsqueeze(dim = 1),
[0241] candidate_news_vector.unsqueeze(dim = 2)).flatten()
[0242] Among them, torch.bmm is the dot product operation, user_vector is the user feature vector, and candidate_news_vector is the news feature vector.
[0243] S3. Train the model: As shown in the appendix Figure 4 as follows:
[0244] S301. Construct the loss function: Adopt the negative sampling technique. Define the news clicked by a user as a positive example, and the news not clicked as a negative example, and calculate the click prediction value p i . The formula is as follows:
[0245]
[0246] Among them, is the click-through rate of the j-th negative example relative to the i-th positive example in the same click sequence, is the i-th positive example, and G is the number of negative examples.
[0247] The loss function of news recommendation is the negative log-likelihood function of all positive examples, and the formula is as follows:
[0248]
[0249] Among them, is the set of positive examples.
[0250] For example: In the PyTorch machine learning framework, the code implementation for the above description is as follows:
[0251] loss = torch.stack([x[0] for x in -F.log_softmax(y_pred, dim = 1)
[0252] ).mean()
[0253] Among them, F.log_softmax is the built-in log_softmax loss function in pytorch, and y_pred is the click prediction value p i 。
[0254] S302. Optimize the model: Select the Adam optimization function as the optimization function of this model. Among them, the learning rate is set to 0.0001, the smoothing constants are set to (0.9, 0.999), eps is set to 1e-8, and the L2 penalty value is set to 0.
[0255] In the experiment, the present invention selects the area under the ROC curve AUC, the mean reciprocal rank MRR, and the cumulative gain nDCG as evaluation indicators.
[0256] For example: The above-described optimization function is represented by the following code in pytorch:
[0257] optimizer = torch.optim.Adam(model.parameters(), lr = learning_rate)
[0258] Among them, torch.optim.Adam is the Adam optimization function embedded in pytorch, model.parameters() is the set of parameters for model training, and learning_rate is the learning rate.
[0259] The model of the present invention has achieved better results than the current model on the MIND public dataset. The comparison of the experimental results is shown in the following table:
[0260]
[0261] The model of the present invention is compared with the existing models, and it can be seen that the performance of the method of the present invention is the best among other methods. Among them, libFM is from the literature "Factorization machines with libfm", and DKN is from the literature "DKN: Deep knowledge-aware network for news recommendation".
[0262] Example 3:
[0263] Build an intelligent news recommendation system based on explicit and implicit interest features based on Example 2. The system includes:
[0264] The training dataset generation unit first obtains the browsing record information of users on online news websites and then performs preprocessing operations on it to obtain user browsing records and their news text content that meet the training requirements. The training dataset generation unit includes,
[0265] The original data acquisition unit is responsible for downloading the publicly available news website dataset on the network and using it as the original data for constructing the training dataset;
[0266] The original data preprocessing unit is responsible for preprocessing each news text in the news dataset, removing stop words and special characters in the news dataset; extracting the key information of each news text, such as the title; thereby constructing the training dataset;
[0267] The news recommendation model construction unit based on explicit and implicit interest features is used to load the training dataset, construct the news encoding module, construct the explicit interest encoding module, construct the TF-IDF algorithm module, construct the implicit interest encoding module, construct the graph neural network module, construct the implicit interest decoding module, and construct the click-through rate predictor module. The news recommendation model construction unit based on explicit and implicit interest features includes,
[0268] The training dataset loading unit is responsible for loading the training dataset;
[0269] The news encoding module construction unit is responsible for training the news feature vectors based on the Glove word vector model in the training dataset and defining all news feature vectors; first encoding the news title vector using a fully connected layer to obtain the hidden layer representation, and finally decoding the hidden layer representation using a fully connected layer to reconstruct the news feature vectors.
[0270] The explicit interest encoding module construction unit is responsible for constructing explicit interest feature vectors according to the user browsing records; among them, the news feature vectors of the user browsing records are obtained by the news encoding module construction unit, and the explicit interest feature vectors are obtained using the Fastformer method;
[0271] The TF-IDF algorithm module construction unit is responsible for extracting the news keywords in the user browsing records using the TF-IDF algorithm, and then mapping each keyword to the same vector space using the word embedding method to obtain the keyword vectors of the news content.
[0272] The implicit interest encoding module construction unit is responsible for extracting the main features of the keyword vectors using a multi-layer perceptron and generating a keyword vector matrix through an aggregation operation, then filtering the possible keywords through a learnable mapping matrix M, and then obtaining the possible keyword vector matrix by calculating the distribution probability of the possible keywords in the keyword vector matrix. This matrix contains the initial implicit interest feature vectors.
[0273] The graph neural network module construction unit is responsible for using the graph neural network to propagate and aggregate the initial implicit interest feature vectors, so as to obtain updated implicit interest feature vectors.
[0274] The implicit interest decoding module construction unit is responsible for using a multi-layer perceptron to decode the updated implicit interest feature vectors, so as to obtain the final implicit interest feature vectors.
[0275] The click-through rate predictor module construction unit first uses a gating network to select important feature information and aggregates the explicit interest feature vectors and the final implicit interest feature vectors to obtain a user feature vector, then fuses the user feature vector and the news feature vector of the candidate news based on the attention network of the candidate news to obtain the final user feature vector, and finally uses the final user feature vector and the news feature vector of the candidate news as inputs, generates the score of each candidate news, that is, the click-through rate, through dot product operation, sorts all candidate news in descending order according to the click-through rate, and recommends the top-K news to the user.
[0276] The model training unit is used to construct the loss function required in the model training process and complete the optimization training of the model; the model training unit includes,
[0277] The loss function construction unit is responsible for calculating the error between the predicted candidate news and the real target news;
[0278] The model optimization unit is responsible for training and adjusting the parameters in the model training to reduce the prediction error.
[0279] Embodiment 4:
[0280] Based on the storage medium of Embodiment 2, which stores multiple instructions that are loaded and executed by a processor to perform the steps of the intelligent news recommendation method based on explicit and implicit interest features in Embodiment 2.
[0281] Embodiment 5:
[0282] Based on the electronic device of Embodiment 4, the electronic device includes: the storage medium of Embodiment 4; and a processor for executing the instructions in the storage medium of Embodiment 4.
[0283] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent news recommendation method based on explicit and implicit interest features, characterized in that This method constructs and trains a news recommendation model composed of a news encoder, an explicit interest encoder, a term frequency-inverse document frequency (TF-IDF) algorithm module, an implicit interest encoder, a graph neural network, an implicit interest decoder, and a click-through rate predictor module. All candidate news are sorted in descending order of click-through rate, and the top-K news are recommended to users. Specifically as follows: Construct a news encoder that takes the title information of the news as input and learns the news feature vector from the above information. Construct a news recommendation model based on explicit and implicit interest features, taking the news feature vector generated by the news encoder as input, and using Fastformer to obtain the explicit interest feature vector; taking the user's browsing record as input, and using the TF-IDF algorithm, multi-layer perceptron, and graph neural network to obtain the implicit interest feature vector. Construct a click-through rate predictor module. First, use a gated network to select important feature information and aggregate the explicit interest feature vector and the final implicit interest feature vector to obtain the user feature vector. Then, based on the attention network of the candidate news, fuse the user feature vector and the news feature vector of the candidate news to obtain the final user feature vector. Finally, take the final user feature vector and the news feature vector of the candidate news as input, and generate the score of each candidate news, that is, the click-through rate, through dot product operation. All candidate news are sorted in descending order of click-through rate, and the top-K news are recommended to users. An intelligent news recommendation method based on explicit and implicit interest features, characterized in that the construction process of the news encoder is specifically as follows: First, construct a word mapping table for each word in the dataset, and map each word in the table to a unique numerical identifier. The mapping rule is: starting from the number 1, and then incrementally sorting in the order of each word being entered into the word mapping table, thus forming a word mapping conversion table; use the Glove pre-trained language model to obtain the word vector representation of each word; in the word embedding layer, convert each news title T = [w1, w2,..., w N into a vector representation, denoted as x = [x1, x2,..., x N , where N represents the length of a news title, x N represents the vector representation of each word, and w represents a word in the news title; Then, using the news title vector x as the input, elements in the input are randomly set to zero with a certain probability to obtain a noisy vector Then, the noisy vector is input into the fully connected layer to obtain the hidden layer representation h, and the formula is as follows: Among them, represents the vector with noise, x represents the news title vector, q(x) represents the random zeroing process, f(·) represents the sigmoid activation function, and U and u are parameters learned from the training process; Finally, taking the hidden layer representation h as input, reconstruct it through a fully connected layer to obtain the news feature vector r, and the formula is as follows: r = f(U'h + u'); where r is the news feature vector, f(·) represents the sigmoid activation function, and U' and u' are parameters learned from the training process.
2. The intelligent news recommendation method based on explicit and implicit interest features according to claim 1, wherein The construction process of the news recommendation model based on explicit and implicit interest features is specifically as follows: Construct an explicit interest encoder: In order to generate the explicit interest feature of the user, the explicit interest encoder uses the Fastformer method to process the user's browsing record and outputs the explicit interest feature vector. Specifically as follows: First, Fastformer converts the input news feature vector into three vector representations of query, key, and value through three linear layers with non-shared parameters. The formula is as follows: q i = W q r i ; k i = W k r i ; v i = W v r i ; Among them, W q , W k and W v are all learnable parameters, r i represents the i-th news feature vector, q i represents the query vector of the i-th news, k i represents the key vector of the i-th news, and v i represents the value vector of the i-th news; Then, use the additive attention mechanism to aggregate and compress the query vector. The formula is expressed as follows: q = Att(q1, q2,..., q N ); Among them, q i represents the query vector of the i-th news, q represents the query vector that aggregates context information, and Att represents the additive attention mechanism; After that, use the additive attention mechanism and bitwise multiplication operation to calculate the interaction information between the key vector and the query vector. The formula is as follows: k = Att(q⊙k1, q⊙k2,..., q⊙k i ,..., q⊙k N ); where k i represents the key vector of the i-th news, k represents the key vector that aggregates the context information, ⊙ represents element-wise multiplication, and Att represents the additive attention mechanism; Then, process the key vector and the value vector through dot product operation and linear layer to obtain the news feature vector of a single attention head. The formula is expressed as follows: Among them, W o is a learnable parameter, ⊙ represents element-wise multiplication, and v i represents the value vector of the i-th news, and is the feature vector of the i-th news output by a single attention head; Finally, based on the output of M attention heads and combined with the user's browsing record, establish the explicit interest feature vector. The formula is expressed as follows: u p = [d1;d2;...;d k ;...;d N ; where [;] represents the concatenation operation, is the k-th news feature vector output by the n-th attention head, M is the number of attention heads, N is the length of the user browsing record, d k is the k-th news feature vector obtained by aggregating and concatenating M attention heads, u p is the explicit interest feature vector; Build the TF-IDF algorithm module: First, input a user's browsing record C u ={v1,...,v i ,...,v t-1} into this module, where v represents each user browsing record; then, use the TF-IDF algorithm to extract keywords from the user browsing record; finally, map the keywords to a keyword vector matrix K through the word embedding layer, where this matrix contains the keyword vectors of this user browsing record; Build an implicit interest encoder, which aims to infer the implicit interests of users from their browsing records, as follows: Build a multi-layer perceptron: Take the keyword vector matrix K of the user browsing record as the input, and use the multi-layer perceptron to encode these vectors. The formula is as follows: C = MLP(W′K + b′); Where K is the keyword vector matrix, W′ represents the learnable parameters of the multi-layer perceptron, b′ is the bias, C represents the keyword vector output after being processed by the multi-layer perceptron, and MLP is the multi-layer perceptron; Build an interest inference module: To infer implicit interests from the keyword vectors of the historical news sequence, first extract the keywords of all news from the news recommendation dataset using the TF-IDF method, map them to a global keyword vector matrix H through the word embedding layer, then filter the possible keywords through a learnable mapping matrix M, and then obtain the possible keyword vector matrix, that is, the initial implicit interest feature vector, by calculating the distribution probability of the possible keywords in the global keyword vector matrix H. The specific process formula is as follows: W p = softmax(HMC); C p = W p H; where softmax represents the softmax normalization function, and W p represents the learnable weight matrix; C p represents the possible keyword vector matrix, which contains all the initial implicit interest feature vectors; Constructing a Graph Neural Network: Using the initial implicit interest feature vector C p as the input, an updated implicit interest feature vector is obtained through the graph neural network; specifically, the operation process of the l-th layer graph neural network is represented as follows: where, σ represents the activation function; H l is the node representation of the l-th layer graph neural network, and W l represents the learnable parameters of the l-th layer graph neural network, D is the degree matrix; A = A + I, where A is the adjacency matrix and I is the identity matrix; specifically, the input of the first layer is C p , then its output is H 0 = C p ; after passing through n layers of graph neural networks, the implicitly interested feature vector updated at time t can be expressed as C t = H n ; Construct an implicit interest decoder: to update the implicit interest feature vector C t As the input, use a multi-layer perceptron as the decoder to generate the final implicit interest feature vector. The formula is as follows: u o = MLP(WC t + b); Among them, C t is the updated implicit interest feature vector, W is the parameter that can be learned by the multi-layer perceptron, b is the bias, and u o is the final implicit interest feature vector, and MLP is the multi-layer perceptron.
3. The intelligent news recommendation method based on explicit and implicit interest features according to claim 1, characterized in that The construction process of the click-through rate predictor is as follows: Construct a gating network: It is designed to select important feature information and aggregate the explicit interest feature vector and the final implicit interest feature vector; use the explicit interest feature vector u generated by the explicit interest encoder p and the final implicit interest feature vector u generated by the implicit interest decoder o as inputs, and generate the user feature vector u through the gating network g ; The formula is as follows: g = ReLU(W g [u o ; u p + b g ); u g = g ⊙ tanh(Vu o + v)+(1 - g) ⊙ u p ; Among them, W g , W b , V and v represent learnable parameters, b g represents the bias, the symbol ; represents the connection operation, u p is the explicit interest feature vector, u o is the final implicit interest feature vector, ReLU and tanh are activation functions, u g is the user feature vector, and g is the gating network; Build an attention network based on candidate news, which is designed to integrate the features of candidate news into the user feature vector to generate the final user feature vector. The formula is as follows: α = Att(W Q n, W K u g ); Among them, W Q and W K are learnable parameters, n is the news feature vector of candidate news generated by the news encoder, u g is the user feature vector, L is the length of a user browsing record, u is the final user feature vector, Att represents the attention mechanism function, and α is the attention weight; Build a prediction module, which takes the news feature vector n of the candidate news generated by the news encoder and the final user feature vector u as the input, and predicts the click-through rate of the candidate news through dot product operation. The formula is as follows: Among them, represents the click-through rate of the candidate news; When the model of this method has not been fully trained, it needs to be trained on the training dataset to optimize the model parameters; when the model training is completed, the click-through rate predictor can predict the recommendation score of each candidate news, and recommend appropriate news to users according to the score.
4. The intelligent news recommendation method based on explicit and implicit interest features according to claim 1, wherein The construction process of the training dataset is as follows: Build a news dataset or select a publicly available news dataset; Preprocess the news dataset: Preprocess each news text in the news dataset, and remove the stop words and special characters in the news dataset; Extract the title, category, subcategory, and summary information of each news text respectively; Build training positive examples: Use the news numbers with label 1 in the historical news sequence and interaction behavior sequence in the user browsing record, that is, the numbers of the news clicked by the user, to build training positive examples; Build training negative examples: Use the news numbers with label 0 in the historical news sequence and interaction behavior sequence in the user browsing record, that is, the numbers of the news not clicked by the user, to build training negative examples; Build a training dataset: Combine all the positive example data and negative example data and shuffle their order to build the final training dataset.
5. The intelligent news recommendation method based on explicit and implicit interest features according to claim 1, characterized in that, After the news recommendation model is built, it is trained and optimized through the training dataset, as follows: Construct the loss function: Using the negative sampling technique, define the news that a user has clicked on as positive examples, and the news that the user has not clicked on as negative examples, and calculate the click prediction value p of the positive examples i ; The formula is as follows: Among them, is the click-through rate of the j-th negative example relative to the i-th positive example in the same click sequence, is the i-th positive example, and G is the number of negative examples; The loss function of news recommendation is the negative log-likelihood function of all positive examples. The formula is as follows: Among them, is the set of positive examples; Optimized training model: The Adam optimization function is selected as the optimization function for this model. Among them, the learning rate is set to 0.001, the smoothing constants are set to (0.9, 0.999), eps is set to 1e-8, and the L2 penalty value is set to 0.
6. An intelligent news recommendation system based on explicit and implicit interest features, which includes A training dataset construction unit that first obtains the browsing record information of users on online news websites and then performs preprocessing operations on it to obtain user browsing records and their news text content that meet the training requirements. The training dataset construction unit includes An original data acquisition unit responsible for downloading the publicly available news website dataset on the network and using it as the original data for constructing the training dataset. An original data preprocessing unit responsible for preprocessing each news text in the news dataset to remove stop words and special characters in the news dataset. Extract the key information of each news text, such as the title, respectively, so as to construct the training dataset. A news recommendation model construction unit based on explicit and implicit interest features is used to load the training dataset, construct a news encoding module, construct an explicit interest encoding module, construct a TF-IDF algorithm module, construct an implicit interest encoding module, construct a graph neural network module, construct an implicit interest decoding module, and construct a click-through rate predictor module. The news recommendation model construction unit based on explicit and implicit interest features includes A training dataset loading unit responsible for loading the training dataset. A news encoding module construction unit responsible for training news feature vectors based on the Glove word vector model in the training dataset and defining all news feature vectors. First, use a fully connected layer to encode the news title vector to obtain a hidden layer representation, and finally use a fully connected layer to decode the hidden layer representation to reconstruct the news feature vector. An explicit interest encoding module construction unit responsible for constructing explicit interest feature vectors according to user browsing records. Among them, the news feature vectors of user browsing records are obtained by the news encoding module construction unit, and the explicit interest feature vectors are obtained using the Fastformer method. A TF-IDF algorithm module construction unit responsible for using the TF-IDF algorithm to extract news keywords in user browsing records, and then using the word embedding method to map each keyword to the same vector space to obtain the keyword vectors of the news content. An implicit interest encoding module construction unit responsible for using a multi-layer perceptron to extract the main features of the keyword vectors and generate a keyword vector matrix through an aggregation operation, then filtering possible keywords through a learnable mapping matrix M, and then calculating the distribution probability of possible keywords in the keyword vector matrix to obtain a possible keyword vector matrix, which contains the initial implicit interest feature vectors. A graph neural network module construction unit responsible for using the graph neural network to propagate and aggregate the initial implicit interest feature vectors to obtain updated implicit interest feature vectors. An implicit interest decoding module construction unit responsible for using a multi-layer perceptron to decode the updated implicit interest feature vectors to obtain the final implicit interest feature vectors. The click-through rate predictor module construction unit first uses a gated network to select important feature information, aggregates the explicit interest feature vector and the final implicit interest feature vector to obtain a user feature vector, then fuses the user feature vector and the news feature vector of the candidate news based on the attention network of the candidate news to obtain the final user feature vector. Finally, the final user feature vector and the news feature vector of the candidate news are used as inputs, and the click-through rate of each candidate news, that is, the score, is generated through dot product operation. All candidate news are sorted from high to low according to the click-through rate, and the top-K news are recommended to the user; The model training unit is used to construct the loss function required in the model training process and complete the optimization training of the model; the model training unit includes, The loss function construction unit is responsible for calculating the error between the predicted candidate news and the real target news; The model optimization unit is responsible for training and adjusting the parameters in the model training to reduce the prediction error.
7. A storage medium in which multiple instructions are stored, characterized in that, The instructions are loaded by the processor and execute the steps of the news recommendation method based on explicit and implicit interest features described in claims 1-5.
8. An electronic device, characterized in that, The electronic device includes: The storage medium according to claim 7; and a processor for executing the instructions in the storage medium.
Citation Information
Patent Citations
Reader preference-based personalized digital book recommendation system and method, computer and storage medium
CN113590970A
Intelligent news recommendation method and system based on multi-interest characteristics of user
CN114896510A