A news recommendation method and system based on gated multi-head self-attention
By introducing gated multi-head self-attention mechanism and pre-trained model BERT in the news recommendation system, the problem of excessive noise information in user interest modeling in the prior art is solved, and the more accurate matching of candidate news and user interests is achieved, and the performance of the recommendation system is improved.
Patent Information
- Application Number
- CN202210867135.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-22
AI Technical Summary
The prior art only considers the relationship between users browsing news when modeling user interests, resulting in the learned user interests containing a large amount of information that is not related to candidate news, making it difficult to accurately match candidate news and specific user interests.
The news recommendation method based on gated multi-head self-attention is adopted, and the unrelated information in the historical browsing news is filtered through the gated mechanism, and combined with the pre-trained model BERT, it enhances the news text representation to achieve the accurate matching of candidate news and user-specific interests.
By reducing noise information, the matching accuracy between candidate news and user interests is improved, and the performance and user experience of news recommendations are enhanced.
Smart Images

Figure CN115098786B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of personalized news recommendation, and particularly relates to a news recommendation method and system based on gated multi-head self-attention. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] News recommendation, as an important branch in the research field of recommendation systems, aims to help users find news that matches their interest preferences as much as possible through news content and user information; different words and different news articles imply different information when representing news and users; the use of the attention mechanism can assign different weights to different words and news to capture the key semantic information of news and the important interest clues of users; for example, An et al. proposed a personalized word-level and news-level attention mechanism to focus on the impact of different words and news on users; Wu et al. proposed an attention-based multi-view learning mechanism to learn the representation of news from multiple components of news (title, category, text); Qi et al. further proposed a multi-head self-attention method to capture the long-distance correlations between words and words, and between news and news; in these methods, the use of the attention mechanism effectively improves the performance of news recommendation.
[0004] However, it may not be optimal to only consider the relationship between news browsed by users during the process of modeling user interests, because users' interests are extensive. For example, if candidate news is not considered during the process of learning user interests through attention, the learned user interests will have more information irrelevant to the candidate news, making it difficult to accurately match the candidate news with specific user interests. Summary of the Invention
[0005] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a news recommendation method and system based on gated multi-head self-attention, which uses a gated multi-head self-attention mechanism to adjust user interests so as to better and accurately match candidate news with specific user interests; moreover, the pre-trained model BERT that enriches language knowledge is applied to news recommendation to enhance news text representation and improve the accuracy of news recommendation.
[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:
[0007] The first aspect of the present invention provides a news recommendation method based on gated multi-head self-attention;
[0008] A news recommendation method based on gated multi-head self-attention includes:
[0009] Obtain historical clicked news and candidate news, and use the pre-trained model BERT to encode the news respectively to obtain historical clicked news features and candidate news features;
[0010] Based on multi-head self-attention, capture the correlation between historical clicked news features, and use candidate news features for feature filtering to obtain user features;
[0011] Combine historical clicked news features, candidate news features and user features to predict the probability that the user browses each candidate news, and recommend candidate news to the user based on the predicted probability.
[0012] Furthermore, the specific steps of the news encoding are as follows:
[0013] Use the pre-trained model BERT to extract the text representation of the news;
[0014] Use Bi-LSTM to capture the bidirectional semantic dependencies of the text representation;
[0015] Based on the bidirectional semantic dependencies, use the attention network to aggregate the output of Bi-LSTM to obtain news features with rich context semantic information.
[0016] Furthermore, the specific steps of capturing the correlation between historical clicked news features are as follows:
[0017] Calculate the query information, key information and value information of each feature in the historical clicked news features;
[0018] Perform multi-round scaled dot-product attention calculations on the query information, key information and value information, and save the calculation results of each round;
[0019] Concatenate all the calculation results and perform a linear transformation to obtain enhanced historical clicked news features.
[0020] Furthermore, the specific steps of using candidate news features for feature filtering are as follows:
[0021] Based on the enhanced historical clicked news features, query information and candidate news features, calculate news internal information and channel modulation gate information;
[0022] Perform a dot-product operation on the news internal information and channel modulation gate information to obtain reconstructed news features;
[0023] Perform attention-weighted aggregation on the reconstructed news features to generate user features.
[0024] Furthermore, the specific steps of predicting the probability that the user browses each candidate news and recommending candidate news to the user based on the predicted probability are as follows:
[0025] Construct a training library consisting of positive and negative samples to train the probability prediction model;
[0026] Input the candidate news features to be predicted into the trained model to obtain the click probability of the candidate news;
[0027] Sort the click probabilities of a group of candidate news and recommend the top several candidate news to the user.
[0028] Furthermore, for the probability prediction model, the input is candidate news features and user features, and the output is the click probability of the candidate news.
[0029] Furthermore, the click probability is the inner product of user features and news features.
[0030] The second aspect of the present invention provides a news recommendation system based on gated multi-head self-attention.
[0031] A news recommendation system based on gated multi-head self-attention includes a news encoding module, a user encoding module, and a probability prediction module:
[0032] The news encoding module is configured to: obtain historical clicked news and candidate news, and use the pre-trained model BERT to perform news encoding respectively to obtain historical clicked news features and candidate news features;
[0033] The user encoding module is configured to: based on multi-head self-attention, capture the correlation between historical clicked news features, and use candidate news features for feature filtering to obtain user features;
[0034] The probability prediction module is configured to: jointly consider historical clicked news features, candidate news features, and user features, predict the probability that the user browses each candidate news, and recommend candidate news to the user based on the predicted probability.
[0035] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in a news recommendation method based on gated multi-head self-attention as described in the first aspect of the present invention.
[0036] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a news recommendation method based on gated multi-head self-attention as described in the first aspect of the present invention.
[0037] The above one or more technical solutions have the following beneficial effects:
[0038] The present invention provides a news recommendation method based on gated multi-head self-attention. Through a gating mechanism, candidate news is used to filter out information irrelevant to the candidate news in the historical browsing news, so as to achieve an accurate match between the candidate news and the user's specific interests.
[0039] The present invention provides a news recommendation method based on gated multi-head self-attention, which uses a pre-trained model to mine the deep semantics of the text, further strengthens the semantic representation of the news, and greatly improves the expressiveness of the model.
[0040] The present invention provides a news recommendation method based on gated multi-head self-attention. By using the relevance between the user's historical clicked news and the candidate news, more relevant user interests are captured, further strengthening the user's interest representation, thereby providing an improved trade-off between accuracy and speed.
[0041] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0043] Figure 1 It is an example diagram of a user's news reading behavior.
[0044] Figure 2 It is a flowchart of the method for the first embodiment.
[0045] Figure 3 It is a system structure diagram of the second embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0047] Note that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0048] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0049] In fact, users' interests are diverse; when a candidate news matches a user, it usually only matches a small part of the user's interests. As Figure 1 shown, the important text fragments are shown in gray. It is inferred from the news browsed by the user that he is interested in news in the fields of animals, health, sports, and politics. The fourth candidate news is a news related to the political theme, which only matches the political news browsed by the user, and its relevance to the news of other themes (animals and sports) browsed by the user is relatively low; in the process of learning the user's interests, the use of attention will focus on establishing connections between news of related themes; however, self-attention learning without relying on candidate news will retain too much information irrelevant to the candidate news; and this is a kind of noise for the accurate matching of candidate news with specific user interests.
[0050] If the candidate news information is not considered in user modeling, it may be difficult to accurately match the candidate news. Therefore, the present invention provides a news recommendation method based on gated multi-head self-attention, and proposes a neural news recommendation framework PGRec with gated multi-head self-attention enhanced by a pre-trained language model. The gated self-attention mechanism uses the candidate news as a channel adjustment gate to filter the output results of the news feature representation browsed by the user in the self-attention learning process, so as to reduce the noise information in the process of accurately matching the candidate news with specific user interests. In order to more effectively understand the semantics of the news, the pre-trained BERT is introduced to enhance the representation of the news; the relevance between the clicked news and the candidate news is also used to weight it to capture more relevant user interests for matching; during the model training process, the negative sampling technique is applied to jointly predict the click score for K + 1 news, that is, the candidate news consists of a positive sample of the user and K negative samples randomly selected from the user, and then jointly predict the positive news and the prediction scores of the negative news.
[0051] Embodiment 1
[0052] This embodiment discloses a news recommendation method based on gated multi-head self-attention;
[0053] As Figure 2 shown, a news recommendation method based on gated multi-head self-attention mainly includes three steps: news encoding, user encoding for candidate participation, and click prediction.
[0054] Step 1: Obtain historical clicked news and candidate news, and use the pre-trained model BERT to perform news encoding respectively to obtain historical clicked news features and candidate news features;
[0055] Step 101: Obtain historical clicked news and candidate news
[0056] Obtain a set of historical clicked news browsed by a given user, denoted as D u =[D1, D2, …, D N , where N is the number of the user's historical clicked news; the goal is to calculate the probability that a given user clicks on a set of candidate news D C The probability D C =[D c1 , D c2 , …, D cM , where M is the number of candidate news, and then sort the probabilities of the candidate news to recommend the best news.
[0057] Step 102: News encoding, the specific steps are as follows:
[0058] (1) Use the pre-trained model BERT to extract the text representation of the news;
[0059] News encoding aims to learn the deep semantic representation of news. Previously, pre-trained word embeddings such as Word2vec and Glove were usually used to initialize the embedding layer of the model. However, most of these pre-trained word embeddings are context-independent, which will cause the semantic information captured by the model not to be sufficient.
[0060] Since the pre-trained model BERT (Devlin et al., 2019) has 12 layers of transform and a large number of parameters, which can enable the model to better model the complex context information in the text, the present invention adopts the pre-trained model BERT as the news encoder.
[0061] The input news features are represented by T tokens, denoted as [w1, w2, …, w T , and after the news text is input into the BRET model and passed through several layers of transform, the tokens representation of the obtained hidden layer is denoted as [e1, e2, …, eT
[0062] (2) Use Bi-LSTM to capture the bidirectional semantic dependencies of text representations;
[0063] To further strengthen the semantic relationship, the output of the hidden layer of the BRET model is sent into the Bi-LSTM to further extract the bidirectional semantic dependencies of text representations.
[0064] (3) Based on the bidirectional semantic dependencies, use the attention network to aggregate the output of the Bi-LSTM to obtain news features with rich context semantic information.
[0065] Connect the output of the Bi-LSTM to the attention network to aggregate the text representations with bidirectional semantic dependencies, and obtain news features h with rich context semantic information, where the user's historical click news features and candidate news features are denoted as [h1, h2, …, h N and h c .
[0066] Step 2: Based on multi-head self-attention, capture the correlation between historical click news features and use candidate news features for feature filtering to obtain user features;
[0067] User encoding aims to learn the representation of the user from the historical click news browsed by the user, including two steps:
[0068] Step 201: Based on the news-level multi-head self-attention network, capture the correlation between historical click news features to obtain a deeper representation of the user. The specific steps are as follows:
[0069] (1) Calculate the query information, key information, and value information of each feature in the historical click news features;
[0070] The historical click news feature matrix is H = [h1, h2, …, h N , where h i is the news feature vector of the i-th historical click of the user. First, convert each feature vector into query information, key information, and value information. The formulas are as follows:
[0071] H query = W query ·H (1)
[0072] H key = W key ·H (2)
[0073] H value = W value ·H (3)
[0074] Among them, Hquery , H key , H value ∈R N×dim represent the query information, key information, and value information of the historical click news features respectively. N represents the number of historical click news, dim represents the unified dimension of the historical click news features, and W query , W key , W value are trainable weight matrices.
[0075] (2) Perform multi-round scaled dot-product attention calculations on the query information, key information, and value information, and save the calculation results of each round;
[0076] The formula for the scaled dot-product attention is:
[0077]
[0078] d k represents the number of hidden units of the neural network.
[0079] H query , H key , H value After a linear transformation, perform a scaled dot-product attention calculation. The formula for the calculation result head i is:
[0080]
[0081] Here, h rounds of scaled dot-product attention calculations need to be performed. Each time, the parameters W for the linear transformation of H query , H key , H value are different.
[0082] (3) Concatenate all the calculation results and perform a linear transformation to obtain the enhanced historical click news features.
[0083] Concatenate the calculation results of h rounds, and then perform a linear transformation. The value obtained is used as the feature after multi-head attention reconstruction. The specific formula is as follows:
[0084] Z news = MultiHead(H query , H key , H value ) = Concat(head1, head2,..., head h )W o (6)
[0085] Among them, W o represents the trainable parameter.
[0086] Record the result as:
[0087] Z news = [Z1, Z2, …, Z N (7)
[0088] S202. Feature filtering is performed on the candidate news features based on the gating adjustment unit and the additional attention unit.
[0089] The gating adjustment unit can filter out information irrelevant to the candidate news from the historical click news features of the user according to different candidate news, so as to achieve the matching of the candidate news with the specific interests of the user.
[0090] (1) Calculate the internal news information and the channel adjustment gate information based on the enhanced historical click news features, query information, and candidate news features;
[0091] The gating adjustment unit includes two input information flows. One is the internal information v of the browsing news composed of the feature Z news reconstructed by multi-head attention and the query information H query . The other is the channel adjustment gate information g composed of the candidate news feature h c and the query information H query . To ensure the same number of features, the candidate news feature vector is first copied N times and linearly transformed, where N is the number of historical click news of the user. The calculation methods of the two information input flows v and g in the gating adjustment unit are as follows:
[0092] h c = W transform repeat(h c ) (8)
[0093]
[0094]
[0095] where is a trainable parameter, and g, v ∈ R N×dim .
[0096] (2) Perform a dot product operation on the internal news information and the channel adjustment gate information to obtain the reconstructed news features;
[0097] To filter out information irrelevant to the candidate news in the user's historical click news, a dot product operation is performed on the internal information v of the browsing news and the channel adjustment gate information g, so as to prevent the self-attention process from containing too much information irrelevant to the candidate news and hindering the matching of the candidate news with the specific information of the user. The gating filter formula is as follows:
[0098]
[0099] Among them, ⊙ represents the click operation of the matrix.
[0100] (3) Perform attention-weighted aggregation on the news features after reconstruction to generate user features;
[0101] Since the correlations between the historical click news features after filtering and reconstruction and the candidate news may be different, use the candidate news to perform attention-weighted aggregation on the news features after reconstruction to generate the user feature u. The formula is as follows:
[0102]
[0103]
[0104]
[0105] Among them, W and b are trainable weight matrices, and φ(·) is a linear network.
[0106] Step 3: Combine the historical click news features, candidate news features, and user features to predict the probability that the user browses each candidate news, and recommend candidate news to the user based on the predicted probability.
[0107] Predict the probability that the user clicks on the candidate news The calculation of the probability is the inner product of the user feature u and the candidate news feature. The formula is as follows:
[0108]
[0109] Predict the click probability and recommend candidate news to the user. The specific steps are as follows:
[0110] Step 301: Construct a training library composed of positive samples and negative samples to train the probability prediction model;
[0111] The probability prediction model is used to predict the click probability of the candidate news. The input is the candidate news feature and the user feature, and the output is the click probability of the candidate news.
[0112] Since the numbers of positive and negative news samples are highly unbalanced, therefore, in model training, apply the negative sampling technique to jointly predict the click probability of k + 1 news. The k + 1 news consists of a positive sample of one user and a negative sample randomly selected from another user. The positive sample is composed of the news clicked by the user in the exposure sequence, and the negative sample is composed of k news randomly selected from the same exposure sequence as the positive sample but not clicked by the user; jointly predict the positive news and k negative news The probability is used to transform the news click prediction problem into a pseudo k+1 classification task. Softmax is used to normalize these probabilities to calculate the click probability of positive samples. The loss function in the model training method is the negative log-likelihood of all positive samples, as follows:
[0113]
[0114] where represents the click probability of the i-th positive news, represents the click probability of the j-th negative news in the same time period as the i-th positive news, and s is the set of positive training samples.
[0115] Step 302: Input the candidate news features to be predicted into the trained model to obtain the click probability of the candidate news;
[0116] Step 303: Sort the click probabilities of a group of candidate news and recommend the top several candidate news to the user.
[0117] Embodiment 2
[0118] This embodiment discloses a news recommendation system based on gated multi-head self-attention;
[0119] As Figure 3 shown, a news recommendation system based on gated multi-head self-attention includes a news encoding module, a user encoding module, and a probability prediction module:
[0120] The news encoding module is configured to: obtain historical click news and candidate news, and use the pre-trained model BERT to perform news encoding respectively to obtain historical click news features and candidate news features;
[0121] The user encoding module is configured to: capture the correlation between historical click news features based on multi-head self-attention, and use candidate news features for feature filtering to obtain user features;
[0122] The probability prediction module is configured to: jointly use historical click news features, candidate news features, and user features to predict the probability that the user browses each candidate news, and recommend candidate news to the user based on the predicted probability.
[0123] Embodiment 3
[0124] The purpose of this embodiment is to provide a computer-readable storage medium.
[0125] The computer-readable storage medium stores a computer program, which when executed by a processor, implements the steps in a news recommendation method based on gated multi-head self-attention as described in Embodiment 1 of the present disclosure.
[0126] Example 4
[0127] The purpose of this embodiment is to provide an electronic device.
[0128] An electronic device includes a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a news recommendation method based on gated multi-head self-attention as described in Embodiment 1 of the present disclosure.
[0129] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0130] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0133] The above are only the preferred embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, various modifications and variations can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
[0134] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A news recommendation method based on gated multi-head self-attention, characterized in that, Including: Obtain historical clicked news and candidate news, and use the pre-trained model BERT to perform news encoding respectively to obtain historical clicked news features and candidate news features; Based on multi-head self-attention, capture the correlation between historical clicked news features, and use candidate news features for feature filtering to obtain user features; The specific steps for capturing the correlation between historical clicked news features are as follows: Calculate the query information, key information, and value information of each feature in the historical clicked news features; Perform multiple rounds of scaled dot-product attention calculations on the query information, key information, and value information, and save the calculation results of each round; Concatenate all the calculation results and perform a linear transformation to obtain enhanced historical clicked news features; The specific steps for using candidate news features for feature filtering are as follows: Based on the enhanced historical clicked news features, query information, and candidate news features, calculate news internal information and channel adjustment gate information; The query information is the query information of each feature in the historical clicked news features; The specific process of calculating news internal information and channel adjustment gate information is as follows: The gating adjustment unit includes the input of two information streams. One is the feature reconstructed through multi-head attention and the query information to form the internal information for browsing news . The other is the channel adjustment gate information composed of the candidate news feature and the query information ; To ensure the same number of features, first copy the candidate news feature vector times and perform a linear transformation . is the number of news items clicked by the user in the historical record. The two information input streams in the gating adjustment unit and are calculated as follows: Among them, is a trainable parameter, ; Perform a dot-product operation on the news internal information and channel adjustment gate information to obtain reconstructed news features; Perform attention-weighted aggregation on the reconstructed news features to generate user features; Combine historical clicked news features, candidate news features, and user features to predict the probability that the user browses each candidate news, and recommend candidate news to the user based on the predicted probability.
2. The news recommendation method based on gated multi-head self-attention according to claim 1, characterized in that The specific steps of the news encoding are as follows: Use the pre-trained model BERT to extract the text representation of the news; Use Bi-LSTM to capture the bidirectional semantic dependencies of the text representation; Based on the bidirectional semantic dependencies, use an attention network to aggregate the output of Bi-LSTM to obtain news features with rich context semantic information.
3. A news recommendation method based on gated multi-head self-attention as described in claim 1, characterized in that The specific steps for predicting the probability that the user browses each candidate news and recommending candidate news to the user based on the predicted probability are as follows: Construct a training library composed of positive samples and negative samples to train the probability prediction model; Input the candidate news features to be predicted into the trained model to obtain the click probability of the candidate news; Sort the click probabilities of a group of candidate news, and select the top several candidate news to recommend to the user.
4. The news recommendation method based on gated multi-head self-attention as described in claim 3, wherein For the probability prediction model, the input is candidate news features and user features, and the output is the click probability of the candidate news.
5. A news recommendation method based on gated multi-head self-attention as described in claim 3, characterized in that The click probability is the inner product of the user features and the news features.
6. A news recommendation system based on gated multi-head self-attention, characterized in that: Including a news encoding module, a user encoding module, and a probability prediction module: The news encoding module is configured to: Obtain historical clicked news and candidate news, and use the pre-trained model BERT to perform news encoding respectively to obtain historical clicked news features and candidate news features; The user encoding module is configured to: Based on multi-head self-attention, capture the correlation between historical clicked news features, and use candidate news features for feature filtering to obtain user features; The steps for capturing the correlation between historical click news features are as follows: calculate the query information, key information, and value information of each feature in the historical click news features; perform multiple rounds of scaled dot-product attention calculations on the query information, key information, and value information, and save the calculation results of each round; splice all the calculation results and perform a linear transformation to obtain the enhanced historical click news features; The specific steps for filtering features with candidate news features are as follows: based on the enhanced historical click news features, query information, and candidate news features, calculate the news internal information and channel adjustment gate information; the query information is the query information of each feature in the historical click news features; the specific process of calculating the news internal information and channel adjustment gate information is as follows: The gating adjustment unit includes the input of two information flows. One is the feature reconstructed through multi-head attention and the query information to form the internal information for browsing news . The other is the channel adjustment gate information composed of the candidate news feature and the query information ; To ensure the same number of features, first copy the candidate news feature vector times and perform a linear transformation . is the number of news items clicked by the user in the history. The calculation methods of the two information input flows in the gating adjustment unit and are as follows: Among them, are trainable parameters, ; Perform a dot-product operation on the news internal information and channel adjustment gate information to obtain the reconstructed news features; perform attention-weighted aggregation on the reconstructed news features to generate user features; A probability prediction module, configured to: jointly use the historical click news features, candidate news features, and user features to predict the probability of the user browsing each candidate news, and recommend candidate news to the user based on the predicted probability.
7. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a news recommendation method based on gated multi-head self-attention as described in any one of claims 1-5.
8. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in a news recommendation method based on gated multi-head self-attention as described in any one of claims 1-5.
Citation Information
Patent Citations
Coarse-grained emotion analysis method based on hierarchical BERT neural network
CN110147452A
Push data acquisition method and device based on attention model, equipment and medium
CN113868542A
Interest activation news recommendation method and system based on multistage matching
CN114201683A
News recommendation method and device based on user portrait, equipment and storage medium
CN114741608A