News recommendation method based on large language model
By building a news tower and a user tower network and using a large language model (LLM) to compress the network, the limitations of traditional news recommendation systems in user interest understanding are solved, personalized and efficient news recommendations are achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510362705.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional news recommendation systems are difficult to quickly and accurately identify user interests in massive news data. The collaborative filtering and content recommendation methods have limitations, and it is difficult to fully understand user's real interests.
Build a news tower and user tower network, use the large language model (LLM) to compress the network, and realize personalized news recommendations through user interaction information learning.
It improves the accuracy and personalization of news recommendations, improves user satisfaction and loyalty, and enhances the intelligence and scalability of the recommendation system.
Smart Images

Figure CN120296250A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of news recommendation, and particularly to a news recommendation method based on a large language model. Background Art
[0002] With the rapid development of the Internet, the way of spreading news and information has undergone a huge change. The channels for users to obtain news are no longer limited to traditional media, but they can obtain a vast amount of news information through various news clients, social media platforms, etc. However, in the face of such a large amount of news data, it is often difficult for users to quickly find the news content they are interested in. Most traditional news recommendation systems are based on the historical behavior data of users and use methods such as collaborative filtering and content recommendation for recommendation, but these methods have certain limitations. For example, the collaborative filtering recommendation system relies on the similarity between users, and when the user data is sparse, the recommendation effect will be greatly reduced; the content recommendation system mainly focuses on the text features of news and it is difficult to fully understand the true interests of users. In recent years, large language models (LLMs) have made remarkable progress in the field of natural language processing, and their powerful language understanding and generation capabilities provide new ideas for news recommendation. However, directly applying large language models to news recommendation is not easy, and a series of problems such as model training, user interest modeling, and news feature extraction need to be solved. Summary of the Invention
[0003] Aiming at the problems existing in the above prior art, the present invention proposes a news recommendation method based on a large language model, which realizes transferable user interaction information learning and accurate personalized news recommendation by constructing news tower and user tower networks and leveraging the content understanding ability of the LLM compression network.
[0004] The technical solution of the present invention is as follows:
[0005] A news recommendation method based on a large language model, and the system for implementing this method includes a data collection module, a data preprocessing module, a model training module, a recommendation algorithm module, and a recommendation generation module. It is characterized in that news recommendation is realized by constructing news tower and user tower networks and leveraging the content understanding ability of the LLM compression network. The specific steps include:
[0006] Step 1: The data collection module obtains user data, the user's news browsing historical data, and news data;
[0007] Step 2: The data preprocessing module performs data cleaning, data standardization, and feature extraction on the user data, the user's news browsing historical data, and the news data obtained in Step 1 to generate structured data inputs, and divides them into a training set, a validation set, and a test set;
[0008] Step 3: The model training module passes the candidate news through the news tower network to obtain news vectors; the news tower network consists of a single large language model (LLM) compression network;
[0009] Step 4: The model training module passes the user browsing history through the user tower network to obtain user vectors; the user tower network consists of k LLM compression networks, a multi-head attention layer, and a user representation layer, where k represents the length of the user historical sequence, that is, the number of news items the user has browsed in the past;
[0010] Step 5: The model training module calculates the loss function through infoNCE using the news vectors and user vectors of the candidate news, performs gradient backpropagation, and updates the model parameters of the news tower and user tower networks;
[0011] Step 6: The recommendation algorithm module validates and evaluates the performance of the model using the validation set described in Step 2, adjusts and optimizes the model based on the validation results until the performance of the recommendation model obtained by the model training module reaches the optimal on the validation set, and finally evaluates the recommendation model using the test set;
[0012] Step 7: The recommendation generation module performs personalized recommendations. According to the data described in Step 1, through the data preprocessing module, combined with the model trained in Steps 2 - 6, when making recommendations, it maps users and news to the corresponding vector spaces, where each user and each news item corresponds to a vector, and generates a recommendation list for recommendation based on the similarity between the user vector and the news vector.
[0013] Furthermore, the specific steps of the data preprocessing in Step 2 are as follows:
[0014] Step 2.1: Data cleaning: Clean the obtained user interaction data, removing noise, errors, and irrelevant data, including abnormal data elimination, inconsistent data processing, null value processing, and illegal value processing;
[0015] Step 2.2: Data standardization: Standardize the numerical features to conform to the standard normal distribution;
[0016] Step 2.3: Feature extraction: Extract user behavior features and news content features;
[0017] Step 2.4: According to the user interaction data, obtain user - news sample pairs of the news items clicked by the user, and use them as positive samples in the training process for training;
[0018] Step 2.5: When constructing negative samples, randomly select samples from the user's exposed but unclicked data and within the batch as negative samples for training, ensuring that the ratio of positive to negative samples is 1:3.
[0019] Further, in step 3, the large language model Llama2 is used to extract the semantic features of news texts, perform in-depth semantic parsing on news titles and contents, and generate news vector representations, which specifically include the following steps:
[0020] Step 3.1: The text features are compressed through the Prompt Layer. The specific content of the Prompt is: [The title of this news is: xx, the content is: xx, and the compression note is one word.];
[0021] Step 3.2: The last non-padding hidden states of the Llama2 model are mapped through a pooling operation to obtain the news vector N t 。
[0022] Further, step 4 specifically includes the following steps:
[0023] Step 4.1: Design a user tower network. The user tower network consists of k LLM compression networks, a multi-head attention layer, and a user representation layer, where k is the length of the user's historical sequence; the LLM compression network is used to perform semantic parsing and compression on news texts and extract key features; the multi-head attention layer is used to analyze the user's historical browsing behavior from multiple perspectives and capture the associations between different news and changes in user interests; the user representation layer is used to further integrate the output of the multi-head attention layer to generate the final user vector;
[0024] Step 4.2: Obtain the text features of each news in the user's historical sequence, compress the text through the Prompt Layer of the LLM compression network. The specific content of the Prompt is: [The title of this news is: xx, the content is: xx, and the compression note is one word.], and map the last non-padding hidden states of Llama2 through a pooling operation to obtain the news vector N i , and the news vectors of each news in the user's historical sequence constitute the input vector E of the multi-head attention layer u =Stack[N1,N2,N3...,N k ;
[0025] Step 4.3: Pass the input vector E u through the multi-head attention layer to obtain the output U context , and the calculation formula is:
[0026] U context =MultiHead(E u ,E u ,E u )
[0027] Step 4.4: Pass Ucontext The user vector U is obtained through the user presentation layer, and the calculation formula is:
[0028] U = AttPool(U context )
[0029] where AttPool is a fully connected layer or a transformation layer.
[0030] Furthermore, step 5 specifically includes the following steps:
[0031] Step 5.1: Calculate the similarity between the user vector and the positive and negative sample pairs of the news vectors of the candidate news. The cosine similarity is used. For the user vector U = (u1, u2,..u i ..,u n ), and the news vector N of the candidate news t = (n1, n2,..n i ..,n n ), where n i represents the value of each dimension of the news vector of the candidate news. The calculation formula for the positive sample is:
[0032]
[0033] The calculation formula for the negative sample is:
[0034]
[0035] Step 5.2: Calculate the InfoNCE loss function based on the similarity between the user vector and the positive and negative samples of the news vectors of the candidate news to adjust the scale of the similarity. The calculation formula is:
[0036]
[0037] where M represents the number of negative samples, and r is the temperature parameter;
[0038] Step 5.3: When the model updates its parameters, the low-rank adaptation (LoRA) method is used to update the model parameters for the LLM compression network. When LoRA acts, the original weight matrix is set as W, and the updated matrix is ΔW. LoRA decomposes ΔW into two low-rank matrices A and B, that is
[0039] ΔW = BA
[0040] where A ∈ R r×d , B ∈ R d×r , where r is the rank of the low-rank matrix, and r << d, where d is the dimension of the original weight matrix. After the transformation, for the input x, the output of the model is expressed as
[0041] h = Wx + BAx
[0042] Among them, BAx is the part updated by LoRA.
[0043] The technical effects of the present invention are as follows:
[0044] A news recommendation method based on a large language model according to the present invention comprehensively utilizes the behavior data and basic information data of users on the news platform, and adopts advanced intelligent recommendation algorithms to improve the accuracy of news recommendations. Combining with the news browsing habits of users, it provides personalized news recommendation services, meets the personalized needs of users, and improves the satisfaction and loyalty of users to the platform. By introducing a large language model and a hybrid recommendation algorithm, it fully explores and utilizes multi-dimensional data resources, and improves the intelligence and refinement level of the recommendation system. Through flexible algorithm design and model training, the recommendation system can quickly adapt to the needs of different user groups and usage scenarios, and enhance the scalability and versatility of the system. Through intelligent data analysis and recommendation strategies, it provides accurate, timely and efficient news recommendation services, and improves the overall experience of users on the news platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of a news recommendation method based on a large language model according to the present invention;
[0046] Figure 2 is a structural diagram of a recommendation model of a news recommendation method based on a large language model according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0047] The present invention will be further clearly and completely described below with reference to the accompanying drawings through specific embodiments.
[0048] A news recommendation method based on a large language model according to the present invention has a flowchart as Figure 1 shown. The system for implementing this method includes a data collection module, a data preprocessing module, a model training module, a recommendation algorithm module, and a recommendation generation module; the data collection module is responsible for collecting user data, the news browsing history data of users, and news data from multiple channels; the data preprocessing module is responsible for cleaning, standardizing, feature extraction, and dimensionality reduction processing of the collected raw data to generate high-quality structured data input; the model training module is responsible for using a large language model (LLM) and a recommendation algorithm to train the preprocessed data to construct a recommendation model; the recommendation algorithm module is responsible for constructing a recommendation model according to the LLM compression network, multi-head attention mechanism, user tower network, and news tower network to perform personalized recommendations for users; the recommendation generation module is responsible for generating a personalized news recommendation list for users according to the output of the recommendation model and displaying it to users.
[0049] The present invention introduces a large language model into the recommendation algorithm in the field of news recommendation. Its structural diagram of the recommendation model is asFigure 2 as shown
[0050] A news recommendation method based on a large language model according to the present invention comprises the following specific steps:
[0051] Step 1: A data collection module obtains user data, the user's news browsing history data, and news data; wherein, the user data includes the user's basic information (such as age, gender, occupation, etc.), the user's news browsing history data (including user click, browse, favorite, etc. behavior data); the news data includes information such as the title, content, release date, category (such as current affairs, finance, sports, entertainment, etc.), author, and source of the news. The data acquisition module obtains the above data from the back-end system of the news platform through API interfaces, database queries, etc., and stores it in a distributed database for subsequent processing.
[0052] Step 2: A data preprocessing module performs data cleaning, data standardization, and feature extraction on the user data, the user's news browsing history data, and the news data obtained in Step 1 to generate structured data inputs, and divides them into a training set, a validation set, and a test set; specifically as follows:
[0053] Step 2.1: Data cleaning. Remove duplicate data: Detect and delete duplicate records through unique identifiers (such as user ID, news ID). Process missing values: For missing user information (such as age, gender, etc.), fill in default values or perform interpolation based on similar user information. Screen out abnormal data: Eliminate data records that are clearly illogical, such as abnormally high browsing durations or unreasonable user behavior sequences.
[0054] Step 2.2: Data standardization. Perform standardization processing on numerical features (such as user age, browsing duration, etc.) to make them conform to the standard normal distribution, where x is the original feature value, μ is the feature mean, and σ is the feature standard deviation, and the formula is:
[0055]
[0056] Step 2.3: Feature extraction. Extract user behavior features: such as the user's browsing frequency, click-through rate, favorite rate, comment activity, etc. Extract news content features: Perform word segmentation on the news title and content, and extract keywords, topic tags, etc.
[0057] Step 2.4: According to the user's interaction data, obtain user-news sample pairs of the news clicked by the user as positive samples during the training process for training.
[0058] Step 2.5: When constructing negative samples, randomly select samples from the user's exposed but unclicked data and within the batch as negative samples for training, ensuring that the ratio of positive and negative samples is approximately 1:3.
[0059] Step 3: The model training module passes the candidate news through the news tower network to obtain the news vector of the candidate news; the news tower network is the core component in the present invention for processing news text features and consists of a single LLM (Large Language Model) compression network. Among numerous large language models, the present invention selects Llama2 as the LLM base. Llama2 has powerful language understanding and generation capabilities and can effectively extract the semantic features of news text. By using Llama2, the news tower network can perform in-depth semantic parsing on news titles and contents and generate high-quality news vector representations;
[0060] Step 3.1: In the news tower network, the processing of text features is carried out through the Prompt Layer. The role of the PromptLayer is to convert news text into a format suitable for large language model processing and at the same time guide the model to extract key information. Specifically, the content of the Prompt is designed as: [The title of this news is: xx, the content is: xx, and the compressed note is a single word.]; this design method can guide the Llama2 model to focus on the core content of the news when processing news text and compress it into a concise keyword or phrase, thereby effectively extracting the key semantic information of the news.
[0061] Step 3.2: After completing the compression of text features, the next step is to convert the output of the Llama2 model into a news vector. Specifically, the last non-padding hidden states of the Llama2 model are used. These hidden states contain the semantic information of the news text after being processed by the model. In order to further extract key features and reduce the dimension, a pooling operation is adopted. The pooling operation generates a fixed-length vector by aggregating the information in the hidden states, and this vector can effectively represent the overall semantic features of the news. Finally, through a mapping operation, the pooled vector is converted into the news vector N t for subsequent recommendation calculations.
[0062] Step 4: The model training module passes the user browsing history through the user tower network to obtain the user vector;
[0063] Step 4.1: The user tower network is the core component in the present invention for processing the user's historical browsing behavior and consists of k LLM compression networks, a multi-head attention layer, and a user representation layer. Among them, k represents the length of the user's historical sequence, that is, the number of news items the user has browsed in the past. The design of the user tower network aims to comprehensively consider the user's historical browsing behavior, capture the changes and dynamic features of the user's interests through the multi-head attention mechanism, and finally generate a user vector that can comprehensively represent the user's interests.
[0064] (1) LLM Compression Network. Each LLM compression network is responsible for processing one news item in the user's historical sequence. The news text is semantically parsed and compressed by an LLM (such as Llama2) to extract key features. These networks can capture the core semantics of each news item and convert it into a vector representation.
[0065] (2) Multi-Head Attention Layer. The multi-head attention mechanism can analyze the user's historical browsing behavior from multiple perspectives simultaneously, capturing the correlations between different news items and changes in the user's interests. In this way, the model can better understand the user's dynamic interests.
[0066] (3) User Representation Layer. The user representation layer is responsible for further integrating the output of the multi-head attention layer to generate the final user vector. This vector can comprehensively represent the user's interest characteristics and is used for subsequent recommendation calculations.
[0067] Step 4.2: Extraction of User Historical News Text Features. In the user tower network, first, it is necessary to obtain the text features of each news item in the user's historical sequence. The specific steps are as follows:
[0068] (1) Text Feature Extraction. For each news item in the user's historical sequence, extract its title and content as text features.
[0069] (2) Text Compression. Compress the text features through the Prompt Layer of the LLM compression network. The specific method is the same as in Step 3.2; the content of the Prompt is designed as: [The title of this news is: xx, the content is: xx, and the compression note is one word.].
[0070] (3) Generation of News Vectors. After pooling the last non-padding hidden states of the Llama2 model, map them to obtain the news vector N for each news item i . These news vectors form the input vector E u =Stack[N1,N2,N3...,N k .
[0071] Step 4.3: Processing of the Multi-Head Attention Layer. Input the news vectors of each news item in the user's historical sequence into the multi-head attention layer, and capture the changes and dynamic features of the user's interests through the multi-head attention mechanism. The specific calculation formula is as follows:
[0072] U context =MultiHead(E u ,E u ,E u )
[0073] The multi-head attention layer processes the input vector simultaneously through multiple attention heads of MultiHead, and can capture the changes in user interests from multiple perspectives. Each attention head learns different features, and finally combines these features to generate a richer representation of user interests.
[0074] Step 4.4: User vector generation. The user representation layer further integrates the output of the multi-head attention layer to generate the final user vector. The specific calculation formula is as follows:
[0075] U = AttPool(U context )
[0076] where U context is the output of the multi-head attention layer, representing the comprehensive features of the user's historical browsing behavior. AttPool is a fully connected layer or transformation layer used to further integrate the output of the multi-head attention layer to generate the final user vector U. The user vector can comprehensively represent the user's interest characteristics and is used for subsequent recommendation calculations. In this way, the model can capture the user's dynamic interests, generate high-quality user representations, and thus achieve personalized news recommendations.
[0077] Step 5: The model training module calculates the loss function through infoNCE for the news vector and user vector of the candidate news, performs gradient backpropagation, and updates the model parameters of the news tower and user tower networks.
[0078] Step 5.1: Similarity calculation. In a news recommendation system, calculating the similarity between the user vector and the news vector of the candidate news is a core part of the recommendation algorithm. The present invention uses cosine similarity as the similarity metric standard to evaluate the matching degree between user interests and news content. For the user vector U = (u1, u2,..u i ..,u n ), the candidate news vector N t = (n1, n2,..n i ..,n n ), where n i represents the value of each dimension of the news vector of the candidate news, and the positive sample calculation formula is:
[0079]
[0080] The negative sample calculation formula is:
[0081]
[0082] Step 5.2: Calculation of the InfoNCE loss function. To train the recommendation model to accurately predict the user's interest in news, the present invention uses the InfoNCE loss function. The InfoNCE loss function is a loss function based on contrastive learning, aiming to maximize the similarity between the positive sample of the news vector of the candidate news and the user vector, while minimizing the similarity between the negative sample of the news vector of the candidate news and the user vector. The specific calculation formula is as follows:
[0083]
[0084] where M represents the number of negative samples, and r is the temperature parameter used to adjust the scale of the similarity. The role of the temperature parameter r is to control the distribution range of the similarity values. A smaller r will make the similarity values more concentrated, thereby enhancing the model's ability to distinguish between positive and negative samples.
[0085] Step 5.3: Update the model parameters using the Low-Rank Adaptation (LoRA) method. During the model training process, to efficiently update the model parameters, the present invention uses the Low-Rank Adaptation (LoRA) method. The LoRA method decomposes the weight matrix into two low-rank matrices, significantly reducing the number of model parameters while maintaining the model's performance. When LoRA takes effect, assuming the original weight matrix is W and the update matrix is ΔW, LoRA decomposes ΔW into two low-rank matrices A and B, that is
[0086] ΔW = BA
[0087] where A ∈ R r×d and B ∈ R d×r , r is the rank of the low-rank matrix, and r << d (d is the dimension of the original weight matrix). In this way, the LoRA method reduces the originally required d×d parameters to d×r + r×d parameters, significantly reducing the computational complexity and memory occupancy. After such a transformation, for the input x, the output of the model can be expressed as
[0088] h = Wx + BAx
[0089] where BAx is the part updated by LoRA. This update method not only reduces the number of parameters but also maintains the model's performance, enabling the model to more efficiently adapt to new data and tasks.
[0090] Step 6: The recommendation algorithm module uses the validation set described in Step 2 to verify and evaluate the performance of the model, adjusts and optimizes the model based on the verification results until the performance of the recommendation model on the validation set reaches a satisfactory level. Finally, the test set is used to evaluate the recommendation model, and the evaluation metrics include recall, AUC, mean reciprocal rank (MRR), and normalized discounted cumulative gain (NDCG) to evaluate the quality of the recommendation list.
[0091] Step 7: The recommendation generation module makes personalized recommendations. According to the data described in Step 1, through the data preprocessing module and in combination with the model trained in Steps 2 - 6, when making recommendations, users and news are mapped to the corresponding vector space, where each user and each piece of news corresponds to a vector. Based on the similarity between the user vector and the news vector, a recommendation list is generated for recommendation.
[0092] The present invention comprehensively utilizes the behavioral data and basic information data of users on the news platform, adopts advanced intelligent recommendation algorithms, combines the content understanding ability of the LLM, and combines the news browsing habits of users to provide personalized news recommendation services. It can better capture the interests and preferences of users, meet the personalized needs of users, and provide more accurate recommendation results. By introducing large language models and intelligent recommendation algorithms, it fully explores and utilizes multi-dimensional data resources to improve the intelligence and refinement level of the recommendation system. Through data cleaning, standardization, and feature extraction, the quality and consistency of the data are ensured, providing a reliable data basis for the recommendation algorithm.
[0093] Finally, it should be noted that the purpose of publishing the embodiments is to help further understand the present invention. However, those skilled in the art can understand that various substitutions and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the embodiments, and the scope of protection claimed by the present invention is defined by the scope of the claims.
Claims
1. A news recommendation method based on large language models. The system for implementing this method includes a data collection module, a data preprocessing module, a model training module, a recommendation algorithm module, and a recommendation generation module, characterized in that, By constructing a news tower and a user tower network and leveraging the content understanding ability of the LLM compression network, news recommendations are realized. The specific steps are as follows: Step 1: The data collection module obtains user data, the user's news browsing history data, and news data; Step 2: The data preprocessing module performs data cleaning, data standardization, and feature extraction on the user data, the user's news browsing history data, and the news data obtained in Step 1, generates structured data inputs, and divides them into a training set, a validation set, and a test set; Step 3: The model training module passes the candidate news through the news tower network to obtain the news vectors of the candidate news; the news tower network consists of a single large language model (LLM) compression network; Step 4: The model training module passes the user browsing history through the user tower network to obtain the user vectors; the user tower network consists of k LLM compression networks, a multi-head attention layer, and a user representation layer, where k represents the length of the user's historical sequence, that is, the number of news items the user has browsed in the past; Step 5: The model training module calculates the loss function through infoNCE for the news vectors and user vectors of the candidate news, performs gradient backpropagation, and updates the model parameters of the news tower and the user tower network; Step 6: The recommendation algorithm module validates and evaluates the performance of the model using the validation set described in Step 2, adjusts and optimizes the model based on the validation results until the performance of the recommendation model obtained by the model training module reaches the optimal on the validation set, and finally evaluates the recommendation model using the test set; Step 7: The recommendation generation module performs personalized recommendations. According to the data described in Step 1, through the data preprocessing module, combined with the model trained in Steps 2 - 6, when making recommendations, the users and news are mapped to the corresponding vector spaces, each user and each news item corresponds to a vector, and a recommendation list is generated for recommendation based on the similarity between the user vectors and the news vectors.
2. The news recommendation method based on a large language model according to claim 1, wherein The specific steps of the data preprocessing in Step 2 are as follows: Step 2.1: Data cleaning: Clean the obtained user interaction data, remove noise, errors, and irrelevant data, including eliminating abnormal data, processing inconsistent data, handling null values, and handling illegal values; Step 2.2: Data standardization: Standardize the numerical features so that they conform to the standard normal distribution; Step 2.3: Feature extraction: Extract user behavior features and news content features; Step 2.4: According to the user's interaction data, obtain user-news sample pairs where the user clicks on the news, and use them as positive samples in the training process for training; Step 2.5: When constructing negative samples, randomly select samples from the user's exposed but unclicked data and within the batch as negative samples for training, ensuring that the ratio of positive to negative samples is 1:
3.
3. The news recommendation method based on a large language model according to claim 1, characterized in that, In Step 3, the large language model Llama2 is used to extract the semantic features of the news text, perform in-depth semantic parsing on the news title and content, and generate news vector representations. The specific steps are as follows: Step 3.1: The text features are compressed through the Prompt Layer. The specific content of the Prompt is: [The title of this news is: xx, the content is: xx, and the compression note is one word.]; Step 3.2: Map the last non-padding hidden states of the Llama2 model to obtain the news vector N through a pooling operation t .
4. The news recommendation method based on a large language model according to claim 1, wherein The specific steps of step 4 are as follows: Step 4.1: Design a user tower network, which consists of k LLM compression networks, a multi-head attention layer, and a user representation layer, where k is the length of the user's historical sequence; the LLM compression network is used to perform semantic parsing and compression on the news text and extract key features; the multi-head attention layer is used to analyze the user's historical browsing behavior from multiple perspectives and capture the associations between different news and changes in user interests; the user representation layer is used to further integrate the output of the multi-head attention layer to generate the final user vector; Step 4.2: Obtain the text features of each news item in the user's historical sequence, and perform text compression through the PromptLayer of the LLM compression network. The specific content of the Prompt is: [The title of this news is: xx, the content is: xx, and the compression note is a single word.]. After pooling the last non-padding hidden states of Llama2, map them to obtain the news vector N i , and the news vectors of each news item in the user's historical sequence form the input vector E of the multi-head attention layer u = Stack[N1, N2, N3..., N k ; Step 4.3: Obtain the output U of the input vector E u through the multi-head attention layer context , and the calculation formula is: U context = MultiHead(E u , E u , E u ) Step 4.4: Obtain the user vector U context through the user presentation layer. The calculation formula is as follows: U = AttPool(U context ) Among them, AttPool is a fully connected layer or a transformation layer.
5. The news recommendation method based on a large language model according to claim 1, wherein The specific steps of step 5 are as follows: Step 5.1: Calculate the similarity between the positive and negative sample pairs of the user vector and the news vector of the candidate news. The cosine similarity is used. For the user vector U = (u1, u2,..u i ..,u n ), and the news vector N t = (n1, n2,..n i ..,n n ), where n i represents the value of each dimension of the news vector of the candidate news. The formula for the positive sample is: The formula for negative samples is: Step 5.2: Calculate the InfoNCE loss function based on the similarity between the user vector and the positive and negative samples of the news vectors of the candidate news to adjust the scale of the similarity. The calculation formula is: Among them, M represents the number of negative samples, and r is the temperature parameter; When the model updates its parameters in step 5.3, the low-rank adaptation (LoRA) method is used to update the model parameters of the LLM compression network. When LoRA takes effect, the original weight matrix is set as W, and the updated matrix is ΔW. LoRA decomposes ΔW into two low-rank matrices A and B, that is ΔW = BA where \(A\in R\) r×d and \(B\in R\) d×r where \(r\) is the rank of the low-rank matrix, and \(r\ll d\), where \(d\) is the dimension of the original weight matrix. After the transformation, for the input \(x\), the output of the model is expressed as h = Wx + BAx Among them, BAx is the part updated by LoRA.