News information processing method and device, electronic equipment and storage medium
By combining external knowledge and multi-task learning methods in the pre-trained language model, the problem of poor accuracy in financial news recommendations is solved, and more accurate recommendation and classification effects are achieved.
Patent Information
- Application Number
- CN202511032899.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-25
AI Technical Summary
The existing technology fails to fully consider the attributes of news when recommending financial news, resulting in poor recommendation accuracy.
The pre-trained language model is adopted, combining external knowledge information and user click history, and through technologies such as naming entity recognition and knowledge graphs, a prompt template is constructed, and the loss function of recommended tasks and attribute classification tasks is comprehensively recommended and the loss function of attribute classification tasks is carried out to perform multi-task learning.
The accuracy and generalization ability of pre-trained language models in financial news recommendations has been improved, collaborative optimization between tasks has been achieved, and more accurate recommendation results and classification results have been provided.
Smart Images

Figure CN120523923A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of news recommendation, and in particular to a news information processing method, device, electronic device and storage medium. Background Art
[0002] Financial news, characterized by high time sensitivity and complex corporate relationships, can help predict the prices of assets like stocks, bonds, and crude oil. Financial news recommendations play a crucial role in helping investors and financial analysts quickly access key market information and investment opportunities. Based on a user's browsing history and market trends, financial news recommendations aim to filter and push the most relevant financial information and analytical reports.
[0003] Currently, existing technologies do not fully consider the attributes of news when recommending financial news, resulting in poor accuracy of news recommendations. Summary of the Invention
[0004] The present invention provides a news information processing method, device, electronic device and storage medium, which are used to solve the technical problem of poor accuracy of news recommendation in the prior art.
[0005] The present invention provides a news information processing method, comprising: Acquiring input information, wherein the input information includes candidate news information; Inputting the input news information into a pre-trained language model to obtain a recommendation result of the recommendation task and a classification result of the attribute classification task output by the pre-trained language model; The loss function of the pre-trained language model combines the loss function of the recommendation task and the loss function of the attribute classification task.
[0006] According to a news information processing method provided by the present invention, the input information also includes external knowledge information and historical news information clicked by the user; Inputting the input information into the pre-trained language model includes: establishing a first prompt template for the recommendation task based on the external knowledge information, the historical news information, and the candidate news information; The first prompt template is input into the pre-trained language model.
[0007] According to a news information processing method provided by the present invention, the external knowledge information includes a knowledge graph; The step of establishing the first prompt template for the recommendation task based on the external knowledge information, the historical news information, and the candidate news information includes: Extracting entities from the candidate news information using named entity recognition technology; Retrieving each triple related to the entity from the knowledge graph; Convert each of the triples into a text sequence; The first prompt template is established based on the text sequence of the triples, the historical news information and the candidate news information.
[0008] According to a news information processing method provided by the present invention, the external knowledge information includes at least one of a knowledge graph, a knowledge base, and online resources.
[0009] According to a news information processing method provided by the present invention, inputting the input information into a pre-trained language model includes: Establishing a second prompt template for the attribute classification task based on the candidate news information; The second prompt template is input into the pre-trained language model.
[0010] According to a news information processing method provided by the present invention, the step of obtaining the classification result includes: Predicting a label word for the mask position in the second prompt template; The classification result corresponding to the label word is determined according to a mapping relationship between the classification result and the label word.
[0011] According to a news information processing method provided by the present invention, the attribute classification task includes at least one of a sentiment analysis task, a topic classification task, and a popularity prediction task.
[0012] The present invention also provides a news information processing device, comprising: An acquisition module, configured to acquire input information, wherein the input information includes candidate news information; A generation module, configured to input the input information into a pre-trained language model to obtain a recommendation result of the recommendation task and a classification result of the attribute classification task output by the pre-trained language model; The loss function of the pre-trained language model combines the loss function of the recommendation task and the loss function of the attribute classification task.
[0013] The present invention also provides an electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-described news information processing methods is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described news information processing methods.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned news information processing methods.
[0016] The news information processing method, device, electronic device and storage medium provided by the present invention enable the loss function of the pre-trained language model to integrate the loss function of the recommendation task and the loss function of the attribute classification task, take into account the attributes of the candidate news information, and utilize the correlation information between different tasks to enhance the performance of the pre-trained language model on each task, improve the generalization ability of the pre-trained language model, realize collaborative optimization between tasks, and make the recommendation results and classification results obtained by the pre-trained language model more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 This is one of the flow charts of the news information processing method provided by the present invention.
[0019] Figure 2 Schematic diagram of the principle of the recommendation task provided by the present invention.
[0020] Figure 3 It is a schematic diagram of the principle of the attribute classification task provided by the present invention.
[0021] Figure 4 It is a schematic diagram of the principle of the news information processing method provided by the present invention.
[0022] Figure 5 It is a schematic diagram of the area under the curve of each model in the recommendation task under different proportions of training sets provided by the present invention; Figure 6 This is a schematic diagram of the accuracy of each model in the sentiment analysis task under different proportions of training sets provided by the present invention; Figure 7 This is a schematic diagram of the accuracy of each model in the topic classification task under different proportions of training sets provided by the present invention; Figure 8 This is a schematic diagram of the accuracy of each model in the popularity prediction classification task under different proportions of training sets provided by the present invention; Figure 9 It is a structural diagram of the news information processing device provided by the present invention.
[0023] Figure 10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0025] It should be noted that, in the description of the present invention, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, the phrase "comprises a..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus comprising the elements. Terms such as "upper" and "lower" indicate positions or relationships based on those shown in the accompanying drawings and are intended solely to facilitate the description of the present invention and simplify the description. They are not intended to indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation, and are therefore not to be construed as limitations on the present invention. Unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be broadly construed, for example, to mean fixed, removable, or integral; mechanical or electrical; direct or indirect through an intermediary; or internal communication between two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0026] The terms "first," "second," and so forth, used herein are used to distinguish similar objects, not to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, allowing embodiments of the present invention to be implemented in an order other than that illustrated or described herein. Furthermore, the terms "first," "second," and so forth generally distinguish objects of a single type, and do not limit the number of objects. For example, the first object may be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the connected objects.
[0027] Generally, news recommendations can also be made using collaborative filtering. Collaborative filtering is a technique that predicts user preferences based on historical user behavior data. It can be categorized as either user-based or item-based, and recommendations are made by analyzing the similarity between users or items. With the rise of deep learning models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), these models have been introduced to the news recommendation field to learn features from user-item interaction data, improving the accuracy of news recommendations. CNNs excel at processing image and text data, while RNNs are particularly well-suited for processing sequential data, such as time series of user behavior.
[0028] In recent years, pre-trained language models such as BERT and GPT have achieved remarkable results in the field of natural language processing (NLP). These models can be adapted to specific downstream tasks, including news recommendation, through fine-tuning to improve the relevance and accuracy of recommendations. The BERT model can learn deeper language features on specific tasks through fine-tuning, thereby improving the performance of the recommendation system. The prompt learning method converts the task of predicting whether a user will click on a candidate news item into a fill-in-the-blank mask prediction task by designing a series of prompt templates. This method leverages the capabilities of pre-trained models to adapt to different tasks by constructing specific input formats without requiring major adjustments to the model architecture. Prompt learning is particularly suitable for processing a small number of samples and can significantly improve the performance of the model on specific tasks.
[0029] The following combination Figures 1-6 The present invention describes a news information processing method, device, electronic device and storage medium.
[0030] Figure 1 This is a flow chart of the news information processing method provided by the present invention. The news information processing method includes: Step S1: obtaining input information, where the input information includes candidate news information.
[0031] In some implementations, the candidate news information may be news information in various fields such as financial news information, current affairs news information, and livelihood news information. The following description will be made using the example of the candidate news information being financial news information.
[0032] In some implementations, candidate news information may be obtained from various news websites.
[0033] Step S2: input the input information into the pre-trained language model to obtain the recommendation result of the recommendation task and the classification result of the attribute classification task output by the pre-trained language model; Among them, the loss function of the pre-trained language model combines the loss function of the comprehensive recommendation task and the loss function of the attribute classification task.
[0034] In some implementations, the pre-trained language model (PLM) can be a RoBERTa model, which can capture rich semantic features in candidate news information. In the RoBERTa model architecture, input data is further refined by MLM heads.
[0035] Among them, the recommendation task is to determine whether to recommend candidate news information. The recommendation results include recommendation and non-recommendation. The loss function of the recommendation task is used to calculate the loss between the predicted label of the news information sample and the true label of the news information sample when training the recommendation function of the pre-trained language model.
[0036] The attribute classification task involves classifying one or more attributes of candidate news information. This task can include at least one of a sentiment analysis task, a topic classification task, and a popularity prediction task. These tasks correspond to the sentiment attributes, topic attributes, and popularity attributes of the candidate news information, respectively. For example, these tasks can include sentiment analysis, topic classification, and popularity prediction.
[0037] Sentiment analysis is used to classify the emotional tendencies expressed by candidate news information. This is crucial for understanding investor sentiment and market trends, and can therefore enhance the sensitivity of recommendation results to user emotional reactions. Topic classification is used to categorize key topics within candidate news information to help understand the core content of the news. The accuracy of topic classification results directly affects the relevance of recommended content, ensuring that users receive news that closely matches their interests and needs. Popularity prediction is used to assess the popularity of candidate news information, which is crucial for predicting the dissemination potential and market influence of news.
[0038] In some embodiments, the classification results of the sentiment analysis task may include positive and negative, the classification results of the topic classification task may include market, finance, policy, corporate management, product operation, and investment, and the classification results of the popularity prediction task may include popular and unpopular.
[0039] Among them, the loss function of the sentiment analysis task is used to calculate the loss between the predicted label of the news information sample and the actual label of the news information sample when training the sentiment analysis function of the pre-trained language model; the loss function of the topic classification task is used to calculate the loss between the predicted label and the actual label when training the topic classification function of the pre-trained language model; the loss function of the popularity prediction task is used to calculate the loss between the predicted label and the actual label when training the popularity prediction function of the pre-trained language model.
[0040] In some embodiments, since the news recommendation task has similarities in knowledge structure with the popularity prediction task, the topic classification task, and the sentiment analysis task, and their data sets can complement each other, a multi-task learning strategy can be used to optimize them. Therefore, the loss function of the pre-trained language model can be the sum of the loss function of the recommendation task and the loss function of each attribute classification task. Exemplarily, if the attribute classification task includes the sentiment analysis task, the topic classification task, and the popularity prediction task, the loss function of the pre-trained language model is the sum of the loss function of the recommendation task, the loss function of the sentiment analysis task, the loss function of the topic classification task, and the loss function of the popularity prediction task.
[0041] In some implementations, the loss function for the recommendation task and the attribute classification task may be a cross-entropy loss function: ; Among them, L is the loss value, N is the number of news information samples, is the true label of the i-th news information sample, is the predicted label of the i-th news information sample.
[0042] It can be understood that since the loss function of the pre-trained language model of the present invention combines the loss function of the recommendation task and the loss function of the attribute classification task, it takes into account the attributes of the candidate news information, and utilizes the correlation information between different tasks to enhance the performance of the pre-trained language model on each task, improve the generalization ability of the pre-trained language model, and realize collaborative optimization between tasks, so that the recommendation results and classification results obtained by the pre-trained language model can be more accurate.
[0043] In the existing technology, pre-trained language models for recommending news mainly rely on general domain data and do not effectively integrate knowledge from specific domains into specific tasks. As a result, pre-trained language models cannot achieve the performance level of other fields when applied in specific fields.
[0044] In order to improve the recommendation performance of the pre-trained language model, for the recommendation task, in some embodiments, the input information may also include external knowledge information and historical news information clicked by the user; in step S2, inputting the input information into the pre-trained language model may include: Establishing a first prompt template for the recommendation task based on external knowledge information, historical news information and candidate news information; The first prompt template is fed into the pre-trained language model.
[0045] Furthermore, after the first prompt template is input into the pre-trained language model, the recommendation result output by the pre-trained language model can be obtained, such as Figure 2 Therefore, the training set of the pre-trained language model needs to include samples of the first prompt template and the real recommended labels of the corresponding news information samples.
[0046] Among them, external knowledge information can include information about companies, industries, economic indicators, etc.
[0047] In some embodiments, the first prompt template It can be expressed as Based on historical news information[ ], recommended candidate news information[ ] is a [MASK] selection [SEP]. Among them, [CLS] and [SEP] are special symbols, indicating the beginning and end of the first prompt template respectively; [MASK] is a hint tag that includes external knowledge information. The pre-trained language model needs to predict the label word at the position of the mask [MASK]. For example, the label word "good" represents a recommendation, and the label word "bad" represents a negative recommendation.
[0048] It's understandable that adding prompt tags containing external knowledge information to a user's reading history to form a longer sequence can enrich the context of the recommendation task. This not only considers the semantic richness of the news content, but also incorporates financial expertise, which can improve model performance. Furthermore, using a structured first prompt template approach, the financial news recommendation task can be transformed into a knowledge-enhanced cloze task. This allows the pre-trained language model to utilize diverse knowledge prompts and learn for the recommendation task. This helps the model better understand and analyze the semantics and context of financial news, and more accurately recommend financial news relevant to the user's interests, thereby improving the quality and personalization of recommendations.
[0049] In some embodiments, external knowledge information may include at least one of a knowledge graph, a knowledge base, and online resources. Exemplarily, external knowledge information may include a knowledge graph, a knowledge base, and online resources. Knowledge graphs, databases, and online resources store information in a variety of forms. Among them, structured knowledge graphs are important resources for knowledge enhancement, describing the relationships between entities in the market. Entities not covered in the knowledge graph can be supplemented by unstructured knowledge bases and online resources to further enrich the knowledge base for recommendation tasks. Knowledge bases may include information such as financial analysis reports, financial terminology dictionaries, financial reports, stock prices, and trading volumes.
[0050] When external knowledge information includes a variety of knowledge graphs, knowledge bases, and online resources, it can enrich the model's domain-specific knowledge, improve the model's efficiency in processing data and recommending news, and reduce the bias of recommendation results.
[0051] In the present invention, when the pre-trained language model predicts the probability distribution of [MASK] on the label word set, the probability of a single label word in the label word set is mapped to the probability of the original label set. The probability P of a certain label word v can be expressed as: ; Among them, g represents the function that converts the predicted probability of each label word in the [MASK] position in the label word set into the predicted probability of the original label, Represents the word set corresponding to a specific category label y in the original label set, Representatives do not recommend, For example, the confidence level of each news recommendation is determined by a comprehensive scoring mechanism, using the variable To measure this confidence, a ranked recommendation list is generated based on the confidence score of each news item.
[0052] Considering the structural differences among knowledge graphs, knowledge bases, and online resources (for example, knowledge graphs are composed of triples, knowledge bases are SQL databases, and online resources are online texts), it is difficult for the model to directly integrate these heterogeneous knowledge sources for effective training. This limits the model's effectiveness in processing and recommending actual financial data, and may lead to deviations in the analysis results.
[0053] In order to integrate the model into the knowledge graph, in some embodiments, the first prompt template for establishing a recommendation task based on external knowledge information, historical news information, and candidate news information may include: Extracting entities from candidate news information through named entity recognition technology; Retrieve the triples related to the entity from the knowledge graph; Convert each triple into a text sequence; A first prompt template is established based on the text sequence of triples, historical news information and candidate news information.
[0054] Financial news is a complex collection of multi-dimensional entities including companies, industries, people, concepts, and industrial chains. To accurately identify these entities, named entity recognition (NER) technology can be used. The knowledge graph describes the relationship between entities in the market and consists of a triple. Representing an entity and relationships The set between them, E is the entity set, and R is the relationship set.
[0055] In a large-scale financial knowledge graph, an entity (such as companies or people) will involve many types of relationships, using Represents the relationship set of entities in financial news. The triples related to these entities show extremely high complexity and diversity. In order to achieve effective prompt enhancement, it is necessary to filter out the triples that are closely related to the news content. Specifically, first determine a refined relationship set based on the relationships mentioned in the news , ; Then combine the entities in the news and refined relationship sets Retrieve all related triples in the knowledge graph , .
[0056] Finally, the triples are converted into text sequences, and entity-related knowledge can be represented in an easy-to-process text format. For example, the text sequence of the triple (Company A, Controller, Ma) can be "The controller of Company A is Ma".
[0057] Of course, for unstructured knowledge bases and online resources, text sequences can be directly integrated into the first prompt template. For example, for the entity "Company A," in a knowledge base, the text sequence might be represented as "Company A's third-quarter financial report for fiscal year 2024 showed an 8% year-on-year revenue growth, but it fell short of expectations, resulting in a decline in stock price and fluctuations in trading volume. The company implemented a stock repurchase plan." In online resources, the text sequence might be represented as "Company A is a globally renowned diversified technology company." Data integration can effectively unify knowledge from different sources and formats, extending the model's knowledge base to nearly all types of financial knowledge.
[0058] For attribute classification tasks, in some embodiments, in step S2, inputting the input information into the pre-trained language model may include: Establishing a second prompt template for attribute classification task based on candidate news information; Feed the second prompt template into the pre-trained language model.
[0059] Furthermore, the second prompt template is input into the pre-trained language model to obtain the classification result output by the pre-trained language model. Therefore, the training set of the pre-trained language model also needs to include samples of the second prompt template and the real classification labels of the corresponding news information samples.
[0060] Among them, the second prompt template is composed of candidate news information , mask [MASK] and prompt words, as shown in Table 1, 、 、 They represent the second prompt templates for sentiment analysis task, topic classification task, and popularity prediction task respectively.
[0061] Table 1
[0062] Using the second prompt template to assist the attribute classification task can enable the model to better understand and distinguish the characteristics and content of different types of news.
[0063] like Figure 3 As shown, in some embodiments, in step S2, the step of obtaining the classification result may include: Predict the label word for the mask position in the second hint template; According to the mapping relationship between the classification result and the label word, the classification result corresponding to the label word is determined.
[0064] For mapping relationships, we can first integrate external knowledge from knowledge graphs, knowledge bases, and online resources to form an initial label word list. We then expand this initial label word list by selecting key words from the training set. We then combine the data characteristics of the training set to remove meaningless and repetitive words and add synonyms. This ensures that each category of words in the label word list is informative and highly relevant to the task, resulting in the final label word list. This process not only enhances the richness of the label word list but also improves its adaptability to specific tasks, providing a precise classification foundation for financial news recommendations.
[0065] Specifically, in sentiment analysis, we extract sentiment-related labels from different dimensions and levels of detail. Positive labels such as "profit growth" and "stock price increase" reflect positive sentiment, while negative labels such as "increasing losses" and "stock price plummet" reflect pessimism. Topic classification tasks are similar to sentiment analysis. For example, for a financial topic, the label vocabulary might include the following: "loss," "declining financial indicators," "insufficient liquidity," "insolvency," "excessive debt-to-asset ratio," and "financial fraud." During model training, the label words in the label vocabulary are mapped to the classification results.
[0066] For popularity prediction tasks, expressing the popularity of news is a core issue. Since there are no explicit features in the news dataset that directly show the popularity of news, an approximate strategy based on click behavior can be adopted. Specifically, the number of times each news appears in the click history of all users is counted and used as an indicator of popularity. If the news is published for a long time but has few views, it indicates that its popularity is low; if the news receives a large number of clicks in a short period of time, it indicates that it has high popularity potential. Construct a financial news popularity The calculation formula is: ; in, represents a natural constant, Represents the time interval from the news release to the present, The click-through rate of a news item is calculated by dividing the number of clicks on the current news item by the total number of clicks. In addition, in order to evaluate its popularity, each news item needs to be labeled according to its content, such as "financial report", "financial policy", "termination of listing", etc. These labels help to identify and predict the popularity of news. News with higher calculated scores, such as news involving major financial decisions or market regulatory policies, will be predicted to have higher popularity; popularity News types with lower scores, such as employee training, recruitment, or outbound investment, will be predicted to be less popular.
[0067] The classification results (labels) corresponding to the label words of each classification task in the present invention are shown in Table 2.
[0068] Table 2
[0069] The label words in the set of label words to be predicted When mapping to the original label, the present invention uses the weighted average of the label scores as the prediction score, and the word with the highest score is the final classification result. Provide a learnable weight , predicted score for: ; in, , Y is the label category set.
[0070] After obtaining the predicted label words, the verbalizer component can be used to map the predicted label words to the actual classification results. This mapping can correspond or map the classification label words in the external knowledge information to the classification labels within the model, thereby enriching the model's ability to classify different types of news.
[0071] Figure 2 and Figure 3 Combining can get Figure 4 The principle diagram is shown.
[0072] In actual experiments, the present invention collected 17,799 entities and 26,798 relationships from the knowledge graph. The knowledge base includes financial analysis reports, a dictionary of financial terms, financial statements, stock prices, and trading volumes. Online resources are from Wikipedia. User click records on financial news were collected from March 4 to March 24, 2024. The dataset contains 373,562 clicks on 16,375 news items by 12,365 users. The test set includes 5,463 articles and 81,245 clicks, and the training set includes 10,912 news items and 292,317 clicks.
[0073] For recommendation tasks, AUC (area under the curve), MRR, NDCG@5, and NDCG@10 are used as evaluation metrics. For attribute classification tasks, Acc (accuracy) and macro F1 are used to measure performance.
[0074] In order to verify the effectiveness of the present invention, two neural network methods (CNN, LSTM), two methods introducing knowledge representation (DKN, MKR), one fine-tuning method (Fine-tuning) and three prompt learning methods (Prompt-tuning, AUTOPrompt-tuning, SOFT Prompt-tuning) were selected for comparison with the present invention.
[0075] Table 3 shows the detailed evaluation results of different models on the recommendation task.
[0076] Table 3
[0077] As can be seen, the accuracy of the neural network models (CNN and LSTM) is relatively low, outperforming the DKN. CNN extracts news features through convolutional layers and max pooling, but fails to fully utilize the knowledge graph and user history information. LSTM captures sequential dependencies by processing user click history and news vectors, but its single historical sequence input limits recommendation accuracy. Both models rely primarily on internal information about the news, without additional auxiliary knowledge, which limits their recommendation performance.
[0078] In addition, the performance of models that introduce external knowledge information is generally better than that of machine learning models. Specifically, DKN uses knowledge graphs to fuse the semantic representation of news, improves the ability to identify positive and negative samples, and performs better than CNN and LSTM. The knowledge graph provides additional context and entity information, enabling news recommendations to better capture the relationships between news and user interests. However, the performance of DKN is still inferior to MKR, indicating that the fusion of a single knowledge graph is limited and the potential of multi-task learning has not been fully utilized. MKR has achieved significant improvements in news recommendation tasks by combining multi-task learning with knowledge graphs. The success of MKR shows that multi-task learning can effectively improve the performance of knowledge-enhanced models.
[0079] Compared to neural networks and knowledge representation models, fine-tuned models demonstrate significant advantages in news recommendation tasks. By performing task-specific optimizations based on pre-training, fine-tuned models can better capture news content and user interests, significantly improving recommendation accuracy.
[0080] Among prompt learning methods, AUTO Prompt-tuning, while not requiring manual label selection, still failed to outperform manual prompt-tuning. This is likely because manual prompt-tuning can more precisely design prompts to match specific tasks. On the other hand, SOFT Prompt-tuning slightly outperformed manual prompt-tuning, primarily due to its prompt optimization through continuous vector representation.
[0081] Our method outperforms all baseline methods, with AUC and NDCG@10 values 1.06% and 0.85% higher than the runner-up, respectively. Experimental results show that the combination of external knowledge integration, cue learning, and multi-task learning significantly improves model performance.
[0082] Table 4 shows the performance comparison results of sentiment analysis task, topic classification task, and popularity prediction task.
[0083] Table 4
[0084] As can be seen, the conclusions from the sentiment analysis, topic classification, and popularity prediction tasks are consistent with those from the recommendation task. While these tasks focus on different areas, they all demonstrate that models that incorporate external knowledge achieve the best performance across all tasks. By integrating external knowledge with a multi-task learning framework, the performance of financial news recommendation and classification tasks has been significantly improved.
[0085] In order to demonstrate the ability of the present invention to process knowledge inputs of different structures, the present invention conducted a comprehensive test on different knowledge sources and their impact on various tasks. The results are detailed in Table 5.
[0086] Table 5
[0087] It can be seen that structured knowledge graph prompts are generally better than unstructured prompts. Although structured knowledge is more difficult to obtain, the information it provides is more accurate and relevant, reducing redundancy; while unstructured knowledge that is easy to obtain, although the amount of information is huge, may contain more interference and non-critical details. In addition, the present invention also examined the use of financial knowledge prompts that are not related to historical news, that is, randomly extracting knowledge items from an external knowledge base. The results show that the effect of this method is far inferior to the aforementioned two methods, and some results are even lower than other prompt learning baselines. These comparative results further verify the effectiveness of the knowledge prompts in the present invention, which can indeed significantly improve the performance of recommendations and news classification.
[0088] The present invention also conducted experiments on news recommendation and classification tasks under low resource conditions, randomly splitting the training set into different proportions (full set, 50%, 30%, 20%, 10%) while keeping the test set unchanged. Figure 5-Figure 8 As shown in Figure 2 (Ours represents our invention), experimental results demonstrate that in small sample size scenarios, the hint-based learning approach significantly outperforms both fine-tuning and traditional machine learning models. These gaps narrow as the sample size increases, demonstrating the superiority of hint-based learning in small sample size scenarios. Even so, our invention demonstrates excellent performance in low-resource environments. In particular, using only 30% of the training data, our model outperforms other baseline models using the full dataset. This is due to the knowledge-enhanced hints and multi-task learning mechanisms that enable the model to efficiently learn contextual semantics even in data-scarce environments.
[0089] The present invention also conducts an ablation study of the model to analyze the impact of four modules: external knowledge information, sentiment analysis task, topic classification task, and popularity prediction task on news recommendation. By gradually removing each module, their independent contribution to the overall recommendation effect is evaluated. The results are shown in Table 6, where w / o represents the removal of the module.
[0090] Table 6
[0091] It can be seen that removing any module leads to a decline in recommendation performance, a finding that emphasizes the importance of each module in improving recommendation quality. In particular, removing external knowledge information resulted in the largest decreases in AUC and NDCG@10, by 0.96% and 0.89%, respectively (the arrows in Table 6 represent decreases, and the numbers following the arrows represent the magnitude of the decrease). This result demonstrates that the introduction of external knowledge information has a significant positive impact on enriching news content and enhancing the accuracy of news recommendations. Furthermore, removing the topic classification and popularity prediction tasks reduced the model's accuracy in capturing news topics. This is because topics are the core of news content and have a direct impact on user interests. Popularity is an important indicator of news appeal, predicting and recommending news that is likely to be widely popular, thereby improving user satisfaction. Finally, removing the sentiment analysis task had a relatively small impact on recommendation performance, suggesting that sentiment analysis may not be the most critical factor for news recommendation.
[0092] like Figure 9 As shown, the news information processing device provided by the present invention includes: An acquisition module is used to acquire input information, including candidate news information; A generation module is used to input the input information into the pre-trained language model to obtain the recommendation results of the recommendation task and the classification results of the attribute classification task output by the pre-trained language model; Among them, the loss function of the pre-trained language model combines the loss function of the comprehensive recommendation task and the loss function of the attribute classification task.
[0093] In some implementations, the input information may also include external knowledge information and historical news information clicked by the user; Generate modules, which can be used specifically for: Establishing a first prompt template for the recommendation task based on external knowledge information, historical news information and candidate news information; The first prompt template is fed into the pre-trained language model.
[0094] In some embodiments, the external knowledge information may include a knowledge graph; Generate modules that can also be used to: Extracting entities from candidate news information through named entity recognition technology; Retrieve the triples related to the entity from the knowledge graph; Convert each triple into a text sequence; A first prompt template is established based on the text sequence of triples, historical news information and candidate news information.
[0095] In some implementations, the external knowledge information may include at least one of a knowledge graph, a knowledge base, and online resources.
[0096] In some embodiments, the generation module may be specifically configured to: Establishing a second prompt template for attribute classification task based on candidate news information; Feed the second prompt template into the pre-trained language model.
[0097] In some embodiments, the generation module may also be used to: Predict the label word for the mask position in the second hint template; According to the mapping relationship between the classification result and the label word, the classification result corresponding to the label word is determined.
[0098] In some implementations, the attribute classification task may include at least one of a sentiment analysis task, a topic classification task, and a popularity prediction task.
[0099] It should be noted that the news information processing device provided by the present invention can execute the news information processing method of any of the above embodiments during specific operation, which will not be described in detail in this embodiment.
[0100] Figure 10 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 10 As shown, the electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may invoke logic instructions in the memory to execute a news information processing method, which includes: obtaining candidate news information; inputting the candidate news information into a pre-trained language model to obtain a recommendation result for a recommendation task and a classification result for an attribute classification task output by the pre-trained language model; wherein the loss function of the pre-trained language model combines the loss function of the recommendation task and the loss function of the attribute classification task.
[0101] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0102] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the news information processing method provided by the above-mentioned embodiments, the method including: obtaining candidate news information; inputting the candidate news information into a pre-trained language model to obtain the recommendation result of the recommendation task and the classification result of the attribute classification task output by the pre-trained language model; wherein the loss function of the pre-trained language model integrates the loss function of the recommendation task and the loss function of the attribute classification task.
[0103] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the news information processing method provided in the above-mentioned embodiments, the method comprising: obtaining candidate news information; inputting the candidate news information into a pre-trained language model to obtain the recommendation result of the recommendation task and the classification result of the attribute classification task output by the pre-trained language model; wherein the loss function of the pre-trained language model integrates the loss function of the recommendation task and the loss function of the attribute classification task.
[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0105] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A news information processing method, characterized in that: include: Acquiring input information, wherein the input information includes candidate news information; Inputting the input information into a pre-trained language model to obtain a recommendation result of the recommendation task and a classification result of the attribute classification task output by the pre-trained language model; The loss function of the pre-trained language model combines the loss function of the recommendation task and the loss function of the attribute classification task.
2. The news information processing method according to claim 1, characterized in that: The input information also includes external knowledge information and historical news information clicked by the user; Inputting the input information into the pre-trained language model includes: establishing a first prompt template for the recommendation task based on the external knowledge information, the historical news information, and the candidate news information; The first prompt template is input into the pre-trained language model.
3. The news information processing method according to claim 2, characterized in that: The external knowledge information includes a knowledge graph; The step of establishing the first prompt template for the recommendation task based on the external knowledge information, the historical news information, and the candidate news information includes: Extracting entities from the candidate news information using named entity recognition technology; Retrieving each triple related to the entity from the knowledge graph; Convert each of the triples into a text sequence; The first prompt template is established based on the text sequence of the triples, the historical news information and the candidate news information.
4. The news information processing method according to claim 2, characterized in that: The external knowledge information includes at least one of a knowledge graph, a knowledge base, and online resources.
5. The news information processing method according to claim 1, characterized in that: Inputting the input information into the pre-trained language model includes: Establishing a second prompt template for the attribute classification task based on the candidate news information; The second prompt template is input into the pre-trained language model.
6. The news information processing method according to claim 5, characterized in that: The step of obtaining the classification result includes: Predicting a label word for the mask position in the second prompt template; The classification result corresponding to the label word is determined according to a mapping relationship between the classification result and the label word.
7. The news information processing method according to claim 1, characterized in that: The attribute classification task includes at least one of a sentiment analysis task, a topic classification task, and a popularity prediction task.
8. A news information processing device, characterized in that: include: An acquisition module, configured to acquire input information, wherein the input information includes candidate news information; A generation module, configured to input the input information into a pre-trained language model to obtain a recommendation result of the recommendation task and a classification result of the attribute classification task output by the pre-trained language model; The loss function of the pre-trained language model combines the loss function of the recommendation task and the loss function of the attribute classification task.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the news information processing method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the news information processing method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Topic recommendation method and device, electronic equipment and storage medium
CN115292460A
News report-oriented multi-scene AI auxiliary manuscript writing method
CN117726826A
Intelligent question and answer method, system and server based on knowledge graph and large language model
CN118446320A
Personalized news recommendation method based on soft prompt tuning
CN118885669A
Intelligent online personal assistant with multi-turn dialog based on visual search
US20180108066A1