News information processing method and device, electronic equipment and storage medium
By combining external knowledge and user history in the pre-trained language model and constructing prompt templates for multi-task learning, the accuracy problem of financial news recommendations is solved, and more accurate news recommendations and classification are achieved.
Patent Information
- Application Number
- CN202511032899.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing technologies fail to fully consider the attributes of news when recommending financial news, resulting in poor recommendation accuracy.
A pre-trained language model is used, combined with external knowledge information and user click history, entity relationships are extracted through named entity recognition and knowledge graph, prompt templates are constructed, and the loss functions of the integrated recommendation task and attribute classification task are used for multi-task learning.
It improves the accuracy and generalization ability of pre-trained language models in financial news recommendations, achieves collaborative optimization between tasks, and improves the accuracy and personalization of recommendation results.
Smart Images

Figure CN120523923B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of news recommendation, and in particular to a news information processing method and device, an electronic device and a storage medium. BACKGROUND
[0002] Financial news has the characteristics of high time sensitivity and complex company relationship, and can help predict the prices of assets such as stocks, bonds and crude oil. Financial news recommendation plays a crucial role in helping investors and financial analysts quickly obtain key market information and investment opportunities. On the basis of given user browsing records and market dynamics, financial news recommendation aims to filter and push the most relevant financial information and analysis reports.
[0003] At present, the existing technology does not comprehensively consider the attributes of news when recommending financial news, resulting in poor accuracy of news recommendation. SUMMARY
[0004] The present application provides a news information processing method, device, electronic device and storage medium to solve the technical problem of poor accuracy of news recommendation in the prior art.
[0005] The present application provides a news information processing method, comprising:
[0006] obtaining input information, the input information comprising candidate news information;
[0007] inputting the input news information into a pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model;
[0008] wherein the loss function of the pre-trained language model integrates the loss function of the recommendation task and the loss function of the attribute classification task.
[0009] According to the news information processing method provided by the present application, the input information further comprises external knowledge information and historical news information clicked by a user;
[0010] The inputting the input information into a pre-trained language model comprises:
[0011] establishing a first prompt template of the recommendation task based on the external knowledge information, the historical news information and the candidate news information;
[0012] inputting the first prompt template into the pre-trained language model.
[0013] According to the news information processing method provided by the present application, the external knowledge information comprises a knowledge graph;
[0014] The first prompt template of the recommendation task is established based on the external knowledge information, the historical news information and the candidate news information, and the first prompt template comprises:
[0015] Each entity in the candidate news information is extracted by a named entity recognition technology;
[0016] Each triple related to the entity is retrieved from the knowledge graph;
[0017] Each triple is converted into a text sequence;
[0018] The first prompt template is established based on the text sequence of the triple, the historical news information and the candidate news information.
[0019] According to the news information processing method provided by the application, the external knowledge information comprises at least one of a knowledge graph, a knowledge base and an online resource.
[0020] According to the news information processing method provided by the application, the input information is input into a pre-trained language model, and the pre-trained language model comprises:
[0021] The second prompt template of the attribute classification task is established based on the candidate news information;
[0022] The second prompt template is input into the pre-trained language model.
[0023] According to the news information processing method provided by the application, the step of obtaining the classification result comprises:
[0024] The label word of the mask position in the second prompt template is predicted;
[0025] According to the mapping relationship between the classification result and the label word, the classification result corresponding to the label word is determined.
[0026] According to the news information processing method provided by the application, the attribute classification task comprises at least one of a sentiment analysis task, a topic classification task and a popularity prediction task.
[0027] The application further provides a news information processing device, comprising:
[0028] An acquisition module is configured to acquire input information, wherein the input information comprises candidate news information;
[0029] A generation module is configured to input the input information into a pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model;
[0030] The loss function of the pre-trained language model synthesizes the loss function of the recommendation task and the loss function of the attribute classification task.
[0031] The application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the news information processing method according to any one of the above when executing the program.
[0032] The application further provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the news information processing method according to any one of the above.
[0033] The application further provides a computer program product, which includes a computer program, and the computer program is executable on a processor to implement the news information processing method according to any one of the above.
[0034] The news information processing method, device, electronic device and storage medium provided by the application make the loss function of the pre-trained language model synthesize the loss function of the recommendation task and the loss function of the attribute classification task, consider the attributes of the candidate news information, enhance the performance of the pre-trained language model on each task by using the associated information between different tasks, improve the generalization ability of the pre-trained language model, realize the collaborative optimization between tasks, and make the recommendation result and the classification result obtained by the pre-trained language model more accurate. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0036] Figure 1 is one of the flowcharts of the news information processing method provided by the application.
[0037] Figure 2 is the principle schematic diagram of the recommendation task provided by the application.
[0038] Figure 3 is the principle schematic diagram of the attribute classification task provided by the application.
[0039] Figure 4 is the principle schematic diagram of the news information processing method provided by the application.
[0040] Figure 5 is the area under the curve of each model on the recommendation task under different proportions of training sets provided by the application.
[0041] Figure 6 is an accuracy diagram of each model in the sentiment analysis task under different proportions of training sets provided by the application;
[0042] Figure 7 is an accuracy diagram of each model in the topic classification task under different proportions of training sets provided by the application;
[0043] Figure 8 is an accuracy diagram of each model in the popularity prediction classification task under different proportions of training sets provided by the application;
[0044] Figure 9 is a structural diagram of the news information processing device provided by the application.
[0045] Figure 10 is a structural diagram of the electronic device provided by the application. DETAILED DESCRIPTION
[0046] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0047] It should be noted that, in the description of the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, method, article or device comprising the element. The terms "upper", "lower" and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. Unless otherwise specified and limited, the terms "mounting", "connecting", "connecting" should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be connected inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0048] The terms "first", "second", and the like used in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a class, and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally means that the front and rear associated objects are in a "or" relationship.
[0049] Generally, the method of collaborative filtering can also be used for news recommendation. Collaborative filtering is a technique that predicts user preferences based on historical behavior data. It can be divided into user-based and item-based. It analyzes the similarity between users or the similarity between items to make recommendations. With the rise of deep learning models such as convolutional neural networks (CNN) and recurrent neural networks (RNN), these models have been introduced into the field of news recommendation to learn features from user and item interaction data, improving the accuracy of news recommendation. CNN performs well on image and text data, while RNN is particularly suitable for processing sequential data such as time-series user behavior.
[0050] In recent years, pre-trained language models such as BERT and GPT have made significant achievements in natural language processing (NLP). These models can be fine-tuned to adapt to specific downstream tasks, including news recommendation, to improve the relevance and accuracy of recommendations. The BERT model can learn deeper language features through fine-tuning on specific tasks, thereby improving the performance of the recommendation system. The prompt learning method converts the task of predicting whether a user will click on a candidate news into a fill-in-the-blank mask prediction task by designing a series of prompt templates. This method leverages the capabilities of pre-trained models by constructing specific input formats to adapt to different tasks without making significant adjustments to the model architecture. Prompt learning is particularly suitable for handling small sample sizes and can significantly improve model performance on specific tasks.
[0051] The following will be described in conjunction with Figures 1-6 The news information processing method, device, electronic equipment and storage medium provided by the present application are described.
[0052] Figure 1 The flowchart of the news information processing method provided by the present application is shown. The news information processing method comprises:
[0053] Step S1, obtaining input information, the input information comprising candidate news information.
[0054] In some embodiments, the candidate news information can be financial news information, political news information, livelihood news information, and news information in various fields. The following will be described with the candidate news information as financial news information as an example.
[0055] In some embodiments, the candidate news information can be obtained from various news websites.
[0056] Step S2, inputting the input information into the pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model;
[0057] The loss function of the pre-trained language model is a combination of the loss function of the recommendation task and the loss function of the attribute classification task.
[0058] In some embodiments, the pre-trained language model (PLM, Pre-trained Language Model) can be a RoBERTa model, which can capture rich semantic features in candidate news information. In the RoBERTa model architecture, the input data is further processed by the MLM heads.
[0059] The recommendation task is to determine whether to recommend the candidate news information, and the recommendation result includes recommendation and non-recommendation. The loss function of the recommendation task is used to calculate the loss between the predicted label of the news information sample and the true label of the news information sample when training the recommendation function of the pre-trained language model.
[0060] The attribute classification task is to classify one or more attributes of the candidate news information. The attribute classification task can include at least one of a sentiment analysis task, a topic classification task, and a popularity prediction task. The sentiment analysis task, the topic classification task, and the popularity prediction task correspond to the sentiment attribute, the topic attribute, and the popularity attribute of the candidate news information, respectively. For example, the attribute classification task can include a sentiment analysis task, a topic classification task, and a popularity prediction task.
[0061] The sentiment analysis task is used to classify the sentiment expressed by the candidate news information, which is crucial for understanding investor sentiment and market trends, and can enhance the sensitivity of the recommendation result to user emotional reactions. The topic classification task is used to classify the key themes in the candidate news information to help understand the core content of the news. The accuracy of the topic classification result directly affects the relevance of the recommended content, ensuring that users can receive news that is highly matched to their interests and needs. The popularity prediction task is used to evaluate the popularity of the candidate news information, which is of great significance for predicting the dissemination potential and market influence of news.
[0062] In some embodiments, the classification result of the sentiment analysis task can include positive and negative, the classification result of the topic classification task can include market, finance, policy, company management, product operation, and investment, and the classification result of the popularity prediction task can include popular and unpopular.
[0063] The loss function of the sentiment analysis task is used to calculate the loss between the predicted label of the news information sample and the real label of the news information sample when training the sentiment analysis function of the pre-trained language model, the loss function of the topic classification task is used to calculate the loss between the predicted label and the real label when training the topic classification function of the pre-trained language model, and the loss function of the popularity prediction task is used to calculate the loss between the predicted label and the real label when training the popularity prediction function of the pre-trained language model.
[0064] In some embodiments, since the news recommendation task has similarity in knowledge structure with the popularity prediction task, the topic classification task and the sentiment analysis task, and the data sets of them can complement each other, a multi-task learning strategy can be used to optimize them. Therefore, the loss function of the pre-trained language model can be the sum of the loss function of the recommendation task and the loss function of each attribute classification task. For example, if the attribute classification task includes the sentiment analysis task, the topic classification task and the popularity prediction task, the loss function of the pre-trained language model is the sum of the loss function of the recommendation task, the loss function of the sentiment analysis task, the loss function of the topic classification task and the loss function of the popularity prediction task.
[0065] In some embodiments, the loss function of the recommendation task and the attribute classification task can be a cross-entropy loss function:
[0066] ;
[0067] wherein L is the loss value, N is the number of news information samples, is the real label of the i-th news information sample, is the predicted label of the i-th news information sample.
[0068] It can be understood that since the loss function of the pre-trained language model of the present application integrates the loss function of the recommendation task and the loss function of the attribute classification task, the attributes of the candidate news information are considered, the correlation information between different tasks is utilized to enhance the performance of the pre-trained language model in each task, the generalization ability of the pre-trained language model is improved, the collaborative optimization between tasks is realized, and the recommendation result and the classification result of the pre-trained language model can be more accurate.
[0069] In the prior art, the pre-trained language model recommends news mainly depending on general field data, and does not effectively integrate specific field knowledge into specific tasks, resulting in that the pre-trained language model cannot achieve the performance level of other fields in the application of specific fields.
[0070] To improve the recommendation performance of the pre-trained language model, for the recommendation task, in some embodiments, the input information can further include external knowledge information and historical news information clicked by the user; in step S2, inputting the input information into the pre-trained language model can include:
[0071] establishing a first prompt template of the recommendation task based on the external knowledge information, the historical news information and the candidate news information;
[0072] inputting the first prompt template into the pre-trained language model.
[0073] Further, after inputting the first prompt template into the pre-trained language model, a recommendation result output by the pre-trained language model can be obtained, as shown in Figure 2 Therefore, the training set of the pre-trained language model needs to include samples of the first prompt template and real recommendation labels of corresponding news information samples.
[0074] The external knowledge information can include information about companies, industries, economic indicators, etc.
[0075] In some embodiments, the first prompt template can be represented as based on the historical news information[ ], the recommendation candidate news information[ ] is a [MASK] selection [SEP]. Wherein [CLS] and [SEP] are special symbols, respectively representing the beginning and end of the first prompt template; is a prompt label including external knowledge information, and [MASK] is a mask. The pre-trained language model needs to predict the label word at the mask [MASK] position, for example, the label word "good" represents recommendation, and the label word "not good" represents non-recommendation.
[0076] It can be understood that adding the prompt label including the external knowledge information to the user's reading history forms a longer sequence, which can enrich the context of the recommendation task, not only considering the semantic richness of the news content, but also integrating professional knowledge in the financial field, which can improve the performance of the model. In addition, through the structured first prompt template, the financial news recommendation task can be converted into a task of knowledge-enhanced cloze, so that the pre-trained language model can utilize diversified knowledge prompts to learn for the recommendation task, which helps the model to better understand and analyze the semantics and background of the financial news, more accurately recommend the financial news related to the user's interest, and thus improve the quality and individualization of the recommendation.
[0077] In some embodiments, the external knowledge information can include at least one of a knowledge graph, a knowledge base, and an online resource. For example, the external knowledge information can include a knowledge graph, a knowledge base, and an online resource. The knowledge graph, the database, and the online resource store information in various forms. Among them, the structured knowledge graph is an important resource for knowledge enhancement, which describes the relationship between entities in the market. For entities not covered in the knowledge graph, non-structured knowledge base and online resources can be used to supplement and further enrich the knowledge base of the recommendation task. The knowledge base can include financial analysis reports, financial term dictionaries, financial reports, stock prices, and trading volumes.
[0078] The external knowledge information includes multiple types of knowledge graphs, knowledge bases, and online resources, which can enrich the domain-specific knowledge of the model, improve the efficiency of the model in processing data and recommending news, and reduce the bias of the recommendation results.
[0079] In the present application, when the pre-trained language model predicts the probability distribution of [MASK] on the label word set, the probability of a single label word in the label word set is mapped to the probability of the original label set. The probability P of a certain label word v can be expressed as:
[0080] ;
[0081] where g represents a function that converts the predicted probability of each label word in the label word set at the [MASK] position to the predicted probability of the original label, represents a word set corresponding to a certain category label y in the original label set, represents not recommended, represents recommended. For example, the recommendation confidence of each news is determined by a comprehensive scoring mechanism, and the variable measures this confidence. According to the confidence score of each news, a ranked recommendation list is generated.
[0082] Considering that the knowledge graph, the knowledge base, and the online resource differ in structure (for example, the knowledge graph consists of triples, the knowledge base is a SQL database, and the online resource is an online text), the model cannot directly integrate these heterogeneous knowledge sources for effective training, which limits the efficiency of the model in processing and recommending actual financial data, and may cause bias in the analysis results.
[0083] In order to enable the model to integrate the knowledge graph, in some embodiments, the first prompt template of the recommendation task is established based on the external knowledge information, the historical news information, and the candidate news information, which can include:
[0084] extracting each entity in the candidate news information through named entity recognition technology;
[0085] Retrieving all triples related to the entity from the knowledge graph;
[0086] Converting all triples into text sequences;
[0087] Establishing a first prompt template based on the text sequence of triples, historical news information, and candidate news information.
[0088] Financial news is a complex collection of entities including companies, industries, people, concepts, and industry chains. To accurately identify these entities, named entity recognition (NER) technology can be used. A knowledge graph describes the relationships between entities in the market, consisting of a triple, representing the set of entities and relationships , E is the set of entities, and R is the set of relationships.
[0089] In a large-scale financial knowledge graph, an entity (such as a company or a person) may involve multiple types of relationships, denoted by , the set of relationships of entities in financial news, and the triples related to these entities exhibit high complexity and diversity. To achieve effective prompt enhancement, we need to filter out triples closely related to news content . Specifically, first, according to the relationships mentioned in the news, determine a refined relationship set , ; then combine the entities in the news and the refined relationship set to retrieve all triples related to them in the knowledge graph , .
[0090] Finally, convert the triples into text sequences. Entity-related knowledge can be represented in an easy-to-process text format, such as the text sequence of the triple (Company A, Controlling Person, Ma). For example, the text sequence of the triple (Company A, Controlling Person, Ma) can be "Company A's controlling person is Ma."
[0091] Of course, for unstructured knowledge bases and online resources, text sequences can be directly integrated into the first prompt template. For example, for the entity "Company A", in the knowledge base, the text sequence may be "Company A's third-quarter 2024 financial report shows an 8% increase in revenue compared to the previous year, but it did not meet expectations, resulting in a decline in stock price, and transaction volume also fluctuated, and the company implemented a stock buyback plan." In online resources, the text sequence may be "Company A is a globally renowned diversified technology company." Through data integration, we can effectively unify knowledge from different sources and formats, expanding the model's knowledge to almost all types of financial knowledge.
[0092] For the attribute classification task, in some embodiments, in step S2, inputting the input information into the pre-trained language model can include:
[0093] establishing a second prompt template for the attribute classification task based on the candidate news information;
[0094] inputting the second prompt template into the pre-trained language model.
[0095] Further, inputting the second prompt template into the pre-trained language model can obtain a classification result output by the pre-trained language model. Therefore, the training set of the pre-trained language model also needs to include samples of the second prompt template and the true classification labels of the corresponding news information samples.
[0096] wherein the second prompt template is composed of the candidate news information , a mask [MASK], and a prompt word, as shown in Table 1, , , represent the second prompt templates of the sentiment analysis task, the topic classification task, and the popularity prediction task, respectively.
[0097] Table 1
[0098]
[0099] Using the second prompt template to assist the attribute classification task can make the model better understand and distinguish the features and content of different types of news.
[0100] As shown in Figure 3 , in some embodiments, in step S2, the step of obtaining the classification result can include:
[0101] predicting a label word at the mask position in the second prompt template;
[0102] determining the classification result corresponding to the label word according to the mapping relationship between the classification result and the label word.
[0103] For the mapping relationship, the external knowledge information of the knowledge graph, knowledge base, and online resources can be integrated first to form an initial label word table, and then the initial label word table can be expanded by selecting key words from the training set, removing some meaningless and repetitive words, and adding synonyms according to the data features of the training set, to ensure that each category of words in the label word table is information-rich and highly relevant to the task, and to obtain the final label word table. This process not only enhances the richness of the label word table, but also improves its adaptability to specific tasks, providing accurate classification basis for financial news recommendation.
[0104] Specifically, in sentiment analysis, we extract sentiment-related labels from different dimensions and levels of detail. Positive labels such as "profit growth" and "stock price increase" reflect positive sentiment, while negative labels such as "increasing losses" and "stock price plummet" reflect pessimism. Topic classification tasks are similar to sentiment analysis. For example, for a financial topic, the label vocabulary might include the following: "loss," "declining financial indicators," "insufficient liquidity," "insolvency," "excessive debt-to-asset ratio," and "financial fraud." During model training, the label words in the label vocabulary are mapped to the classification results.
[0105] For popularity prediction tasks, expressing the popularity of news is a core issue. Since there are no explicit features in the news dataset that directly show the popularity of news, an approximate strategy based on click behavior can be adopted. Specifically, the number of times each news appears in the click history of all users is counted and used as an indicator of popularity. If the news is published for a long time but has few views, it indicates that its popularity is low; if the news receives a large number of clicks in a short period of time, it indicates that it has high popularity potential. Construct a financial news popularity The calculation formula is:
[0106] ;
[0107] in, represents a natural constant, Represents the time interval from the news release to the present, The click-through rate of a news item is calculated by dividing the number of clicks on the current news item by the total number of clicks. In addition, in order to evaluate the popularity, each news item needs to be labeled according to its content, such as "financial report", "financial policy", "termination of listing", etc. These labels help to identify and predict the popularity of news. Combined with the popularity calculation method, popularity News with higher calculated scores, such as news involving major financial decisions or market regulatory policies, will be predicted to have higher popularity; popularity News types with lower scores, such as employee training, recruitment, or outbound investment, will be predicted to be less popular.
[0108] The classification results (labels) corresponding to the label words of each classification task in the present invention are shown in Table 2.
[0109] Table 2
[0110]
[0111] The label words in the set of label words to be predicted When mapping to original labels, the application adopts a weighted average of label scores as a prediction score, and the word with the highest score is the final classification result. In addition, for the label word set provides a learnable weight , the prediction score is:
[0112] ;
[0113] wherein, Y is a set of label categories.
[0114] After obtaining the predicted label words, the predicted label words can be mapped to the actual classification result through the verbalizer component. Through mapping, the classification label words in the external knowledge information can be corresponded or mapped to the classification labels inside the model, thereby enriching the classification ability of the model for different types of news.
[0115] Figure 2 and Figure 3 can be obtained in combination as shown in the principle schematic diagram. Figure 4
[0116] In actual experiments, the application collects 17,799 entities and 26,798 relationships in the knowledge graph, and the knowledge base includes financial analysis reports, financial term dictionaries, financial statements, stock prices and trading volumes, and online resources are resources on Wikipedia. Collect user click records on financial news, data covering March 4, 2024 to March 24, 2024, the data set includes 12,365 user click operations on 16,375 news 373,562 times. The test set includes 5,463 articles and 81,245 clicks, and the training set includes 10,912 news and 292,317 clicks.
[0117] For the recommendation task, AUC (area under the curve), MRR, NDCG@5 and NDCG@10 are used as evaluation indexes. For the attribute classification task, Acc (accuracy) and macro F1 are used to measure performance.
[0118] In order to verify the effectiveness of the application, 2 neural network methods (CNN, LSTM), 2 knowledge representation methods (DKN, MKR), 1 fine-tuning method (Fine-tuning) and 3 prompt learning methods (Prompt-tuning, AUTOPrompt-tuning, SOFT Prompt-tuning) are selected for comparison with the application.
[0119] Table 3 shows the detailed evaluation results of different models on the recommendation task.
[0120] Table 3
[0121]
[0122] As can be seen, the accuracy of the neural network models (CNN and LSTM) is relatively low, outperforming the DKN. CNN extracts news features through convolutional layers and max pooling, but fails to fully utilize the knowledge graph and user history information. LSTM captures sequential dependencies by processing user click history and news vectors, but its single historical sequence input limits recommendation accuracy. Both models rely primarily on internal information about the news, without additional auxiliary knowledge, which limits their recommendation performance.
[0123] In addition, the performance of models that introduce external knowledge information is generally better than that of machine learning models. Specifically, DKN uses knowledge graphs to fuse the semantic representation of news, improves the ability to identify positive and negative samples, and performs better than CNN and LSTM. The knowledge graph provides additional context and entity information, enabling news recommendations to better capture the relationships between news and user interests. However, the performance of DKN is still inferior to MKR, indicating that the fusion of a single knowledge graph is limited and the potential of multi-task learning has not been fully utilized. MKR has achieved significant improvements in news recommendation tasks by combining multi-task learning with knowledge graphs. The success of MKR shows that multi-task learning can effectively improve the performance of knowledge-enhanced models.
[0124] Compared to neural networks and knowledge representation models, fine-tuned models demonstrate significant advantages in news recommendation tasks. By performing task-specific optimizations based on pre-training, fine-tuned models can better capture news content and user interests, significantly improving recommendation accuracy.
[0125] Among prompt learning methods, AUTO Prompt-tuning, while not requiring manual label selection, still failed to outperform manual prompt-tuning. This is likely because manual prompt-tuning can more precisely design prompts to match specific tasks. On the other hand, SOFT Prompt-tuning slightly outperformed manual prompt-tuning, primarily due to its prompt optimization through continuous vector representation.
[0126] Our method outperforms all baseline methods, with AUC and NDCG@10 values 1.06% and 0.85% higher than the runner-up, respectively. Experimental results show that the combination of external knowledge integration, cue learning, and multi-task learning significantly improves model performance.
[0127] Table 4 shows the performance comparison results of sentiment analysis task, topic classification task, and popularity prediction task.
[0128] Table 4
[0129]
[0130] As can be seen, the conclusions from the sentiment analysis, topic classification, and popularity prediction tasks are consistent with those from the recommendation task. While these tasks focus on different areas, they all demonstrate that models that incorporate external knowledge achieve the best performance across all tasks. By integrating external knowledge with a multi-task learning framework, the performance of financial news recommendation and classification tasks has been significantly improved.
[0131] In order to demonstrate the ability of the present invention to process knowledge inputs of different structures, the present invention conducted a comprehensive test on different knowledge sources and their impact on various tasks. The results are detailed in Table 5.
[0132] Table 5
[0133]
[0134] It can be seen that structured knowledge graph prompts are generally better than unstructured prompts. Although structured knowledge is more difficult to obtain, the information it provides is more accurate and relevant, reducing redundancy; while unstructured knowledge that is easy to obtain, although the amount of information is huge, may contain more interference and non-critical details. In addition, the present invention also examined the use of financial knowledge prompts that are not related to historical news, that is, randomly extracting knowledge items from an external knowledge base. The results show that the effect of this method is far inferior to the aforementioned two methods, and some results are even lower than other prompt learning baselines. These comparative results further verify the effectiveness of the knowledge prompts in the present invention, which can indeed significantly improve the performance of recommendations and news classification.
[0135] The present invention also conducted experiments on news recommendation and classification tasks under low resource conditions, randomly splitting the training set into different proportions (full set, 50%, 30%, 20%, 10%) while keeping the test set unchanged. Figures 5-8 As shown in Figure 2 (Ours represents our invention), experimental results demonstrate that in small sample size scenarios, the hint-based learning approach significantly outperforms both fine-tuning and traditional machine learning models. These gaps narrow as the sample size increases, demonstrating the superiority of hint-based learning in small sample size scenarios. Even so, our invention demonstrates excellent performance in low-resource environments. In particular, using only 30% of the training data, our model outperforms other baseline models using the full dataset. This is due to the knowledge-enhanced hints and multi-task learning mechanisms that enable the model to efficiently learn contextual semantics even in data-scarce environments.
[0136] The ablation study of the model is also performed, and the influence of the four modules of external knowledge information, sentiment analysis task, topic classification task and popularity prediction task on news recommendation is analyzed. The independent contribution of each module to the overall recommendation effect is evaluated by removing each module step by step, and the results are shown in Table 6, and w / o represents removing the module.
[0137] Table 6
[0138]
[0139] It can be seen that the removal of any one module will cause the decline of the recommendation performance, and this finding emphasizes the importance of each module in improving the recommendation quality. In particular, after removing the external knowledge information, the AUC and NDCG@10 decrease by 0.96% and 0.89% respectively (the arrow in Table 6 represents the decrease, and the number after the arrow represents the decrease amplitude). This result shows that the introduction of external knowledge information has a significant positive impact on enriching news content and enhancing the accuracy of news recommendation. In addition, after removing the topic classification task and the popularity prediction task, the accuracy of the model in capturing news topics decreases, which is due to the fact that the topic is the core of the news content and has a direct impact on the user's interest point, and the popularity is an important indicator to measure the attractiveness of news, which can predict and recommend those news that may be widely welcomed, and improve the user satisfaction. Finally, the removal of the sentiment analysis task has relatively small influence on the recommendation performance, which means that the sentiment analysis task may not be the most critical factor for news recommendation.
[0140] As shown in Figure 9 The news information processing device provided by the application comprises:
[0141] The acquisition module is configured to acquire input information, and the input information comprises candidate news information.
[0142] The generation module is configured to input the input information into a pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model.
[0143] The loss function of the pre-trained language model is a combination of the loss function of the recommendation task and the loss function of the attribute classification task.
[0144] In some embodiments, the input information can further comprise external knowledge information and historical news information clicked by a user.
[0145] The generation module can be specifically configured to:
[0146] establish a first prompt template of the recommendation task based on the external knowledge information, the historical news information and the candidate news information.
[0147] input the first prompt template into the pre-trained language model.
[0148] In some embodiments, the external knowledge information can include a knowledge graph.
[0149] The generation module can also be configured to:
[0150] extract each entity in the candidate news information through a named entity recognition technology;
[0151] retrieve each triple related to the entity from the knowledge graph;
[0152] convert each triple into a text sequence;
[0153] establish a first prompt template based on the text sequence of the triple, the historical news information and the candidate news information.
[0154] In some embodiments, the external knowledge information can include at least one of a knowledge graph, a knowledge base and an online resource.
[0155] In some embodiments, the generation module can be specifically configured to:
[0156] establish a second prompt template of an attribute classification task based on the candidate news information;
[0157] input the second prompt template into the pre-trained language model.
[0158] In some embodiments, the generation module can also be configured to:
[0159] predict a label word at a mask position in the second prompt template;
[0160] determine the classification result corresponding to the label word according to the mapping relationship between the classification result and the label word.
[0161] In some embodiments, the attribute classification task can include at least one of a sentiment analysis task, a topic classification task and a popularity prediction task.
[0162] It should be noted that the news information processing apparatus provided by the present application can execute the news information processing method of any of the above embodiments when it is actually running, and therefore the present embodiment will not be described here.
[0163] Figure 10 is a structural schematic diagram of an electronic device provided by the present application, as Figure 10As shown, the electronic device can include a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory complete communication with each other through the communications bus. The processor can invoke a logical instruction in the memory to execute a news information processing method, which includes: obtaining candidate news information; inputting the candidate news information into a pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model; and wherein a loss function of the pre-trained language model is a combination of a loss function of the recommendation task and a loss function of the attribute classification task.
[0164] In addition, the logical instructions in the memory described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0165] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, when the program instructions are executed by a computer, the computer can execute the news information processing method provided by the above-mentioned embodiments, which includes: obtaining candidate news information; inputting the candidate news information into a pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model; and wherein a loss function of the pre-trained language model is a combination of a loss function of the recommendation task and a loss function of the attribute classification task.
[0166] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the news information processing method provided by any of the above embodiments, and the method comprises: obtaining candidate news information; inputting the candidate news information into a pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model; and wherein a loss function of the pre-trained language model is a combination of a loss function of the recommendation task and a loss function of the attribute classification task.
[0167] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment according to actual needs. Those skilled in the art can understand and implement it without creative labor.
[0168] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software plus necessary universal hardware platforms, and of course, can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0169] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A news information processing method, characterized in that: include: Acquiring input information, wherein the input information includes candidate news information; Inputting the input information into a pre-trained language model to obtain a recommendation result of a recommendation task and a classification result of an attribute classification task output by the pre-trained language model; the attribute classification task includes a sentiment analysis task, a topic classification task, and a popularity prediction task; The popularity prediction task includes calculating the popularity , popularity The calculation formula is: ; represents a natural constant, Represents the time interval from the news release to the present, Represents the click rate of news; Among them, the recommendation task has similarities in knowledge structure with the popularity prediction task, the topic classification task, and the sentiment analysis task, and their data sets can complement each other. The recommendation task and the attribute classification task are optimized using a multi-task learning strategy; the loss function of the pre-trained language model integrates the loss function of the recommendation task and the loss function of the attribute classification task. The loss function of the pre-trained language model is the sum of the loss function of the recommendation task, the loss function of the sentiment analysis task, the loss function of the topic classification task, and the loss function of the popularity prediction task.
2. The news information processing method according to claim 1, characterized in that: The input information also includes external knowledge information and historical news information clicked by the user; Inputting the input information into the pre-trained language model includes: establishing a first prompt template for the recommendation task based on the external knowledge information, the historical news information, and the candidate news information; The first prompt template is input into the pre-trained language model.
3. The news information processing method according to claim 2, characterized in that: The external knowledge information includes a knowledge graph; The step of establishing the first prompt template for the recommendation task based on the external knowledge information, the historical news information, and the candidate news information includes: Extracting entities from the candidate news information using named entity recognition technology; Retrieving each triple related to the entity from the knowledge graph; Convert each of the triples into a text sequence; The first prompt template is established based on the text sequence of the triples, the historical news information and the candidate news information.
4. The news information processing method according to claim 2, characterized in that: The external knowledge information includes at least one of a knowledge graph, a knowledge base, and online resources.
5. The news information processing method according to claim 1, characterized in that: Inputting the input information into the pre-trained language model includes: Establishing a second prompt template for the attribute classification task based on the candidate news information; The second prompt template is input into the pre-trained language model.
6. The news information processing method according to claim 5, characterized in that: The step of obtaining the classification result includes: Predicting a label word for the mask position in the second prompt template; The classification result corresponding to the label word is determined according to a mapping relationship between the classification result and the label word.
7. A news information processing device, characterized in that: include: An acquisition module, configured to acquire input information, wherein the input information includes candidate news information; A generation module is used to input the input information into a pre-trained language model to obtain the recommendation results of the recommendation task and the classification results of the attribute classification task output by the pre-trained language model; the attribute classification task includes a sentiment analysis task, a topic classification task, and a popularity prediction task; The popularity prediction task includes calculating the popularity , popularity The calculation formula is: ; represents a natural constant, Represents the time interval from the news release to the present, Represents the click rate of news; Among them, the recommendation task has similarities in knowledge structure with the popularity prediction task, the topic classification task, and the sentiment analysis task, and their data sets can complement each other. The recommendation task and the attribute classification task are optimized using a multi-task learning strategy; the loss function of the pre-trained language model integrates the loss function of the recommendation task and the loss function of the attribute classification task. The loss function of the pre-trained language model is the sum of the loss function of the recommendation task, the loss function of the sentiment analysis task, the loss function of the topic classification task, and the loss function of the popularity prediction task.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the news information processing method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the news information processing method according to any one of claims 1 to 6 is implemented.