News event classification method and device

By introducing a combination of search model and generative model in the news event classification, the problems of low efficiency and large subjective deviation in the existing technology are solved, and more efficient and accurate news event classification is achieved.

CN119988632APending Publication Date: 2025-05-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510122570.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology has problems such as low efficiency and prone to subjective deviations in the classification of news events, especially when processing massive news data, it is difficult to ensure consistency and accuracy.

Method used

A news event classification method based on the search model and the generation model is proposed. By generating sample input query, pairwise matching and calculating matching scores, building positive and negative samples, and training the search model; at the same time, using the news event knowledge base to generate prompt information, adjust the parameters of the large language model, train the generation model, and finally combining the search model and the generation model for news event classification.

Benefits of technology

It improves the efficiency and accuracy of news event classification, reduces subjective bias, and better understands the semantics and context of news texts, thereby improving the accuracy of classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988632A_ABST
    Figure CN119988632A_ABST
Patent Text Reader

Abstract

The invention provides a news event classification method and device, and relates to the technical field of artificial intelligence, in particular to the technical field of natural language processing, deep learning and large language models. A specific embodiment of the method comprises the steps of generating an input query based on a news text; inputting the input query and the news event knowledge base into a retrieval model to obtain reference news event knowledge; generating prompt information based on the reference news event knowledge; and inputting the news text and the prompt information into the generation model to obtain a news event category of the news text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the field of natural language processing, deep learning and large language model technology. Background Art

[0002] With the popularization of the Internet and the diversification of media channels, the speed and amount of news information are increasing rapidly, forming an information explosion phenomenon. How to effectively organize, manage and retrieve this news information has become an important issue. News event classification technology can help media companies, search engines and data analysis agencies to structure the management of massive news and improve the efficiency and accuracy of information retrieval.

[0003] Manually classifying and organizing news is not only time-consuming and labor-intensive, but also prone to subjective bias. Automated news classification systems can provide consistency and efficiency in large-scale data processing. Automated news classification technology can be applied to personalized recommendations, public opinion monitoring, market analysis and other fields to provide users with more accurate services. The development of natural language processing, machine learning, deep learning and other technologies has provided new methods and tools for news event classification. These technologies can better understand the semantics and context of news texts, thereby improving the accuracy of classification.

[0004] At present, news event classification mainly relies on technologies such as natural language processing, machine learning, and deep learning. Common implementation solutions are as follows: First, keyword-based classification: This method relies on predefined keywords or phrases to classify news texts, and determines the category to which it belongs by matching keywords in the text.

[0005] Second, the traditional machine learning method: use feature engineering to convert text into vector representation, and then apply traditional machine learning algorithms such as naive Bayes, support vector machine, logistic regression, etc. for classification.

[0006] Third, classification based on topic models: Use topic models (such as LDA (Latent Dirichlet Allocation)) to extract topics from texts, and then classify them according to the topic distribution.

[0007] Fourth, deep learning methods: use deep neural networks (such as convolutional neural networks, recurrent neural networks, and long short-term memory networks) to classify news texts. Summary of the invention

[0008] The embodiments of the present disclosure provide a news event classification method, apparatus, device, storage medium, and program product.

[0009] In a first aspect, an embodiment of the present disclosure proposes a retrieval model training method, comprising: generating a sample input query set based on a sample news text set; matching sample input queries in the sample input query set in pairs to obtain related input queries of the sample input queries; calculating matching scores of the sample input queries and the related input queries; constructing positive samples and negative samples based on the matching scores; and training a retrieval model based on the positive samples and the negative samples.

[0010] In a second aspect, an embodiment of the present disclosure proposes a generative model training method, comprising: generating a sample input query based on a sample news text; searching a news event knowledge base based on the sample input query to obtain sample reference news event knowledge; generating sample prompt information based on the sample reference news event knowledge; inputting the sample news text and the sample prompt information into a large language model to obtain a predicted news event category of the sample news text; adjusting the parameters of the large language model based on the difference between the predicted news event category of the sample news text and the real news event category to obtain a generative model.

[0011] In a third aspect, an embodiment of the present disclosure proposes a news event classification method, including: generating an input query based on news text; inputting the input query and a news event knowledge base into a retrieval model to obtain reference news event knowledge; generating prompt information based on the reference news event knowledge; inputting the news text and the prompt information into a generation model to obtain a news event category of the news text.

[0012] In a fourth aspect, an embodiment of the present disclosure proposes a retrieval model training device, comprising: a generation module, configured to generate a sample input query set based on a sample news text set; a matching module, configured to match sample input queries in the sample input query set in pairs to obtain related input queries of the sample input queries; a calculation module, configured to calculate the matching scores of the sample input queries and the related input queries; a construction module, configured to construct positive samples and negative samples based on the matching scores; and a training module, configured to train a retrieval model based on the positive samples and the negative samples.

[0013] In a fifth aspect, an embodiment of the present disclosure proposes a generative model training device, comprising: a first generation module, configured to generate a sample input query based on a sample news text; a retrieval module, configured to search a news event knowledge base based on the sample input query to obtain sample reference news event knowledge; a second generation module, configured to generate sample prompt information based on the sample reference news event knowledge; a classification module, configured to input the sample news text and the sample prompt information into a large language model to obtain a predicted news event category of the sample news text; an adjustment module, configured to adjust the parameters of the large language model based on the difference between the predicted news event category of the sample news text and the real news event category to obtain a generative model.

[0014] In the sixth aspect, the embodiment of the present disclosure proposes a news event classification device, including: a first generation module, configured to generate an input query based on a news text; a retrieval module, configured to input the input query and a news event knowledge base into a retrieval model to obtain reference news event knowledge; a second generation module, configured to generate prompt information based on the reference news event knowledge; a third generation module, configured to input the news text and the prompt information into the generation model to obtain a news event category of the news text.

[0015] In the seventh aspect, an embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the first aspect, the second aspect, or the third aspect.

[0016] In an eighth aspect, an embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to execute the method described in the first aspect, the second aspect, or the third aspect.

[0017] In a ninth aspect, an embodiment of the present disclosure proposes a computer program product, including a computer program, which implements the method described in the first aspect, the second aspect, or the third aspect when executed by a processor.

[0018] The key or important features of the embodiments of the present disclosure are not intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Other features, purposes and advantages of the present disclosure will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings. The drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. Among them: Figure 1 is a flowchart of an embodiment of a retrieval model training method according to the present disclosure; Figure 2 is a flow chart of an embodiment of a generative model training method according to the present disclosure; Figure 3 is a flow chart of an embodiment of a news event classification method according to the present disclosure; Figure 4 It is the overall framework diagram of news event classification; Figure 5 It is a framework diagram of model retrieval and re-ranking; Figure 6 It is a flow chart of news event classification; Figure 7 It is a flowchart for constructing a retrieval model fine-tuning dataset; Figure 8 It is the online knowledge base update flow chart; Fig. 9 is a structural schematic diagram of an embodiment of a retrieval model training device according to the present disclosure; Fig.10 is a structural schematic diagram of an embodiment of a generative model training device according to the present disclosure; Fig.11 is a structural schematic diagram of an embodiment of a news event classification device according to the present disclosure; Fig.12 It is a block diagram of an electronic device used to implement the news event classification method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0021] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0022] The main goal of news event classification is to accurately assign news articles to one or more predefined event categories based on their content. These categories are usually related to specific topics, events, or activities. Therefore, news event classification is usually defined as a text classification task.

[0023] News event classification can be built based on the big model RAG (Retrieval-Augmented Generation) technology, the core idea of ​​which is to enhance the ability of the generation model by retrieving relevant information from a knowledge base, so that it can generate more accurate and informative texts. The big model RAG can mainly include the retrieval model and the generation model.

[0024] The retrieval model can find text snippets related to the input query from a predefined vertical domain knowledge base, which can provide additional information support for the generation model.

[0025] At present, commonly used retrieval models can be roughly divided into two categories: one is a dense vector retrieval model based on deep learning models such as BERT (Bidirectional Encoder Representation from Transformers) and GPT (Generative Pre-trained Transformer); the other is a sparse vector retrieval model represented by BM25 (Best Matching 25) and TF-IDF (Term Frequency-Inverse Document Frequency).

[0026] The dense vector retrieval model is better at extracting semantic information from text, while the sparse vector retrieval model is better at extracting keyword information. Both the input query and the news event knowledge base text are encoded as vectors, and the retrieval process is achieved by calculating the similarity between the query vector and the knowledge base vector. By introducing an external knowledge base, the retrieval model can provide real-time updated information and more extensive background knowledge, which is especially important for generation tasks that require the latest information.

[0027] The task of the generative model can be to generate coherent and meaningful text based on the given input (including the original input and the retrieved related knowledge). Models with Transformer architecture are usually used, such as GPT or BART (Bidirectional and Auto-Regressive Transformers). The generative model can combine the input context information with the retrieved text fragments to generate the final output. By combining the information provided by the retrieval model, the generative model can better understand the context and generate more accurate and detailed content. This combination also helps to reduce the probability of generating errors or hallucinations (i.e. generating information that does not conform to the facts). At present, generative models based on large models are generally used. Commonly used large model bases may include but are not limited to: Llama, Qwen, ChatGLM, Baichuan, etc.

[0028] Figure 1 A process 100 of an embodiment of a retrieval model training method according to the present disclosure is shown. The retrieval model training method comprises the following steps: Step 101: Generate a sample input query set based on a sample news text set.

[0029] In this embodiment, the execution subject of the retrieval model training method can generate a sample input query set based on a sample news text set.

[0030] The execution subject of the retrieval model training method is usually a server. The server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0031] The sample news text set may include a large number of sample news texts. The sample news text may be any news text. The sample input query set may include a large number of sample input queries. The sample input query may be key information of the sample news text. A sample news text in the sample news text set may generate a sample input query in the sample input query set.

[0032] In some embodiments, in order to improve the accuracy of the search, the sample news text may be preprocessed, wherein the preprocessing may include but is not limited to at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, correcting spelling, etc.

[0033] In some embodiments, since the sample news text is generally a long text, sentence segmentation is required. By segmenting the sample news text, several related sample natural sentences can be obtained. Based on these sample natural sentences, a sample input query can be formed. For example, the sample news text is first segmented based on punctuation marks to obtain a sequence of sample natural sentences; then, for sample natural sentences whose length is less than the maximum sentence length, adjacent sample natural sentences can be used to splice them to obtain sample spliced ​​natural sentences; finally, based on sample natural sentences and sample spliced ​​natural sentences whose length is not less than the maximum sentence length, a sample input query can be generated. Among them, when splicing, the length of the sample spliced ​​natural sentence can be made as close to the maximum sentence length as possible. And, ensure that some overlapping sentences are contained between the paragraphs after splicing. In this way, the contextual information between paragraphs or sentences can be retained to the maximum extent without exceeding the specified text block size.

[0034] Optionally, after the sample news text set is preprocessed and sentence segmented to obtain the sample input query set, the sample input query set may be deduplicated.

[0035] Step 102: Match sample input queries in the sample input query set in pairs to obtain related input queries of the sample input queries.

[0036] In this embodiment, the execution entity may perform pairwise matching of the sample input queries in the sample input query set to obtain related input queries of the sample input queries.

[0037] In some embodiments, the sample input queries in the sample input query set can be matched in pairs using the BM25 algorithm, and the most likely related input queries are matched for each sample input query. The BM25 algorithm is an algorithm for information retrieval and text mining, and is widely used in search engines and related fields. BM25 is based on the idea of ​​TF-IDF, but is improved to take into account factors such as the length of the document.

[0038] Step 103: Calculate the matching scores between the sample input query and the related input query.

[0039] In this embodiment, the execution entity may calculate the matching score between the sample input query and the related input query. The matching score may be used to characterize the matching degree between the sample input query and the related input query. The higher the matching degree, the more the sample input query matches the related input query; conversely, the less the sample input query matches the related input query.

[0040] In some embodiments, the BGE (BAAI General Embedding) model can be used to calculate the matching scores of sample input queries and related input queries. BGE is a general semantic vector model with leading multilingual and cross-lingual retrieval capabilities, which fully and high-quality supports input texts of different granularities such as "sentences", "paragraphs", "chapters", and "documents", and integrates three retrieval functions, namely dense retrieval, sparse retrieval, and multi-vector retrieval, in one stop, achieving the best level in multiple evaluation benchmarks.

[0041] Step 104: construct positive samples and negative samples based on the matching scores.

[0042] In this embodiment, the execution entity can construct positive samples and negative samples based on the matching scores. The data format of the positive samples and negative samples can be, for example, {"query": str, "pos": List[str], "neg": List[str]}. For each sample input query, several positive samples (such as 2-5 positive samples) and several negative samples (such as 10 negative samples) can be provided.

[0043] In some embodiments, if the matching score is lower than a preset low score threshold, it can be considered that the two sample input queries are likely to be unrelated sentence pairs, and can be used to construct negative samples. If the matching score is higher than a preset high score threshold, it can be considered that the correlation between the two sample input queries is large, and can be used to construct positive samples. If the matching score is between the preset low score threshold and the preset high score threshold, a large model (such as GPT-4) can be used for annotation.

[0044] It should be noted that when constructing positive samples and negative samples, two samples with matching scores between a preset low score threshold and a preset high score threshold can be preferentially selected to input queries to construct positive samples and negative samples. If the amount of negative sample data is insufficient, random text can be extracted as negative samples.

[0045] Step 105: Based on the positive samples and the negative samples, a retrieval model is trained.

[0046] In this embodiment, the execution subject can train a retrieval model based on positive samples and negative samples.

[0047] In some embodiments, the retrieval model is initialized, and the sample is input into the retrieval model to obtain the retrieval result. Based on the retrieval result and the label of the sample, the loss value is calculated. Based on the loss value, the parameters of the retrieval model are adjusted until the loss value is small enough. The label of the sample can be used to characterize a positive sample or a negative sample.

[0048] The disclosed embodiment provides a retrieval model training method, which uses domain data to fine-tune the retrieval model, thereby improving the reasoning effect of the model in the news field.

[0049] Figure 2 A process 200 of an embodiment of a generative model training method according to the present disclosure is shown. The generative model training method comprises the following steps: Step 201, generating a sample input query based on a sample news text.

[0050] In this embodiment, the execution subject of the generation model training method can generate a sample input query based on the sample news text.

[0051] The execution subject of the generation model training method is usually a server. The server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0052] The sample news text may be any news text, and the sample input query may be key information of the sample news text.

[0053] In some embodiments, in order to improve the accuracy of the search, the sample news text may be preprocessed, wherein the preprocessing may include but is not limited to at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, correcting spelling, etc.

[0054] In some embodiments, since the sample news text is generally a long text, sentence segmentation is required. By segmenting the sample news text, several related sample natural sentences can be obtained. Based on these sample natural sentences, a sample input query can be formed. For example, the sample news text is first segmented based on punctuation marks to obtain a sequence of sample natural sentences; then, for sample natural sentences whose length is less than the maximum sentence length, adjacent sample natural sentences can be used to splice them to obtain sample spliced ​​natural sentences; finally, based on sample natural sentences and sample spliced ​​natural sentences whose length is not less than the maximum sentence length, a sample input query can be generated. Among them, when splicing, the length of the sample spliced ​​natural sentence can be made as close to the maximum sentence length as possible. And, ensure that some overlapping sentences are contained between the paragraphs after splicing. In this way, the contextual information between paragraphs or sentences can be retained to the maximum extent without exceeding the specified text block size.

[0055] Optionally, after the sample news text set is preprocessed and sentence segmented to obtain the sample input query set, the sample input query set may be deduplicated.

[0056] Step 202, based on the sample input query, search in the news event knowledge base to obtain sample reference news event knowledge.

[0057] In this embodiment, the execution subject may search the news event knowledge base based on the sample input query to obtain sample reference news event knowledge.

[0058] The news event knowledge base may include news event knowledge of various categories. When constructing the news event knowledge base, multiple news event categories may be preset, positive examples and negative examples may be set for each news event category, and positioning words may be filtered for each news event category.

[0059] Define the news event categories that need to be classified and recalled, and set several positive and negative examples for each news event category. For example, select the news event categories that need to be monitored, and domain experts can write definitions and related examples based on the news event categories, with 10 initial positive and negative examples for each category. Among them, the definition can specify the cases that belong to positive examples and the cases that belong to negative examples as needed. For the parts of the news event categories that are not monitored, "other" can be used as the category label.

[0060] To improve the efficiency of the RAG model, when constructing a news event knowledge base, relevant positioning words can be screened out for each news event category based on the characteristics of each news event category, so as to quickly exclude irrelevant text paragraphs.

[0061] In some embodiments, when generating a sample input query, the positioning words may be used to match the sample natural sentences in the sample natural sentence sequence, and the sample natural sentences that do not contain the positioning words may be filtered.

[0062] Step 203, generating sample prompt information based on sample reference news event knowledge.

[0063] In this embodiment, the above-mentioned execution subject can generate sample prompt information based on sample reference news event knowledge.

[0064] In the RAG classification scenario, the large model can be used to classify the recalled samples based on the knowledge of news events. Therefore, the prompt information template can be constructed as follows: prompt template You are an expert in the field of news. For the input news text, refer to the definitions, judgment criteria and examples of the relevant categories provided below to determine the category to which it belongs.

[0065] Output format description : Output in json format. The output result is the "label" field, which indicates the category to which the text belongs. The value range is {candidate category label list}. Please refer to the definition and judgment basis provided below for details.

[0066] # Definition of relevant categories and their judgment criteria: {Category definitions and criteria} # Example: {Example} --- enter: {Enter text} Output: {Output Category} Among them, {candidate category label list} is the category label set involved in all sample reference news event knowledge retrieved, and the category "other" is added at the end as a fallback category. {category definition and its judgment criteria} is the pre-defined examples and possible positive and negative examples. The input of {example} is the first k' sample reference news event knowledge retrieved, and the output is the category to which it belongs, expressed in json.

[0067] In order to ensure the scalability of the generative model for unknown news events, a small amount of classified data can be constructed to perform LoRA (Low-Rank Adaptation) fine-tuning on the output format of the large language model. The training data can include input text and output labels. The input text can be a sample prompt information constructed using the prompt template, and the output annotation format is the json output format. The output category label can annotate the real news event category of the sample news text. Using training data to train the generative model can make the output format of the generative model meet the requirements of the task.

[0068] It should be noted that when there are sufficient annotated datasets in the field, large quantities of annotated corpora in the field can also be used for fine-tuning to further improve the prediction accuracy of the generative model.

[0069] Step 204: input the sample news text and the sample prompt information into the large language model to obtain the predicted news event category of the sample news text.

[0070] In this embodiment, the execution entity may input the sample news text and the sample prompt information into the large language model to obtain the predicted news event category of the sample news text.

[0071] Step 205, based on the difference between the predicted news event category and the real news event category of the sample news text, the parameters of the large language model are adjusted to obtain a generation model.

[0072] In this embodiment, the execution subject may adjust the parameters of the large language model based on the difference between the predicted news event category and the real news event category of the sample news text to obtain a generation model.

[0073] In some embodiments, a loss value is calculated based on the predicted news event category and the real news event category of the sample news text. Based on the loss value, the parameters of the large language model are adjusted until the loss value is small enough to obtain a generative model.

[0074] The disclosed embodiment provides a generative model training method, which uses domain data to fine-tune the retrieval model, thereby improving the reasoning effect of the model in the news field.

[0075] Figure 3 A process 300 of an embodiment of a news event classification method according to the present disclosure is shown. The news event classification method comprises the following steps: Step 301, generating an input query based on the news text.

[0076] In this embodiment, the execution subject of the news event classification method can generate an input query based on the news text.

[0077] The execution subject of the news event classification method is usually a server. The server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0078] The news text may be any news text input by the user, and the input query may be key information of the news text.

[0079] In some embodiments, in order to improve the accuracy of the search, the news text may be preprocessed, wherein the preprocessing may include but is not limited to at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, correcting spelling, etc.

[0080] In some embodiments, since news text is generally a long text, sentence segmentation is required. By segmenting the news text, several related natural sentences can be obtained. Based on these natural sentences, an input query can be formed. For example, first, the news text is segmented based on punctuation marks to obtain a sequence of natural sentences; then, for natural sentences whose length is less than the maximum sentence length, adjacent natural sentences can be used to splice them to obtain spliced ​​natural sentences; finally, based on natural sentences whose length is not less than the maximum sentence length and spliced ​​natural sentences, an input query can be generated. Among them, when splicing, the length of the spliced ​​natural sentence can be made as close to the maximum sentence length as possible. And, ensure that some overlapping sentences are contained between the paragraphs after splicing. In this way, the contextual information between paragraphs or sentences can be retained to the maximum extent without exceeding the specified text block size.

[0081] Optionally, after the news text set is preprocessed and sentence segmented to obtain the input query set, the input query set may be deduplicated.

[0082] Step 302, input the input query and the news event knowledge base into the retrieval model to obtain reference news event knowledge.

[0083] In this embodiment, the above-mentioned execution entity can input the input query and the news event knowledge base into the retrieval model to obtain reference news event knowledge.

[0084] The news event knowledge base may include news event knowledge of various categories. When constructing the news event knowledge base, multiple news event categories may be preset, positive examples and negative examples may be set for each news event category, and positioning words may be filtered for each news event category.

[0085] Define the news event categories that need to be classified and recalled, and set several positive and negative examples for each news event category. For example, select the news event categories that need to be monitored, and domain experts can write definitions and related examples based on the news event categories, with 10 initial positive and negative examples for each category. Among them, the definition can specify the cases that belong to positive examples and the cases that belong to negative examples as needed. For the parts of the news event categories that are not monitored, "other" can be used as the category label.

[0086] To improve the efficiency of the RAG model, when constructing a news event knowledge base, relevant positioning words can be screened out for each news event category based on the characteristics of each news event category, so as to quickly exclude irrelevant text paragraphs.

[0087] In some embodiments, when generating an input query, the natural sentences in the natural sentence sequence may be matched using the positioning words, and the natural sentences that do not contain the positioning words may be filtered out.

[0088] The retrieval model can pre-process the input news text to obtain a set of related natural sentences to form an input query. Use the set retrieval strategies to perform retrieval and recall in the news event knowledge base to obtain a number of reference news event knowledge related to the input query.

[0089] After the retrieval model vectorizes the input query, it can determine the similarity of the text by calculating the distance between it and the news event knowledge in the news event knowledge base, and finally select the k news event knowledge with the highest similarity for recall. and , the distance measurement method generally uses cosine similarity or Euclidean distance, and the specific calculation formula can be as follows: Cosine Similarity: ; Euclidean distance: .

[0090] In the recall phase of the large model RAG, the sparse vector retrieval model and the dense vector retrieval model can be combined. The sparse vector retrieval model can be retrieved using, for example, the BM25 model, and the dense vector retrieval model can be vectorized using, for example, the bge-large-zh-v1.5 model, and then retrieved by similarity sorting.

[0091] In some embodiments, the input query and the news event knowledge base are input into a sparse vector retrieval model to obtain first reference news event knowledge. The input query and the news event knowledge base are input into a dense vector retrieval model to obtain second reference news event knowledge. The first reference news event knowledge and the second reference news event knowledge are combined and deduplicated to obtain reference news event knowledge. Among them, the sparse vector retrieval model is good at finding related texts based on keywords, while the dense vector retrieval model is good at finding texts based on semantic information, and the two complement each other. Since different retrieval models have different score distributions, the m news event knowledge with the highest scores in each retrieval model can be selected for recall, and k news event knowledge can be obtained after deduplication. Where k≤lm, l is the number of retrieval models.

[0092] It should be noted that the retrieval model can be used Figure 1 The embodiment shown is obtained by training and will not be described in detail here.

[0093] In some embodiments, the input query and reference news event knowledge are input into the re-ranking model, and the reference news event knowledge can be filtered based on the score. Using the re-ranking model, the input query and the reference news event knowledge can be ranked based on their relevance to obtain a refined ranking set.

[0094] Since multiple retrieval models are used in the retrieval stage, and the score distribution of each retrieval model is different, it is difficult to directly compare and rank them. Therefore, a reranking model is used to further refine and rank the results obtained from the preliminary retrieval to improve the relevance of the results. Among them, the reranking model can use the cross entropy loss function to perform binary classification task learning to determine whether the two input sentences match. Through secondary sorting by the reranking model, the k' news event knowledge with the highest scores can be selected from the k news event knowledge for subsequent model generation. The reranking model can be, for example, bge-reranker-large.

[0095] Step 303, generating prompt information based on the reference news event knowledge.

[0096] In this embodiment, the above-mentioned execution subject can generate prompt information based on the reference news event knowledge.

[0097] In the RAG classification scenario, using the large model, classification can be performed based on the recalled reference news event knowledge.

[0098] In some embodiments, reference news event categories are extracted from reference news event knowledge; based on the reference news event categories, a list of candidate category labels in a prompt information template is filled in; the definition, positive examples, and negative examples of the reference news event categories are extracted from the reference news event knowledge; based on the definition, positive examples, and negative examples of the reference news event categories, the category definition and judgment criteria in the prompt information template are filled in; based on the reference news event knowledge, examples in the prompt information template are filled in.

[0099] Typically, the examples filled in the prompt information template may be partial reference news event knowledge.

[0100] In some embodiments, a matching score between an input query and reference news event knowledge is calculated; based on the matching score, example news event knowledge is selected from the reference news event knowledge and written into the example in the prompt information template.

[0101] In some embodiments, an input query and reference news event knowledge are input into a re-ranking model to obtain a matching score between the input query and the reference news event knowledge.

[0102] By using the re-ranking model, we can sort based on the relevance of the input query and the reference news event knowledge to obtain a refined ranking set.

[0103] Since multiple retrieval models are used in the retrieval stage, and the score distribution of each retrieval model is different, it is difficult to directly compare and rank them. Therefore, a reranking model is used to further refine and rank the results obtained from the preliminary retrieval to improve the relevance of the results. Among them, the reranking model can use the cross entropy loss function to perform binary classification task learning to determine whether the two input sentences match. Through secondary sorting by the reranking model, the k' news event knowledge with the highest scores can be selected from the k news event knowledge for subsequent model generation. The reranking model can be, for example, bge-reranker-large.

[0104] The prompt information template can be constructed as follows: prompt template You are an expert in the field of news. For the input news text, refer to the definitions, judgment criteria and examples of the relevant categories provided below to determine the category to which it belongs.

[0105] Output format description : Output in json format. The output result is the "label" field, which indicates the category to which the text belongs. The value range is {candidate category label list}. Please refer to the definition and judgment basis provided below for details.

[0106] # Definition of relevant categories and their judgment criteria: {Category definitions and criteria} # Example: {Example} --- enter: {Enter text} Output: {Output Category} Among them, {candidate category label list} is the category label set involved in all sample reference news event knowledge retrieved, and the category "other" is added at the end as a fallback category. {category definition and its judgment criteria} is the pre-defined examples and possible positive and negative examples. The input of {example} is the first k' sample reference news event knowledge retrieved, and the output is the category to which it belongs, expressed in json.

[0107] Step 304: input the news text and prompt information into the generation model to obtain the news event category of the news text.

[0108] In this embodiment, the execution subject may input the news text and the prompt information into the generation model to obtain the news event category of the news text.

[0109] Using the supervised fine-tuned generative model, the news event categories of the news text can be predicted, and finally all the news event categories about the news text can be output.

[0110] In the RAG classification task, the generative model can generate category labels based on the retrieved news event knowledge, and a large model is usually used as the generative model, such as Qwen2.5-7B-Instruct.

[0111] It should be noted that the generative model can be used Figure 2 The embodiment shown is obtained by training and will not be described in detail here.

[0112] The disclosed embodiment provides a news event classification method, which is mainly used in the news field to classify events contained in news texts so as to identify existing and new news events. Timely grasping and analyzing news hot events can help enterprises and research institutions use news classification data to conduct market trend analysis, competitor monitoring and consumer behavior research, so as to make more informed business decisions. Government agencies and non-governmental organizations can also understand hot issues and social dynamics of public concern by analyzing news events of different categories.

[0113] Accurate news classification can help readers quickly obtain the information they need and improve user experience. In the business field, companies can make more informed decisions by analyzing news data for market forecasts and competition analysis. At the social level, news event classification technology can be used to monitor social hot spots in real time and assist government departments in public opinion management and emergency response.

[0114] In summary, news event classification has important background and research significance in information management, technology development, social application, etc. Researching and developing an efficient news classification system can not only improve the efficiency of information processing, but also help promote the innovation and development of related industries.

[0115] In addition, the algorithm is also applicable to other fields that require background knowledge, such as medicine and technology, and can effectively perform text classification tasks in these vertical fields.

[0116] Figure 4 The overall framework diagram of news event classification is shown.

[0117] 1. News event knowledge base construction: define the news event categories that need to be classified and recalled, and set several positive and negative examples for each news event category. The definitions and positive and negative examples 401 can be encoded into sentence vectors 402 and stored in the news event knowledge base 403.

[0118] 2. Input query formation: Encoding the input news text 404 can form a sentence vector 404.

[0119] 3. Model retrieval: Use the set retrieval strategies to perform retrieval and recall in the news event knowledge base 404, and output several retrieval results 405 related to the sentence vector 404.

[0120] 4. Prompt information construction: Based on the search results 405 , prompt information 406 can be generated.

[0121] 5. Model generation: Input the prompt information 406 into the supervised fine-tuned generation model 407 and output the prediction result 408.

[0122] 6. Analyze the prediction result 408 to obtain the classification result 409.

[0123] Figure 5 A framework diagram of model retrieval and re-ranking is shown.

[0124] In the recall phase of the large model RAG, the sparse vector retrieval model 501 and the dense vector retrieval model 502 can be combined. The input query 503 is input into the sparse vector retrieval model 501 and the dense vector retrieval model 502 for mixed retrieval to obtain a candidate set 504. The candidate set 504 is input into the re-ranking model 505 for sorting to obtain a refined ranking set 506.

[0125] Figure 6 A flow chart of news event classification is shown.

[0126] Step 601, input news text.

[0127] Step 602, preprocessing the news text to obtain the preprocessed news text.

[0128] Step 603, segment the preprocessed news text to obtain a natural sentence set.

[0129] Step 604: Use the positioning words to filter the natural sentence set, and remove the natural sentences that do not contain the positioning words.

[0130] Step 605, based on the filtered natural sentence set, model retrieval and re-ranking are performed on the news event knowledge base to obtain prediction results.

[0131] Step 606: remove sentences whose top-k prediction results are all negative to obtain samples.

[0132] Step 607: sort the matched samples according to the search scores.

[0133] Step 608, select the positive sample of the news event knowledge base with the top-1 retrieval score.

[0134] Step 609: Select the negative sample with the highest retrieval score of the corresponding category.

[0135] Step 610, selecting the first k samples and corresponding labels.

[0136] Step 611, combining the selected results into prompt information.

[0137] Step 612, perform large model prediction on the prompt information to obtain a prediction result.

[0138] Step 613, output the result.

[0139] Figure 7 The flowchart of constructing the retrieval model fine-tuning dataset is shown.

[0140] Step 701, output sample news text.

[0141] Step 702, preprocessing the sample news text to obtain the preprocessed sample news text.

[0142] Step 703, segment the preprocessed sample news text to obtain a set of sample natural sentences.

[0143] Step 704, perform BM25 matching on the sample natural sentence set to obtain matching results.

[0144] Step 705, select top-20 relevant texts from the matching results.

[0145] Step 706, using the BGE model to calculate the matching score.

[0146] Step 707: For low scores, the two sentences are considered to be a partially irrelevant sentence pair.

[0147] Step 708: For high scores, the two sentences are considered to be a partially related sentence pair.

[0148] Step 709: For the intermediate scores, use GPT-4 to determine whether they are relevant.

[0149] Step 710: Sample the partially irrelevant sentence pairs to obtain negative samples.

[0150] Step 711, output negative samples.

[0151] Step 712: Sample partially related sentence pairs to obtain positive samples.

[0152] Step 713: output positive samples.

[0153] Step 714: For sentences that cannot be judged, manual annotation is performed.

[0154] Figure 8 The flowchart of online knowledge base update is shown.

[0155] When used online, the news event knowledge base 801 can be continuously added and cleaned based on manual feedback to improve the recall rate and accuracy of classification. On the one hand, new event categories 802 in the news field continue to emerge, and domain experts can continuously add definitions and examples 803 in the news event knowledge base; on the other hand, users can provide feedback on the accuracy of online recall categories, and count high false call examples 804 with low accuracy in the news event knowledge base 801, which can be removed in a timely manner after manual evaluation.

[0156] Further references Fig. 9 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a retrieval model training device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0157] like Fig. 9 As shown, the retrieval model training device 900 of this embodiment may include: a generation module 901, a matching module 902, a calculation module 903, a construction module 904 and a training module 905. The generation module 901 is configured to generate a sample input query set based on a sample news text set; the matching module 902 is configured to match the sample input queries in the sample input query set in pairs to obtain related input queries of the sample input queries; the calculation module 903 is configured to calculate the matching score of the sample input query and the related input query; the construction module 904 is configured to construct positive samples and negative samples based on the matching score; the training module 905 is configured to train the retrieval model based on the positive samples and the negative samples.

[0158] In this embodiment, in the retrieval model training device 900, the specific processing of the generation module 901, the matching module 902, the calculation module 903, the construction module 904 and the training module 905 and the technical effects thereof can be referred to respectively. Figure 1 The relevant descriptions of steps 101-105 in the corresponding embodiment are not repeated here.

[0159] In some optional implementations of this embodiment, the retrieval model training device 900 also includes: a preprocessing module, configured to preprocess the sample news text, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

[0160] In some optional implementations of the present embodiment, the generation module 901 is further configured to: segment the sample news texts in the sample news text set based on punctuation marks to obtain a sample natural sentence sequence; for sample natural sentences in the sample natural sentence sequence whose length is less than the maximum sentence length, use adjacent sample natural sentences to splice the sample natural sentences whose length is less than the maximum sentence length to obtain sample spliced ​​natural sentences; based on the sample natural sentences and sample spliced ​​natural sentences in the sample natural sentence sequence whose length is not less than the maximum sentence length, generate sample input queries in the sample input query set.

[0161] Further references Fig.10 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a generation model training device, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0162] like Fig.10 As shown, the generative model training device 1000 of this embodiment may include: a first generation module 1001, a retrieval module 1002, a second generation module 1003, a classification module 1004 and an adjustment module 1005. The first generation module 1001 is configured to generate a sample input query based on a sample news text; the retrieval module 1002 is configured to search in a news event knowledge base based on the sample input query to obtain sample reference news event knowledge; the second generation module 1003 is configured to generate sample prompt information based on the sample reference news event knowledge; the classification module 1004 is configured to input the sample news text and the sample prompt information into a large language model to obtain a predicted news event category of the sample news text; the adjustment module 1005 is configured to adjust the parameters of the large language model based on the difference between the predicted news event category of the sample news text and the real news event category to obtain a generative model.

[0163] In this embodiment, in the generation model training device 1000, the specific processing of the first generation module 1001, the retrieval module 1002, the second generation module 1003, the classification module 1004 and the adjustment module 1005 and the technical effects thereof can be referred to respectively. Figure 2 The relevant descriptions of steps 201 - 205 in the corresponding embodiment are not repeated here.

[0164] In some optional implementations of the present embodiment, the generation model training device 1000 also includes: a setting module, which is configured to preset multiple news event categories, set positive examples and negative examples for each news event category, and filter positioning words for each news event category to obtain a news event knowledge base.

[0165] In some optional implementations of the present embodiment, the generation model training device 1000 also includes: a preprocessing module, configured to preprocess the sample news text, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

[0166] In some optional implementations of the present embodiment, the first generation module 1001 is further configured to: segment the sample marked news text based on punctuation marks to obtain a sample natural sentence sequence; for sample natural sentences in the sample natural sentence sequence whose length is less than the maximum sentence length, use adjacent sample natural sentences to splice the sample natural sentences whose length is less than the maximum sentence length to obtain sample spliced ​​natural sentences; generate a sample input query based on the sample natural sentences and sample spliced ​​natural sentences in the sample natural sentence sequence whose length is not less than the maximum sentence length.

[0167] In some optional implementations of this embodiment, the generation model training device 1000 also includes: a matching module, which is configured to use positioning words to match sample natural sentences in the sample natural sentence sequence and filter sample natural sentences that do not contain positioning words.

[0168] Further references Fig.11 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a news event classification device. Figure 3 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0169] like Fig.11 As shown, the news event classification device 1100 of this embodiment may include: a first generation module 1101, a retrieval module 1102, a second generation module 1103 and a third generation module 1104. The first generation module 1101 is configured to generate an input query based on the news text; the retrieval module 1102 is configured to input the input query and the news event knowledge base into the retrieval model to obtain reference news event knowledge; the second generation module 1103 is configured to generate prompt information based on the reference news event knowledge; the third generation module 1104 is configured to input the news text and the prompt information into the generation model to obtain the news event category of the news text.

[0170] In this embodiment, in the news event classification device 1100, the specific processing of the first generation module 1101, the search module 1102, the second generation module 1103 and the third generation module 1104 and the technical effects thereof can be referred to in Figure 3 The relevant descriptions of steps 301 - 304 in the corresponding embodiment are not repeated here.

[0171] In some optional implementations of the present embodiment, the second generation module 1103 includes: a first extraction submodule, configured to extract reference news event categories from reference news event knowledge; a first filling submodule, configured to fill in the candidate category label list in the prompt information template based on the reference news event categories; a second extraction submodule, configured to extract the definition, positive examples and negative examples of the reference news event categories from the reference news event knowledge; a second filling submodule, configured to fill in the category definition and judgment criteria in the prompt information template based on the definition, positive examples and negative examples of the reference news event categories; and a third filling submodule, configured to fill in the examples in the prompt information template based on the reference news event knowledge.

[0172] In some optional implementations of this embodiment, the third filling sub-module includes: a calculation unit, configured to calculate the matching score between the input query and the reference news event knowledge; a selection unit, configured to select example news event knowledge from the reference news event knowledge based on the matching score, and write the example into the prompt information template.

[0173] In some optional implementations of this embodiment, the computing unit is further configured to: input the input query and the reference news event knowledge into the re-ranking model to obtain a matching score between the input query and the reference news event knowledge. In some optional implementations of this embodiment, the news event classification device 1100 further includes: a setting module configured to preset multiple news event categories, set positive examples and negative examples for each news event category, and filter positioning words for each news event category to obtain a news event knowledge base.

[0174] In some optional implementations of this embodiment, the news event classification device 1100 also includes: a preprocessing module, configured to preprocess the news text, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, word form restoration, full-width and half-width conversion, and spelling correction.

[0175] In some optional implementations of the present embodiment, the first generation module 1101 is further configured to: segment the news text based on punctuation marks to obtain a natural sentence sequence; for natural sentences in the natural sentence sequence whose length is less than the maximum sentence length, use adjacent natural sentences to splice the natural sentences whose length is less than the maximum sentence length to obtain a spliced ​​natural sentence; generate an input query based on the natural sentences in the natural sentence sequence whose length is not less than the maximum sentence length and the spliced ​​natural sentences.

[0176] In some optional implementations of this embodiment, the news event classification device 1100 also includes: a matching module, which is configured to use positioning words to match natural sentences in the natural sentence sequence and filter natural sentences that do not contain positioning words.

[0177] In some optional implementations of the present embodiment, the retrieval module 1102 is further configured to: input the input query and the news event knowledge base into a sparse vector retrieval model to obtain first reference news event knowledge; input the input query and the news event knowledge base into a dense vector retrieval model to obtain second reference news event knowledge; merge the first reference news event knowledge and the second reference news event knowledge and remove duplication to obtain reference news event knowledge.

[0178] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0179] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0180] Fig.12 A schematic block diagram of an example electronic device 1200 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0181] like Fig.12 As shown, the device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.

[0182] A number of components in the device 1200 are connected to the I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 1208, such as a disk, an optical disk, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the device 1200 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0183] The computing unit 1201 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1201 performs the various methods and processes described above, such as the news event classification method. For example, in some embodiments, the news event classification method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the news event classification method described above may be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute the news event classification method in any other appropriate manner (eg, by means of firmware).

[0184] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0185] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0186] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0188] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0189] A computer system may include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises through computer programs running on respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.

[0190] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions provided by this disclosure can be achieved, and this document does not limit this.

[0191] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A news event classification method, comprising: Based on the news text, generate input query; Inputting the input query and the news event knowledge base into the retrieval model to obtain reference news event knowledge; Based on the reference news event knowledge, generating prompt information; The news text and the prompt information are input into a generation model to obtain a news event category of the news text.

2. The method according to claim 1, wherein: The generating of prompt information based on the reference news event knowledge includes: Extracting reference news event categories from the reference news event knowledge; Based on the reference news event category, fill in the candidate category label list in the prompt information template; Extracting the definition, positive examples and negative examples of the reference news event category from the reference news event knowledge; Based on the definition, positive examples and negative examples of the reference news event category, fill in the category definition and judgment criteria in the prompt information template; Based on the reference news event knowledge, fill in the example in the prompt information template.

3. The method according to claim 2, wherein: The example of filling in the prompt information template based on the reference news event knowledge includes: Calculating a matching score between the input query and the reference news event knowledge; Based on the matching score, example news event knowledge is selected from the reference news event knowledge and written into the example in the prompt information template.

4. The method according to claim 3, wherein: The calculating the matching score between the input query and the reference news event knowledge includes: The input query and the reference news event knowledge are input into a re-ranking model to obtain a matching score between the input query and the reference news event knowledge.

5. The method according to claim 1, wherein: The method further comprises: Multiple news event categories are preset, positive examples and negative examples are set for each news event category, and positioning words are filtered for each news event category to obtain a news event knowledge base.

6. The method according to claim 1, wherein: The method further comprises: The news text is preprocessed, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

7. The method according to claim 1, wherein: The step of generating an input query based on the news text includes: Segmenting the news text based on punctuation marks to obtain a natural sentence sequence; For natural sentences in the natural sentence sequence whose length is less than the maximum sentence length, concatenating the natural sentences whose length is less than the maximum sentence length using adjacent natural sentences to obtain a concatenated natural sentence; The input query is generated based on the natural sentences in the natural sentence sequence whose length is not less than the maximum sentence length and the concatenated natural sentences.

8. The method according to claim 7, wherein: The method further comprises: The natural sentences in the natural sentence sequence are matched using the positioning words, and the natural sentences that do not contain the positioning words are filtered out.

9. The method according to claim 1, wherein: The step of inputting the input query and the news event knowledge base into a retrieval model to obtain reference news event knowledge includes: Inputting the input query and the news event knowledge base into a sparse vector retrieval model to obtain first reference news event knowledge; Inputting the input query and the news event knowledge base into a dense vector retrieval model to obtain second reference news event knowledge; The first reference news event knowledge and the second reference news event knowledge are combined and then deduplicated to obtain the reference news event knowledge.

10. A retrieval model training method, comprising: Based on the sample news text set, generate a sample input query set; Matching sample input queries in the sample input query set in pairs to obtain related input queries of the sample input queries; Calculating a matching score between the sample input query and the related input query; Based on the matching scores, construct positive samples and negative samples; A retrieval model is trained based on the positive samples and the negative samples.

11. The method according to claim 10, wherein: The method further comprises: The sample news text is preprocessed, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

12. The method according to claim 10, wherein: The step of generating a sample input query set based on the sample news text set includes: Segmenting the sample news texts in the sample news text set based on punctuation marks to obtain a sample natural sentence sequence; For a sample natural sentence in the sample natural sentence sequence whose length is less than the maximum sentence length, concatenate the sample natural sentences whose length is less than the maximum sentence length using adjacent sample natural sentences to obtain a sample concatenated natural sentence; Based on the sample natural sentences in the sample natural sentence sequence whose length is not less than the maximum sentence length and the sample concatenated natural sentences, a sample input query in the sample input query set is generated.

13. A generative model training method, comprising: Based on the sample news text, generate sample input queries; Based on the sample input query, searching in the news event knowledge base to obtain sample reference news event knowledge; Generate sample prompt information based on the sample reference news event knowledge; Inputting the sample news text and the sample prompt information into a large language model to obtain a predicted news event category of the sample news text; Based on the difference between the predicted news event category and the real news event category of the sample news text, the parameters of the large language model are adjusted to obtain a generation model.

14. The method according to claim 13, wherein: The method further comprises: Multiple news event categories are preset, positive examples and negative examples are set for each news event category, and positioning words are filtered for each news event category to obtain a news event knowledge base.

15. The method according to claim 13, wherein: The method further comprises: The sample news text is preprocessed, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

16. The method according to claim 13, wherein: The step of generating a sample input query based on the sample news text includes: Segmenting the sample marked news text based on punctuation marks to obtain a sample natural sentence sequence; For a sample natural sentence in the sample natural sentence sequence whose length is less than the maximum sentence length, concatenate the sample natural sentences whose length is less than the maximum sentence length using adjacent sample natural sentences to obtain a sample concatenated natural sentence; The sample input query is generated based on the sample natural sentences whose length is not less than the maximum sentence length and the sample concatenated natural sentences in the sample natural sentence sequence.

17. The method according to claim 16, wherein: The method further comprises: The sample natural sentences in the sample natural sentence sequence are matched using the positioning words, and the sample natural sentences that do not contain the positioning words are filtered out.

18. A news event classification device, comprising: A first generation module is configured to generate an input query based on the news text; A retrieval module is configured to input the input query and the news event knowledge base into a retrieval model to obtain reference news event knowledge; A second generating module is configured to generate prompt information based on the reference news event knowledge; The third generation module is configured to input the news text and the prompt information into a generation model to obtain a news event category of the news text.

19. The device according to claim 18, wherein: The second generation module comprises: A first extraction submodule is configured to extract reference news event categories from the reference news event knowledge; A first filling submodule is configured to fill in a candidate category label list in a prompt information template based on the reference news event category; A second extraction submodule is configured to extract the definition, positive examples and negative examples of the reference news event category from the reference news event knowledge; The second filling submodule is configured to fill in the category definition and judgment criteria in the prompt information template based on the definition, positive examples and negative examples of the reference news event category; The third filling submodule is configured to fill in the examples in the prompt information template based on the reference news event knowledge.

20. The device according to claim 19, wherein The third filling submodule includes: A calculation unit, configured to calculate a matching score between the input query and the reference news event knowledge; The selection unit is configured to select example news event knowledge from the reference news event knowledge based on the matching score and write the example into the prompt information template.

21. The device according to claim 20, wherein: The computing unit is further configured to: The input query and the reference news event knowledge are input into a re-ranking model to obtain a matching score between the input query and the reference news event knowledge.

22. The device according to claim 18, wherein The device also includes: The setting module is configured to preset multiple news event categories, set positive examples and negative examples for each news event category, and filter positioning words for each news event category to obtain a news event knowledge base.

23. The device according to claim 18, wherein The device also includes: The preprocessing module is configured to preprocess the news text, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

24. The device according to claim 18, wherein The first generation module is further configured to: Segmenting the news text based on punctuation marks to obtain a natural sentence sequence; For natural sentences in the natural sentence sequence whose length is less than the maximum sentence length, concatenating the natural sentences whose length is less than the maximum sentence length using adjacent natural sentences to obtain a concatenated natural sentence; The input query is generated based on the natural sentences in the natural sentence sequence whose length is not less than the maximum sentence length and the concatenated natural sentences.

25. The apparatus of claim 18, wherein: The device also includes: The matching module is configured to match the natural sentences in the natural sentence sequence using the positioning words and filter the natural sentences that do not contain the positioning words.

26. The device according to claim 18, wherein The retrieval module is further configured to: Inputting the input query and the news event knowledge base into a sparse vector retrieval model to obtain first reference news event knowledge; Inputting the input query and the news event knowledge base into a dense vector retrieval model to obtain second reference news event knowledge; The first reference news event knowledge and the second reference news event knowledge are combined and then deduplicated to obtain the reference news event knowledge.

27. A retrieval model training device, comprising: A generation module is configured to generate a sample input query set based on the sample news text set; A matching module, configured to perform pairwise matching on the sample input queries in the sample input query set to obtain related input queries of the sample input queries; A calculation module, configured to calculate a matching score between the sample input query and the related input query; A construction module, configured to construct positive samples and negative samples based on the matching scores; The training module is configured to train a retrieval model based on the positive samples and the negative samples.

28. The device according to claim 27, wherein The device also includes: The preprocessing module is configured to preprocess the sample news text, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

29. The apparatus of claim 276, wherein: The generation module is further configured to: Segmenting the sample news texts in the sample news text set based on punctuation marks to obtain a sample natural sentence sequence; For a sample natural sentence in the sample natural sentence sequence whose length is less than the maximum sentence length, concatenate the sample natural sentences whose length is less than the maximum sentence length using adjacent sample natural sentences to obtain a sample concatenated natural sentence; Based on the sample natural sentences in the sample natural sentence sequence whose length is not less than the maximum sentence length and the sample concatenated natural sentences, a sample input query in the sample input query set is generated.

30. A generative model training device, comprising: A first generation module is configured to generate a sample input query based on the sample news text; A retrieval module is configured to search in a news event knowledge base based on the sample input query to obtain sample reference news event knowledge; A second generating module is configured to generate sample prompt information based on the sample reference news event knowledge; A classification module, configured to input the sample news text and the sample prompt information into a large language model to obtain a predicted news event category of the sample news text; The adjustment module is configured to adjust the parameters of the large language model based on the difference between the predicted news event category and the real news event category of the sample news text to obtain a generation model.

31. The device according to claim 30, wherein The device also includes: The setting module is configured to preset multiple news event categories, set positive examples and negative examples for each news event category, and filter positioning words for each news event category to obtain a news event knowledge base.

32. The device according to claim 30, wherein: The device also includes: The preprocessing module is configured to preprocess the sample news text, wherein the preprocessing includes at least one of the following: removing stop words, removing non-printing characters, removing garbled characters, converting traditional Chinese to simplified Chinese, restoring word forms, converting full-width to half-width, and correcting spelling.

33. The device according to claim 30, wherein: The first generation module is further configured to: Segmenting the sample marked news text based on punctuation marks to obtain a sample natural sentence sequence; For a sample natural sentence in the sample natural sentence sequence whose length is less than the maximum sentence length, concatenate the sample natural sentences whose length is less than the maximum sentence length using adjacent sample natural sentences to obtain a sample concatenated natural sentence; The sample input query is generated based on the sample natural sentences whose length is not less than the maximum sentence length and the sample concatenated natural sentences in the sample natural sentence sequence.

34. The device according to claim 30, wherein: The device also includes: The matching module is configured to match the sample natural sentences in the sample natural sentence sequence with the positioning words, and filter the sample natural sentences that do not contain the positioning words.

35. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-9 or 10-12 or 13-17.

36. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method of any one of claims 1-9 or 10-12 or 13-17.

37. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-9 or 10-12 or 13-17.

Citation Information

Cited By

  • Topic screening method and device, equipment and medium

    CN122412618A