A news retrieval question-answering method and system based on a large language model of RAG

By leveraging RAG-based large language models for intent understanding, question classification, and multi-source retrieval, the system addresses the issues of incomplete and inaccurate answers in news retrieval question-answering systems, achieving logically clear answers and event context display, thereby enhancing the user experience.

CN120144700BActive Publication Date: 2025-12-02MEMORY TENSOR (SHANGHAI) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510164244.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-12-02
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

Existing news retrieval question-answering systems based on large language models suffer from incomplete, inaccurate, and unclear answers. In particular, when faced with redundant or messy search content, they struggle to accurately identify relevant information and generate logically coherent answers, and they lack a clear display of the event context.

Method used

Using a large language model based on RAG, the intent understanding module determines the question type, classifies and breaks down the questions, recalls relevant materials using multi-source retrieval, and generates progressively layered answers by combining text similarity and large model validation, and displays the development of events through a panoramic timeline.

Benefits of technology

It improves the comprehensiveness and accuracy of the answers, ensures that the generated content is logically clear, can intuitively show the development of events, and enhances users' comprehension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144700B_ABST
    Figure CN120144700B_ABST
Patent Text Reader

Abstract

This invention discloses a news retrieval question-answering method and system based on a large language model (RAG), relating to the field of large language model text generation technology. It performs intent understanding on user questions and provides a news retrieval question-answering method based on a large language model. This method enhances the comprehensiveness of search results and final answers by increasing data sources and expanding the query scope; accurately grasps user intent and correctly identifies and uses parts of the search results relevant to the user's question to improve the accuracy of the answer; answers questions in a progressive manner and provides an intuitive display of the event timeline to help users analyze the development of events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model text generation technology, and in particular to a news retrieval question-answering method and system based on a large language model (RAG). Background Technology

[0002] Large Language Models (LLMs) are deep learning-based artificial intelligence models capable of processing and generating natural language text. Their core technology is neural networks, particularly the Transformer architecture, which, through training on large amounts of text data, can understand and generate human-like language.

[0003] The development of large language models has a long history, from early reliance on rules and statistical methods, to the rise of language models based on Hidden Markov Models, and then to the emergence of Transformer architecture which greatly improves natural language processing capabilities. Models such as Qwen and GPT have made significant progress on multiple natural language processing tasks.

[0004] The mainstream approach for news retrieval and question-answering systems based on large language models involves using a Transformer architecture to further train a base model on the news corpus, building upon an open-source large model. This process includes pre-training and fine-tuning phases, followed by fine-tuning based on the base model to achieve the retrieval and question-answering functionality. However, existing solutions suffer from incomplete, inaccurate, and unclear answers. Retrieval Augmentation (RAG) technology can address these shortcomings. It generates queries based on user questions, retrieves relevant documents from search engines or external knowledge bases, integrates the information, and generates an answer.

[0005] Existing solutions often fail to fully utilize retrieved content, particularly when the retrieved content is excessive or complex, leading to omissions of key information. Due to weak fundamental capabilities, the models' responses heavily rely on the retrieved content, failing to identify irrelevant or unsupported portions, resulting in chronological errors and misattributions. Furthermore, the responses lack logical flow, and the lengthy text generated in extended analysis reports can confuse users. Moreover, most existing models fail to establish a clear timeline of events, hindering users' ability to intuitively analyze the development of events. Summary of the Invention

[0006] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is to provide a news retrieval question-and-answer method based on a large language model, which increases the data source and expands the query scope to enhance the comprehensiveness of the retrieval results and the final answer; accurately grasps the user's intent and correctly identifies and uses the parts of the retrieval results that are relevant to the user's question to improve the accuracy of the answer; answers questions in a progressive manner and provides an intuitive display of the event timeline to help users analyze the development process of the event.

[0007] To achieve the above objectives, this invention provides a news retrieval and question-answering method based on a large language model (RAG), characterized by the following steps:

[0008] The intent of the user's question is understood, and the output of the intent understanding module is divided into three categories: refusal to answer, supplementary information, and direct answer. The model is trained to determine whether the user's question belongs to the news field and whether it is clear and specific enough. Questions that do not belong to the news field are refused to be answered, and supplementary information is provided for questions that are not clear and specific to guide the user to improve the question.

[0009] Questions that are not rejected are categorized to determine whether they are simple questions that can be answered with a single search or complex questions that require intermediate results to answer. Complex questions are further divided into complex questions that can be broken down into multiple independent sub-questions and complex questions with dependent sub-questions. For the sub-questions after they are broken down, they are summarized after being answered to finally answer the user's question. Mind maps are used to clarify the dependencies between questions. If the simple questions after they are broken down are too general, they are further broken down into multiple specific questions through question augmentation.

[0010] After completing the problem decomposition and enhancement, a unified retrieval and recall process is performed for each decomposed simple problem and its enhanced problem. A multi-source retrieval approach is adopted to retrieve more comprehensive search materials from multiple data sources. The most similar search results are selected by text similarity model or keyword filtering. The retrieved text is subjected to strict relevance judgment. In addition to using the results output by the text similarity model, a separate large model for judging relevance is built for further verification. Only when the text similarity reaches the threshold and is judged as relevant by the large model can the search result be added to the candidate pool.

[0011] The search results in the candidate pool are distributed to the questions on the mind map in order of similarity, so that each question has its own reference materials to assist in the answer. These questions and their reference materials are then input into the large question-answering model in order, and the answer of each downstream question is used as the reference material for the upstream question to complete the answer to the original question input by the user.

[0012] For the search results in the candidate pool, the large model is called to extract the event information and add it to the event pool. Then, the similarity discrimination model is used to determine whether these events are related to the user's question. Irrelevant and duplicate events are removed, and the remaining event information is input into the large model to complete the classification and summary of events. The summarized event information is then displayed to the user to help the user intuitively understand the development process of the event.

[0013] Preferably, the intent understanding module determines the domain and specificity of the user's question by training a model, specifically:

[0014] The training model learns the characteristics of the news field and the criteria for judging the clarity of a question;

[0015] The user inputs their question into the trained model, and the model outputs the category to which the question belongs.

[0016] Preferably, the multi-source retrieval method retrieves information from multiple data sources, specifically as follows:

[0017] Simultaneously send retrieval requests to both online databases and self-built databases;

[0018] Receive search results returned by various data sources;

[0019] The search results are integrated and initially screened.

[0020] Preferably, the relevance determination is accomplished jointly by a text similarity model and a separately constructed large model, specifically as follows:

[0021] Calculate the text similarity between the search results and the question;

[0022] The search results are input into a separately constructed large model to determine their relevance.

[0023] Only when the text similarity reaches the threshold and is judged as relevant by the large model, will the search result be added to the candidate pool.

[0024] Preferably, the specific steps for the panoramic timeline module to display the development of events are as follows:

[0025] Event information is extracted from the search results in the candidate pool and added to the event pool;

[0026] A similarity-based discrimination model is used to remove events from the event pool that are irrelevant to the user's question.

[0027] Duplicate events in the event pool are removed using a similarity discrimination model;

[0028] The remaining event information is input into a large model for classification and summarization;

[0029] The summarized event information is displayed to the user.

[0030] Another aspect of the present invention is a news retrieval and question-answering system based on a large language model (RAG), characterized in that it includes:

[0031] The intent understanding module is used to understand the intent of user questions and categorizes its output into three types: refuse to answer, supplement, and direct answer. By training the model, it determines whether the user's question belongs to the news field and whether it is clear and specific enough. For questions that do not belong to the news field, it refuses to answer and provides supplementary information for questions that are not clear and specific.

[0032] The question classification module is used to categorize unrejected questions, determine whether they are simple or complex, further subdivide complex questions, and summarize the answers to user questions after answering the sub-questions. Mind maps are used to clarify the dependencies between questions, and question enhancement is performed for general simple questions.

[0033] The retrieval and recall module is used to perform multi-source retrieval and recall for each of the simple questions and their enhanced questions after the question is split and enhanced. It retrieves information from multiple data sources, selects the most similar search results through text similarity model and keyword filtering, and performs strict relevance judgment on the search results, adding the search results that meet the conditions to the candidate pool.

[0034] The question-answering module is used to distribute the search results in the candidate pool to the questions on the mind map, input the questions and their reference materials into the large model responsible for question answering, and use the answers to downstream questions as reference materials for upstream questions to complete the answer to the original question.

[0035] The panoramic timeline module is used to extract event information from the candidate pool retrieval results. After relevance and repetition judgment, the remaining event information is input into the large model for classification and summarization, and the development of the event is displayed to the user.

[0036] Preferably, the training model in the intent understanding module learns news domain features and question clarity judgment criteria to determine the category to which the user's question belongs.

[0037] Preferably, the multi-source retrieval in the retrieval and recall module simultaneously retrieves information from online databases and self-built databases and integrates and filters it.

[0038] Preferably, the relevance determination in the retrieval and recall module is accomplished jointly by a text similarity model and a separately constructed large model, ensuring the relevance of the retrieval results added to the candidate pool.

[0039] Preferably, the panoramic timeline module visually displays the development process of events by extracting event information, determining relevance and repetition, and classifying and summarizing them.

[0040] Beneficial technical effects of the present invention:

[0041] This invention, through question splitting, question enhancement, and multi-path recall, can more comprehensively recall relevant materials, and the split sub-questions also make the answers more detailed.

[0042] The intent understanding module of this invention provides differentiated processing for different types of questions, improving the accuracy of material usage; the distribution and relevance judgment of retrieval results, as well as citation generation and hallucination mitigation algorithms, further improve the accuracy of responses.

[0043] This invention uses mind maps to break down user questions and answers them in a progressive manner. It also introduces a panoramic timeline generation module, allowing users to more intuitively understand the development of events.

[0044] The following will further explain the concept, specific structure, and technical effects of the present invention in conjunction with the accompanying drawings, so as to fully understand the purpose, features, and effects of the present invention. Attached Figure Description

[0045] Figure 1 This is a system flowchart of a preferred embodiment of the present invention;

[0046] Figure 2 This is a mind map question-and-answer flowchart of the present invention;

[0047] Figure 3 This is a flowchart of the multi-source retrieval process of the present invention;

[0048] Figure 4 This is a flowchart of the correlation determination process of the present invention;

[0049] Figure 5 This is a flowchart of the timeline generation process of the present invention. Detailed Implementation

[0050] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.

[0051] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.

[0052] This invention provides a news retrieval question and answer system based on a large language model (RAG), which increases data sources and expands the query scope to enhance the comprehensiveness of search results and final answers; accurately grasps user intent and correctly identifies and uses the parts of the search results relevant to the user's question to improve the accuracy of the answer; answers questions in a progressive manner and provides an intuitive display of the event timeline to help users analyze the development of events.

[0053] This invention proposes a framework for news retrieval and question answering, as shown in the attached figure. In this framework, the user's question is first analyzed to understand their intent. The output of the intent understanding module is divided into three categories: rejection, supplementary information, and direct answer. By training the model, it is made capable of determining whether the user's question belongs to the news domain and whether it is sufficiently specific. For questions that do not belong to the news domain, the answer will be rejected; for questions that are not specific, supplementary information will be provided to guide the user to refine their question.

[0054] Next, the questions that were not rejected are categorized into simple questions that can be answered in a single search, and complex questions that require intermediate results to answer. Complex questions are further divided into complex questions that can be broken down into multiple independent sub-questions and complex questions with dependent sub-questions. After answering the decomposed sub-questions, the model summarizes these answers to finally answer the user's question. A Graph of Thought (GoT) is used to clarify the dependencies between these questions. In addition, sometimes the decomposed simple questions are too general, such as summarizing the sequence of events or summarizing a document, so question augmentation is used to further decompose them into multiple specific questions.

[0055] After problem decomposition and enhancement, a unified retrieval and recall process is performed for each decomposed simple problem and its enhanced version. A multi-source retrieval approach is adopted, searching from multiple data sources including Baidu, Bing, and a self-built database to retrieve more comprehensive search materials. The most similar search results are selected using text similarity models such as BGE and Qwen, as well as keyword filtering. However, if the retrieved results are irrelevant to the questions on the mind map, the model may exhibit significant misleading behavior; therefore, strict relevance assessment of the retrieved text is necessary. In this step, in addition to using the results output by the aforementioned BGE text similarity model, a separate large-scale model for relevance assessment is constructed for further validation. Only search results whose BGE text similarity reaches a threshold and are judged as relevant by the large-scale model are added to the candidate pool.

[0056] At this point, we have several search results in the candidate pool and several questions on the mind map. Next, the search results in the candidate pool are distributed to the questions on the mind map according to their similarity, so that each question has its own reference materials to aid in the answer. These questions and their reference materials are then sequentially input into the large question-answering model, and the answer to each downstream question is used as reference material for the upstream question, thus completing the answer to the original question input by the user.

[0057] At this point, the user's question has been answered. However, for users who want a clear understanding of the event's development, displaying the information in long text format is clearly not intuitive enough. To address this issue, a panoramic timeline module is introduced. This module presents events related to the user's question in a panoramic timeline format. Specifically, for the search results in the candidate pool, a large model is used to extract event information (including but not limited to time, location, people, and event descriptions) and add it to the event pool. Then, a similarity discrimination model is used to determine whether these events are relevant to the user's question, removing irrelevant events from the event pool. A similarity discrimination model is then used again to remove duplicate events from the event pool. Finally, the remaining event information is input into the large model to classify and summarize the events, and the summarized event information is displayed to the user, helping them intuitively understand the event's development.

[0058] The present invention will now be illustrated through specific embodiments:

[0059] Example 1: Single Event News Retrieval Question and Answer

[0060] Example Scenario

[0061] Suppose a user is interested in the news of a company releasing a new product and wants to know detailed information about the product, its market reaction, and its impact on the company's future development.

[0062] Intention understanding stage

[0063] The user inputs the question: "How is the new product released by Company X? What impact will it have on the company's development?" The intent understanding module in the system inputs the question into the trained model. After learning the characteristics of the news domain and the criteria for judging question clarity, the model determines that the question belongs to the news domain and is clear and specific, so it can be answered directly without involving refusal to answer or supplementary information.

[0064] Problem classification stage

[0065] This problem was deemed complex because it involves both the product itself and its impact on the company's development, and these two aspects are somewhat interdependent (the product's situation may affect the company's development). The model uses a GoT (Go of Thought) to clarify the relationship between the problems, breaking them down into two sub-problems: "Detailed information about a company's newly released product (including features, characteristics, etc.)" and "The impact of this product's release on the company's development (such as changes in market share, financial impact, etc.)."

[0066] Search and recall phase

[0067] For the two sub-questions after decomposition, the retrieval and recall module employs a multi-source retrieval approach. Relevant information is searched from Baidu, Bing, and a self-built database. Assuming 100 relevant news articles are retrieved from Baidu, 80 from Bing, and 50 from the self-built database, the similarity between these search results and the sub-questions is calculated using a BGE text similarity model, combined with keyword filtering, such as setting keywords like "company name," "new product name," "product function," "market reaction," and "company development." After initial screening, Baidu retains 30 of the most similar results, Bing retains 25, and the self-built database retains 20. These initially filtered results are then input into a separately constructed large-scale relevance-based model for further validation. With a BGE text similarity threshold of 0.7, the large-scale model ultimately adds 20 search results from Baidu, 18 from Bing, and 15 from the self-built database to the candidate pool.

[0068] Q&A Phase

[0069] The search results in the candidate pool are distributed to the corresponding two sub-questions based on their similarity. For example, search results highly relevant to product details are assigned to the first sub-question, while those highly relevant to the product's impact on the company's development are assigned to the second sub-question. These questions and their reference materials are then sequentially input into the large-scale question-answering model. The large-scale model answers the first sub-question based on the reference materials, such as "A company's newly released product has innovative features X and Y, and performs well." This answer serves as one of the reference materials for answering the second sub-question. It then answers the second sub-question, such as "Due to the product's innovative features and good performance, it has attracted significant consumer attention, which is expected to increase the company's market share, positively impact the company's future financial situation, and drive further development." Finally, the answers to both sub-questions are combined to answer the user's original question: "A company's newly released product has innovative features X and Y, and performs well. Because it has attracted significant consumer attention, it is expected to increase the company's market share, positively impact the company's future financial situation, and drive further development."

[0070] Panoramic Timeline Stage

[0071] For the search results in the candidate pool, the panoramic timeline module calls the large model to extract event information, such as product launch time and location, and the speech content of relevant company personnel at the launch event. Assuming 10 event information items are extracted, a similarity discrimination model is used to determine the relevance of these events to the user's question, eliminating 3 irrelevant events. Then, another similarity discrimination model is used to eliminate 2 duplicate events. The remaining 5 event information items are input into the large model for classification and summarization, for example, arranged in chronological order, and displayed to the user, such as "[Time 1] A company launched a new product at [Location 1], and [Responsible Person's Name] introduced the product's innovative feature X. [Time 2] Market research institutions began evaluating the product. [Time 3] Consumers showed strong interest in the product, and pre-orders exceeded expectations, etc.", helping users intuitively understand the development process of the events.

[0072] Example 2: News Search Q&A for a Series of Events

[0073] Example Scenario

[0074] Taking a series of events in a sports event as an example, users want to know about an athlete's performance in this event, including the results of each game, key events, and their impact on their career.

[0075] Intention understanding stage

[0076] The user inputs the question: "How did a certain athlete perform in this competition? What impact will it have on their career?" The intent understanding module inputs the question into the model, and the model determines that the question belongs to the news field and is specific and can be answered directly.

[0077] Problem classification stage

[0078] This is a complex problem involving an athlete's performance in multiple matches within a competition and its impact on their career. The model uses a GoT (Go of Thinking) approach to break it down into multiple sub-problems, such as "An athlete's performance and key events in the first match of the competition," "An athlete's performance and key events in the second match of the competition," and so on, "The impact of an athlete's overall performance in this competition on their career (e.g., ranking changes, increased commercial value)," etc. Furthermore, since the question "Impact of overall performance on career" is relatively general, it is further broken down into specific questions such as "How does competition performance affect their ranking in the sports world?" and "How does competition performance increase their commercial endorsement value?" through question augmentation.

[0079] Search and recall phase

[0080] For each sub-question and its augmentation question, the retrieval and recall module performs multi-source searches. Relevant information is retrieved from Baidu, Bing, and a self-built database. Assuming that for each sub-question, an average of 120 results are retrieved from Baidu, 100 from Bing, and 60 from the self-built database. Using a BGE text similarity model and keyword filtering, Baidu initially retains 40 results, Bing retains 30, and the self-built database retains 25. After further large-scale model validation, with a BGE text similarity threshold of 0.65, an average of 30 results from Baidu, 25 from Bing, and 20 from the self-built database are added to the candidate pool.

[0081] Q&A Phase

[0082] The search results in the candidate pool are distributed to various sub-questions based on similarity. For example, search results related to the first match are assigned to the sub-question "An athlete's performance and key events in the first match of the event." The main model answers based on reference materials, such as "In the first match, an athlete achieved [specific results], and at a key moment [description of key events]." After answering each sub-question in turn, these answers are summarized to answer the user's question about the athlete's performance in the event. For the augmented question "How does event performance affect their ranking in the sports world?", the main model answers based on previous answers about match results, combined with reference materials, such as "Due to their outstanding performance in this event, their ranking in the sports world is expected to rise by [X] places." Other augmented questions are answered in the same way, and finally, the answers to the user's question about the impact of event performance on their career are summarized to answer the user's original question as a whole.

[0083] Panoramic Timeline Stage

[0084] Event information is extracted from the candidate pool search results, such as the time, location, opponent, and highlights of each match. Assuming 20 event information items are extracted, a similarity discrimination model is used to remove 5 irrelevant events, and then 3 duplicate events are removed. The remaining 12 event information items are input into a larger model for classification and aggregation, and displayed to the user in chronological order of the matches, such as "[Time 1] An athlete played the first match against [Opponent 1] at [Location 1], achieving [Score 1], and [Highlight 1] during the match. [Time 2] The second match took place...", helping users clearly understand the development of the events.

[0085] Example 3: Comprehensive News Event Retrieval Questions and Answers

[0086] Example Scenario

[0087] Consider a comprehensive news event involving economic, political, and social factors, such as a new policy introduced in a certain region. Users want to know the specific content of the policy, its impact on local economic development, the reactions from all sectors of society, and its connection with other relevant policies.

[0088] Intention understanding stage

[0089] The user inputs the question: "What are the specific details of a newly introduced policy in a certain region, its impact on the economy and society, and its relationship with other policies?" The intent understanding module inputs the question into the model, and the model determines that the question belongs to the news field and is clear and specific, so it can be answered directly.

[0090] Problem classification stage

[0091] This is a complex problem. The model uses a GoT (Go-of-Touch) mind map to break it down into the following sub-problems: "The specific content (clauses, objectives, etc.) of a newly introduced policy in a certain region," "The impact of this policy on local economic development (such as industrial restructuring, changes in employment, etc.)," ​​"The reactions of various sectors of society (businesses, the public, etc.) to this policy," and "The relationship between this policy and other related policies (such as previous similar policies or supporting policies)." The sub-problem "The reactions of various sectors of society to this policy" is relatively general, so it is further broken down into specific questions such as "The attitudes and responses of the business community to this policy" and "The views and expectations of the public regarding this policy."

[0092] Search and recall phase

[0093] For each sub-question and its enhancement, the retrieval and recall module performs multi-source searches. Information is searched from Baidu, Bing, and a self-built database. Assuming 200 relevant news articles are retrieved from Baidu, 150 from Bing, and 80 from the self-built database, using a BGE text similarity model and keyword filtering, Baidu initially retains 60, Bing retains 50, and the self-built database retains 30. After further large-scale model validation, with a BGE text similarity threshold of 0.75, 45 results from Baidu, 40 from Bing, and 25 from the self-built database are added to the candidate pool.

[0094] Q&A Phase

[0095] The search results in the candidate pool are distributed to corresponding sub-questions based on similarity. For example, search results related to the specific content of a policy are assigned to the sub-question "the specific content (clauses, objectives, etc.) of a newly introduced policy in a certain region," and the large model answers as "the main clauses of this policy include [listing clauses], and the objectives are [elaborating on the objectives]." Other sub-questions are answered sequentially. For example, for "the attitudes and countermeasures of the business community towards this policy," the large model answers based on reference materials as "the business community generally supports this policy, and some companies plan to [list corporate countermeasures]." All sub-question answers are summarized to answer the user's original question, such as "the specific content of a newly introduced policy in a certain region is [detailed content], its impact on local economic development will be [description], the business community's [attitudes and measures], the public's [views and expectations], and the [association description] of this policy with other policies."

[0096] Panoramic Timeline Stage

[0097] Event information is extracted from the candidate pool search results, such as the policy release date, background information of the press conference, and the timing of feedback from various sectors of society at different stages. Assuming 15 event information items are extracted, a similarity discrimination model is used to remove 4 irrelevant events and 2 duplicate events. The remaining 9 event information items are input into a larger model for classification and summarization, and displayed to the user in chronological order, such as "[Time 1] A local government held a press conference to announce a new policy, [Key points of the official's speech]. [Time 2] Some enterprise representatives expressed their views on the policy...", helping users intuitively grasp the development process of the events.

[0098] Example 4: Example of a News Retrieval and Question-Answering System for Emerging Events

[0099] Example Scenario

[0100] Suppose a natural disaster, such as an earthquake, occurs in a certain location. Users want to quickly obtain detailed information about the earthquake, including the time of occurrence, magnitude, affected area, rescue progress, and impact on the lives of local residents.

[0101] Intention understanding stage

[0102] The user inputs the question: "What are the specific details of the earthquake that occurred in a certain area? How is the rescue work progressing? What impact has it had on the lives of local residents?" The intent understanding module inputs this question into the trained model. Based on the learned news domain characteristics and question clarity judgment criteria, the model quickly determines that the question belongs to the news domain and is sufficiently clear and specific, requiring no refusal to answer or supplementary information, and can proceed directly to subsequent processing.

[0103] Problem classification stage

[0104] This problem was identified as complex because it encompasses multiple aspects of the earthquake itself and its subsequent impacts, with certain logical connections between these aspects. The model uses a GoT (Go of Thinking) approach to break it down into the following sub-problems: "Precise time of the earthquake," "Earthquake magnitude," "Specific affected area (including which regions, towns, etc.)," ​​"Current progress of rescue efforts (deployment of rescue teams, transportation of relief supplies, etc.)," ​​and "Impact of the earthquake on local residents' lives (housing, water and electricity supply, daily life, etc.)." Among these, the questions "Specific affected area" and "Impact of the earthquake on local residents' lives" are relatively general and are further refined through question enhancement into specific questions such as "List of major affected towns and villages," "Damage to houses and resettlement of residents in affected areas," and "Areas with water and electricity supply disruptions and estimated restoration time."

[0105] Search and recall phase

[0106] For each sub-question and its enhanced questions, the retrieval and recall module initiates a multi-source retrieval mechanism. Retrieval requests are sent to Baidu, Bing, and a self-built database to obtain relevant information. For example, for the sub-question "the precise time of the earthquake," 50 relevant news reports are retrieved from Baidu, 40 from Bing, and 20 from the self-built database. The similarity between these search results and the sub-question is calculated using a BGE text similarity model, and filtered using keywords such as "location," "earthquake," and "time of occurrence." After initial screening, Baidu retains 20 of the most similar results, Bing retains 15, and the self-built database retains 10. These initially filtered results are then input into a specially constructed large-scale relevance-based model for further validation. With a BGE text similarity threshold of 0.8, the large-scale model determines that 15 results from Baidu, 12 from Bing, and 8 from the self-built database are added to the candidate pool. The same retrieval and recall process is applied to other sub-questions to ensure that the most relevant references are obtained for each question.

[0107] Q&A Phase

[0108] The search results in the candidate pool are distributed to the corresponding sub-questions according to their similarity. For example, the search results with the highest relevance to the earthquake's occurrence time are assigned to the sub-question "the exact time of the earthquake," and the large model provides an accurate answer based on these reference materials, such as "the earthquake occurred at [specific time]." For the sub-question "the current progress of rescue work," the large model answers by comprehensively considering the reference materials, "[X] rescue teams have been dispatched to the disaster area, and relief supplies are being transported to the severely affected [specific areas], and [X] temporary resettlement sites have been set up." This process continues, answering each sub-question in turn, and then summarizing the answers to each sub-question. For example, "The earthquake had a magnitude of [X], occurred at [specific time], and affected areas include [list the main affected towns and villages]. Rescue work is currently proceeding in an orderly manner, multiple rescue teams have been dispatched, material transportation is progressing, and some residents in the affected areas have been resettled. The earthquake caused extensive damage to houses, and water and electricity supplies were interrupted in some areas; the estimated recovery time is [estimated]," thus providing a complete answer to the user's original question.

[0109] Panoramic Timeline Stage

[0110] The panoramic timeline module extracts various event information from the candidate pool's search results, including the immediate situation of the earthquake, the arrival time of rescue teams, and the phased progress of material distribution. Assuming 25 event information items are extracted, a similarity discrimination model is used to eliminate 6 irrelevant events, such as earthquake events unrelated to other regions or other information unrelated to this earthquake. Next, the similarity discrimination model eliminates 4 duplicate events, such as repeated reports of the same rescue team's departure. The remaining 15 event information items are input into a larger model for classification and summarization, clearly displayed to the user in chronological order, such as "[Time 1] Earthquake occurs, with noticeable tremors felt in many areas. [Time 2] Local government activates emergency response, rescue teams begin to assemble. [Time 3] The first rescue team arrives at the most severely affected [area name] and begins rescue work..." This helps users intuitively understand the entire development process of the earthquake event from its occurrence to the rescue efforts, allowing for a better grasp of the overall picture.

[0111] Intent Understanding Module Algorithm

[0112] Model selection and training

[0113] Fine-tuning is performed using a neural network model based on the Transformer architecture, such as the Qwen model.

[0114] The training data consists of a large number of labeled questions in both the news and non-news domains. The labels include the domain to which the question belongs (news / non-news) and the specificity of the question (specific / unspecific).

[0115] The training objective is to minimize the cross-entropy loss function between the predicted and labeled results. The Adam optimizer can be used, with a learning rate of 0.0001 and 10 training epochs.

[0116] Intent determination process

[0117] The system preprocesses user-input questions, including word segmentation and stop word removal.

[0118] The preprocessed question is input into the trained Qwen model to obtain the vector representation of the question.

[0119] The vector representation is classified through a fully connected layer to determine the category of the question (e.g., refusing to answer if the question belongs to mathematics or coding, and providing supplementary information if details are missing).

[0120] Problem classification module algorithm

[0121] Problem classification model

[0122] Construct a classification model based on decision trees.

[0123] Feature selection includes the keywords of the question, the length of the question, and the sentence structure of the question.

[0124] The training data learns the classification patterns of questions under different feature combinations. The training data consists of a large number of classified news questions (simple questions, complex questions, and sub-types of complex questions).

[0125] Problem decomposition and augmentation algorithms

[0126] For complex problems, use rule-based decomposition methods. For example, break the problem down into subproblems based on specific keywords (such as "and", "as well as", "impact on", etc.).

[0127] For general and simple questions, an enhanced approach using template matching is employed. For example, for questions like "outlining the sequence of events," a preset template is matched to generate specific questions such as "the time, place, and main characters of the event" and "key milestones in the development of the event."

[0128] Retrieval and Recall Module Algorithm

[0129] Multi-source retrieval algorithm

[0130] For search engines such as Baidu and Bing, use their provided API interfaces to send search requests. Request parameters include keywords (generated from the question), search scope (news category), and time range (set according to the timeliness of the news, such as the past week).

[0131] For self-built databases, use full-text search technology, such as the Lucene-based search engine, to search for news document titles, body texts, and other fields in the database based on keywords in the question.

[0132] Correlation discrimination algorithm

[0133] First, the BGE text similarity model is used to calculate the similarity between the search results and the question, using the formula: where is the question vector and is the search result vector. A similarity threshold of 0.7 is set.

[0134] Simultaneously, the search results are input into a separately constructed large-scale model based on Qwen for relevance assessment. The input to the large-scale model is the search results and the question, and the output is a relevance or irrelevance judgment. Only search results that simultaneously meet the BGE text similarity threshold and are judged as relevant by the large-scale model are added to the candidate pool.

[0135] Question answering module algorithm

[0136] Answer generation model

[0137] The Qwen2-72B-Insturct model was used as the base model for fine-tuning.

[0138] The training data consists of pairs of questions and their corresponding correct answers. Through supervised learning, the model learns to generate accurate answers based on the questions.

[0139] During fine-tuning, the optimization objective is to minimize the edit distance between the generated answer and the reference answer. The optimizer uses Adagrad with a learning rate of 0.001 and 5 training epochs.

[0140] Answering process

[0141] After sorting the search results in the candidate pool according to similarity, they are combined with the question in turn to form an input sequence, which is then fed into the fine-tuned Qwen2-72B-Insturct model.

[0142] The model generates answers based on the input sequence. For sub-questions with dependencies, the answers to downstream questions are used as additional input information for the answers to upstream questions to generate more accurate final answers.

[0143] Panoramic Timeline Module Algorithm

[0144] Event extraction model

[0145] An event extraction model was built based on Qwen2-72B-Insturct.

[0146] The training data consists of news documents and manually annotated event information (including time, location, people, event description, etc.). The system learns the ability to identify and extract event-related information from text through the annotated data.

[0147] The model uses cross-entropy loss as its loss function, RMS Prop as its optimizer, and a learning rate of 0.0005. Training continues until the loss function converges.

[0148] Event handling and display algorithms

[0149] After extracting event information from the candidate pool retrieval results, an irrelevant event is removed using a cosine similarity-based discrimination model, with a threshold set to 0.6.

[0150] For duplicate events, deduplication is performed by comparing key information about the event (such as time, location, and core content of the event).

[0151] The processed event information is input into a hierarchical clustering-based classification and summarization algorithm, which classifies the events according to their time, theme, and other attributes. The information is then displayed to the user in a timeline format: [Time] - [Event Description].

[0152] Example 5: Example of a Web-based News Retrieval and Question Answering Software

[0153] Example Scenario

[0154] Develop a web-based news retrieval and question-and-answer software. Users can access the software through a web browser, enter questions about various news events, such as new product launches in the technology field, celebrity news in the entertainment industry, and changes in the international political situation. The software can quickly and accurately provide comprehensive, clear, and accurate answers, and present the event timeline in an intuitive way.

[0155] Software architecture and module design

[0156] Front-end Interface: The user interface is built using HTML, CSS, and JavaScript. A simple and intuitive input box is designed for easy question input. Areas are reserved for displaying answer results and a comprehensive timeline, presented to the user with clear layout. For example, answer results are presented in paragraph format, the comprehensive timeline is displayed as a visual timeline chart, and event nodes are indicated by different colors to indicate their importance.

[0157] Backend server: The backend server is built using the Flask framework in Python. It is responsible for receiving user questions sent from the frontend, processing them by calling the intent understanding module, question classification module, retrieval and recall module, question answering module, and panoramic timeline module, and returning the final result to the frontend.

[0158] Database: MongoDB is chosen to store the self-built news corpus, user question records, and intermediate data from model training. Elasticsearch is used to create an index to improve the search efficiency of the retrieval module.

[0159] Implementation of the intent understanding module in software

[0160] The trained intent understanding model (fine-tuned based on Qwen) is deployed on the backend server. When a user enters a question on the frontend, the backend sends the question to the intent understanding model. After preprocessing (word segmentation, stop word removal, etc.), the model outputs the intent classification result of the question (refusal to answer, supplementary information, direct answer). If it is a supplementary information request, the backend generates a corresponding prompt and returns it to the frontend to guide the user to complete the question; if it is a refusal to answer, a friendly rejection prompt is returned; if it is a direct answer, the subsequent process continues.

[0161] Implementation of the problem classification module in the software

[0162] The decision tree-based question classification model is deployed on the backend. Upon receiving a question that can be answered directly, the backend inputs the question into the classification model. The model determines the question type (simple, complex, or sub-types of complex questions) based on features such as keywords, length, and sentence structure. For complex questions, they are broken down according to rules, and the sub-questions are numbered and their dependencies are recorded (using a GoT mind map data structure stored in memory). For general simple questions, template matching is used to enhance the question, generating specific sub-questions, which are also numbered and their relationships recorded.

[0163] Implementation of the retrieval and recall module in the software

[0164] For the multi-source retrieval part, the backend sends requests to the news search APIs of Baidu and Bing using Python network request libraries (such as requests), while simultaneously performing searches in a self-built database using Elasticsearch. After obtaining the search results, the BGE text similarity model (implemented using Python deep learning libraries such as PyTorch) is used to calculate similarity, combined with keyword filtering. The initially filtered results are then input into a large model based on Qwen (deployed on a backend server) for relevance verification. Search results that meet the criteria (BGE text similarity reaches a threshold and is judged as relevant by the large model) are added to the candidate pool and stored in memory in a specific data structure (such as a list nested dictionary, where the dictionary contains the search result content, similarity score, etc.).

[0165] Implementation of the question-and-answer module in the software

[0166] The finely tuned Qwen2-72B-Instruct model is deployed on the backend server. After sorting the search results in the candidate pool by similarity, the backend sequentially combines each sub-question and its corresponding search results into an input sequence and sends it to the GPT-3.5 model. After the model generates an answer, the backend uses the answers to downstream questions as reference information for the answers to upstream questions, based on the dependencies between questions, to progressively construct the final answer. Finally, the complete answer is returned to the frontend and presented to the user in the answer display area of ​​the frontend interface.

[0167] Implementation of the panoramic timeline module in the software

[0168] An event extraction model based on Qwen2-72B-Instruct is deployed on the backend. The backend extracts text content from the candidate pool, inputs it into the event extraction model, and obtains event information (time, location, people, event description, etc.). A cosine similarity discrimination model (implemented using BGE text similarity model) is used to eliminate irrelevant events, and then duplicates are removed by comparing key information. The processed event information is then input into a hierarchical clustering-based classification algorithm (implemented using a Python data analysis library such as Scikit-learn) for classification. Finally, the backend returns the classified and summarized event information to the frontend in a timeline data format (e.g., a JSON array, where each element contains attributes such as time and event description). The frontend uses a JavaScript visualization library (e.g., D3.js) to draw and display the timeline in a panoramic timeline display area, allowing users to interact with the event details using the mouse.

[0169] Software operation flow example

[0170] Users enter the question on the front end: "What new features does a certain mobile phone company's latest phone have? What impact will it have on the mobile phone market?"

[0171] After receiving the question, the backend first uses the intent understanding module to determine if the question can be answered directly.

[0172] The problem classification module identifies it as a complex problem, breaks it down into two sub-problems: "New features of a mobile phone released by a certain mobile phone company" and "Impact of the release of this mobile phone on the mobile phone market", and records the dependencies.

[0173] The retrieval and recall module performs multi-source searches, obtaining relevant news reports from Baidu, Bing, and its own databases. After relevance assessment, suitable search results are added to the candidate pool.

[0174] The question-and-answer module generates answers to sub-questions based on the candidate pool, such as "A mobile phone company's latest mobile phone has new features A, B, C, etc." and "After the release of this mobile phone, its market share is expected to change to some extent, putting pressure on competitors, etc." The final answers are then compiled and returned to the front end.

[0175] The panoramic timeline module extracts event information, such as "[Release Time] A mobile phone company releases a mobile phone, [Release Location], [Speaker introduces new features]", etc., and displays it on the front end in the form of a timeline, so that users can clearly see the development of the event.

[0176] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A news retrieval and question-answering method based on a large language model (RAG), characterized in that, Includes the following steps: The intent of the user's question is understood, and the output of the intent understanding module is divided into three categories: refusal to answer, supplementary information, and direct answer. The model is trained to determine whether the user's question belongs to the news field and whether it is clear and specific enough. Questions that do not belong to the news field are refused to be answered, and supplementary information is provided for questions that are not clear and specific to guide the user to improve the question. Questions that are not rejected are categorized to determine whether they are simple questions that can be answered with a single search or complex questions that require intermediate results to answer. Complex questions are further divided into complex questions that can be broken down into multiple independent sub-questions and complex questions with dependent sub-questions. For the sub-questions after they are broken down, they are summarized after being answered to finally answer the user's question. Mind maps are used to clarify the dependencies between questions. If the simple questions after they are broken down are too general, they are further broken down into multiple specific questions through question augmentation. After completing the problem decomposition and enhancement, a unified retrieval and recall process is performed for each decomposed simple problem and its enhanced problem. A multi-source retrieval approach is adopted to retrieve more comprehensive search materials from multiple data sources. The most similar search results are selected by text similarity model or keyword filtering. The retrieved text is subjected to strict relevance judgment. In addition to using the results output by the text similarity model, a separate large model for judging relevance is built for further verification. Only when the text similarity reaches the threshold and is judged as relevant by the large model can the search result be added to the candidate pool. The search results in the candidate pool are distributed to the questions on the mind map in order of similarity, so that each question has its own reference materials to assist in the answer. These questions and their reference materials are then input into the large question-answering model in order, and the answer of each downstream question is used as the reference material for the upstream question to complete the answer to the original question input by the user. For the search results in the candidate pool, the large model is called to extract the event information and add it to the event pool. Then, the similarity discrimination model is used to determine whether these events are related to the user's question. Irrelevant and duplicate events are removed, and the remaining event information is input into the large model to complete the classification and summary of events. The summarized event information is then displayed to the user to help the user intuitively understand the development process of the event.

2. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that, The intent understanding module determines the domain and specificity of the user's question by training a model, specifically: The training model learns the characteristics of the news field and the criteria for judging the clarity of a question; The user inputs their question into the trained model, and the model outputs the category to which the question belongs.

3. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that, The multi-source retrieval method retrieves information from multiple data sources, specifically: Simultaneously send retrieval requests to both online databases and self-built databases; Receive search results returned by various data sources; The search results are integrated and initially screened.

4. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that, The relevance determination is accomplished jointly by a text similarity model and a separately constructed large model, specifically as follows: Calculate the text similarity between the search results and the question; The search results are input into a separately constructed large model to determine their relevance. Only when the text similarity reaches the threshold and is judged as relevant by the large model, will the search result be added to the candidate pool.

5. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that, The specific steps for displaying the development of events using the panoramic timeline module are as follows: Event information is extracted from the search results in the candidate pool and added to the event pool; Use a similarity discrimination model to remove events from the event pool that are irrelevant to the user's question; Duplicate events in the event pool are removed using a similarity discrimination model; Input the remaining event information into the large model for classification and summarization; The summarized event information is displayed to the user.

6. A news retrieval and question-answering system based on a large language model using RAG, characterized in that, include: The intent understanding module is used to understand the intent of user questions and categorizes its output into three types: refuse to answer, supplement, and direct answer. By training the model, it determines whether the user's question belongs to the news field and whether it is clear and specific enough. For questions that do not belong to the news field, it refuses to answer and provides supplementary information for questions that are not clear and specific. The question classification module is used to categorize unrejected questions, determine whether they are simple or complex, further subdivide complex questions, and summarize the answers to user questions after answering the sub-questions. Mind maps are used to clarify the dependencies between questions, and question enhancement is performed for general simple questions. The retrieval and recall module is used to perform multi-source retrieval and recall for each of the simple questions and their enhanced questions after the question is split and enhanced. It retrieves information from multiple data sources, selects the most similar search results through text similarity model and keyword filtering, and performs strict relevance judgment on the search results, adding the search results that meet the conditions to the candidate pool. The question-answering module is used to distribute the search results in the candidate pool to the questions on the mind map, input the questions and their reference materials into the large model responsible for question answering, and use the answers to downstream questions as reference materials for upstream questions to complete the answer to the original question. The panoramic timeline module is used to extract event information from the candidate pool retrieval results. After relevance and repetition judgment, the remaining event information is input into the large model for classification and summarization, and the development of the event is displayed to the user.

7. The RAG-based large language model news retrieval and question-answering system according to claim 6, characterized in that, The training model in the intent understanding module learns news domain features and question clarity judgment criteria to determine the category to which the user's question belongs.

8. The RAG-based large language model news retrieval and question-answering system according to claim 6, characterized in that, The multi-source retrieval module of the retrieval and recall module simultaneously retrieves information from online databases and self-built databases and integrates and filters it.

9. The RAG-based large language model news retrieval and question-answering system according to claim 6, characterized in that, The relevance determination in the retrieval and recall module is accomplished jointly by a text similarity model and a separately constructed large model, ensuring the relevance of the retrieval results added to the candidate pool.

10. The RAG-based large language model news retrieval and question-answering system according to claim 6, characterized in that, The panoramic timeline module extracts event information, determines relevance and repetition, and categorizes and summarizes it to intuitively display the development process of events.

Citation Information

Patent Citations

  • Question and answer method and device, equipment and storage medium

    CN118035415A

  • Knowledge retrieval enhancement generation method and system based on large language model

    CN118394890A