News retrieval question and answer method and system of big language model based on RAG

By introducing RAG-based methods into the large language model news retrieval question and answer system, the existing system has solved the problems of incomplete, inaccurate and unclear answers, and more comprehensive, accurate and logically clear answers are achieved, and an intuitive display of event development context is provided.

CN120144700AActive Publication Date: 2025-06-13MEMORY TENSOR (SHANGHAI) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510164244.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-13
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The existing large language model news retrieval question and answer system has problems such as incomplete, inaccurate and unclear answer results. Especially when the search content is too much or too complicated, it is difficult to effectively utilize the retrieved content, resulting in lack of logic and unclear event context.

Method used

The RAG-based news search question-and-answer method is used to judge the domain and clarity of user problems through the intention understanding module, classify and split the questions, and use multi-source search to recall relevant materials, ensure the relevance of answers through text similarity model and large-model verification, and use the panoramic timeline module to display the event development context.

Benefits of technology

It improves the comprehensiveness and accuracy of the answers, ensures that the answer logic is clear, can intuitively display the event development process, and enhances users' understanding of event development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144700A_ABST
    Figure CN120144700A_ABST
Patent Text Reader

Abstract

The invention discloses a news retrieval question-answering method and system of a big language model based on RAG, relates to the technical field of big language model text generation, and provides the news retrieval question-answering method of the big language model for performing intention understanding on questions of users. Data sources are increased, and the query range is expanded to enhance the comprehensiveness of retrieval results and final answers; the intention of the user is accurately grasped, and the part related to the question of the user in the retrieval result is correctly identified and used to improve the accuracy of answering; questions are answered in a layer-by-layer progressive mode, and an intuitive display mode of event veins is provided to help a user to analyze the development process of events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language model text generation, and in particular to a news retrieval and question answering method and system based on RAG for large language models. Background Art

[0002] Large Language Model (LLM) is an artificial intelligence model based on deep learning that can process and generate natural language text. Its core technology is neural networks, especially the Transformer architecture. By training on a large amount of text data, it can understand and generate human-like language.

[0003] The development of large language models has a long history. From relying on rule-based and statistical methods in the early days, to the rise of language models based on Hidden Markov Models, and then to the emergence of the Transformer architecture which has greatly improved natural language processing capabilities. Models such as Qwen and GPT have made remarkable progress in multiple natural language processing tasks.

[0004] For news retrieval and question answering systems based on large language models, the mainstream solution is to further train news corpora on the basis of open-source large models using the Transformer architecture to form a base model, which is divided into pre-training and fine-tuning stages, and then perform instruction fine-tuning based on the base model to achieve the retrieval and question answering function. However, the existing solutions have problems such as incomplete, inaccurate, and unclear answer results. The Retrieval-Augmented Generation (RAG) technology can make up for these deficiencies. It can generate queries according to user questions, retrieve relevant documents through a retrieval engine or an external knowledge base, and generate answers after integrating the information.

[0005] In the existing solutions, the model cannot fully utilize the retrieved content. Especially when the retrieved content is too much and too miscellaneous, some retrieved content summaries will be missed. Due to poor basic capabilities, the answers given by the model are extremely dependent on the retrieved content and cannot identify the parts in the retrieved content that are irrelevant to the question or lack information support, resulting in errors such as chronological disorder and misattribution. The answer content lacks logic. When generating a long analysis report, the long text generated by the model may confuse users. Moreover, most existing models do not sort out the context of events, and users cannot intuitively analyze the development of events. Summary of the Invention

[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is to provide a news retrieval and question answering method for large language models, which increases data sources, expands the query scope to enhance the comprehensiveness of retrieval results and final answers; accurately grasps the user's intention, and correctly identifies and uses the parts related to the user's question in the retrieval results to improve the accuracy of answers; answers questions in a progressive manner and provides an intuitive display method of the event context to help users analyze the development process of events.

[0007] To achieve the above object, the present invention provides a news retrieval and question - answering method for large - language models based on RAG, which is characterized by including the following steps:

[0008] Understand the intention of the user's question, and classify the output of the intention understanding module into three categories: refusal to answer, supplementation, and direct answer. By training the model, it is enabled to judge whether the user's question belongs to the news field and whether it is clear and specific enough. For questions that do not belong to the news field, answer refusal is provided, and for questions that are not specific and clear, supplementary information is provided to guide the user to improve the question;

[0009] Classify the questions that have not been refused to answer, and judge whether they are simple questions that can be answered by single - retrieval or complex questions that require intermediate results to answer. Complex questions are further divided into complex questions that can be split into multiple independent sub - questions and complex questions with dependent sub - questions. For the split sub - questions, after answering, they are summarized to finally answer the user's question. Use a mind map to clarify the dependency relationship between questions. When the split simple questions are too general, they are further decomposed into multiple specific questions through question enhancement;

[0010] After completing question splitting and enhancement, for each split simple question and its enhanced question, unified retrieval and recall are performed. Adopt a multi - source retrieval method to retrieve from multiple data sources to recall more comprehensive retrieval materials, and select the most similar retrieval results through a text similarity model or keyword filtering. Conduct strict relevance discrimination on the retrieved and recalled text. In addition to using the results output by the text similarity model, a large model for judging relevance is separately constructed for further verification. Only when the text similarity reaches the threshold and the retrieval results are judged to be relevant by the large model can they be added to the candidate pool;

[0011] Distribute the retrieved content in the candidate pool to each question on the mind map in order of similarity, so that each question has its own reference materials to assist in answering. Input these questions and their reference materials into the large model responsible for question - answering in turn, and use the answer to each downstream question as the reference material for the upstream question to complete the answer to the original question input by the user;

[0012] For the retrieval results in the candidate pool, call the large model to extract the event information therein and add it to the event pool, and then use the similarity discrimination model to judge whether these events are relevant to the user's question, eliminate irrelevant and duplicate events, input the remaining event information into the large model, complete the classification and summary of the events, and display the summarized event information to the user to help the user intuitively understand the development process of the event.

[0013] Preferably, the intention understanding module judges the field and clarity of the user's question through a trained model, specifically:

[0014] The training model learns the features in the news field and the judgment criteria for question clarity;

[0015] The user's question is input into the trained model, and the model outputs the category to which the question belongs.

[0016] Preferably, the multi-source retrieval method retrieves information from multiple data sources, specifically:

[0017] Send retrieval requests to the network database and the self-built database simultaneously;

[0018] Receive the retrieval results returned by each data source;

[0019] Integrate and preliminarily screen the retrieval results.

[0020] Preferably, the relevance discrimination is jointly completed by a text similarity model and a separately constructed large model, specifically:

[0021] Calculate the text similarity between the retrieval result and the question;

[0022] Input the retrieval result into the separately constructed large model to judge its relevance;

[0023] Only when the text similarity reaches the threshold and is judged as relevant by the large model, the retrieval result is added to the candidate pool.

[0024] Preferably, the specific steps for the panoramic timeline module to display the development context of events are:

[0025] Extract event information from the retrieval results in the candidate pool and add it to the event pool;

[0026] Use the similarity discrimination model to eliminate the events in the event pool that are not relevant to the user's question;

[0027] Use the similarity discrimination model to eliminate duplicate events in the event pool;

[0028] Input the remaining event information into the large model for classification and summarization;

[0029] Display the summarized event information to the user.

[0030] On the other hand, the present invention also provides a news retrieval and question-answering system based on a RAG large language model, which is characterized by including:

[0031] An intent understanding module, which is used to understand the intent of the user's question, and classify its output into three categories: refusal to answer, supplementation, and direct answer. It judges whether the user's question belongs to the news field and whether it is specific and clear enough through a training model. For questions that do not belong to the news field, it refuses to answer, and for questions that are not specific and clear, it provides supplementary information;

[0032] A problem classification module, which is used to classify the questions that have not been rejected, determine whether they are simple questions or complex questions, further subdivide the complex questions, and summarize and answer the user's questions after answering the split sub-questions. It uses a mind map to clarify the problem dependency relationship and enhance the general simple questions.

[0033] A retrieval and recall module, which is used to perform multi-source retrieval and recall on each split simple question and its enhanced question after question splitting and enhancement, retrieve information from multiple data sources, select the most similar retrieval results through a text similarity model and keyword filtering, and perform strict relevance discrimination on the retrieval results, and add the qualified retrieval results to the candidate pool.

[0034] An answer module, which is used to distribute the retrieved content in the candidate pool to the questions on the mind map, input the questions and their reference materials into the large model responsible for answering, and use the downstream question answers as reference materials for the upstream questions to complete the answer to the original question.

[0035] A panoramic timeline module, which is used to extract event information from the retrieval results in the candidate pool, and after discrimination of relevance and repeatability, input the remaining event information into the large model for classification and summary, and display the development context of the events to the user.

[0036] Preferably, the training model in the intention understanding module learns the characteristics of the news field and the judgment criteria for question clarity to determine the category to which the user's question belongs.

[0037] Preferably, the multi-source retrieval in the retrieval and recall module retrieves information from both the network database and the self-built database and integrates and filters it.

[0038] Preferably, the relevance discrimination in the retrieval and recall module is jointly completed by a text similarity model and a separately constructed large model to ensure the relevance of the retrieval results added to the candidate pool.

[0039] Preferably, the panoramic timeline module intuitively displays the development process of the events by extracting event information, discriminating relevance and repeatability, and classifying and summarizing.

[0040] The beneficial technical effects of the present invention:

[0041] Through question splitting, question enhancement and multi-channel recall, the present invention can recall relevant materials more comprehensively, and the split sub-questions also make the answer content more detailed.

[0042] The intention understanding module of the present invention differentially processes different types of questions to improve the accuracy of material use; the distribution and relevance discrimination of the retrieval and recall results, as well as the citation generation and hallucination mitigation algorithms, further improve the answer accuracy.

[0043] The present invention uses mind maps to decompose user questions and answers the questions in a progressive manner. A panoramic timeline generation module is introduced to enable users to more intuitively feel the development context of events.

[0044] The following will further illustrate the concept, specific structure and technical effects of the present invention with reference to the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a system flow chart of a preferred embodiment of the present invention;

[0046] Figure 2 is a mind map Q&A flow chart of the present invention;

[0047] Figure 3 is a multi-source retrieval flow chart of the present invention;

[0048] Figure 4 is a relevance discrimination flow chart of the present invention;

[0049] Figure 5 is a timeline generation flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0051] In the accompanying drawings, components with the same structure are denoted by the same numeral labels, and components with similar structures or functions are denoted by similar numeral labels. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. In order to make the drawings clearer, the thickness of some components is appropriately exaggerated in the drawings.

[0052] The present invention provides a news retrieval Q&A based on a large language model of RAG, which increases data sources and expands the query scope to enhance the comprehensiveness of retrieval results and final answers; accurately grasps the user's intention, and correctly identifies and uses the parts related to the user's question in the retrieval results to improve the accuracy of the answer; answers questions in a progressive manner, and provides an intuitive display method of the event context to help users analyze the development process of events.

[0053] The present invention proposes a framework for news retrieval and question answering, as shown in the accompanying drawings. In this framework, first, the intention of the user's question is understood. The output of the intention understanding module is divided into three categories: refusal to answer, supplementation, and direct answer. By training the model, it is enabled to determine whether the user's question belongs to the news field and whether it is specific and clear enough. Questions that do not belong to the news field will be refused to answer; for questions that are not specific and clear, supplementary information will be provided to guide the user to refine the question.

[0054] Next, the questions that have not been refused to answer are classified to determine whether they are simple questions that can be answered by a single retrieval or complex questions that require intermediate results to answer. Complex questions are further divided into complex questions that can be split into multiple independent sub-questions and complex questions with dependent sub-questions. After answering the split sub-questions, the model summarizes these answers to finally answer the user's question. A Graph of Thought (GoT) is used to clarify the dependencies between these questions. In addition, sometimes the split simple questions are too general, such as questions about sorting out the context of an event or summarizing a manuscript, so they are further decomposed into multiple specific questions through question enhancement.

[0055] After completing question splitting and enhancement, for each split simple question and its enhanced question, retrieval and recall are uniformly performed. A multi-source retrieval method is adopted to retrieve from multiple data sources including Baidu, Bing, and the self-built database to recall more comprehensive retrieval materials, and the most similar retrieval results are selected through text similarity models such as BGE and Qwen and keyword filtering. However, if the retrieved results are irrelevant to the questions on the thought map, the model may produce serious hallucinations, so strict relevance discrimination needs to be performed on the retrieved text. In this step, in addition to using the results output by the above BGE text similarity model, a large model for judging relevance is separately constructed for further verification. Only when the BGE text similarity reaches the threshold and the retrieved results are judged to be relevant by the large model can they be added to the candidate pool.

[0056] At this time, there are several retrieved results in the candidate pool and several questions on the thought map. Next, the retrieved content in the candidate pool is distributed to each question on the thought map in order of similarity, so that each question has its own reference materials to assist in answering. These questions and their reference materials are sequentially input into the large model responsible for question answering, and the answer to each downstream question is used as the reference material for the upstream question, thus completing the answer to the original question input by the user.

[0057] So far, the answer to the user's question has been completed. However, for users who wish to clearly understand the context of events, the form of long-text display is obviously not intuitive enough. To solve this problem, a panoramic timeline module is introduced. This module presents the events related to the user's question to the user in the form of a panoramic timeline. Specifically, for the retrieval results in the candidate pool, the large model is called to extract the event information (including but not limited to time, location, person, event description, etc.) and add it to the event pool. Then, the similarity discrimination model is used to determine whether these events are relevant to the user's question, and the irrelevant events are removed from the event pool. The similarity discrimination model is also used to remove duplicate events from the event pool. Finally, the remaining event information is input into the large model to complete the classification and summary of events, and the summarized event information is presented to the user to help the user intuitively understand the development process of events.

[0058] Next, the present invention will be described through specific embodiments:

[0059] Embodiment 1: Single-event news retrieval and Q&A

[0060] Embodiment scenario

[0061] Suppose a user is interested in the news event of a certain company launching a new product and wants to know the detailed information about the product, the market reaction, and the impact on the company's future development.

[0062] Intention understanding stage

[0063] The user enters the question: "What about the newly launched product of a certain company? What impact does it have on the company's development?" The intention understanding module in the system inputs the question into the trained model. The model, after learning the characteristics of the news field and the criteria for judging question clarity, determines that this question belongs to the news field and is clear and specific, and can be directly answered without involving the situation of refusal to answer or supplementing information.

[0064] Question classification stage

[0065] This question is determined to be a complex question because it involves two aspects: the situation of the product itself and the impact on the company's development, and there is a certain dependency relationship between these two aspects (the product situation may affect the company's development). The model uses a mind map (GoT) to clarify the relationship between the questions and splits it into two sub-questions: "Detailed information about the newly launched product of a certain company (including functions, features, etc.)" and "The impact on the company's development after the product is launched (such as changes in market share, impact on financial status, etc.)."

[0066] Retrieval and recall stage

[0067] For the two sub-questions after splitting, the retrieval and recall module adopts a multi-source retrieval method. Search for relevant information from Baidu, Bing, and the self-built database. Suppose 100 relevant news are retrieved from Baidu, 80 from Bing, and 50 from the self-built database. Calculate the similarity between these retrieval results and the sub-questions through the BGE text similarity model, and combine keyword filtering. For example, set keywords such as "a company name", "new product name", "product function", "market reaction", "company development", etc. After preliminary screening, 30 most similar results are retained from Baidu, 25 from Bing, and 20 from the self-built database. Then, input these preliminarily screened results into a separately constructed large model for judging relevance for further verification. Set the BGE text similarity threshold to 0.7. After being judged by the large model, finally, 20 retrieval results from Baidu, 18 from Bing, and 15 from the self-built database are added to the candidate pool.

[0068] Question and Answer Stage

[0069] Distribute the retrieved content in the candidate pool to the corresponding two sub-questions in order of similarity. For example, distribute the retrieval results with a higher relevance to the product details to the first sub-question, and those with a higher relevance to the impact of the product on the company's development to the second sub-question. Then input these questions and their reference materials into the large model responsible for answering questions in turn. The large model answers the first sub-question based on the reference materials, such as "The newly released product of a certain company has innovative function X, features Y, and good performance, etc.", and this answer will be used as one of the reference materials for answering the second sub-question. Then answer the second sub-question, such as "Due to the innovative function and good performance of the product, it has attracted the attention of a large number of consumers, is expected to increase the company's market share, have a positive impact on the company's future financial situation, and promote the further development of the company, etc.". Finally, summarize the answers to the two sub-questions and answer the user's original question: "The newly released product of a certain company has innovative function X, features Y, and good performance. Because it has attracted the attention of a large number of consumers, it is expected to increase the company's market share, have a positive impact on the company's future financial situation, and promote the further development of the company."

[0070] Panoramic Timeline Stage

[0071] For the retrieval results in the candidate pool, the panoramic timeline module calls a large model to extract event information from them, such as the product release time, location, and the content of the company's relevant person in charge's speech at the press conference. Suppose 10 pieces of event information are extracted. The similarity discrimination model is used to judge the relevance of these events to the user's question, and 3 irrelevant events are removed. Then, the similarity discrimination model is used to remove 2 duplicate events. The remaining 5 pieces of event information are input into the large model for classification and summary, for example, arranged in chronological order and presented to the user, such as "[Time 1] A certain company releases a new product at [Location 1], and [Name of the person in charge] introduces that the product has innovative function X. [Time 2] A market research institution begins to evaluate the product. [Time 3] Consumers show strong interest in the product, and the pre-order volume exceeds expectations, etc.", to help the user intuitively understand the event development process.

[0072] Example 2: Series of Event News Retrieval and Q&A

[0073] Example Scenario

[0074] Taking a series of events of a certain sports event as an example, the user wants to know the performance of a certain athlete in this event, including the results of each game, key events, and the impact on his / her career.

[0075] Intention Understanding Stage

[0076] The user inputs the question: "What is the performance of a certain athlete in this event? What is the impact on his / her career?" The intention understanding module inputs the question into the model, and the model judges that this question belongs to the news field and is clear and specific, and can be directly answered.

[0077] Question Classification Stage

[0078] This question is a complex question, involving the performance of the athlete in multiple games of the event and the impact on his / her career. The model uses a mind map (GoT) to split it into multiple sub-questions, such as "The results and key events of a certain athlete in the first game of the event", "The results and key events of a certain athlete in the second game of the event"... "The impact of the overall performance of a certain athlete in this event on his / her career (such as ranking changes, commercial value improvement, etc.)". At the same time, since the question of "the impact of the overall performance on the career" is relatively general, it is further decomposed into specific questions such as "How does the event performance affect his / her ranking in the sports world" and "How does the event performance improve his / her commercial endorsement value" through question enhancement.

[0079] Retrieval and Recall Stage

[0080] For each sub-question and its enhanced question, the retrieval and recall module conducts multi-source retrieval. Relevant information is retrieved from Baidu, Bing, and the self-built database. Suppose for each sub-question, on average 120 results are retrieved from Baidu, 100 from Bing, and 60 from the self-built database. Through the BGE text similarity model and keyword filtering, initially 40 results are retained from Baidu, 30 from Bing, and 25 from the self-built database. After verification by the large model, setting the BGE text similarity threshold to 0.65, finally on average 30 results from Baidu, 25 from Bing, and 20 from the self-built database are added to the candidate pool.

[0081] Question and Answer Phase

[0082] The retrieved content in the candidate pool is distributed to each sub-question according to similarity. For example, the retrieval results related to the first game are assigned to the sub-question "The performance and key events of a certain athlete in the first game of the event". The large model answers based on the reference materials, such as "In the first game, a certain athlete achieved [specific performance], and at the crucial moment [description of the key event]". After answering each sub-question in turn, these answers are summarized to answer the user's questions about the athlete's performance in the event. For the enhanced question "How does the event performance affect their ranking in the sports world", the large model answers based on the previous answers about game results, etc., combined with the reference materials, such as "Due to the excellent performance in this event, their ranking in the sports world is expected to rise by [X] positions". Similarly, other enhanced questions are answered, and finally the answers to the user's questions about the impact of event performance on the career are summarized to answer the user's original question as a whole.

[0083] Panoramic Timeline Phase

[0084] Event information is extracted from the retrieval results in the candidate pool, such as the time, location, opponent, and exciting moments of each game. Suppose 20 pieces of event information are extracted, 5 irrelevant events are removed using the similarity discrimination model, and then 3 duplicate events are removed. The remaining 12 pieces of event information are input into the large model for classification and summary, and presented to the user in chronological order of the games, such as "[Time 1] A certain athlete played the first game against [Opponent 1] at [Location 1], achieving [Performance 1], and [Exciting Moment 1] during the game. [Time 2] The second game was played......", helping the user clearly understand the context of the event development.

[0085] Example 3: Comprehensive News Event Retrieval and Question Answering

[0086] Example Scenario

[0087] Consider a comprehensive news event involving economic, political, and social factors, such as a new policy being introduced in a certain region. The user wants to know the specific content of the policy, its impact on the local economic development, the reactions of all sectors of society, and its association with other relevant policies.

[0088] Intention Understanding Stage

[0089] The user inputs the question: "What are the specific details of the newly introduced policy in a certain area, its impact on the economy and society, and its relationship with other policies?" The intention understanding module inputs the question into the model, and the model determines that this question belongs to the news field and is clear and specific, and can be answered directly.

[0090] Question Classification Stage

[0091] This question is a complex question. The model uses a mind map (GoT) to break it down into the following sub-questions: "The specific content (clauses, goals, etc.) of the newly introduced policy in a certain area", "The impact of this policy on the local economic development (such as industrial structure adjustment, changes in employment situation, etc.)", "The reactions of all sectors of society (enterprises, the public, etc.) to this policy", "The relationship between this policy and other relevant policies (such as previous similar policies or supporting policies)". For the sub-question of "The reactions of all sectors of society to this policy", since it is relatively general, it is decomposed into specific questions such as "The attitudes and coping measures of the business community towards this policy" and "The views and expectations of the public towards this policy" through question enhancement.

[0092] Retrieval and Recall Stage

[0093] For each sub-question and its enhanced questions, the retrieval and recall module conducts multi-source retrieval. Search for information from Baidu, Bing, and the self-built database. Suppose 200 relevant news articles are retrieved from Baidu, 150 from Bing, and 80 from the self-built database. Through the BGE text similarity model and keyword filtering, 60 articles are initially retained from Baidu, 50 from Bing, and 30 from the self-built database. After verification by the large model, setting the BGE text similarity threshold to 0.75, 45 articles from Baidu, 40 from Bing, and 25 from the self-built database are finally added to the candidate pool.

[0094] Question Answering Stage

[0095] Distribute the retrieved content in the candidate pool to the corresponding sub-questions according to similarity. For example, distribute the retrieval results related to the specific content of the policy to the sub-question of "The specific content (clauses, goals, etc.) of the newly introduced policy in a certain area". The large model answers such as "The main clauses of this policy include [list the clauses], and the goal is [elaborate the goal]". Answer other sub-questions in turn. For example, for "The attitudes and coping measures of the business community towards this policy", the large model answers according to the reference materials "The business community generally expresses support for this policy, and some enterprises plan to [list the enterprise coping measures]". Summarize the answers to all sub-questions and answer the user's original question, such as "The specific content of the newly introduced policy in a certain area is [detailed content], it will [describe the impact] on the local economic development, the business community [attitudes and measures], the public [views and expectations], and the relationship between this policy and other policies [describe the relationship]".

[0096] Panoramic Timeline Phase

[0097] Extract event information from the retrieved results in the candidate pool, such as the policy release time, the situation of the release background briefing, the feedback time points of all sectors of society at different stages, etc. Suppose 15 pieces of event information are extracted. Use a similarity discrimination model to eliminate 4 irrelevant events, and then eliminate 2 duplicate events. Input the remaining 9 pieces of event information into a large model for classification and summarization, and display them to the user in chronological order, such as "[Time 1] The government of a certain region held a press conference to announce a new policy, [Highlights of the official's speech]. [Time 2] Some enterprise representatives expressed their views on the policy...", to help the user intuitively grasp the development process of the event.

[0098] Example 4: Example of an Emergency News Retrieval Q&A System

[0099] Example Scenario

[0100] Suppose a natural disaster, such as an earthquake, occurred in a certain place. The user hopes to quickly obtain detailed information about this earthquake, including the occurrence time, magnitude, affected area, rescue progress, and the impact on the lives of local residents, etc.

[0101] Intention Understanding Phase

[0102] The user inputs the question: "What is the specific situation of the earthquake that occurred in a certain place? How is the rescue work progressing? What is the impact on the lives of local residents?" The intention understanding module inputs this question into a trained model. Based on the learned news domain features and question clarity judgment criteria, the model quickly determines that this question belongs to the news domain and is clear and specific enough, without the need to reject the answer or supplement information, and can directly proceed with subsequent processing.

[0103] Question Classification Phase

[0104] This question is identified as a complex question because it covers multiple aspects of the earthquake itself and subsequent impacts, etc., and there is a certain logical relationship among various aspects. The model uses a Graph of Thoughts (GoT) to decompose it into the following sub-questions: "The exact time of the earthquake", "What is the magnitude of the earthquake", "The specific affected area (including which regions, towns, etc.)", "The progress of the current rescue work (rescue team deployment, rescue supply transportation, etc.)", "The impact of the earthquake on the lives of local residents (housing, water and electricity supply, daily life, etc.)". Among them, the questions of "The specific affected area" and "The impact of the earthquake on the lives of local residents" are relatively general and are further refined into specific questions such as "List of the main towns and villages affected", "Damage situation of houses in the affected area and resettlement situation of residents", "Regions where water and electricity supply is interrupted and the estimated recovery time", etc. through question enhancement.

[0105] Retrieval and Recall Phase

[0106] For each sub-question and its enhanced question, the retrieval and recall module activates a multi-source retrieval mechanism. Retrieval requests are sent to Baidu, Bing, and the self-built database respectively to obtain relevant information. For example, for the sub-question of "the exact time of the earthquake occurrence", 50 relevant news reports are retrieved from Baidu, 40 from Bing, and 20 from the self-built database. The similarity between these retrieval results and the sub-question is calculated through the BGE text similarity model, and filtering is carried out in combination with the keywords "a certain place", "earthquake", and "occurrence time". After preliminary screening, Baidu retains 20 most similar results, Bing retains 15, and the self-built database retains 10. Then, these preliminarily screened results are input into a specially constructed large model for judging relevance for further verification. Setting the BGE text similarity threshold to 0.8, after being judged by the large model, finally 15 retrieval results from Baidu, 12 from Bing, and 8 from the self-built database are added to the candidate pool. The retrieval and recall process for other sub-questions is carried out in the same way to ensure obtaining the most relevant reference materials for each question.

[0107] Question and Answer Phase

[0108] The retrieved content in the candidate pool is distributed to the corresponding sub-questions in order of similarity. For example, the retrieval results with the highest relevance to the earthquake occurrence time are assigned to the sub-question of "the exact time of the earthquake occurrence", and the large model gives an accurate answer based on these reference materials, such as "The earthquake occurred at [specific time]". For the sub-question of "the progress of the current rescue work", the large model comprehensively answers based on the reference materials, "Currently, [X] rescue teams have been dispatched to the disaster area, rescue supplies are being continuously transported to the severely affected [specific area], and [X] temporary resettlement sites have been set up". And so on, each sub-question is answered in turn, and the answers to each sub-question are summarized. Such as "The earthquake magnitude is [X], it occurred at [specific time], and the affected areas involve [list the main affected towns and villages]. Currently, the rescue work is in an orderly manner. Multiple rescue teams have been dispatched, the transportation of supplies is advancing, and some of the residents in the affected areas have been resettled. The earthquake caused a large number of houses to be damaged, and the water and electricity supply was interrupted in some areas. It is expected to be restored at [estimated restoration time]", thus completely answering the user's original question.

[0109] Panoramic Timeline Phase

[0110] The panoramic timeline module extracts various types of event information from the retrieval results of the candidate pool, including the situation at the moment of the earthquake, the time nodes when the rescue teams arrive, the phased progress of material distribution, etc. Suppose 25 pieces of event information are extracted, and 6 irrelevant events are removed using the similarity discrimination model, such as earthquake events unrelated to other regions or other information not related to this earthquake. Then, 4 duplicate events are removed through the similarity discrimination model, such as the departure information of the same rescue team reported repeatedly. The remaining 15 pieces of event information are input into the large model for classification and summary, and clearly presented to the user in chronological order, such as "[Time 1] The earthquake occurred, and obvious tremors were felt in many places. [Time 2] The local government activated the emergency response, and the rescue teams began to assemble. [Time 3] The first rescue team arrived at the most severely affected [region name] and carried out rescue work...", helping users intuitively understand the development process of the entire earthquake event from occurrence to rescue, so as to better grasp the overall situation of the event.

[0111] Intention understanding module algorithm

[0112] Model selection and training

[0113] A neural network model based on the Transformer architecture, such as the Qwen model, is used for fine-tuning.

[0114] The training data is a large number of labeled questions in the news field and non-news field, and the labeled content includes the field to which the question belongs (news / non-news) and the clarity of the question (clear / unclear).

[0115] The training objective is to minimize the cross-entropy loss function between the prediction result and the labeled result. The Adam optimizer can be selected as the optimizer, the learning rate is set to 0.0001, and the number of training epochs is 10.

[0116] Intention judgment process

[0117] Preprocess the question input by the user, including operations such as word segmentation and stop word removal.

[0118] Input the preprocessed question into the trained Qwen model to obtain the vector representation of the question.

[0119] Classify the vector representation through the fully connected layer to judge the category to which the question belongs (such as rejecting the answer when the question belongs to a mathematics or code question, and supplementing details when there are missing details, etc.).

[0120] Question classification module algorithm

[0121] Question classification model

[0122] Build a classification model based on a decision tree.

[0123] Feature selection includes keywords of the question, the length of the question, the sentence structure of the question, etc.

[0124] Learn the classification patterns of questions under different feature combinations through training data, which are a large number of classified news questions (simple questions, complex questions and sub-types of complex questions).

[0125] Question splitting and enhancement algorithm

[0126] For complex questions, use a rule-based splitting method. For example, split the question into sub-questions according to specific keywords (such as "and", "as well as", "the impact on...", etc.).

[0127] For general simple questions, adopt a template matching enhancement method. For questions like "Sort out the context of the event", match the preset template to generate specific questions such as "The time, place, and main characters when the event started" and "The key nodes in the development process of the event".

[0128] Retrieval and recall module algorithm

[0129] Multi-source retrieval algorithm

[0130] For search engines such as Baidu and Bing, use the API interfaces provided by them to send retrieval requests. The request parameters include keywords (generated from the question), retrieval scope (news category), time range (set according to news timeliness, such as the recent week), etc.

[0131] For the self-built database, use full-text retrieval technology, such as a retrieval engine based on Lucene, to retrieve in fields such as the news document title and body in the database according to the keywords in the question.

[0132] Relevance discrimination algorithm

[0133] First, use the BGE text similarity model to calculate the similarity between the retrieval result and the question. The formula is: where is the question vector and is the retrieval result vector. Set the similarity threshold to 0.7.

[0134] At the same time, input the retrieval result into a separately constructed large model based on Qwen for relevance judgment. The input of the large model is the retrieval result and the question, and the output is a judgment result of relevant or not relevant. Only the retrieval results that simultaneously meet the BGE text similarity threshold and are judged as relevant by the large model are added to the candidate pool.

[0135] Question and answer module algorithm

[0136] Answer generation model

[0137] Adopt the Qwen2-72B-Insturct model as the base model for fine-tuning.

[0138] The training data consists of question and corresponding correct answer pairs. Through supervised learning, the model learns the ability to generate accurate answers based on questions.

[0139] During fine-tuning, the optimization objective is to minimize the edit distance between the generated answer and the reference answer. The optimizer used is Adagrad, with a learning rate of 0.001 and 5 training epochs.

[0140] Answer process

[0141] After sorting the retrieved content in the candidate pool by similarity, it is combined with the question into an input sequence and input into the fine-tuned Qwen2-72B-Insturct model in turn.

[0142] The model generates an answer based on the input sequence. For sub-questions with dependencies, the answer to the downstream question is used as additional input information for the answer to the upstream question to generate a more accurate final answer.

[0143] Panoramic timeline module algorithm

[0144] Event extraction model

[0145] An event extraction model is constructed based on Qwen2-72B-Insturct.

[0146] The training data is news documents and manually annotated event information (including time, location, people, event description, etc.). By learning from the annotated data, the ability to identify and extract event-related information from text is learned.

[0147] The loss function of the model uses cross-entropy loss, the optimizer is RMS Prop, the learning rate is 0.0005, and training is performed until the loss function converges.

[0148] Event processing and display algorithm

[0149] After extracting event information from the retrieval results in the candidate pool, a discriminant model based on cosine similarity is used to eliminate irrelevant events, and the threshold is set to 0.6.

[0150] For duplicate events, deduplication is performed by comparing the key information of the events (such as time, location, core content of the event).

[0151] The processed event information is input into a classification and summarization algorithm based on hierarchical clustering, classified according to attributes such as the time and theme of the event, and then presented to the user in the form of a timeline. The display format is: [time] - [event description].

[0152] Example 5: Example of news retrieval and question-answering software based on a Web platform

[0153] Example scenario

[0154] Develop a news retrieval and Q&A software based on the Web platform. Users can access this software through a web browser and input questions about various news events, such as new product releases in the technology field, celebrity dynamics in the entertainment industry, changes in the international political situation, etc. The software can quickly and accurately provide comprehensive, clear, and accurate answers and display the event context in an intuitive way.

[0155] Software Architecture and Module Design

[0156] Front-end Interface: Use HTML, CSS, and JavaScript technologies to build the user interaction interface. Design a simple and intuitive input box to facilitate users to input questions. At the same time, reserve areas for displaying the answer results and the panoramic timeline, and display them to users in a clear layout. For example, the answer results are presented in paragraphs, the panoramic timeline is displayed as a visual timeline chart, and the event nodes are marked with different colors to indicate the importance level.

[0157] Back-end Server: Use the Flask framework of Python to build the back-end server. It is responsible for receiving the user questions sent by the front-end, calling the intent understanding module, question classification module, retrieval and recall module, Q&A module, and panoramic timeline module for processing, and returning the final results to the front-end.

[0158] Database: Select MongoDB database to store the self-built news corpus, user question records, and intermediate data during model training, etc. At the same time, use Elasticsearch to build an index to improve the search efficiency of the retrieval and recall module.

[0159] Implementation of the Intent Understanding Module in the Software

[0160] The trained intent understanding model (fine-tuned based on Qwen) is deployed on the back-end server. When the user inputs a question on the front-end, the back-end sends the question to the intent understanding model. After preprocessing (word segmentation, stop word removal, etc.), the model outputs the intent classification result of the question (reject answer, supplement, direct answer). In the case of supplementing information, the back-end generates corresponding prompt information and returns it to the front-end to guide the user to perfect the question; in the case of rejecting an answer, a friendly rejection prompt is returned; in the case of a direct answer, the subsequent process continues.

[0161] Implementation of the Question Classification Module in the Software

[0162] The decision tree-based question classification model is deployed on the backend. After receiving a question that can be directly answered, the backend inputs the question into the classification model. The model determines the question type (simple question, complex question, and sub-types of complex questions) based on features such as the keywords, length, and sentence structure of the question. For complex questions, they are split according to rules, and the split sub-questions are numbered and the dependency relationships are recorded (stored in memory using the data structure of the Mind Map GoT). For general simple questions, question enhancement is performed through template matching to generate specific sub-questions and record the numbers and relationships in the same way.

[0163] Implementation of the retrieval and recall module in the software

[0164] In the multi-source retrieval part, the backend sends requests to the news search APIs of Baidu and Bing respectively through Python's network request library (such as requests), and at the same time performs retrieval in the self-built database using Elasticsearch. After obtaining the retrieval results, the BGE text similarity model (implemented through Python's deep learning library such as PyTorch) is used to calculate the similarity, and keyword filtering is combined. Then the preliminarily screened results are input into the large model based on Qwen (implemented through the deployment of the backend server) for relevance verification. The retrieval results that meet the conditions (the BGE text similarity reaches the threshold and are judged as relevant by the large model) are added to the candidate pool and stored in memory in a specific data structure (such as a list nested dictionary, where the dictionary contains information such as the content of the retrieval result and the similarity score).

[0165] Implementation of the question and answer module in the software

[0166] The fine-tuned Qwen2-72B-Instruct model is deployed on the backend server. After sorting the retrieval content in the candidate pool by similarity, the backend sequentially combines each sub-question and its corresponding retrieval content into an input sequence and sends it to the GPT-3.5 model. After the model generates an answer, the backend uses the answer to the downstream question as a reference for answering the upstream question according to the dependency relationship between the questions, and gradually constructs the final answer. Finally, the complete answer is returned to the front end and presented to the user in the answer display area of the front-end interface.

[0167] Implementation of the panoramic timeline module in the software

[0168] The event extraction model based on Qwen2-72B-Instruct is deployed on the backend. The backend extracts the text content from the retrieval results of the candidate pool, inputs it into the event extraction model, and obtains event information (such as time, location, people, event description, etc.). The cosine similarity discriminant model (implemented through the BGE text similarity model) is used to eliminate irrelevant events, and then duplicate events are removed by comparing key information. Then, the processed event information is input into a classification and summarization algorithm based on hierarchical clustering (implemented through Python data analysis libraries such as Scikit-learn) for classification. Finally, the backend returns the classified and summarized event information to the frontend in the format of timeline data (such as a JSON array, where each element contains attributes such as time and event description). The frontend uses a JavaScript visualization library (such as D3.js) to draw and display the timeline in the panoramic timeline display area, and users can view the event details through mouse interaction.

[0169] Example of software operation process

[0170] The user enters the question at the frontend: "What are the new features of the latest mobile phones released by a certain mobile phone company? What impact will it have on the mobile phone market?"

[0171] After receiving the question, the backend first determines through the intent understanding module that the question can be answered directly.

[0172] The question classification module identifies it as a complex question, splits it into two sub-questions: "The new features of the latest mobile phones released by a certain mobile phone company" and "The impact of the release of this mobile phone on the mobile phone market", and records the dependency relationship.

[0173] The retrieval and recall module conducts multi-source retrieval, obtains relevant news reports from Baidu, Bing, and the self-built database, and adds the appropriate retrieval results to the candidate pool after relevance discrimination.

[0174] The Q&A module generates answers to the sub-questions based on the content of the candidate pool, such as "The latest mobile phones released by a certain mobile phone company have new features A, B, C, etc." and "After the release of this mobile phone, the market share is expected to change to a certain extent, putting pressure on competitors, etc.", and summarizes the final answer and returns it to the frontend.

[0175] The panoramic timeline module extracts event information, such as "[Release time] A certain mobile phone company releases a mobile phone, [Press conference location], [Speaker introduces new features]", etc., and displays it in the form of a timeline on the frontend after processing. Users can clearly see the development context of the events.

[0176] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art shall fall within the protection scope determined by the claims.

Claims

1. A large language model news retrieval question answering method based on RAG, characterized in that: The following steps are involved: Understand the intention of the user's question and divide the output of the intention understanding module into three categories: refusal to answer, supplementation and direct answer. Through training, the model is able to judge whether the user's question belongs to the news field and whether it is clear and specific enough. It refuses to answer questions that do not belong to the news field, and provides supplementary information for unclear questions to guide users to improve their questions. Classify the questions that have not been rejected to determine whether they are simple questions that can be answered with a single search or complex questions that require intermediate results to answer. Complex questions are further divided into complex questions that can be split into multiple independent sub-questions and complex questions with dependent sub-questions. For the split sub-questions, summarize them after answering them to finally answer the user's question. Use mind maps to clarify the dependencies between questions. If the split simple questions are too general, they can be broken down into multiple specific questions through question enhancement. After question splitting and enhancement, for each split simple question and its enhanced question, unified retrieval recall is performed, using multi-source retrieval to retrieve from multiple data sources to recall more comprehensive retrieval materials, and the most similar retrieval results are selected through text similarity model or keyword filtering, and strict relevance judgment is performed on the retrieved texts. In addition to using the results output by the text similarity model, a separate large model for judging relevance is constructed for further verification. Only when the text similarity reaches the threshold and is judged as a relevant retrieval result by the large model can it be added to the candidate pool; Distribute the search content in the candidate pool to each question on the mind map in order according to similarity, so that each question has its own reference material to assist in answering. Input these questions and their reference materials into the large model responsible for question answering in turn, and use the answer to each downstream question as a reference for the upstream question to complete the answer to the original question input by the user. For the search results in the candidate pool, the big model is called to extract the event information and add it to the event pool. Then, the similarity discrimination model is used to determine whether these events are related to the user's questions, and irrelevant and repeated events are eliminated. The remaining event information is input into the big model to complete the classification and summary of the events. The summarized event information is displayed to the user to help the user intuitively understand the development process of the event.

2. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that: The intent understanding module determines the domain and clarity of the user's question by training the model, specifically: The training model learns the characteristics of the news field and the criteria for judging the clarity of issues; The user's question is input into the trained model, and the model outputs the category to which the question belongs.

3. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that: The multi-source retrieval method retrieves information from multiple data sources, specifically: Send search requests to network databases and self-built databases at the same time; Receive search results returned by various data sources; Integrate and preliminarily screen the search results.

4. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that: The relevance determination is accomplished by a text similarity model and a separately constructed large model, specifically: Calculate the text similarity between the search results and the question; The search results are fed into a separately constructed large model to determine their relevance; The retrieval results are added to the candidate pool only when the text similarity reaches the threshold and is judged to be relevant by the large model.

5. The RAG-based large language model news retrieval question answering method according to claim 1, characterized in that: The specific steps of the panoramic timeline module to display the development of events are: Extract event information from the search results of the candidate pool and add it to the event pool; Use the similarity discrimination model to remove events in the event pool that are irrelevant to the user's question; Use the similarity discrimination model to remove duplicate events in the event pool; Input the remaining event information into the large model for classification and aggregation; Display the aggregated event information to the user.

6. A large language model news retrieval question answering system based on RAG, characterized in that: include: The intention understanding module is used to understand the intention of the user's question and divide its output into three categories: refusal to answer, supplement and direct answer. The training model is used to determine whether the user's question belongs to the news field and whether it is clear and specific enough. For questions that do not belong to the news field, the module refuses to answer and provides supplementary information for questions that are not specific and clear. The question classification module is used to classify questions that have not been rejected and determine whether they are simple or complex questions. Complex questions are further subdivided and summarized after answering the sub-questions. Mind maps are used to clarify question dependencies and to enhance general simple questions. The retrieval and recall module is used to perform multi-source retrieval and recall on each simple question and its enhanced question after question splitting and enhancement, retrieve information from multiple data sources, select the most similar search results through text similarity model and keyword filtering, and perform strict relevance judgment on the search results, and add the search results that meet the conditions to the candidate pool; The question-answering module is used to distribute the search content in the candidate pool to the questions on the mind map, input the questions and their reference materials into the large model responsible for question-answering, use the answers to downstream questions as reference materials for upstream questions, and complete the answers to the original questions; The panoramic timeline module is used to extract event information from the candidate pool retrieval results. After correlation and repeatability judgment, the remaining event information is input into the large model for classification and summary, showing the development context of the event to the user.

7. The RAG-based large language model news retrieval question answering system according to claim 6, characterized in that: The training model in the intention understanding module learns news domain characteristics and question clarity judgment criteria to determine the category to which the user's question belongs.

8. The RAG-based large language model news retrieval question answering system according to claim 6, characterized in that: The multi-source search in the search and recall module simultaneously searches for information from network databases and self-built databases and integrates and screens them.

9. The RAG-based large language model news retrieval question answering system according to claim 6, characterized in that: The relevance determination in the retrieval recall module is accomplished by a text similarity model and a separately constructed large model, thereby ensuring the relevance of the retrieval results added to the candidate pool.

10. The RAG-based large language model news retrieval question answering system according to claim 6, characterized in that: The panoramic timeline module intuitively displays the development process of events by extracting event information, determining relevance and repetition, and classifying and summarizing.

Citation Information

Patent Citations

  • Question and answer method and device, equipment and storage medium

    CN118035415A

  • Knowledge retrieval enhancement generation method and system based on large language model

    CN118394890A

  • A system for generating answers to multiple questions using rag-based generative artificial intelligence technology

    KR102710159B1

  • Supervised Summarization and Structuring of Unstructured Documents

    US20240012842A1

Cited By

  • Generative AI real-time live broadcast method

    CN121193959A

  • Key value cache scheduling method and system for language model reasoning

    CN122019408A

  • A key-value cache scheduling method and system for language model inference

    CN122019408B