Large language model question and answer method, device and equipment based on retrieval enhancement and medium
By combining knowledge graphs and search engine multi-path retrieval techniques, the large language model is assisted in generating question-and-answer results, which solves the illusion problem of large language models in factual question-and-answer scenarios and achieves more accurate and richer question-and-answer results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2026-04-14
AI Technical Summary
Large language models are prone to hallucination problems in question-answering scenarios with strict requirements of factuality or real-time performance, which are difficult to solve effectively with existing technologies.
By combining internal retrieval of knowledge graph data and external retrieval of search engines, a retrieval plan is generated. Target retrieval results are obtained through multi-path retrieval and then used to assist a large language model in generating question-and-answer results.
It improves the accuracy and realism of large language model question answering, solves the illusion problem, and generates richer and more accurate results.
Smart Images

Figure CN121858685A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for question answering based on a large language model with retrieval enhancement. Background Technology
[0002] Large language models (MLMs) are artificial intelligence models designed to understand and generate human language. They typically refer to language models containing hundreds of billions or more parameters, trained on massive amounts of text data to gain a deep understanding of language. MLMs have demonstrated excellent semantic understanding capabilities, making question-answering systems based on them a hot research topic. However, the content generated by MLM responses is sometimes irrelevant to the question or inaccurate, a phenomenon known as the "large model illusion." In question-answering scenarios with stringent requirements for factual accuracy or real-time performance, mitigating the large language model illusion is a crucial issue that urgently needs to be addressed. Summary of the Invention
[0003] This invention provides a retrieval-enhanced question-answering method, apparatus, device, and medium for large language models, enabling more accurate resolution of the illusion problem in large language models.
[0004] Firstly, this embodiment provides a retrieval-enhanced question-answering method based on a large language model, the method comprising:
[0005] Receive user query information and perform intent understanding on the user query information to obtain intent understanding results;
[0006] Based on the intent understanding results and the determined current business scenario, and combined with the pre-built search template library, a search plan is generated relative to the user query information;
[0007] The retrieval is performed according to the retrieval plan to obtain the target retrieval results relative to the user query information. The retrieval plan includes an internal retrieval plan based on knowledge graph data and an external retrieval plan based on a search engine.
[0008] The target retrieval results and the user query information are input into the large language model to obtain the target question-answering results relative to the user query information.
[0009] Secondly, this embodiment provides a large language model question-answering device based on retrieval enhancement, the device comprising:
[0010] The intent understanding module is used to receive user query information and perform intent understanding on the user query information to obtain intent understanding results;
[0011] The plan generation module is used to generate a retrieval plan relative to the user query information based on the intent understanding results and the determined current business scenario, combined with a pre-built retrieval template library.
[0012] The retrieval module is used to perform retrieval according to the retrieval plan and obtain target retrieval results relative to the user query information. The retrieval plan includes an internal retrieval plan based on knowledge graph data and an external retrieval plan based on a search engine.
[0013] The result generation module is used to input the target retrieval result and the user query information into the large language model to obtain the target question answer result relative to the user query information.
[0014] Thirdly, this embodiment provides an electronic device, including:
[0015] At least one processor; and
[0016] A memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the retrieval-enhanced large language model question answering method according to any embodiment of the present invention.
[0018] Fourthly, this embodiment provides a computer-readable storage medium, wherein the computer program is executed by the at least one processor to enable the at least one processor to perform the retrieval-enhanced large language model question-answering method according to any embodiment of the present invention.
[0019] Fifthly, embodiments of the present invention also provide a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the retrieval-enhanced large language model question-answering method as described in any embodiment of the present invention.
[0020] This invention provides a retrieval-enhanced question-answering method, apparatus, device, and medium for a large language model. The method includes: receiving user query information and performing intent understanding on the query information to obtain intent understanding results; generating a retrieval plan relative to the user query information based on the intent understanding results and the determined current business scenario, combined with a pre-built retrieval template library; performing a retrieval according to the retrieval plan to obtain target retrieval results relative to the user query information, wherein the retrieval plan includes an internal retrieval plan based on knowledge graph data and an external retrieval plan based on a search engine; and inputting the target retrieval results and the user query information into a large language model to obtain target question-answering results relative to the user query information. This technical solution combines internal retrieval based on knowledge graph data and external retrieval based on a search engine to perform multi-path retrieval to obtain target retrieval results. These target retrieval results are then used to assist in the generation of responses by the large language model. The external knowledge acquisition addresses the illusion problem of large language models. By combining internal and external retrieval, the generated target retrieval results are richer and more accurate, more precisely solving the illusion problem of large language models.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating a large language model question-answering method based on retrieval enhancement provided in Embodiment 1 of the present invention;
[0024] Figure 2 This is a flowchart illustrating another large language model question-answering method based on retrieval enhancement provided in Embodiment 2 of the present invention;
[0025] Figure 3 This is a schematic diagram of the structure of a large language model question-answering device based on retrieval enhancement provided in Embodiment 3 of the present invention;
[0026] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] It is understood that before using the technical solutions disclosed in the various embodiments of the present invention, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in the present invention and their authorization should be obtained in accordance with relevant laws and regulations through appropriate means.
[0030] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application program, server, or storage medium executing the operation of this invention, based on the prompt message.
[0031] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0032] It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of the present invention. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present invention.
[0033] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0034] Example 1
[0035] Figure 1 This is a flowchart illustrating a large language model-based question answering method with retrieval enhancement provided in Embodiment 1 of the present invention. This method is applicable to question answering based on a large language model. The method can be executed by a large language model-based question answering device with retrieval enhancement. This large language model-based question answering device with retrieval enhancement can be implemented in hardware and / or software and is generally integrated into an electronic device.
[0036] like Figure 1 As shown, the large language model question answering method based on retrieval enhancement provided in this embodiment may specifically include the following steps:
[0037] S101. Receive user query information and perform intent understanding on the user query information to obtain intent understanding results.
[0038] In this embodiment, when a user has a question-and-answer requirement, they can ask the electronic device integrated with this solution. The information to be asked is recorded as the user's question information. There are no specific restrictions on the input format of the user's question information; for example, it can be in voice or text format. For example, the user's question information may involve vehicle control commands, entertainment needs, or question-and-answer related needs. After receiving the user's question information, this step first performs intent understanding to determine the user's current needs, and records the understanding result as the intent understanding result.
[0039] The intent understanding of user queries can be based on multiple types of pre-trained models. First, the intent of the query is determined. For example, if a user says "open the car window" or "set the air conditioning to XX degrees," the current business scenario can be considered a vehicle control intent. If a user says "play music," the current business scenario can be considered an entertainment intent. If a user says "what are the tourist attractions in XX city," the current business scenario can be considered a question-and-answer intent. In addition to intent determination, the system also generates search requests related to the user query and identifies entities contained within the query.
[0040] It should be noted that only after recognizing the current business scenario as a question-and-answer intent will a series of subsequent processes be performed to obtain the question-and-answer result relative to the user's query and then fed back to the user. Unlike existing technologies that directly use a pre-trained large language model to generate question-and-answer results from user queries, leading to the illusion problem, this embodiment employs retrieval-augmented generation (RAG) technology to retrieve auxiliary information relevant to the user's query before inputting the question into the large language model. This auxiliary information then assists the large language model in generating the question-and-answer result. Specifically, RAG technology first retrieves the most relevant information to the input query from a large text database, and then uses this information to assist in generating the response or text, thereby solving the large model illusion problem.
[0041] This step first involves understanding the user's query information. Once the current business scenario is identified as a question-and-answer intent, the results identified by different models are aggregated to generate the information needed for the retrieval enhancement process, such as the category of the question to be returned, the query, and the slot. This information is recorded as the intent understanding result.
[0042] S102. Based on the intent understanding results and the determined current business scenario, and combined with the pre-built search template library, generate a search plan for the user's query information.
[0043] In this embodiment, several search templates are pre-built. Different business scenarios and search methods correspond to different search templates. The library storing different types of search templates is denoted as the search template library. For example, different search templates are set for different business scenarios. For instance, what search templates are available for searching encyclopedia information about a person? What search templates are available for searching news articles? What search templates are available for searching automotive knowledge within a company? Each business scenario's search template may include knowledge graph-based search templates, semantic full-text search templates, and search templates for third-party external engine searches, each with its own corresponding search template for different business scenarios.
[0044] When the current business scenario is determined to be a question-and-answer intent, based on the determined intent understanding result and the current business scenario, a search template corresponding to the current business scenario can be obtained from a pre-built search template library. The intent understanding result is then filled into the obtained search template to generate a search plan relative to the user's query information. In this embodiment, multiple types of search plans are generated to perform multi-path retrieval based on different types of search plans. For example, if the user's query information is "What are the tourist attractions in city XX?", a search plan relative to this question will be generated. The search plan includes: a search plan based on knowledge graph retrieval, a semantic full-text search plan, and an external search plan for third-party search engines.
[0045] S103. Conduct a search according to the search plan to obtain the target search results relative to the user's query information.
[0046] The retrieval plan includes an internal retrieval plan based on knowledge graph data and an external retrieval plan based on search engines. In this embodiment, in addition to the internal retrieval based on knowledge graph data, external retrieval based on third-party search engines is also supported, thereby obtaining richer retrieval results for subsequent assistance to the large language model in generating question-and-answer results.
[0047] In this embodiment, internal retrieval is performed based on knowledge graph data, which can include knowledge graph retrieval and semantic full-text retrieval. Internal retrieval based on the internal knowledge graph data retrieves content related to the user's query. External retrieval is performed using third-party search engines to find richer content related to the user's query.
[0048] It's important to understand that when performing external searches using search engines, the retrieved web pages require web crawling technology to automatically capture the pages, which are then parsed to obtain the content from the retrieved pages. Since there may be multiple relevant web pages, and the content within them is often extensive, with some content highly relevant to the user's query and others less so, this embodiment further optimizes the retrieved content by summarizing it to select the most relevant summary content as the final result.
[0049] As described above, fusing internal and external multi-source search results can yield richer information related to the user's query. Furthermore, it allows for the completion of response dependency information. For example, response dependency information can include the source of the response information, related images, etc. The results of internal and external multi-source searches and the completed information are then used as the target search results.
[0050] S104. Input the target retrieval results and user query information into the large language model to obtain the target question-and-answer results relative to the user query information.
[0051] In this embodiment, the target retrieval results and user query information need to be customized in conjunction with the current business scenario to generate instructions or prompts submitted to the large language model. These instructions or prompts contain richer information than the user query information and follow a prompt customization engineering strategy. Specifically, the target retrieval results are injected into the prompt, which, together with the user query information, serves as the prompt information. This prompt information is then input into the large language model, which performs summary understanding and reasoning on the content of the target retrieval results, thereby generating high-quality question responses. For example, prompt customization includes general instructions and control instructions; user query information includes user information, user context, planning signals, etc.; and the target retrieval results are obtained based on the user query. These serve as inputs to the large language model. It should be noted that, in addition to generating high-quality question responses, this embodiment also completes the question responses with information. For example, the completed information may include the webpage source of the question response, image links related to the question response, and, if multimedia information is included, playback links. This completed information and the question response can be used together as the answer to the user query information, denoted as the target answer result.
[0052] Understandably, the execution steps described above can be integrated as a plugin into the large language model. When the current business scenario of a user's inquiry is a question-and-answer intent, the plugin is used to enhance the retrieval of the user's inquiry. When the current business scenario of a user's inquiry is not a question-and-answer intent, but rather some other meaningless dialogue, it does not need to be processed by the plugin; instead, a fallback response is directly provided based on the large language model, achieving a more user-friendly approach.
[0053] This invention provides a retrieval-enhanced question-answering method for large language models. It combines internal retrieval based on knowledge graph data with external retrieval based on search engines to perform multi-path retrieval and obtain target retrieval results. These results are then used to assist in generating responses for the large language model. The external knowledge acquisition addresses the illusion problem of large language models. By combining internal and external retrieval, the generated target retrieval results are richer and more accurate, thus more precisely resolving the illusion problem of large language models.
[0054] As an optional embodiment of the present invention, the method can be further optimized based on the above embodiments by including: displaying the target question-and-answer result on the human-computer interaction interface.
[0055] The target question-and-answer result includes not only the response text to the user's query but also hyperlinks associated with the response text, displayed on the human-computer interaction interface. Based on the target answer result displayed on the interface, users can not only see the response text to their query but also perform related operations based on the displayed hyperlinks. For example, the displayed hyperlinks can include webpage links, image links, and playback links. When a user wants to access the source webpage of the response text, they can trigger the webpage link to navigate to and browse the webpage. When a user wants to access an image associated with the response text, they can trigger the image link to navigate to and browse the image page associated with the response text. If the response text includes multimedia information such as music or video, the user can trigger the displayed playback link to play the multimedia.
[0056] The above technical solution adds the function of displaying the results of the target question. When displaying the results, it includes not only the reply text to the user's query information, but also related hyperlinks. Based on the hyperlinks, users can easily and quickly access related pages, which improves the convenience for users.
[0057] Example 2
[0058] Figure 2 This is a flowchart illustrating another retrieval-enhanced large language model question-answering method provided in Embodiment 2 of the present invention. This embodiment is a further optimization of the above embodiment. In this embodiment, the limitations of "performing intent understanding on the user query information and obtaining intent understanding results" are further optimized, as are the limitations of "performing retrieval according to the retrieval plan and obtaining target retrieval results relative to the user query information" and "inputting the target retrieval results and the user query information into the large language model to obtain target question-answering results relative to the user query information" are further optimized.
[0059] like Figure 2 As shown in the figure, this embodiment 2 provides a large language model question answering method based on retrieval enhancement, which specifically includes the following steps:
[0060] S201. Receive user query information and perform intent understanding on the user query information based on at least one type of model to obtain the model output result corresponding to the user query information.
[0061] The model output should include at least the current business scenario of the user's inquiry, the search call request, and the entities involved.
[0062] This step receives user-input queries. Upon receiving the queries, the system performs intent understanding based on different models to obtain the current business context, search request, and involved entities. This understanding is then used as the model's output. For example, if the user query is "What are the tourist attractions in XX city?", the system first determines that the query is a question-and-answer intent, with entities including XX city and tourist attractions. A corresponding search request is then generated to invoke the relevant functions.
[0063] S202. If the current business scenario is a question-and-answer intent, then process the model output to obtain the retrieval information used for retrieval as the intent understanding result.
[0064] In this embodiment, if the current business scenario is a question-and-answer intent, the model output will be processed to obtain retrieval information such as the category, query command, and slots related to the user's query information. This retrieval information is used for subsequent retrieval and is denoted as the intent understanding result. The category of the user's query information can be divided into primary categories, secondary categories, etc. For example, primary categories include vehicle control domain, navigation domain, etc., while secondary categories are more detailed lower-level categories. The category of the user's query information is equivalent to a multi-level analysis of the question-and-answer intent. The query command is mainly for multi-round queries. Based on the understanding of the user's query information, a corresponding query command is generated. For example, if the first round question is "What is the capital of XX province?", and the second round question is "What is its area?", then the query command for the second round question needs to be rewritten to generate "What is the area of the capital of XX province?" as the query command for this round. A slot refers to a location used to fill specific information. These slots usually correspond to the specific needs of the question or a specific part of the answer. By filling these slots, the system can better understand and generate answers, thereby improving the accuracy and relevance of the question and answer. For example, "XX province" and "capital" are slots.
[0065] S203. Based on the intent understanding results and the determined current business scenario, and in conjunction with the pre-built search template library, generate a search plan for the user's query information.
[0066] S204. Based on the retrieval plan, perform internal retrieval based on knowledge graph data and external retrieval based on search engines to obtain multi-path retrieval results.
[0067] In this embodiment, internal retrieval is performed based on internal knowledge graph data to retrieve content related to the user's query. For example, the knowledge graph data may include internal domain knowledge, external domain knowledge, and general knowledge graphs. External retrieval is then performed using third-party search engines to find richer content related to the user's query. During external retrieval, since multiple relevant web pages may exist, and each page contains a large amount of content, some highly relevant to the user's query and others less so, this embodiment further refines the externally retrieved content by summarizing it to select the most relevant summary content as the final external retrieval result. Furthermore, the final result obtained from external retrieval can be standardized using a knowledge graph approach to obtain a standard-formatted retrieval content. The results of both internal and external retrieval are then used as the multi-path retrieval result.
[0068] As a specific implementation method, the steps of performing internal retrieval based on knowledge graph data and external retrieval based on search engines according to the retrieval plan to obtain multi-path retrieval results can be optimized, including:
[0069] a1) Based on the internal retrieval plan in the retrieval plan, perform knowledge graph retrieval and semantic full-text retrieval on the internal knowledge graph data to obtain internal retrieval results.
[0070] In this embodiment, the generated retrieval plan includes an internal retrieval plan, which is based on internal retrieval of internal knowledge graph data. The internal retrieval plan includes knowledge graph retrieval and semantic full-text retrieval. The results obtained from the internal retrieval are recorded as internal retrieval results.
[0071] b1) Perform external searches on the search engine according to the external search plan in the search plan, obtain external search results, and optimize the external search results to obtain optimized external search results.
[0072] In this embodiment, when performing external searches based on a search engine, the relevant web pages found need to be automatically crawled using web crawler technology, and then their content parsed to obtain the content from the retrieved relevant web pages. After obtaining the content from the web pages, the retrieved content can be stored in document form, and the retrieved content can also be standardized according to knowledge graph data to obtain external search content in a standard format. During external searches, since there may be multiple relevant web pages, and the content contained in each page is also substantial, with some content highly relevant to the user's query and others less so, this embodiment further optimizes the externally retrieved content, selecting the summary content with high relevance to the user's query as the final result of the external search, denoted as the optimized external search result.
[0073] Furthermore, the steps for optimizing external search results to obtain optimized external search results can be optimized, including:
[0074] b11) Segment each document contained in the external search results into paragraphs, and for each document, input the paragraph vector of each paragraph obtained after the document segmentation and the question vector of the user query information into a preset first relevance model to obtain the first relevance score of each paragraph in the document and the user query information.
[0075] In this embodiment, the external search results may contain several documents. When optimizing the documents, each document is first segmented into paragraphs. Paragraph segmentation can be performed using a pre-trained paragraph segmentation model. The document is input into the paragraph segmentation model, and the output is the segmented paragraphs.
[0076] The model used to calculate the relevance between paragraphs and user queries is denoted as the first relevance model. For example, the first relevance model can employ a dual-tower model (Base Graph Embedding, BGE). For each document, after segmentation, the paragraphs are converted into paragraph vectors, and the user queries are converted into question vectors. Then, the paragraph vectors of each paragraph and the question vectors of the user queries are input into the first relevance model to obtain the relevance score between each paragraph and the user queries, denoted as the first relevance score.
[0077] For example, suppose the question is "In what year was XX car launched?". An external search engine finds three relevant web pages. The content of these three pages is obtained through web crawling and content parsing, and stored as three documents, denoted as Document 0, Document 1, and Document 2. Each document is then segmented into paragraphs. Taking Document 0 as an example, it is segmented into five paragraphs: paragraph 0, paragraph 1, paragraph 2, paragraph 3, and paragraph 4. The paragraph vectors of these five paragraphs, along with the question vector of the user's query, are input into a first relevance model. The relevance scores are: paragraph 0: 0.5; paragraph 1: 0.6; paragraph 2: 0.9; paragraph 3: 0.7; and paragraph 4: 0.4.
[0078] b12) For each document, based on the first relevance score, obtain the paragraphs whose ranking is greater than or equal to the first top ranking threshold, and use them as the target paragraphs after document optimization.
[0079] The first top ranking threshold can be set according to actual needs. Specifically, for each document, paragraphs are ranked from highest to lowest based on the first relevance score. Paragraphs with a ranking greater than or equal to the first top ranking threshold are selected as target paragraphs for document optimization. These target paragraphs can be considered the most relevant to the user's query, thereby removing paragraphs with weak relevance to the user's query and achieving a concise summary. Summary optimization is performed on each document obtained from an external search, resulting in an optimized summary for each document.
[0080] For example, continuing with the above example, assuming the first top ranking threshold is 3, paragraphs 2, 3, and 1 are obtained as the optimized summary of document 1.
[0081] b13) For each document, input the paragraph vector and question vector of the top-ranked target paragraph into the preset second relevance model to obtain the second relevance score of each document to the user's query information.
[0082] Another model used to calculate the relevance between paragraphs and user queries is denoted as the second relevance model. For example, the second relevance model can employ a single-tower model. Specifically, the paragraph vector and question vector of the top-ranked target paragraph are input into the second relevance model to obtain a relevance score between the target paragraph and the user query. This score is then used as the document's relevance score to the user query and is denoted as the second relevance score. It is important to understand that a second relevance score is calculated for each document.
[0083] For example, continuing with the above example, paragraph 2 in document 0 is the paragraph with the highest relevance ranking. The paragraph vector of paragraph 2 and the question vector of the user's query information are input into the second relevance model, so that the relevance score of paragraph 2 and the question vector of the user's query information is 0.7. 0.7 is used as the relevance score of document 0 and the question to be answered.
[0084] b14) Based on each second relevance score, obtain the target document with a ranking greater than or equal to the second top ranking threshold, and use the optimized target paragraph of the target document as the external search result.
[0085] The second top ranking threshold can be set according to actual needs. Specifically, based on the second relevance score of each document to the user's query information, documents are ranked from high to low. Target documents with a ranking greater than or equal to the second top ranking threshold are obtained, and then the optimized summaries of these target documents are used as external search results. For example, assuming document 0 has a relevance score of 0.7 to the user's query information, document 1 has a relevance score of 0.9, document 2 has a relevance score of 0.6, and the second top ranking threshold is 2, then the optimized summaries of document 1 and document 0 are taken as external search results.
[0086] The above technical solution specifies the steps for summary optimization, and achieves more accurate and refined external retrieval results by optimizing externally retrieved documents and document paragraphs, thus providing basic data for solving the large model illusion problem.
[0087] c1) Combine internal search results and optimized external search results as multiple search results.
[0088] Specifically, internal search results and optimized external search results are combined to form multi-path search results.
[0089] The above technical solution utilizes knowledge graph data for knowledge graph retrieval and semantic full-text retrieval, and performs external retrieval and summary optimization through a search engine, thereby obtaining richer, more comprehensive, and refined search results. This provides more accurate auxiliary information for subsequent large language model inference.
[0090] S205. Perform multi-factor sorting on the multi-path search results to obtain the target search results.
[0091] When ranking multiple search results, factors such as relevance, timeliness, authority, and content quality of the search results to the user's query can be considered. A comprehensive ranking of the search results is then performed based on these factors. A predetermined number of search results are selected as the target search results. For example, optimized search results can be input into a large language model in JsonResponse format.
[0092] As a specific implementation method, the steps of multi-factor ranking of multi-way search results to obtain the target search results can be optimized, including:
[0093] a2) Determine the relevance, timeliness, and authority of the multi-path search results to the user's query information.
[0094] This step is used to determine the relevance of each search result in the multi-path search results to the user query. The relevance determination can be based on a trained relevance model to evaluate the relevance between the search results and the user query: for each search result, if the search result includes both a title and content, a relevance score can be calculated for the title and the user query, and a relevance score for the content and the user query can be calculated separately. The combined score of these two relevance scores is then used as the relevance score between the search result and the user query. The relevance of each search result to the user query can be represented based on the relevance score. For example, values in the range [0,1] are used as the relevance score; a larger value indicates a higher relevance between the search result and the user query, and a smaller value indicates a lower relevance.
[0095] Timeliness is used to assess whether search results are up-to-date. This can be determined based on the current query time and the publication time of the content related to the search results; the closer the publication time is to the current query time, the higher the timeliness. Authority is used to assess whether search results originate from authoritative websites, portals, etc. The publication time and content quality of search results can be obtained by standardizing the search content based on a knowledge graph. It should be noted that evaluation factors can include not only relevance, timeliness, and authority, but also content quality to assess the quality of search results. Evaluation factors can be set according to actual needs; no specific restrictions are imposed here.
[0096] b2) Based on the relevance, timeliness, and authority of the multi-source search results and the user's query information, the search results in the multi-source search results are sorted by multiple factors to determine the comprehensive ranking of the multi-source search results.
[0097] In this embodiment, when sorting the search results in the multi-way search results, it is necessary to consider the relevance, timeliness, and authority of the search results to the user's query information, and to perform multi-factor sorting on each search result in the multi-way search results to obtain a comprehensive sort of the multi-way search results.
[0098] c2) Obtain the top search results from the comprehensive ranking that have a ranking greater than or equal to the third top ranking threshold.
[0099] The third top ranking threshold can be set according to actual needs. This step is used to obtain search results with a ranking greater than or equal to the third top ranking threshold from the overall ranking as the top search results. For example, assuming the third top ranking threshold is 20, the top 20 search results from the overall ranking are taken as the top search results.
[0100] d2) Complete the source citations for the top search results and use the completed top search results as the target search results.
[0101] In this embodiment, the determined top search results will also be supplemented with source citations, such as links to the source webpages of the search results, webpage links to related images of the search results, and webpage links to videos played in the search results. The supplemented top results will be used as the target search results.
[0102] The aforementioned technical solution evaluates search results from multiple perspectives, including relevance, timeliness, authority, and content quality, to obtain the optimal search results that best match the user's query. This ensures the relevance, timeliness, authority, and quality of the relevant content, resulting in more accurate search results that serve as auxiliary information for subsequent large language model inference, thus guaranteeing more accurate question-and-answer results. Simultaneously, information completion of the search results ensures richer results and provides foundational data for displaying relevant links on the human-computer interaction interface. Multi-dimensional search result filtering guarantees higher quality and more complete information in the search results.
[0103] S206. Input the target retrieval results and user query information into the large language model to obtain the initial question and answer results relative to the user query information.
[0104] Specifically, the user query and preferred search terms are input into a large language model. The large language model performs summary understanding and reasoning on the content of the target search results, thereby generating a high-quality question response. This question response serves as the initial question-and-answer result relative to the user query. For example, the initial question-and-answer result can be output in Markdown text format.
[0105] S207. Based on the response source citations in the target search results, generate response source links relative to the initial question and answer results.
[0106] The reply source link is a hyperlink used to access content related to the initial question and answer result.
[0107] In this embodiment, the target search results include not only internal and external search results but also response source references. Response source references can be understood as indicating which web pages, databases, or knowledge graph data these search results were retrieved from. Response source links can also serve as a basis for judging the credibility of the initial question results. Based on the response source references in the target search results, corresponding web page links to access these contents, image links to related images, and even playback links can be generated for each search result; these links are called response source links. Response source links can be seen as hyperlinks used to access content related to the initial results.
[0108] For example, assuming a user asks "What are some good movies?", in addition to generating a reply text related to good movies, we can also generate a link to access that reply text, a link to the movie cover image, and a link to play the movie. The reply text can be considered the initial question-and-answer result, and the links to access that reply text, the movie cover image, and the movie playback link can be considered the source links of the reply.
[0109] S208. Use the initial question and answer results and the link to the source of the response as the target question and answer results.
[0110] In this embodiment, the initial question-and-answer result and the link to the source of the reply associated with the initial question-and-answer result are considered together as the target question-and-answer result. The target question-and-answer result can be considered to include not only the reply text, but also the webpage link to access that content, links to access related images, and even playback links, etc.
[0111] The above technical solution specifies the steps of intent understanding, retrieval, and obtaining target question-and-answer results relative to user query information. It solves the illusion problem of large models, realizes the integration of internal knowledge graph and external third-party search capabilities, and optimizes summary generation, thereby improving the authenticity and relevance of large model question-and-answer responses and enabling real-time information acquisition.
[0112] Example 3
[0113] Figure 3 This is a schematic diagram of a retrieval-enhanced large language model question-answering device provided in Embodiment 3 of the present invention. This device is applicable to question-answering based on a large language model. The retrieval-enhanced large language model question-answering device can be implemented in hardware and / or software, and is generally integrated into an electronic device. For example... Figure 3 As shown, the system includes: an intent understanding module 31, a plan generation module 32, a retrieval module 33, and a result generation module 34; wherein,
[0114] The intent understanding module 31 is used to receive user query information and perform intent understanding on the user query information to obtain intent understanding results;
[0115] The plan generation module 32 is used to generate a retrieval plan relative to the user's query information based on the intent understanding results and the determined current business scenario, combined with a pre-built retrieval template library.
[0116] The retrieval module 33 is used to perform retrieval according to the retrieval plan and obtain the target retrieval results relative to the user's query information. The retrieval plan includes an internal retrieval plan based on knowledge graph data and an external retrieval plan based on a search engine.
[0117] The result generation module 34 is used to input the target retrieval results and user query information into the large language model to obtain the target question-answering results relative to the user query information.
[0118] The above technical solution combines internal retrieval based on knowledge graph data with external retrieval based on search engines to obtain target retrieval results through multi-path retrieval. These results are then used to assist in the response generation of the large language model, thus solving the illusion problem of the large language model by acquiring external knowledge. By combining internal and external retrieval, the generated target retrieval results are richer and more accurate, enabling a more precise solution to the illusion problem of the large language model.
[0119] Optionally, the intent understanding module 31 is specifically used for:
[0120] Based on at least one type of model, the intent of the user query information is understood, and the model output result corresponding to the user query information is obtained. The model output result includes at least the current business scenario of the user query information, the search call request, and the entities involved.
[0121] If the current business scenario is a question-and-answer intent, the model output is processed to obtain retrieval information for use in the search, which is then used as the intent understanding result.
[0122] Optionally, the retrieval module 33 may include:
[0123] The retrieval unit is used to perform internal retrieval based on knowledge graph data and external retrieval based on search engines according to the retrieval plan, and obtain multi-path retrieval results;
[0124] The retrieval and sorting unit is used to sort the results from multiple sources by multiple factors to obtain the target retrieval results.
[0125] Optionally, the retrieval unit is specifically used for:
[0126] Based on the internal retrieval plan in the retrieval plan, knowledge graph retrieval and semantic full-text retrieval are performed on the internal knowledge graph data to obtain internal retrieval results;
[0127] Perform external searches on search engines according to the external search plan in the search plan, obtain external search results, and optimize the summary of the external search results to obtain optimized external search results;
[0128] The internal search results and the optimized external search results are used as multiple search results.
[0129] Optionally, the retrieval unit is used to perform the step of optimizing external retrieval results to obtain optimized external retrieval results, including:
[0130] Each document contained in the external search results is segmented into paragraphs. For each document, the paragraph vectors of each segment and the question vector of the user query information are input into the preset first relevance model to obtain the first relevance score of each paragraph in the document and the user query information.
[0131] For each document, based on the first relevance score, the paragraphs with a ranking greater than or equal to the first top ranking threshold are selected as the document's optimized summary.
[0132] For each document, the paragraph vector and question vector of the top-ranked target paragraph are input into a pre-defined second relevance model to obtain a second relevance score for each document to the user's query information;
[0133] Based on each second relevance score, target documents with a ranking greater than or equal to the second top ranking threshold are obtained, and the optimized target paragraphs of the target documents are used as external search results.
[0134] Optionally, the retrieval sorting unit is specifically used for:
[0135] Determine the relevance, timeliness, and authority of multi-path search results to user queries;
[0136] Based on the relevance, timeliness, and authority of the multi-source search results and the user's query information, the search results in the multi-source search results are ranked by multiple factors to determine the comprehensive ranking of the multi-source search results.
[0137] Obtain top search results with a ranking greater than or equal to the third-highest ranking threshold from the comprehensive ranking;
[0138] The source citations of the top search results are completed, and the completed top search results are used as the target search results.
[0139] Optionally, the result generation module 34 is specifically used for:
[0140] The user query information and the target retrieval results are input into the large language model to obtain the initial question and answer results relative to the user query information;
[0141] Based on the source citations in the target search results, a source link for the response relative to the initial question and answer results is generated. The source link for the response is used as a hyperlink to access content related to the initial question and answer results.
[0142] Use the initial question and answer results and the link to the source of the response as the target question and answer results.
[0143] Optionally, the device further includes a result display module for:
[0144] The target question and answer results are displayed on the human-computer interaction interface.
[0145] The retrieval-enhanced large language model question answering device provided in the embodiments of the present invention can execute the retrieval-enhanced large language model question answering method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0146] Example 4
[0147] Figure 4 This is a schematic diagram of an electronic device according to Embodiment 4 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0148] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded into the RAM 43 from storage unit 48. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0149] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0150] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as a retrieval-enhanced large language model question answering method.
[0151] In some embodiments, the retrieval-enhanced large language model question-answering method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the retrieval-enhanced large language model question-answering method described above can be performed. Alternatively, in other embodiments, processor 41 can be configured to perform the retrieval-enhanced large language model question-answering method by any other suitable means (e.g., by means of firmware).
[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0153] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0154] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0155] To provide interaction with the user, the systems and technologies described herein can be implemented in a vehicle having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the vehicle. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0156] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0157] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0158] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the retrieval-enhanced large language model question-answering method as provided in any embodiment of this invention.
[0159] In implementing a computer program product, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0160] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0161] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A question-answering method for a large language model based on retrieval enhancement, characterized in that, include: Receive user query information and perform intent understanding on the user query information to obtain intent understanding results; Based on the intent understanding results and the determined current business scenario, and combined with the pre-built search template library, a search plan is generated relative to the user query information; The retrieval is performed according to the retrieval plan to obtain the target retrieval results relative to the user query information. The retrieval plan includes an internal retrieval plan based on knowledge graph data and an external retrieval plan based on a search engine. The target retrieval results and the user query information are input into the large language model to obtain the target question-answering results relative to the user query information.
2. The method according to claim 1, characterized in that, The step of performing intent understanding on the user query information to obtain intent understanding results includes: The intent of the user query information is understood based on at least one type of model to obtain the model output result corresponding to the user query information. The model output result includes at least the current business scenario of the user query information, the search call request, and the entities involved. If the current business scenario is a question-and-answer intent, then the output results of each model are processed to obtain retrieval information used for retrieval as the intent understanding result.
3. The method according to claim 1, characterized in that, The step of performing a search according to the search plan to obtain target search results relative to the user query information includes: According to the retrieval plan, internal retrieval based on knowledge graph data and external retrieval based on search engine are performed to obtain the multi-path retrieval results; The multi-factor sorting of the multi-path search results is performed to obtain the target search results.
4. The method according to claim 3, characterized in that, The step of performing internal retrieval based on knowledge graph data and external retrieval based on search engines according to the retrieval plan to obtain the multi-path retrieval results includes: According to the internal retrieval plan in the retrieval plan, knowledge graph retrieval and semantic full-text retrieval are performed on the internal knowledge graph data to obtain internal retrieval results; According to the external retrieval plan in the retrieval plan, an external retrieval is performed on the search engine to obtain external retrieval results, and the external retrieval results are optimized to obtain optimized external retrieval results; The internal search results and the optimized external search results are used as the multi-path search results.
5. The method according to claim 4, characterized in that, The optimization of the external search results to obtain optimized external search results includes: Each document contained in the external search results is segmented into paragraphs. For each document, the paragraph vectors of each segment obtained after the document segmentation and the question vector of the user query information are input into a preset first relevance model to obtain the first relevance score between each paragraph in the document and the user query information. For each document, based on each of the first relevance scores, obtain the paragraphs whose ranking is greater than or equal to the first top ranking threshold, and use them as the target paragraphs after the document optimization. For each document, the paragraph vector of the top-ranked target paragraph and the question vector are input into a preset second relevance model to obtain a second relevance score for each document and the user's query information; Based on each of the second relevance scores, target documents with a ranking greater than or equal to the second top ranking threshold are obtained, and the optimized target paragraphs of the target documents are used as the external search results.
6. The method according to claim 3, characterized in that, The step of performing multi-factor sorting on the multi-path retrieval results to obtain the target retrieval results includes: Determine the relevance, timeliness, and authority of the multi-path retrieval results to the user query information; Based on the relevance, timeliness, and authority of the multi-path retrieval results and the user query information, the retrieval results in the multi-path retrieval results are sorted by multiple factors to determine the comprehensive ranking of the multi-path retrieval results; Obtain the top search results with a ranking greater than or equal to the third top ranking threshold from the comprehensive ranking; The top search results are supplemented with source citations, and the supplemented top search results are used as the target search results.
7. The method according to claim 1, characterized in that, The step of inputting the target retrieval result and the user query information into a large language model to obtain a target question-answering result relative to the user query information includes: The target retrieval results and the user query information are input into the large language model to obtain the initial question and answer results relative to the user query information; Based on the response source reference in the target search results, a response source link relative to the initial question and answer results is generated. The response source link is used to access hyperlinks of content associated with the initial question and answer results. The initial question-and-answer result and the link to the source of the response are used as the target question-and-answer result.
8. The method according to claim 1, characterized in that, Also includes: The target question-and-answer results are displayed on the human-computer interaction interface.
9. A large language model question-answering device based on retrieval enhancement, characterized in that, include: The intent understanding module is used to receive user query information and perform intent understanding on the user query information to obtain intent understanding results; The plan generation module is used to generate a retrieval plan relative to the user query information based on the intent understanding results and the determined current business scenario, combined with a pre-built retrieval template library. The retrieval module is used to perform retrieval according to the retrieval plan and obtain target retrieval results relative to the user query information. The retrieval plan includes an internal retrieval plan based on knowledge graph data and an external retrieval plan based on a search engine. The result generation module is used to input the target retrieval result and the user query information into the large language model to obtain the target question answer result relative to the user query information.
10. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the retrieval-enhanced large language model question answering method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the retrieval-enhanced large language model question-answering method as described in any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the retrieval-enhanced large language model question-answering method as described in any one of claims 1-8.