Intelligent question answering method and system based on question index retrieval enhancement generation technology

By dividing the document into pieces and generating hypothetical problem indexes and generating answers in combination with large language models, the problem of insufficient information coverage when processing semantic complex queries is solved, and the retrieval performance and accuracy are improved.

CN120124640AActive Publication Date: 2025-06-10BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510179300.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

When existing search enhancement generation (RAG) technologies deal with semantic complex or ambiguity, there is a problem of insufficient information coverage, resulting in limited improvement in retrieval performance.

Method used

By dividing the documents in the knowledge base into multiple document shards, extracting abstract questions as hypothetical questions, and building a question index query library, combining large language models to generate answers, improve retrieval performance.

Benefits of technology

By generating an ordered set of hypothetical questions, the semantic coverage and diversity of the search module are enhanced, the semantic gap between the question and the answer is reduced, and the hit rate and accuracy of the search are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124640A_ABST
    Figure CN120124640A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent question answering method and system based on a question index retrieval enhancement generation technology, and the method comprises the steps: segmenting a knowledge base document into document fragments, extracting a plurality of abstract questions from each document fragment as hypothetical questions, and taking an ordered set of hypothetical question vectors of the document fragments as an index to construct a query library; a user question is rewritten and vectorized to generate a user question vector, then document fragments related to the user question are obtained through index retrieval, and the process is as follows: the similarity between the user question vector and each hypothetical question vector of each document fragment is calculated, and the hypothetical question vectors with the similarity greater than a threshold value are selected; then selecting document fragments to which a plurality of hypothetical problem vectors with large similarity belong to form a related text set; and forming the user question and the related text set into cue words, and generating answers through a large language model. The invention relates to the technical field of natural language processing, and can effectively improve the retrieval performance of the RAG technology and ensure the accuracy of large model generation answers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an intelligent question - answering method and system based on question - index retrieval - enhanced generation technology, and relates to the technical field of natural language processing. Background Art

[0002] In the Retrieval - Augmented Generation (RAG) framework, the performance of the retrieval module highly depends on the semantic similarity between the input question and the document or fragment. However, due to the diversity and complexity of natural language expressions, there is often a semantic gap between the input question and the semantics in the document, resulting in limited retrieval effects. In international research, common optimization means include question rewriting and hypothetical document embedding, mainly based on semantic adjustment of user input. These methods focus on optimization from the query side, assuming that generation depends on the model's understanding of the query, but do not fully consider the mining of potential information on the document side. They perform well in dealing with semantically complex or polysemous queries, but when faced with long texts or multi - topic documents, there may still be problems of insufficient information coverage, thus limiting the improvement of retrieval performance.

[0003] Therefore, how to effectively improve the retrieval performance of RAG technology to ensure the accuracy of the answers generated by large models has become a key technical issue that technicians focus on. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an intelligent question - answering method and system based on question - index retrieval - enhanced generation technology, which can effectively improve the retrieval performance of RAG technology, thereby ensuring the accuracy of the answers generated by large models.

[0005] To achieve the above - mentioned purpose, the present invention provides an intelligent question - answering method based on question - index retrieval - enhanced generation technology, including:

[0006] Step 1: Split all documents in the knowledge base into multiple document shards, then extract several summary questions from each document shard as the hypothetical questions of the document shard, then vectorize each hypothetical question to form an ordered set of hypothetical question vectors of the document shard, and finally use the ordered set of hypothetical question vectors of each document shard as an index to construct a query library;

[0007] Step 2: Rewrite the question input by the user, vectorize it to generate a user question vector, and then retrieve relevant document shards from the query library through index retrieval. The specific process of index retrieval is as follows: Calculate the similarity between the user question vector and each hypothetical question vector in the ordered set of hypothetical question vectors of each document shard in the query library, select all hypothetical question vectors with similarities greater than the threshold, then sort the selected hypothetical question vectors in descending order of similarity, select multiple hypothetical question vectors ranked at the front, and finally extract the document shards to which the selected hypothetical question vectors belong to form a text set related to the user question;

[0008] Step 3: Combine the user question and the text set related to the user question to form a prompt, and generate an answer corresponding to the user question through a large language model.

[0009] To achieve the above object, the present invention also provides an intelligent question-answering system based on a question index retrieval enhancement generation technology, including:

[0010] An index generation device, configured to split all documents in the knowledge base into multiple document shards, then extract several summary questions from each document shard as hypothetical questions of the document shard, vectorize each hypothetical question to form an ordered set of hypothetical question vectors of the document shard, and finally use the ordered set of hypothetical question vectors of each document shard as an index to build a query library;

[0011] A document selection device, configured to rewrite the question input by the user, vectorize it to generate a user question vector, and then retrieve relevant document shards from the query library through index retrieval. The specific process of index retrieval is as follows: Calculate the similarity between the user question vector and each hypothetical question vector in the ordered set of hypothetical question vectors of each document shard in the query library, select all hypothetical question vectors with similarities greater than the threshold, then sort the selected hypothetical question vectors in descending order of similarity, select multiple hypothetical question vectors ranked at the front, and finally extract the document shards to which the selected hypothetical question vectors belong to form a text set related to the user question;

[0012] An answer generation device, configured to combine the user question and the text set related to the user question to form a prompt, and generate an answer corresponding to the user question through a large language model.

[0013] To achieve the above object, the present invention also provides a computing device, including:

[0014] A memory and a processor;

[0015] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the intelligent question-answering method based on the problem index retrieval enhancement generation technology are implemented.

[0016] To achieve the above object, the present invention also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the intelligent question-answering method based on the problem index retrieval enhancement generation technology are implemented.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: Since the expression mode of the original question may be inconsistent with the document content, and semantic loss will occur during the vectorization process of long documents, the present invention can supplement diverse semantic perspectives, mine potential information, reduce the semantic gap between the question and the answer by automatically generating hypothetical questions by the large language model. The constructed metadata (such as paragraph summaries and hypothetical questions) can provide more semantic anchors for the retrieval module, enabling it to find relevant content within a wider semantic range. Specifically, the present invention uses the large model to generate questions that the document can answer, and then calculates the similarity between the original question and the hypothetical question during retrieval to reduce the semantic gap between the question and the answer. At the same time, two indicators of semantic coverage and semantic diversity are applied to improve the quality of the ordered set of hypothetical questions and enhance the hit rate and accuracy of retrieval; the present invention effectively compresses the information volume of the source document through the abstract question extraction method, and only extracts several core semantic information from it. Specifically, several key hypothetical questions are generated from the source document, and the answers to these questions can be directly obtained from the document or through reasoning. These key questions constitute a compact and efficient semantic space. Compared with directly retrieving in the original semantic space, the constructed semantic space enables the query question to be retrieved within a semantic range closer to it, thus greatly reducing the scale of the query semantic space; in the process of using the embedding model to map the text to the high-dimensional semantic space, each text (such as a question or a document) is represented by an embedding vector. However, since the embedding is usually a compressed representation of the content (such as a sentence vector or a paragraph vector), information loss may occur when capturing the detailed information of the text. The present invention can effectively mitigate the impact of detail loss by extracting short text questions from the long text and embedding them into the high-dimensional semantic space, and ensure that the ordered set of hypothetical questions can fully cover the important semantic units in the document and have semantic differences. Description of the Drawings

[0018] Figure 1 It is a flowchart of an intelligent question-answering method based on the problem index retrieval enhancement generation technology shown in an exemplary embodiment of the present invention.

[0019] Figure 2as shown in an exemplary embodiment of the present invention Figure 1 The specific step flowchart of extracting several summary questions from each document fragment as the hypothetical questions of the document fragment in Step 1.

[0020] Figure 3 The structural schematic diagram of an intelligent question-answering system based on the question index retrieval enhanced generation technology as shown in an exemplary embodiment of the present invention.

[0021] Figure 4 The structural schematic diagram of a computer device as shown in an exemplary embodiment of the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.

[0023] As Figure 1 shown, an intelligent question-answering method based on the question index retrieval enhanced generation technology of the present invention includes:

[0024] Step 1: Cut all documents in the knowledge base into multiple document fragments, then extract several summary questions from each document fragment as the hypothetical questions of the document fragment, then vectorize each hypothetical question to form an ordered set of hypothetical question vectors of the document fragment, and finally use the ordered set of hypothetical question vectors of each document fragment as an index to construct a query library;

[0025] Step 2: Rewrite the question input by the user, vectorize it to generate a user question vector, and then obtain the document fragments related to the user question from the query library through index retrieval. The specific process of index retrieval is as follows: calculate the similarity between the user question vector and each hypothetical question vector in the ordered set of hypothetical question vectors of each document fragment in the query library, and select all hypothetical question vectors with a similarity greater than the threshold, then sort the selected hypothetical question vectors in descending order of similarity, select multiple hypothetical question vectors ranked at the front, and finally extract the document fragments to which the selected hypothetical question vectors belong to form a text set related to the user question;

[0026] The COT (Chain of Thought) method can be used to rewrite the question input by the user;

[0027] Step 3: Use the user question and the text set related to the user question as prompt words, and generate an answer corresponding to the user question through a large language model.

[0028] In step one, it is also possible to extract the content of the documents in the knowledge base and convert non-text data into text data. When the uploaded documents are in formats such as HTML, PDF, docx, etc., corresponding extractors are used to extract the document content. If there is image data, the specific processing process is as follows:

[0029] First, extract the caption or annotation of the image, then perform OCR (i.e., Optical Character Recognition) on the image to extract character information, and then input the caption or annotation of the image and the character information into a multi-modal large model to generate image description information. Finally, use the caption or annotation of the image, the character information, and the image description information to replace the image data in the knowledge base.

[0030] In step one, it is possible to generate summary questions for each document shard through a large language model, and then, based on semantic coverage and diversity, select several hypothetical questions from the summary questions. As Figure 2 shown, taking the document shard D as an example, Figure 1 In step one, extracting several summary questions from each document shard as the hypothetical questions of the document shard may include:

[0031] Step 11: Generate the corresponding summary content and summary question set for the document shard D through a large language model. The summary question set contains n summary questions, where n is the preset number of summary questions, and its value can be set according to actual business needs;

[0032] It is possible to require the large language model to extract summary questions from multiple perspectives such as "usage", "definition", "facts", etc., and the n extracted summary questions form the summary question set;

[0033] Step 12: Construct an ordered set Q of hypothetical questions and initialize it as an empty set. At the same time, initialize the set semantic coverage Cov 0 and the set semantic diversity value Div 0 to 0;

[0034] Step 13: Select each summary question from the summary question set one by one and calculate the information gain value corresponding to each summary question: Info_Gain(q i ) = α * (Cov(q i ) - Cov 0 ) + (1 - α) * (Div(q i ) - Div 0 ), where q i is the i-th summary question in the summary question set, and Info_Gain(q i ), Cov(q i ), and Div(q i ) are the information gain, semantic coverage, and semantic diversity of qi The corresponding information gain value, the contribution value to semantic coverage, and the contribution value to semantic diversity. α is an important weight parameter for adjusting the importance of coverage and diversity, and its value can be set according to actual business needs. Then, select the maximum value from all the information gain values corresponding to the summary questions, write the summary question corresponding to the maximum value into the ordered set Q of hypothetical questions, and delete the summary question from the set of summary questions. At the same time, update the set semantic coverage Cov 0 and the set semantic diversity value Div 0 : Cov 0 ’ = Cov 0 + Cov(q u ), Div 0 ’ = Div 0 + Div(q u ), Cov 0 ’ 、Div 0 ’ are the updated set semantic coverage and set semantic diversity values, q u is the summary question corresponding to the maximum value;

[0035] Step 14: Determine whether the number of hypothetical questions in the ordered set Q of hypothetical questions is less than the preset threshold m of the number of hypothetical questions. If so, go to Step 13; if not, vectorize each hypothetical question in the ordered set Q of hypothetical questions to generate an ordered set of hypothetical question vectors. m can be set according to actual business needs.

[0036] In Step 13, the calculation process of Cov(q i ) is as follows:

[0037] First, decompose the document shard into multiple semantic units such as paragraphs, sentences, phrases, or keywords, and use the embedding model to convert q i and each semantic unit into vector representations. Then, calculate the cosine similarity between q i and the vector representation of each semantic unit one by one. Finally, calculate Cov(q i ): Among them, Total_Units is the number of semantic units in the document shard, and NUM_Units(q i ) is the total number of cosine similarities between q i and the vector representations of all semantic units in the document shard that are greater than the preset threshold, which is used to represent the number of semantic units covered by the summary question.

[0038] The calculation formula of Div(q i ) is as follows: Among them, Q_NUM is the number of hypothetical questions in the ordered set Q of hypothetical questions, and v qj is the vector representation of the j-th hypothetical question in the ordered set Q of hypothetical questions, is q i 's vector representation, is the cosine similarity between.

[0039] To more clearly explain the technical effects of the method of the present invention, the following shows an embodiment of applying the indexing and retrieval method of the present invention:

[0040] Document fragment D:

[0041] Dujiangyan is one of the most famous ancient water conservancy projects in China and is located on the Minjiang River in Sichuan. It was built under the auspices of Li Bing and his son, and its main function is flood control and irrigation. Dujiangyan diverts the Minjiang River into the Inner River and the Outer River through the fish mouth structure, thus effectively reducing the threat of floods to downstream farmland. At the same time, the project has significantly improved the agricultural productivity in the Minjiang River Basin, forming a rich situation of the "Land of Abundance".

[0042] The completion of Dujiangyan is not only a milestone in the development of Chinese water conservancy technology, but also has a profound impact on the regional social economy. The project has promoted the development of farming culture, supported population growth and the urbanization process. Through the expansion of the irrigation network, a large number of farmlands have been efficiently utilized, providing a stable food supply for successive regimes....

[0043] The ordered set of hypothetical questions constructed for the document fragment D is as follows:

[0044] [1] How does Dujiangyan improve agricultural productivity?

[0045] [2] What is the contribution of Dujiangyan to regional social development?

[0046] [3] How do ancient water conservancy projects help with flood control?

[0047] User question: How did ancient societies use technology to improve production and maintain social stability?

[0048] In the above embodiment, since the semantics expressed by the user question are too abstract or generalized, it is impossible to directly hit the keywords or semantics related to the document. The ordered set of hypothetical questions in the present invention can connect the user question with the document content by refining the core semantics of the document, making up for the defect of direct retrieval. Although the matching degree between the user question and the document itself is not high, it can match the content in Q. Through experimental verification, when the document scale is large, the user question is fuzzy or the semantic expression is complex, the indexing and retrieval method in the present invention is particularly effective.

[0049] When the index retrieval fails to retrieve enough document shards, the present invention can also use hybrid retrieval, and the index retrieval and the hybrid retrieval can be used complementarily. In this way, when the number of document shards in the text set related to the user's question in step two is less than or equal to 1, hybrid retrieval can still be performed. The specific steps of the hybrid retrieval include:

[0050] Calculate the keyword scores of the user's question and each document shard in the knowledge base through the BM25 algorithm, and select the top K document shards with the highest keyword scores. At the same time, calculate the similarity between the user's question and each document shard in the knowledge base by calculating the cosine similarity of the embedding vectors, and select the top K document shards with the highest similarity. Finally, write the K document shards selected according to the keyword scores and similarities into the text set related to the user's question, and K can be set according to actual business needs.

[0051] Since the retrieved content of the hybrid retrieval has insufficient relevance to the question itself: Embedding and BM25 mainly focus on the semantic similarity and keyword matching between texts, but cannot ensure that the retrieved content is truly helpful for answering the question, thus affecting the system performance. For example:

[0052] User's question: What are the main causes of global warming? The retrieved content of the hybrid retrieval is as follows:

[0053] Document shard 1: In recent years, the trend of global warming has been very obvious (semantically relevant but without information value).

[0054] Document shard 2: Greenhouse gases are the main cause of global warming (directly answering the question).

[0055] In the above example, although document shard 1 is semantically relevant to the question, it does not contain the key information required to answer the question and may even interfere with the generation process.

[0056] Another example:

[0057] User's question: What is the impact of the Amazon rainforest on the global climate? The retrieved content of the hybrid retrieval is as follows:

[0058] Document shard 1: The Amazon River is the river with the largest flow in the world, and its length is about 6,400 kilometers. (Noise information)

[0059] Document shard 2: The Amazon rainforest plays a key role in the global carbon cycle and climate regulation by absorbing carbon dioxide and releasing oxygen.

[0060] In the above example, although the content of the retrieved document shard is related to "Amazon", it completely deviates from the core requirements of the question (about the climate impact of the rainforest), and may even mislead the large language model to discuss the Amazon River rather than the Amazon rainforest.

[0061] To improve the effectiveness and accuracy of retrieving documents, the present invention can also evaluate and screen the retrieval effect of hybrid retrieval through an evaluation model, and further includes:

[0062] Extract each document fragment obtained by hybrid retrieval one by one from the user question-related text set, then use the evaluation model to calculate the relevance score between the extracted document fragment and the user question. The relevance score is used to characterize whether the document fragment is useful for answering the question. Then, generate a dynamic label for identifying the relationship between the user question and the document fragment based on the relevance score, and filter accordingly to avoid interference from useless information. The evaluation model can be implemented using GPT-4, the language model MonoT5, etc. For example, when implemented using GPT-4, a method based on instruction fine-tuning can be used to score the relevance. When implemented using MonoT5, the relevance scoring method can be adopted, and an end-to-end generative scoring method can be used to directly judge Relevant or Irrelevant. The dynamic label can include: Relevant (highly relevant), that is, the document fragment directly answers the question and the information is complete; Partially Relevant (partially relevant), that is, the document fragment may contain some relevant information, but it is not direct or comprehensive enough; Irrelevant (irrelevant), that is, the document fragment has nothing to do with the question or contains useless information. If the numerical range of the relevance score is [0,1], the corresponding interval for Relevant can be set as: [0.7,1], the corresponding interval for Partially Relevant can be set as: (0.4,0.7), and the corresponding interval for Irrelevant can be set as: [0,0.4]. In this way, when the calculated relevance score belongs to the corresponding interval of Irrelevant, delete the document fragment from the user question-related text set.

[0063] The ordered set of hypothetical questions is generated to improve the retrieval efficiency and quality, but the semantic range it covers may not be precisely matched with the real user questions, which may lead to retrieval deviation. Therefore, the present invention also designs a dynamic update mechanism for the ordered set of hypothetical questions. When the document fragment retrieved by hybrid retrieval and the real user question are determined to be highly relevant (i.e., Relevant) through the evaluation model, part of the hypothetical questions can be replaced with this user question, thereby optimizing the index retrieval. It further includes:

[0064] Step A1: Extract the first document fragment from the user question-related text set;

[0065] Step A2: Determine whether the dynamic label identifying the relationship between the user question and the extracted document fragment is Relevant. If it is, continue to Step A3; if not, turn to Step A7;

[0066] Step A3: Determine whether the retrieval method of the extracted document shard is a hybrid retrieval. If so, proceed to Step A4; if not, go to Step A6;

[0067] Step A4: Calculate the similarity between the user question and the extracted document shard, and the BM25 keyword score between the user question and the extracted document shard. Then determine whether the similarity is less than a preset similarity threshold and the BM25 keyword score is less than a preset keyword score threshold. If so, proceed to Step A5; if not, go to Step A7;

[0068] Step A5: Determine whether the number of hypothetical questions in the ordered set of hypothetical questions of the extracted document shard is less than m. If so, add the user question to the end position of the ordered set of hypothetical questions of the extracted document shard, that is, the user question is the last hypothetical question in the ordered set of hypothetical questions of the extracted document shard, and then go to Step A7; if not, delete the last hypothetical question in the ordered set of hypothetical questions of the extracted document shard, and add the user question to the first position of the ordered set of hypothetical questions of the extracted document shard, that is, the user question is the first hypothetical question in the ordered set of hypothetical questions of the extracted document shard, and then go to Step A7;

[0069] Step A6: Determine whether the retrieval method of the extracted document shard is an index retrieval. If so, obtain the hypothetical question with the highest similarity to the user question from the ordered set of hypothetical questions of the extracted document shard, then move the obtained hypothetical question to the first position of the ordered set of hypothetical questions of the extracted document shard, and then proceed to Step A7; if not, proceed to Step A7;

[0070] Step A7: Determine whether all document shards in the user question-related text set have been extracted. If so, this process ends; if not, continue to extract the next document shard in the user question-related text set and go to Step A2.

[0071] As Figure 3 shown, an intelligent question-answering system based on the problem index retrieval enhancement generation technology of the present invention includes:

[0072] An index generation device, configured to slice all documents in the knowledge base into multiple document shards, then extract several summary questions from each document shard as the hypothetical questions of the document shard, then vectorize each hypothetical question to form an ordered set of hypothetical question vectors of the document shard, and finally use the ordered set of hypothetical question vectors of each document shard as an index to construct a query library;

[0073] A document selection device is used to rewrite the problem input by the user, generate a user problem vector after vectorization, and then obtain document shards related to the user problem from the query library through index retrieval. The specific process of index retrieval is as follows: Calculate the similarity between the user problem vector and each hypothetical problem vector in the ordered set of hypothetical problem vectors of each document shard in the query library, select all hypothetical problem vectors with similarities greater than the threshold, then sort the selected hypothetical problem vectors in descending order of similarity, select multiple hypothetical problem vectors ranked at the front, and finally extract the document shards to which the selected hypothetical problem vectors belong to form a text set related to the user problem;

[0074] An answer generation device is used to form a prompt word with the user problem and the text set related to the user problem, and generate an answer corresponding to the user problem through a large language model.

[0075] See Figure 4 , Figure 4 is a structural block diagram of a computing device 400 shown in an exemplary embodiment of this specification. The components of the computing device 400 include but are not limited to a memory 410 and a processor 420. The processor 420 is connected to the memory 410 through a bus 430, and a database 450 is used to store data.

[0076] The computing device 400 further includes an access device 440, and the access device 440 enables the computing device 400 to communicate via one or more networks 460. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 440 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0077] In one embodiment of this specification, the above components of the computing device 400 and Figure 4 other components not shown therein may also be connected to each other, for example, via a bus. It should be understood that Figure 4 the block diagram of the computing device shown is merely for illustrative purposes and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0078] The computing device 400 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 400 can also be a mobile or stationary server or cloud server, etc.

[0079] The processor 420 is used to execute the following computer-executable instructions, which when executed by the processor implement the steps of the intelligent question-answering method based on the above problem index retrieval enhancement generation technology.

[0080] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the intelligent question-answering method based on the above problem index retrieval enhancement generation technology belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the intelligent question-answering method based on the above problem index retrieval enhancement generation technology.

[0081] One embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions that, when executed by a processor, implement the steps of the intelligent question-answering method based on the above problem index retrieval enhancement generation technology.

[0082] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the intelligent question-answering method based on the above problem index retrieval enhancement generation technology belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the intelligent question-answering method or system based on the above problem index retrieval enhancement generation technology.

[0083] One embodiment of this specification also provides a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the intelligent question-answering method based on the above problem index retrieval enhancement generation technology.

[0084] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above intelligent question-answering method based on the problem index retrieval enhancement generation technology belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above intelligent question-answering method or system based on the problem index retrieval enhancement generation technology.

[0085] The above has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0086] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0087] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.

[0088] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An intelligent question answering method based on question index retrieval enhancement generation technology, characterized in that: Included are: Step 1: Divide all documents in the knowledge base into multiple document fragments, then extract several summary questions from each document fragment as hypothetical questions of the document fragment, then vectorize each hypothetical question to form an ordered set of hypothetical question vectors of the document fragment, and finally use the ordered set of hypothetical question vectors of each document fragment as an index to build a query library; Step 2: rewrite the question input by the user, and generate a user question vector after vectorization, and then obtain the document fragments related to the user question from the query library through index retrieval. The specific process of index retrieval is as follows: calculate the similarity between the user question vector and each hypothetical question vector in the ordered set of hypothetical question vectors of each document fragment in the query library, and select all hypothetical question vectors with similarities greater than a threshold, and then sort the selected hypothetical question vectors in descending order of similarity, select multiple hypothetical question vectors ranked first, and finally extract the document fragments to which the selected hypothetical question vectors belong to form a text set related to the user question; Step 3: The user question and the text set related to the user question are combined into prompt words, and the answer corresponding to the user question is generated through the large language model.

2. The method according to claim 1, characterized in that In step 1, content is extracted from the documents in the knowledge base to convert non-text data into text data. If there is image data, the specific processing process is as follows: First, the caption or annotation of the image is extracted, and then the image is subjected to OCR recognition to extract character information. The caption or annotation and character information of the image are then input into a multimodal large model to generate image description information. Finally, the caption or annotation, character information and image description information of the image are used to replace the image data in the knowledge base.

3. The method according to claim 1, characterized in that In step 1, several summary questions are extracted from each document segment as hypothetical questions for the document segment, including: Step 11: Generate summary content and summary question set corresponding to document segment D through the large language model, where the summary question set includes n summary questions, where n is the preset number of summary questions; Step 12: construct an ordered set Q of hypothetical questions and initialize it to an empty set. At the same time, initialize the set semantic coverage Cov0 and the set semantic diversity value Div0 to 0; Step 13: Select each summary question from the summary question set one by one, and calculate the information gain value corresponding to each summary question: Info_Gain(q i )=α*(Cov(q i )-Cov0)+(1-α)*(Div(q i )-Div0),q i is the i-th summary question in the summary question set, Info_Gain(q i )、Cov(q i )、Div(q i ) are q i The corresponding information gain value, contribution value to semantic coverage and contribution value to semantic diversity, α is the importance weight parameter used to adjust coverage and diversity, and then select the maximum value from the information gain values ​​corresponding to all summary questions, write the summary question corresponding to the maximum value into the ordered set of hypothetical questions Q, and delete the summary question from the summary question set. At the same time, update the set semantic coverage Cov0 and the set semantic diversity value Div0: Cov0 ’ =Cov0+Cov(q u ), Div0 ’ = Div0+Div(q u ), Cov0 ’ 、Div0 ’ is the updated set semantic coverage and set semantic diversity value, q u It is the summary problem corresponding to the maximum value; Step 14: determine whether the number of hypothetical questions in the ordered set of hypothetical questions Q is less than the preset threshold value m of the number of hypothetical questions. If yes, go to step 13; if not, vectorize each hypothetical question in the ordered set of hypothetical questions Q to generate an ordered set of hypothetical question vectors.

4. The method according to claim 3, characterized in that In step 13, Cov(q i ) is calculated as follows: First, the document fragment is decomposed into multiple semantic units, and the q i Each semantic unit is converted into a vector representation, and then q is calculated one by one i The cosine similarity between the vector representation of each semantic unit is calculated, and finally Cov(q i ): Among them, Total_Units is the number of semantic units in the document fragment, NUM_Units(q i ) is q i The total number of vector representations of all semantic units in the document fragment whose cosine similarity is greater than a preset threshold. Div(q i ) is calculated as follows: Where Q_NUM is the number of hypothetical questions in the ordered set Q of hypothetical questions, v qj is the vector representation of the jth hypothetical question in the ordered set of hypothetical questions Q, v qi Yes i The vector representation of Sim(v qi ,v qj ) is v qi 、v qj The cosine similarity between .

5. The method according to claim 1, characterized in that When the number of document fragments in the user question-related text set in step 2 is less than or equal to 1, a hybrid search is also performed. The specific steps of the hybrid search include: The keyword scores of the user question and each document fragment in the knowledge base are calculated through the BM25 algorithm, and the K document fragments with the highest keyword scores are selected. At the same time, the similarity between the user question and each document fragment in the knowledge base is calculated by calculating the cosine similarity of the embedded vector, and the K document fragments with the highest similarity are selected. Finally, the K document fragments selected according to the keyword score and similarity are written into the text set related to the user question.

6. The method according to claim 5, characterized in that Also included are: Each document fragment obtained by hybrid retrieval is extracted one by one from the text set related to the user question, and then the evaluation model is used to calculate the relevance score between the extracted document fragment and the user question. The relevance score is used to characterize whether the document fragment is useful for answering the question. Then, a dynamic label for identifying the relationship between the user question and the document fragment is generated based on the relevance score, and the dynamic labels are screened accordingly. The dynamic labels include: Relevant, Partially Relevant, and Irrelevant.

7. The method according to claim 6, characterized in that Also included are: Step A1: extract the first document fragment from the user question related text set; Step A2: determine whether the dynamic tag identifying the relationship between the user question and the extracted document segment is Relevant. If so, proceed to Step A3; if not, proceed to Step A7; Step A3: determine whether the retrieval method of the extracted document fragment is hybrid retrieval. If yes, proceed to step A4; if no, proceed to step A6; Step A4, calculate the similarity between the user question and the extracted document fragment, the BM25 keyword score of the user question and the extracted document fragment, and determine whether the similarity is less than a preset similarity threshold and the BM25 keyword score is less than a preset keyword score threshold. If yes, proceed to step A5; if not, go to step A7; Step A5, determine whether the number of hypothetical questions in the ordered set of hypothetical questions of the extracted document segment is less than m. If so, add the user question to the end of the ordered set of hypothetical questions of the extracted document segment, that is, the user question is the last hypothetical question in the ordered set of hypothetical questions of the extracted document segment, and then turn to step A7; if not, delete the last hypothetical question in the ordered set of hypothetical questions of the extracted document segment, and add the user question to the first position of the ordered set of hypothetical questions of the extracted document segment, that is, the user question is the first hypothetical question in the ordered set of hypothetical questions of the extracted document segment, and then turn to step A7; Step A6: determine whether the retrieval method of the extracted document fragment is index retrieval. If yes, obtain the hypothetical question with the greatest similarity to the user question from the ordered set of hypothetical questions of the extracted document fragment, and then move the obtained hypothetical question to the first position of the ordered set of hypothetical questions of the extracted document fragment, and then proceed to step A7; if no, proceed to step A7; Step A7: determine whether all document segments in the user question related text set have been extracted. If yes, this process ends; if not, continue to extract the next document segment in the user question related text set and go to step A2.

8. An intelligent question answering system based on question index retrieval enhancement generation technology, characterized in that: Included are: An index generation device is used to divide all documents in the knowledge base into multiple document fragments, then extract a number of summary questions from each document fragment as hypothetical questions of the document fragment, then vectorize each hypothetical question to form an ordered set of hypothetical question vectors of the document fragment, and finally use the ordered set of hypothetical question vectors of each document fragment as an index to construct a query library; The document selection device is used to rewrite the question input by the user, generate a user question vector after vectorization, and then obtain the document fragments related to the user question from the query library through index retrieval. The specific process of index retrieval is as follows: calculate the similarity between the user question vector and each hypothetical question vector of the ordered set of hypothetical question vectors of each document fragment in the query library, and select all hypothetical question vectors with similarities greater than a threshold, and then sort the selected hypothetical question vectors in descending order of similarity, and finally select multiple document fragments ranked first to form a text set related to the user question; The answer generation device is used to form prompt words from user questions and text sets related to user questions, and generate answers corresponding to user questions through a large language model.

9. A computing device, characterized in that include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the intelligent question-answering method based on question index retrieval enhanced generation technology as described in any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that: It stores computer executable instructions, which, when executed by a processor, implement the steps of the intelligent question answering method based on question index retrieval enhancement generation technology as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • push-button type telephone call device.

    BE594786A

  • RAG knowledge question-answering method and device based on fusion vector and keyword retrieval

    CN117951274A

  • Large model RAG method based on hierarchical indexing and hybrid retrieval

    CN118779437A

  • Enterprise document query multi-mode question-answering system generated by combining retrieval enhancement

    CN119003724A

  • RAG text processing method and device based on multi-path recall and medium

    CN119167921A

Cited By

  • Information retrieval method and device based on large language model, medium and electronic equipment

    CN120763306A

  • Intelligent knowledge base management system and method based on large model

    CN120872995A