Retrieval enhancement generation method and system based on large language model
By adopting a search enhancement generation method based on a large language model in the information retrieval system, using a multi-source knowledge base to rewrite and extract query information, identify user intentions, and perform multi-channel hybrid searches, the problem of insufficient efficiency and accuracy in handling diversified and cross-domain queries is solved, and high-precision information retrieval and high-quality reply are achieved.
Patent Information
- Application Number
- CN202411937157.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing information retrieval systems are insufficient in efficiency and accuracy when handling diversified and cross-domain queries, and it is difficult to deeply understand the user's complex query intentions, resulting in the search results often deviating from the user's actual needs.
A search-enhanced generation method based on a large language model is adopted. By deploying a multi-source knowledge base, rewrite the query information entered by users, extract key information, identify user intentions and classify them, perform mixed searches and result rearrangements, and finally generate replies.
It realizes high-precision identification of user query needs and intentions in open domains, improves the immediacy, accuracy and professionalism of the information retrieval system, and improves the quality of system reply.
Smart Images

Figure CN120045750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a retrieval augmented generation method and system based on a large language model. Background Art
[0002] Today, with the explosive growth of digital information, it is particularly crucial to build an efficient information retrieval system that can quickly and accurately filter out the required information from massive data. However, traditional retrieval systems often rely on a single knowledge source, such as an Internet search engine or a professional database, which greatly limits their efficiency and accuracy in processing diverse and cross-domain queries. In addition, most of these systems are based on keyword matching and are difficult to deeply understand the complex query intentions of users, resulting in retrieval results often deviating from the actual needs of users.
[0003] In order to improve the accuracy of retrieval, in recent years, language models based on deep learning have emerged, such as GPT and BERT. These models utilize semantic understanding technology and significantly improve the precision of retrieval. Nevertheless, these single-model-dependent solutions still face challenges in integrating multi-source data and adapting to specific domain knowledge, and the retrieval results in the open domain often contain a large amount of noise, which limits the efficiency of knowledge services.
[0004] Therefore, how to accurately identify the query needs and intentions of users, achieve high-precision information retrieval in the open domain, and improve the immediacy, accuracy, and professionalism of the responses of information retrieval systems has become an urgent technical problem to be solved currently. Summary of the Invention
[0005] In view of the above problems, the present invention provides a retrieval augmented generation method and system based on a large language model. By deploying a multi-source knowledge base, using the large language model to rewrite the input query information of users, extract the key information of the input query information of users, identify and classify the intentions of users, then perform multi-way hybrid retrieval and re-rank the retrieval results, and finally generate a response. The present invention can accurately identify the needs and intentions of user queries, retrieve relevant information with high precision in the open domain, improve the immediacy, accuracy, and professionalism of the information obtained by the information retrieval system, and improve the response quality of the system.
[0006] The present invention provides a retrieval-enhanced generation method based on a large language model, including: obtaining input query information of a user; respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information; identifying the user intention and performing intention classification based on the large language model according to the input query information to obtain an intention classification result; performing multi-channel hybrid retrieval through a preset multi-source knowledge base according to the first query instruction and the second query instruction to obtain an initial retrieval result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; rearranging and screening the initial retrieval result according to the intention classification result to obtain an enhanced prompt; inputting the input query information and the enhanced prompt into the large language model to obtain a retrieval-enhanced generation result output by the large language model.
[0007] According to the retrieval-enhanced generation method based on a large language model provided by the present invention, the step of respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction includes: performing natural language processing on the input query information to identify the grammatical structure and semantic content in the query; according to the identified grammatical structure and semantic content, using the large language model to generate multiple rewritten versions; evaluating the semantic similarity and information retention degree of each rewritten version, and selecting the rewritten version with the highest semantic similarity and information retention degree to be integrated into the first query instruction; using the large language model to extract query key information from the input query information to generate the second query instruction according to the query information; the query key information includes entities and concepts.
[0008] According to the retrieval-enhanced generation method based on a large language model provided by the present invention, the step of identifying the user intention and performing intention classification based on the large language model according to the input query information to obtain an intention classification result includes: using the large language model to perform in-depth semantic analysis on the input query information to understand the user's needs and query purposes to obtain an in-depth semantic analysis result; matching the in-depth semantic analysis result with preset intention categories to obtain the intention classification result.
[0009] A retrieval augmented generation method based on a large language model provided by the present invention, according to the first query instruction and the second query instruction, performing multi-way hybrid retrieval through a preset multi-source knowledge base to obtain an initial retrieval result, including: segmenting the preset multi-source knowledge base into multiple text segments; converting the multiple text segments into vector representations based on a preset vector model to obtain multiple text segment vectors; converting the first query instruction and the second query instruction into vector representations respectively based on the preset vector model to obtain a first query instruction vector and a second query instruction vector; performing multi-way hybrid retrieval according to the multiple text segment vectors, the first query instruction vector and the second query instruction vector, evaluating similarity and relevance, and obtaining the initial retrieval result.
[0010] A retrieval augmented generation method based on a large language model provided by the present invention, performing multi-way hybrid retrieval according to the multiple text segment vectors, the first query instruction vector and the second query instruction vector, evaluating similarity and relevance, and obtaining the initial retrieval result, including: calculating the cosine distance between the first query instruction vector and all the text segment vectors to obtain a first cosine distance; the first cosine distance is used to evaluate the similarity and relevance between the first query instruction vector and all the text segment vectors in multi-way hybrid retrieval; calculating the cosine vector between the second query instruction vector and all the text segment vectors to obtain a second cosine distance; the second cosine distance is used to evaluate the similarity and relevance between the second query instruction vector and all the text segment vectors in multi-way hybrid retrieval; adding the first cosine distance and the second cosine distance to obtain a third cosine distance; obtaining the initial retrieval result according to the third cosine distance.
[0011] A retrieval augmented generation method based on a large language model provided by the present invention, re-ranking and screening the initial retrieval result according to the intent classification result to obtain an enhanced prompt word, including: adjusting the parameters of the initial retrieval result according to the intent classification result to obtain an adjusted retrieval result; re-ranking the retrieval result according to the parameters of the adjusted retrieval result, and screening the results ranked within a preset number as the enhanced prompt word.
[0012] The present invention also provides a retrieval enhanced generation system based on a large language model, including: an acquisition module for acquiring input query information of a user; a query instruction module for respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including key information; an intent classification module for identifying and classifying the user intent based on the large language model according to the input query information to obtain an intent classification result; a retrieval module for performing multi-way hybrid retrieval through a preset multi-source knowledge base according to the first query instruction and the second query instruction to obtain an initial retrieval result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; an enhancement prompt module for re-ranking and filtering the initial retrieval result according to the intent classification result to obtain an enhanced prompt word; a result generation module for inputting the input query information and the enhanced prompt word into the large language model to obtain a retrieval enhanced generation result output by the large language model.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, and when the processor executes the computer program, it implements the retrieval enhanced generation method based on the large language model as described in any one of the above.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the retrieval enhanced generation method based on the large language model as described in any one of the above.
[0015] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the retrieval enhanced generation method based on the large language model as described in any one of the above.
[0016] A retrieval enhanced generation method and system based on a large language model provided by the present invention, the method comprising: obtaining input query information of a user; respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction; identifying the user intention and performing intention classification based on the large language model according to the input query information to obtain an intention classification result; performing multi-channel hybrid retrieval through a preset multi-source knowledge base according to the first query instruction and the second query instruction to obtain an initial retrieval result; performing result rearrangement and screening on the initial retrieval result according to the intention classification result to obtain an enhanced prompt; and inputting the input query information and the enhanced prompt into the large language model to obtain a retrieval enhanced generation result. The present invention can accurately identify the requirements and intentions of user queries, retrieve relevant information with high precision in the open domain, and improve the timeliness, accuracy, professionalism and reply quality of information retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic flowchart of a retrieval enhanced generation method based on a large language model provided by the present invention.
[0019] Figure 2 is a schematic structural diagram of a retrieval enhanced generation system based on a large language model provided by the present invention.
[0020] Figure 3 is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0022] Please refer to Figure 1 , Figure 1 is a schematic flowchart of a retrieval enhanced generation method based on a large language model provided by the present invention.
[0023] The present invention provides a retrieval enhanced generation method based on a large language model, comprising: 101: Obtain the input query information of the user.
[0024] 102: Based on the large language model, rewrite and extract key information from the input query information respectively to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information.
[0025] As a preferred embodiment, based on the large language model, rewrite and extract key information from the input query information respectively to obtain a first query instruction and a second query instruction, including: performing natural language processing on the input query information to identify the syntactic structure and semantic content in the query; according to the identified syntactic structure and semantic content, using the large language model to generate multiple rewritten versions; evaluating the semantic similarity and information retention degree of each rewritten version, and selecting the rewritten version with the highest semantic similarity and information retention degree to integrate into the first query instruction; using the large language model to extract query key information from the input query information to generate a second query instruction according to the query information; the query key information includes entities and concepts.
[0026] In this embodiment, the large language model first receives and analyzes the input query information in an oral form of the user, and through deep learning and natural language processing capabilities, identifies the syntactic structure and semantic content in the query. Key information is extracted according to the identified syntactic structure and semantic content. The large language model automatically modifies the query according to these key information to obtain multiple rewritten versions. These rewritten versions are designed to express the same query intention from different perspectives to increase the coverage of retrieval. Then, cosine similarity can be used to calculate the semantic similarity of each rewritten version. At the same time, the information retention degree can be evaluated by comparing the matching degree of keywords and concepts in the rewritten version and the original query. In order to generate a query instruction with a clearer structure, accurate expression and no ambiguity, select the rewritten version with the highest semantic similarity and information retention degree and integrate it into the first query instruction. This transformation not only enhances the clarity of the query, but also optimizes the relevance and efficiency of the subsequent retrieval process. In this way, it is ensured that accurate requirements are extracted from the user's informal expression, providing a solid foundation for the hybrid retrieval process.
[0027] Use a large language model to deeply analyze the input query information of the user to extract the direct and implicit query key information contained, so as to generate a second query instruction. This process involves identifying important entities in the query text, such as names of people, places, events, and proper nouns, etc. The large language model, through its advanced semantic understanding ability, can not only capture the obvious keywords, but also insight into those implicit information that may not be directly expressed but are crucial to the query intention. The generated second query instruction thus includes all key entities and concepts, which provides a precise orientation for subsequent hybrid retrieval and data processing, ensuring the relevance and comprehensiveness of the retrieval results.
[0028] 103: According to the input query information, identify the user intention and perform intention classification based on the large language model to obtain the intention classification result.
[0029] As a preferred embodiment, according to the input query information, identify the user intention and perform intention classification based on the large language model to obtain the intention classification result, including: using the large language model to perform in-depth semantic analysis on the input query information to understand the user's needs and query purposes, and obtaining the in-depth semantic analysis result; matching the in-depth semantic analysis result with the preset intention categories to obtain the intention classification result.
[0030] In this embodiment, use the large language model to perform in-depth semantic analysis on the input query information. This analysis process involves parsing the natural language text input by the user to identify the user's true needs and query purposes. The process of in-depth semantic analysis can include text preprocessing, feature extraction, and semantic understanding. Subsequently, the large language model matches the in-depth semantic analysis result (the information understood) with the preset intention categories, and classifies the query into one of the three preset categories: Internet category, Wikipedia category, or academic literature category, to obtain the intention classification result. This classification is based on the characteristics of the query content and the type of information that the user may seek. For example, a query about current events may be classified into the Internet category, a query about historical or definitional content may be classified into the Wikipedia category, and a query with academic or research depth is classified into the academic literature category. This step not only speeds up the retrieval process, but also ensures the accuracy and relevance of the retrieval results, effectively matching the user's retrieval needs with the most suitable data source.
[0031] 104: According to the first query instruction and the second query instruction, perform multi-channel hybrid retrieval through a preset multi-source knowledge base to obtain an initial retrieval result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base.
[0032] As a preferred embodiment, according to the first query instruction and the second query instruction, multi-channel hybrid retrieval is performed through a preset multi-source knowledge base to obtain an initial retrieval result, including: segmenting the preset multi-source knowledge base into multiple text segments; converting the multiple text segments into vector representations based on a preset vector model to obtain multiple text segment vectors; converting the first query instruction and the second query instruction into vector representations respectively based on the preset vector model to obtain a first query instruction vector and a second query instruction vector; performing multi-channel hybrid retrieval according to the multiple text segment vectors, the first query instruction vector, and the second query instruction vector, evaluating similarity and relevance, and obtaining an initial retrieval result.
[0033] As a preferred embodiment, multi-channel hybrid retrieval is performed according to the multiple text segment vectors, the first query instruction vector, and the second query instruction vector, evaluating similarity and relevance, and obtaining an initial retrieval result, including: calculating the cosine distance between the first query instruction vector and all text segment vectors to obtain a first cosine distance; the first cosine distance is used to evaluate the similarity and relevance between the first query instruction vector and all text segment vectors in multi-channel hybrid retrieval; calculating the cosine vector between the second query instruction vector and all text segment vectors to obtain a second cosine distance; the second cosine distance is used to evaluate the similarity and relevance between the second query instruction vector and all text segment vectors in multi-channel hybrid retrieval; adding the first cosine distance and the second cosine distance to obtain a third cosine distance; obtaining an initial retrieval result according to the third cosine distance.
[0034] In this embodiment, the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base.
[0035] The Internet knowledge base refers to a knowledge base that uses the API (Application Programming Interface) of the Bing search engine as an interface. The open interface provided by the Bing search engine is used to query data on the Internet in real time. This deployment method allows access to the widest range of information sources, including the latest web content and dynamic data. Through API calls, the user's query can be directly sent to the Bing search engine, and search results are received, which are then used for subsequent processing and analysis.
[0036] The Wikipedia knowledge base refers to a local knowledge base obtained by crawling all entries on Wikipedia. There is a crawler program that runs regularly, which is responsible for downloading new and updated entry content from the Wikipedia website and storing it on a local server. The advantage of this method is that it allows access to the extensive information on Wikipedia without a real-time Internet connection, while also reducing the request burden on the Wikipedia server. The existence of the local knowledge base ensures the fast response and high reliability of retrieval operations.
[0037] An academic literature knowledge base refers to a knowledge base that uses the API of Semantic Scholar as an interface. Semantic Scholar is a widely used academic paper search engine that provides access to a large number of academic literatures. Through the API, one can directly query the Semantic Scholar database to obtain detailed information about academic papers, including paper abstracts, citation information, and download links. This deployment method can utilize the latest academic research results to support in-depth analysis and understanding of scientific issues.
[0038] Set the maximum number of characters to 1024, and segment the content in the Internet knowledge base, Wikipedia knowledge base, and academic literature knowledge base to obtain multiple text fragments. This process ensures that each text fragment is of a manageable and processable size, which helps improve the efficiency and effectiveness of subsequent processing. Then, based on a preset vector model (such as word embedding or BERT model), convert the text fragments in the Internet knowledge base, Wikipedia knowledge base, and academic literature knowledge base into vector representations to obtain multiple text fragment vectors, enabling each text fragment to have its mathematical expression in the vector space for easy mathematical comparison and calculation. Based on the preset vector model, convert the first query instruction into a vector representation to obtain the first query instruction vector. Use the same vector model to ensure that the query instruction and text fragments are in the same vector space, making the subsequent comparison operations feasible and accurate. Calculate the cosine distance between the first query instruction vector and all text fragment vectors to obtain the first cosine distance. The cosine distance is used to measure the similarity between vectors and is used here to determine the degree of relevance of each text fragment to the query instruction. Based on the vector model, convert the second query instruction into a vector representation to obtain the second query instruction vector. The second query instruction is also converted into a vector representation to ensure that this query instruction can effectively compare similarity with text fragments. Calculate the cosine distance between the second query instruction vector and all text fragment vectors to obtain the second cosine distance, which is also used to evaluate similarity and relevance. Add the first cosine distance and the second cosine distance to obtain a comprehensive cosine distance metric (the third cosine distance), which takes into account the relevance evaluation of the same text fragments by two different dimensional query instructions. Take the reciprocal of the third cosine distance as the score of the data fragment (the initial retrieval result). Such a scoring method makes the text fragments with smaller distances (i.e., more relevant) have higher scores and are thus preferentially selected.
[0039] Add the first cosine distance and the second cosine distance. Specifically, the sum of the weights of the first cosine distance and the second cosine distance is 1, and the general values are 0.7 and 0.3 respectively. The specific values need to be determined through specific experiments.
[0040] 105: Rearrange and filter the initial retrieval results according to the intent classification results to obtain enhanced prompt words.
[0041] As a preferred embodiment, rearranging and filtering the initial retrieval results according to the intent classification results to obtain enhanced prompt words includes: adjusting the parameters of the initial retrieval results according to the intent classification results to obtain the adjusted retrieval results; reordering the retrieval results according to the parameters of the adjusted retrieval results, and filtering the results ranked within a preset number as enhanced prompt words.
[0042] In this embodiment, first, the parameters of the initial retrieval results are adjusted according to the intent classification results to obtain the adjusted retrieval results. Specifically, if a data segment belongs to the knowledge base corresponding to the identified query category (Internet category, Wikipedia category, or academic literature category), the score of this data segment will be doubled. This score adjustment is to give priority to data sources that are more relevant to the user's query intent. Subsequently, the retrieval results are reordered according to the parameters of the adjusted retrieval results, and the results ranked within a preset number are filtered as enhanced prompt words. Specifically, all the adjusted scores are reordered, and the top N data segments with the highest scores are selected as enhanced prompt words. The default value of N is set to 5, but this parameter can be adjusted according to the needs of experiments or actual applications. This step ensures the high relevance and accuracy of the retrieval results, thus better meeting the user's information needs.
[0043] 106: Input the input query information and the enhanced prompt words into the large language model to obtain the retrieval enhancement generation result output by the large language model.
[0044] In this embodiment, the user's input query information and N segments (enhanced prompt words) are assembled into an input prompt word. The input prompt word is designed to be concise and unambiguous, containing all key information to ensure that it can effectively guide the large language model to understand and respond to the user's query needs. Then, the input prompt word is input into the large language model. The large language model uses this information for in-depth semantic processing and generates a detailed answer to the user's query or an output providing relevant information to obtain the retrieval enhancement generation result. This process makes full use of the high-quality data segments obtained by hybrid retrieval in order to achieve the optimal answer quality and user satisfaction.
[0045] The large language model is an open-source large language model. The large language model (LLM), also known as the large-scale language model, is an artificial intelligence model designed to understand and generate human language, that is, natural language. The large language model is trained on a large amount of text data and can perform a wide range of tasks, including dialogue, question and answer, text classification, text summarization, translation, sentiment analysis, and so on. The large language model is characterized by its large scale, containing billions of parameters, which helps it learn complex patterns in language data. The large language model is usually based on deep learning architectures such as transformers, which helps it perform various natural language processing tasks.
[0046] The large-scale language model can also be any open-source model of ChatGLM-6B, ChatGPT series, StableVicuna, PaLM, Galactica, or LLaMA series, which can be determined according to specific circumstances. In the open-source large language model, the code is open-source, the dataset is open-source, and there is a license.
[0047] In the present invention, by deploying a multi-source knowledge base, using the large language model to rewrite the input query information of the user, extract the key information of the input query information of the user, identify the intention of the user and classify it, then perform multi-way hybrid retrieval and re-rank the retrieval results, and finally generate a reply. The present invention can accurately identify the needs and intentions of the user's query, retrieve relevant information with high precision in the open domain, improve the timeliness, accuracy, and professionalism of the information obtained by the information retrieval system, and improve the reply quality of the system.
[0048] The retrieval augmented generation system based on the large language model provided by the present invention will be described below. The retrieval augmented generation system based on the large language model described below can be mutually referred to with the retrieval augmented generation method based on the large language model described above.
[0049] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a retrieval augmented generation system based on the large language model provided by the present invention.
[0050] The present invention also provides a retrieval augmented generation system based on a large language model, including: an acquisition module 201 for acquiring the input query information of a user; a query instruction module 202 for respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information; an intention classification module 203 for identifying and classifying the user intention based on the large language model according to the input query information to obtain an intention classification result; a retrieval module 204 for performing multi-way hybrid retrieval through a preset multi-source knowledge base according to the first query instruction and the second query instruction to obtain an initial retrieval result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; an enhancement prompt module 205 for re-ranking and filtering the initial retrieval result according to the intention classification result to obtain an enhanced prompt word; and a result generation module 206 for inputting the input query information and the enhanced prompt word into the large language model to obtain a retrieval augmented generation result output by the large language model.
[0051] The retrieval augmented generation system based on the large language model of the present invention includes an acquisition module, a query instruction module, an intention classification module, a retrieval module, an enhancement prompt module, and a result generation module. The functions of each module are closely connected. Through the collaborative work of multiple specially designed modules, it aims to retrieve relevant information with high precision in the open domain and improve the immediacy, accuracy, and professionalism of the responses generated by the information retrieval system.
[0052] Figure 3 Illustrated is a schematic structural diagram of an electronic device, such as Figure 3As shown in the figure, the electronic device may include: a processor 301, a communications interface 302, a memory 303, and a communication bus 304. Among them, the processor 301, the communications interface 302, and the memory 303 complete communication with each other through the communication bus 304. The processor 301 may call the logical instructions in the memory 303 to execute a retrieval augmented generation method based on a large language model. The method includes: obtaining input query information of a user; respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information; identifying the user intention and performing intention classification based on the large language model according to the input query information to obtain an intention classification result; performing multi-channel hybrid retrieval through a preset multi-source knowledge base according to the first query instruction and the second query instruction to obtain an initial retrieval result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; rearranging and screening the initial retrieval result according to the intention classification result to obtain an enhanced prompt; inputting the input query information and the enhanced prompt into the large language model to obtain a retrieval augmented generation result output by the large language model.
[0053] In addition, when the logical instructions in the above-mentioned memory 303 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0054] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the retrieval-enhanced generation method based on a large language model provided by the above-mentioned various methods. The method includes: obtaining input query information of a user; respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information; based on the input query information, identifying the user's intention and performing intention classification based on the large language model to obtain an intention classification result; according to the first query instruction and the second query instruction, performing multi-way hybrid retrieval through a preset multi-source knowledge base to obtain an initial retrieval result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; according to the intention classification result, rearranging and screening the initial retrieval result to obtain an enhanced prompt word; inputting the input query information and the enhanced prompt word into the large language model to obtain a retrieval-enhanced generation result output by the large language model.
[0055] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the retrieval-enhanced generation method based on a large language model provided by the above-mentioned various methods. The method includes: obtaining input query information of a user; respectively rewriting and extracting key information from the input query information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information; based on the input query information, identifying the user's intention and performing intention classification based on the large language model to obtain an intention classification result; according to the first query instruction and the second query instruction, performing multi-way hybrid retrieval through a preset multi-source knowledge base to obtain an initial retrieval result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; according to the intention classification result, rearranging and screening the initial retrieval result to obtain an enhanced prompt word; inputting the input query information and the enhanced prompt word into the large language model to obtain a retrieval-enhanced generation result output by the large language model.
[0056] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0057] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A retrieval enhancement generation method based on a large language model, characterized in that: include: Get the user's input query information; Rewriting the input query information and extracting key information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information; According to the input query information, identifying the user intent based on the large language model and performing intent classification to obtain an intent classification result; According to the first query instruction and the second query instruction, a multi-path mixed search is performed through a preset multi-source knowledge base to obtain an initial search result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; Rearranging and screening the initial search results according to the intention classification results to obtain enhanced prompt words; The input query information and the enhanced prompt word are input into the large language model to obtain a retrieval enhancement generation result output by the large language model.
2. The retrieval enhancement generation method based on a large language model according to claim 1, characterized in that: The rewriting and key information extraction of the input query information based on the large language model to obtain the first query instruction and the second query instruction include: Performing natural language processing on the input query information to identify grammatical structure and semantic content in the query; Generate multiple rewritten versions using the large language model according to the identified grammatical structure and the semantic content; Evaluate the semantic similarity and information retention of each of the rewritten versions, and select the rewritten version with the highest semantic similarity and information retention to integrate into the first query instruction; The large language model is used to extract query key information from the input query information to generate the second query instruction according to the query information; the query key information includes entities and concepts.
3. The retrieval enhancement generation method based on a large language model according to claim 1, characterized in that: The step of identifying the user intent and classifying the intent based on the large language model according to the input query information to obtain the intent classification result includes: Using the large language model to perform deep semantic analysis on the input query information to understand the user's needs and query purpose, and obtain a deep semantic analysis result; The deep semantic analysis result is matched with the preset intent category to obtain the intent classification result.
4. The retrieval enhancement generation method based on a large language model according to claim 1, characterized in that: The step of performing a multi-path mixed search through a preset multi-source knowledge base according to the first query instruction and the second query instruction to obtain an initial search result includes: Performing text segmentation on the preset multi-source knowledge base to obtain multiple text segments; Convert the plurality of text segments into vector representations based on a preset vector model to obtain a plurality of text segment vectors; Based on the preset vector model, the first query instruction and the second query instruction are respectively converted into vector representations to obtain a first query instruction vector and a second query instruction vector; A multi-way mixed search is performed based on the plurality of text segment vectors, the first query instruction vector and the second query instruction vector, similarity and relevance are evaluated, and the initial search result is obtained.
5. The retrieval enhancement generation method based on a large language model according to claim 4, characterized in that: The performing of multi-path mixed retrieval according to the plurality of text segment vectors, the first query instruction vector and the second query instruction vector, evaluating similarity and relevance, and obtaining the initial retrieval result includes: Calculating the cosine distance between the first query instruction vector and all the text segment vectors to obtain a first cosine distance; the first cosine distance is used to evaluate the similarity and correlation between the first query instruction vector and all the text segment vectors in a multi-way mixed search; Calculating the cosine vectors of the second query instruction vector and all the text segment vectors to obtain a second cosine distance; the second cosine distance is used to evaluate the similarity and correlation between the second query instruction vector and all the text segment vectors in multi-way mixed retrieval; Adding the first cosine distance to the second cosine distance to obtain a third cosine distance; The initial search result is obtained according to the third cosine distance.
6. The retrieval enhancement generation method based on a large language model according to any one of claims 1 to 5, characterized in that: The step of rearranging and screening the initial search results according to the intention classification results to obtain enhanced prompt words includes: Adjusting the parameters of the initial search result according to the intention classification result to obtain an adjusted search result; The search results are re-sorted according to the adjusted search result parameters, and the results sorted within a preset number are selected as the enhanced prompt words.
7. A retrieval enhancement generation system based on a large language model, characterized in that: include: The acquisition module is used to obtain the user's input query information; A query instruction module, used to rewrite the input query information and extract key information based on the large language model to obtain a first query instruction and a second query instruction; the first query instruction is an instruction including the rewritten query information, and the second query instruction is an instruction including the key information; An intention classification module, used to identify user intentions and perform intention classification based on the large language model according to the input query information to obtain an intention classification result; A retrieval module, configured to perform a multi-path mixed search through a preset multi-source knowledge base according to the first query instruction and the second query instruction to obtain an initial search result; the multi-source knowledge base includes an Internet knowledge base, a Wikipedia knowledge base, and an academic literature knowledge base; An enhanced prompt module, used to rearrange and screen the initial search results according to the intention classification results to obtain enhanced prompt words; The result generation module is used to input the input query information and the enhanced prompt word into the large language model to obtain the retrieval enhancement generation result output by the large language model.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the retrieval enhancement generation method based on a large language model is implemented as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the retrieval enhancement generation method based on a large language model is implemented as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the retrieval enhancement generation method based on a large language model is implemented as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Knowledge retrieval method and device, storage medium and server
CN111339239A
Knowledge retrieval enhancement-based large language model question and answer method and device
CN118113836A
Information retrieval method and device, storage medium and computer program product
CN118939763A
Database AI question answering system and method based on large language model and retrieval enhancement
CN118981522A
Multi-source and multi-mode fused knowledge reasoning method, system and device and medium
CN119005340A
Cited By
Method, computing program product and device for identifying specific type of user address
CN120499150A
Electric power scientific research information text retrieval optimization method and system
CN120508607A
Private domain question and answer method and system based on automatic scheduling
CN120804316A
Intelligent semantic matching distributed knowledge base query method, device and equipment
CN120832368A
Database natural language query autoregression intention classification method and system
CN120995178A