Retrieval enhancement generation method and device, electronic equipment and storage medium
By using retrieval-enhanced generation methods and algorithms in intelligent question-answering robots, combined with vector databases and user features to optimize large language models, the problem of insufficient accuracy in generated content is solved, and more accurate and targeted answers are achieved.
Patent Information
- Application Number
- CN202510877439.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-10
AI Technical Summary
The large language models in existing intelligent question-answering robots cannot provide more targeted answers based on users' historical interactions and preferences, resulting in insufficient accuracy in the generated content.
By generating user query requests, using the retrieval enhancement generation algorithm to perform vector retrieval in the preset vector database, combining the user's question characteristics and the large speech model to generate target prompt information, and optimizing the large language model to improve the accuracy of the generated content.
It improves the large language model's ability to understand user question information and the targetedness of generated content, thereby enhancing the accuracy of generated content.
Smart Images

Figure CN120763294A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a retrieval enhancement generation method, device, electronic device, and storage medium. Background Art
[0002] Currently, large language models are being applied to intelligent question-and-answer robots, allowing them to generate decisions based on their decisions. For example, large language models can be used to generate responses to questions from employees in the financial technology sector during their learning and training on enterprise information security. In financial technology business systems, intelligent voice robots built using large language models can use these models to generate responses to user questions regarding financial business knowledge, financial business details, or how to operate financial business systems. In healthcare and elderly care business systems, intelligent voice robots built using large language models can use these models to generate responses to user questions regarding healthcare or elderly care. However, the large language models in existing intelligent question-and-answer robots cannot provide more targeted responses based on users' historical interactions and preferences, hindering the accuracy of generated content. Summary of the Invention
[0003] In view of the above problems, the embodiments of the present application provide a retrieval enhancement generation method, device, electronic device and storage medium to solve the above technical problems that are not conducive to improving the user-specificity and accuracy of generated content.
[0004] In a first aspect, an embodiment of the present application provides a search enhancement generation method, comprising:
[0005] Generate corresponding query requests based on the user's question information;
[0006] Based on the search enhancement generation algorithm, performing a vector search on the query request in a preset vector database to obtain a first search result;
[0007] Target prompt information is generated based on the question information and the user question characteristics, the target prompt information and the first retrieval result are input into a large speech model, and a first output result of the large speech model is obtained, wherein the user question characteristics are obtained based on the user's historical question information.
[0008] Optionally, the retrieval enhancement generation method further includes:
[0009] Vectorizing at least one of the historical question information of the user to obtain a historical question information vector, and constructing a user's question information set based on the historical question information vector;
[0010] The large language model is used to perform feature learning on the question information set to obtain the corresponding user question features.
[0011] Optionally, generating target prompt information according to the question information and user question characteristics includes:
[0012] receiving system prompt information set by a user, wherein the system prompt information is used to represent how the large language model understands the question information;
[0013] The question information, the user question feature and the system prompt information are spliced together to obtain the target prompt information.
[0014] Optionally, the retrieval enhancement generation method further includes:
[0015] Obtaining first evaluation information input by a user for the first output result;
[0016] The large language model is optimized according to the first evaluation information.
[0017] Optionally, the retrieval enhancement generation method further includes:
[0018] receiving feedback information input by a user regarding the first output result;
[0019] Target prompt information is regenerated according to the question information, the user question characteristics and the feedback information, the regenerated target prompt information and the first retrieval result are input into the large speech model to obtain a second output result of the large speech model.
[0020] Optionally, the retrieval enhancement generation method further includes:
[0021] Acquire the second evaluation information according to the first output result and the second output result;
[0022] The large language model is optimized according to the second evaluation information.
[0023] Optionally, the retrieval enhancement generation method further includes:
[0024] Performing data processing on the original knowledge document corresponding to the information security knowledge base to obtain the preset vector database;
[0025] The search enhancement generation algorithm is based on performing a vector search on the query request in a preset vector database to obtain a first search result, including:
[0026] Based on a search enhancement generation algorithm, matching and searching the query request with the preset vector database to obtain at least one context information;
[0027] The relevance between each piece of context information and the query request is calculated respectively, and the context information with the relevance greater than or equal to a preset relevance threshold is taken as the first search result.
[0028] In a second aspect, an embodiment of the present application provides a search enhancement generation device, comprising:
[0029] A request generation module is used to generate a corresponding query request based on the user's question information;
[0030] A retrieval module, configured to perform a vector search on the query request in a preset vector database based on a retrieval enhancement generation algorithm to obtain a first retrieval result;
[0031] An inference module is used to generate target prompt information based on the question information and user question characteristics, input the target prompt information and the first retrieval result into a large speech model, and obtain a first output result of the large speech model, wherein the user question characteristics are obtained based on the user's historical question information.
[0032] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory coupled to the processor, wherein the memory stores program instructions that can be executed by the processor; when the processor executes the program instructions stored in the memory, the above-mentioned retrieval enhancement generation method is implemented.
[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, and when the program instructions are executed by a processor, the above-mentioned retrieval enhancement generation method can be implemented.
[0034] The retrieval enhancement generation method, device, electronic device and storage medium provided in the embodiments of the present application generate a corresponding query request based on the user's question information; based on the retrieval enhancement generation algorithm, perform a vector search on the query request in a preset vector database to obtain a first retrieval result; generate target prompt information based on the question information and the user's question characteristics, input the target prompt information and the first retrieval result into the large voice model to obtain the first output result of the large voice model, wherein the user question characteristics are obtained based on the user's historical question information; through the above method, the preset vector database provides a reference framework for the large language model, so that the content generated by the large language model is more accurate; the combination of the preset vector database and the retrieval enhancement generation algorithm improves the adaptability of the large language model to application scenarios, so that the large language model can better understand the user's question information; and the user's question characteristics can enable the large language model to understand the user's question information more specifically, and can make the content generated by the large language model more user-specific, which is conducive to improving the accuracy of the generated content.
[0035] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A flowchart of the retrieval enhancement generation method provided in an embodiment of the present application is shown.
[0037] Figure 2 A structural diagram of a search enhancement generation device provided in an embodiment of the present application is shown.
[0038] Figure 3 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown.
[0039] Figure 4 A schematic diagram of the structure of the storage medium provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0040] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.
[0041] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0042] In the embodiments of the present application, it should be noted that, in this document, relational terms such as first and second, etc., are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0043] Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0044] In the description of the embodiments of this application, words such as "example" or "for example" are used to indicate an example, illustration, or description. Any embodiment or design described as "for example" or "for example" in the embodiments of this application is not to be construed as being preferred or having more advantages than another embodiment or design. The use of words such as "example" or "for example" is intended to clearly present relative concepts.
[0045] In addition, in the embodiments of the present application, "plurality" refers to two or more. In view of this, in the embodiments of the present application, "plurality" can also be understood as "at least two". "At least one" can be understood as one or more, for example, one, two, or more. For example, "including at least one" means including one, two, or more, and does not limit which ones are included. For example, "including at least one of A, B, and C" means including A, B, C, A and B, A and C, B and C, or A, B, and C.
[0046] It should be noted that in the embodiments of the present application, "and / or" describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / ", unless otherwise specified, generally indicates that the associated objects are in an "or" relationship.
[0047] Large Language Model (LLM), through deep learning technology, can process and generate natural language text. Large language models can include the following types: models based on Transformer architecture, models based on encoder-decoder architecture, hybrid architectures based on autoregressive and autoregressive models, and multimodal models. Exemplarily, models based on Transformer architecture can include the GPT series (Generative Pre-trained Transformer), the BERT series (Bidirectional Encoder Representations from Transformers), and the T5 series (Text-to-Text Transfer Transformer).
[0048] Retrieval-Augmented Generation (RAG): A model that combines information retrieval and generation techniques. This technique generates answers or content by referencing information from external knowledge bases and can be used in a variety of natural language processing tasks.
[0049] An embodiment of the present application provides a search enhancement generation method, see Figure 1As shown, the search enhancement generation method includes the following steps S11 to S13:
[0050] Step S11: Generate a corresponding query request according to the user's question information.
[0051] The question information can be directly vectorized to obtain the corresponding query request. Specifically, if the question information is in text form, the question information is mapped to a vector (embedding). Alternatively, keywords can be extracted from the question information first, and then the question information and the extracted keywords are vectorized to obtain the corresponding query request.
[0052] For example, a user uses an intelligent voice robot in the field of medical care, health care and elderly care, and asks a question such as "I often feel tired and dizzy. Is this related to an imbalance in sugar metabolism?"
[0053] For example, a user uses an intelligent voice robot in the financial technology field and asks a question such as "I am 30 years old and want to purchase a critical illness insurance policy with a coverage of 500,000 yuan. How much does it cost per year?"
[0054] Step S12: Based on the search enhancement generation algorithm, a vector search is performed on the query request in a preset vector database to obtain a first search result.
[0055] Among them, the preset vector database is obtained by vectorizing the text knowledge base, and the query request is matched with the data in the preset vector database for retrieval. The retrieval algorithm adopted is based on the retrieval enhancement generation algorithm, which can generate multiple similar retrieval requests according to the query request, and use multiple parallel channels to retrieve the query request from multiple dimensions respectively. For example, the multiple dimensions can be overall content, keywords, semantic understanding, time, space and user type. By understanding the query request from different angles at the same time, relevant data can be retrieved more comprehensively.
[0056] For example, if a user uses an intelligent voice robot in the healthcare and elderly care field, the preset vector database is obtained by vectorizing the text knowledge base in this field. If a user uses an intelligent voice robot in the financial technology field, the preset vector database is obtained by vectorizing the text knowledge base in this field.
[0057] Step S13: Generate target prompt information based on the question information and the user question feature, input the target prompt information and the first search result into the large speech model, and obtain the first output result of the large speech model, wherein the user question feature is obtained based on the user's historical question information.
[0058] Among them, the user question features are obtained by extracting features from the user's historical question information. The user question features are used to describe the characteristics of the questions asked by the user. For example, the user question features may include the topic, complexity, length, language style used, etc. of the user's questions; the user question features can be used to characterize the way the large language model understands the question information. For example, when asking questions, the user tends to get to the point simply and directly, or the user's thinking is relatively confused and the topic is unclear when asking questions.
[0059] The user question features may be obtained by extracting features from the user's historical question information using a large language model, or by extracting features from the user's historical question information using a trained feature extraction model.
[0060] The first search result is used as context information, and the large language model performs reasoning based on the target prompt information within the framework of the first search information to obtain the first output result.
[0061] For example, the question information (the vector of the question information) and the user question feature may be concatenated to obtain the target prompt information.
[0062] For example, a user uses an intelligent voice robot in the field of medical care, health and elderly care, and asks the following question: "I have been under a lot of work pressure recently, often working overtime, feeling extremely tired, and often dizzy." Based on the user's historical question information, the user's historical question topic "blood sugar problem" is extracted. Although the user did not explicitly state the blood sugar problem in the question information, the large language model can take the blood sugar problem into consideration when making inferences based on the user's question characteristics, making the content of the first output result more accurate.
[0063] In this embodiment, the preset vector database provides a reference framework for the large language model, making the content generated by the large language model more accurate; the combination of the preset vector database and the retrieval enhancement generation algorithm improves the adaptability of the large language model to application scenarios, enabling the large language model to better understand user question information; and the user question features can enable the large language model to understand the user's question information more targeted, and can make the content generated by the large language model more user-targeted, which is conducive to improving the accuracy of the generated content.
[0064] For example, in the field of enterprise information security applications, the documents in the information security knowledge base can be constructed as a preset vector database, and the user can complete the learning or query of enterprise information security knowledge by inputting question information and obtaining the first output result.
[0065] As an embodiment, the search enhancement generation method further includes the following steps:
[0066] Step S21: vectorize the at least one historical question information of the user to obtain a historical question information vector, and construct a question information set of the user according to the historical question information vector;
[0067] Step S22: learning features of the question information set by using a large language model to obtain corresponding user question features.
[0068] Among them, feature learning refers to automatically discovering pattern features in the question information set through a machine learning algorithm, and taking the pattern features as user question features. Specifically, by using the powerful semantic analysis capability of the large language model, the question information set is feature extracted through deep learning, the question mode law contained in the question information set is extracted, and the user question features are obtained.
[0069] As an implementation manner, in step S13, generating the target prompt information according to the question information and the user question features specifically includes the following steps:
[0070] Step S31: receiving system prompt information set by the user, the system prompt information being used to represent a way in which the large language model understands the question information.
[0071] Among them, the system prompt information is a way of thinking of the large language model. Different ways of thinking may have differences in understanding the question information input by the user. Illustratively, the system prompt information can be a chain-of-thought (CoT) or a cumulative reasoning or a multimodal chain-of-thought or a reflective mechanism or a structured reasoning, etc.
[0072] Step S32: concatenating the question information, the user question features, and the system prompt information to obtain the target prompt information.
[0073] In this embodiment, by selecting the system prompt information, a first output result with higher adaptation degree can be obtained.
[0074] As an implementation manner, the retrieval enhancement generation method further includes the following steps:
[0075] Step S41: obtaining first evaluation information input by the user for the first output result;
[0076] Among them, the first evaluation information can be a feedback of the user on the accuracy of the first output result. The first evaluation information can represent that the first output result is accurate or that the first output result is inaccurate.
[0077] Step S42: Optimize the large language model according to the first evaluation information.
[0078] The corresponding question information, the first output result and the first evaluation information may be formed into a first training data, and by collecting a certain amount of the first training data, parameter optimization training may be performed on the large language model.
[0079] In this embodiment, the large language model is continuously optimized through the first evaluation information, which is conducive to further improving the accuracy of the generated content.
[0080] As an embodiment, the search enhancement generation method further includes the following steps:
[0081] Step S51: receiving feedback information input by a user regarding a first output result;
[0082] Step S52: regenerate target prompt information according to the question information, user question characteristics and feedback information, input the regenerated target prompt information and the first search result into the large speech model, and obtain a second output result of the large speech model.
[0083] After receiving the first output result, if the user believes the content of the first output result is incomplete or the logical order needs to be optimized, they can supplement the query information. The user's supplementary content is the feedback information. The target prompt information is regenerated based on the feedback information. The large language model infers the new target prompt information within the framework of the first search information to obtain the second output result.
[0084] In this embodiment, when the first output result is relatively accurate but needs slight adjustment, it can be re-output through feedback information, which is conducive to further improving the accuracy of the generated content.
[0085] In some embodiments, the search enhancement generation method further comprises the following steps:
[0086] Step S61: Obtaining second evaluation information according to the first output result and the second output result;
[0087] Step S62: Optimize the large language model according to the second evaluation information.
[0088] Among them, the second output result is better than the second output result, and the second evaluation information can represent the difference between the second output result and the first output result. The large language model is optimized according to the above difference, so that the optimized large language model can generate the second output result without the need for user feedback information to be supplemented, which is conducive to further improving the accuracy of the generated content.
[0089] As an embodiment, the search enhancement generation method further includes the following steps:
[0090] Step S71: performing data processing on the original knowledge document corresponding to the information security knowledge base to obtain a preset vector database.
[0091] The original knowledge document is preprocessed to convert it into a format suitable for vectorization processing. The preprocessing process specifically includes:
[0092] The first step is tokenization, which divides the text into words or subword units. For example, for Chinese text, you can use a word segmentation tool such as jieba; for English text, you can use NLTK or spaCy; the second step is stop words removal, which removes common meaningless words in the text, such as "的", "是", "和", etc. (Chinese) or "the", "is", "and", etc. (English); the third step is stemming and lemmatization, which restores words to their basic form. For example, "running" can be restored to "run"; the fourth step is text cleaning, which removes noise in the text, such as HTML tags, special characters, numbers, etc.
[0093] After preprocessing, the preprocessed knowledge documents can be vectorized using the M3E (Moka Massive Mixed Embedding) vector model to obtain a preset vector database.
[0094] Exemplarily, the above steps may also be used when question information is vectorized.
[0095] Accordingly, step S12 specifically includes the following steps:
[0096] Step S81: Based on the search enhancement generation algorithm, a matching search is performed on the query request and the preset vector database to obtain at least one context information;
[0097] Among them, the large language model can generate multiple similar retrieval requests based on the query request, and use multiple parallel channels to search the query request from multiple dimensions to obtain at least one context information.
[0098] Step S82: Calculate the relevance of each piece of context information with the query request, and take the context information with a relevance greater than or equal to a preset relevance threshold as the first search result.
[0099] Among them, highly matching context information is used as the first retrieval result for the large language model to reason, which is conducive to further improving the accuracy of the generated content.
[0100] An embodiment of the present application provides a search enhancement generation device 200, see Figure 2 As shown, the retrieval enhancement generation device 200 includes: a request generation module 21, which is used to generate a corresponding query request based on the user's question information; a retrieval module 22, which is used to perform vector retrieval on the query request in a preset vector database based on a retrieval enhancement generation algorithm to obtain a first retrieval result; an inference module 23, which is used to generate target prompt information based on the question information and the user's question characteristics, input the target prompt information and the first retrieval result into a large speech model, and obtain a first output result of the large speech model, wherein the user's question characteristics are obtained based on the user's historical question information.
[0101] In some optional embodiments, the reasoning module 23 is further used to: vectorize at least one of the user's historical question information to obtain a historical question information vector, and construct a user's question information set based on the historical question information vector; use the large language model to perform feature learning on the question information set to obtain the corresponding user question features.
[0102] In some optional embodiments, the reasoning module 23 is also used to: receive system prompt information set by the user, wherein the system prompt information is used to characterize how the large language model understands the question information; and splice the question information, the user question features, and the system prompt information to obtain the target prompt information.
[0103] In some optional implementations, the reasoning module 23 is further configured to: obtain first evaluation information input by a user for the first output result; and optimize the large language model according to the first evaluation information.
[0104] In some optional embodiments, the reasoning module 23 is also used to: receive feedback information input by the user regarding the first output result; regenerate target prompt information based on the question information, the user question characteristics and the feedback information, input the regenerated target prompt information and the first retrieval result into the large speech model, and obtain the second output result of the large speech model.
[0105] In some optional implementations, the reasoning module 23 is further configured to: obtain the second evaluation information according to the first output result and the second output result; and optimize the large language model according to the second evaluation information.
[0106] In some optional implementations, the request generation module 21 is further configured to: perform data processing on the original knowledge document corresponding to the information security knowledge base to obtain the preset vector database;
[0107] Correspondingly, the reasoning module 23 is also used to: based on the retrieval enhancement generation algorithm, match and search the query request with the preset vector database to obtain at least one context information; calculate the correlation between each context information and the query request respectively, and take the context information whose correlation is greater than or equal to the preset correlation threshold as the first retrieval result.
[0108] In this embodiment, the preset vector database provides a reference framework for the large language model, making the content generated by the large language model more accurate; the combination of the preset vector database and the retrieval enhancement generation algorithm improves the adaptability of the large language model to application scenarios, enabling the large language model to better understand user question information; and the user question features can enable the large language model to understand the user's question information more targeted, and can make the content generated by the large language model more user-targeted, which is conducive to improving the accuracy of the generated content.
[0109] Figure 3 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 3 As shown, the electronic device 30 includes a processor 31 and a memory 32 coupled to the processor 31 .
[0110] The memory 32 stores program instructions for implementing the search enhancement generation method according to any one of the above embodiments.
[0111] The processor 31 is configured to execute program instructions stored in the memory 32 to perform retrieval enhancement generation.
[0112] The processor 31 may also be referred to as a CPU (Central Processing Unit). The processor 31 may be an integrated circuit chip having signal processing capabilities. The processor 31 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor.
[0113] In this embodiment, the preset vector database provides a reference framework for the large language model, making the content generated by the large language model more accurate; the combination of the preset vector database and the retrieval enhancement generation algorithm improves the adaptability of the large language model to application scenarios, enabling the large language model to better understand user question information; and the user question features can enable the large language model to understand the user's question information more targeted, and can make the content generated by the large language model more user-targeted, which is conducive to improving the accuracy of the generated content.
[0114] See Figure 4 , Figure 4 Schematic diagram of the structure of a computer-readable storage medium of an embodiment of the present application. The computer-readable storage medium 40 of the embodiment of the present application stores program instructions 41 that can implement all the above methods, wherein the program instructions 41 can be stored in the above storage medium in the form of a software product, including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, or terminal devices such as a computer, a server, a mobile phone, and a tablet.
[0115] In this embodiment, the preset vector database provides a reference framework for the large language model, making the content generated by the large language model more accurate; the combination of the preset vector database and the retrieval enhancement generation algorithm improves the adaptability of the large language model to application scenarios, enabling the large language model to better understand user question information; and the user question features can enable the large language model to understand the user's question information more targeted, and can make the content generated by the large language model more user-targeted, which is conducive to improving the accuracy of the generated content.
[0116] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0117] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the content of the description and drawings of this application, or directly or indirectly used in other related technical fields, is also included in the patent protection scope of the present application.
[0118] The above is only an implementation method of the present application. It should be pointed out that for ordinary technicians in this field, improvements can be made without departing from the creative concept of the present application, but these all fall within the scope of protection of the present application.
Claims
1. A search enhancement generation method, characterized in that: include: Generate corresponding query requests based on the user's question information; Based on the search enhancement generation algorithm, performing a vector search on the query request in a preset vector database to obtain a first search result; Target prompt information is generated based on the question information and the user question characteristics, the target prompt information and the first retrieval result are input into a large speech model, and a first output result of the large speech model is obtained, wherein the user question characteristics are obtained based on the user's historical question information.
2. The search enhancement generation method according to claim 1, characterized in that: The retrieval enhancement generation method further includes: Vectorizing at least one of the historical question information of the user to obtain a historical question information vector, and constructing a user's question information set based on the historical question information vector; The large language model is used to perform feature learning on the question information set to obtain the corresponding user question features.
3. The search enhancement generation method according to claim 1, characterized in that: Generating target prompt information according to the question information and user question characteristics includes: receiving system prompt information set by a user, wherein the system prompt information is used to represent how the large language model understands the question information; The question information, the user question feature and the system prompt information are spliced together to obtain the target prompt information.
4. The search enhancement generation method according to claim 1, characterized in that: The retrieval enhancement generation method further includes: Obtaining first evaluation information input by a user for the first output result; The large language model is optimized according to the first evaluation information.
5. The search enhancement generation method according to claim 1, characterized in that: The retrieval enhancement generation method further includes: receiving feedback information input by a user regarding the first output result; Target prompt information is regenerated according to the question information, the user question characteristics and the feedback information, the regenerated target prompt information and the first retrieval result are input into the large speech model to obtain a second output result of the large speech model.
6. The search enhancement generation method according to claim 5, characterized in that: The retrieval enhancement generation method further includes: Acquire the second evaluation information according to the first output result and the second output result; The large language model is optimized according to the second evaluation information.
7. The search enhancement generation method according to claim 1, characterized in that: The retrieval enhancement generation method further includes: Performing data processing on the original knowledge document corresponding to the information security knowledge base to obtain the preset vector database; The search enhancement generation algorithm is based on performing a vector search on the query request in a preset vector database to obtain a first search result, including: Based on a search enhancement generation algorithm, matching and searching the query request with the preset vector database to obtain at least one context information; The relevance between each piece of context information and the query request is calculated respectively, and the context information with the relevance greater than or equal to a preset relevance threshold is taken as the first search result.
8. A search enhancement generation device, characterized in that: include: A request generation module is used to generate a corresponding query request based on the user's question information; A retrieval module, configured to perform a vector search on the query request in a preset vector database based on a retrieval enhancement generation algorithm to obtain a first retrieval result; An inference module is used to generate target prompt information based on the question information and user question characteristics, input the target prompt information and the first retrieval result into a large speech model, and obtain a first output result of the large speech model, wherein the user question characteristics are obtained based on the user's historical question information.
9. An electronic device, characterized in that: It comprises a processor and a memory coupled to the processor, wherein the memory stores program instructions that can be executed by the processor; when the processor executes the program instructions stored in the memory, the retrieval enhancement generation method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and when the program instructions are executed by a processor, the search enhancement generation method according to any one of claims 1 to 7 can be implemented.
Citation Information
Cited By
Customs declaration information auditing method, device and equipment and storage medium
CN121684071A