Retrieval enhancement generation method fusing dynamic adaptive prompt engineering and semantic enhancement

By semantically enhancing and vectorizing the knowledge base in the field of college admissions consultation, and combining it with dynamic adaptive prompting engineering, the shortcomings of existing RAG technology in query understanding, data processing and system robustness are solved, and more accurate and logically coherent intelligent question answering is achieved.

CN121166733AActive Publication Date: 2025-12-19浙江航大科技开发有限公司
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511695194.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2025-12-19
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Existing RAG technology suffers from problems such as insufficient query understanding, rough data processing, inadequate system robustness, and rigid reasoning guidance in specific fields such as college admissions consultation, resulting in untargeted search results, semantic misalignment, and incorrect answers.

Method used

This paper adopts a method that integrates dynamic adaptive prompting engineering and semantic enhancement to perform offline semantic enhancement preprocessing and vectorization processing on the target domain knowledge base. Natural language queries are converted into structured subqueries through query rewriting nodes, and the processing path is selected based on the high-frequency word judgment results to construct the context and finally generate the target answer.

Benefits of technology

It significantly improves search accuracy, reduces the error rate of high-frequency word queries, enhances system robustness, and generates more accurate and logically coherent answers that meet the needs of professional fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121166733A_ABST
    Figure CN121166733A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval enhancement generation method fusing dynamic adaptive prompt engineering and semantic enhancement, and relates to the technical field of intelligent question answering. The method comprises the steps of performing semantic enhancement preprocessing on original data in a target domain knowledge base; vectorizing the enhanced data and constructing a corresponding vector knowledge base; after receiving a natural language query of a user, converting the natural language query into a structured sub-query through a preset prompt project; judging whether the structured sub-query contains a preset high-frequency word or not, and selecting a corresponding processing path to construct a context; a target answer is generated based on the context and the original query. According to the method, the problem of semantic sparsity of data is solved through semantic enhancement, accurate intention understanding is realized through query rewriting, and high-frequency word retrieval prejudice is avoided through dynamic path selection, so that the intelligent question and answer quality in specific fields such as college enrollment and consultation is remarkably improved, and the robustness and the adaptive capacity of an intelligent question and answer system are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent question answering, in particular to a retrieval enhancement generation method fusing dynamic adaptive prompting engineering and semantic enhancement, and is especially suitable for intelligent question answering scenarios in professional and data-intensive fields such as college enrollment consultation. BACKGROUND

[0002] In recent years, large language models (LLMs) have shown excellent capabilities in natural language processing. However, due to their core capabilities being derived from pre-training on massive static data sets, there are two major limitations in internal knowledge: on the one hand, they cannot learn new information after the training cutoff date; on the other hand, the knowledge learned from general-purpose corpora cannot meet the requirements of specific domains for information accuracy and depth. In addition, without external fact constraints, large language models are prone to produce "hallucination" content that is inconsistent with facts.

[0003] To solve the above problems, retrieval-augmented generation (RAG) technology has emerged as the mainstream paradigm for adapting large language models to specific tasks. The core idea of RAG is to retrieve relevant information segments from an external knowledge base as instant context and provide them to the language model before model inference and generation. By combining the generation capabilities of large language models with the real-time and accuracy of external knowledge sources, RAG improves the consistency of factual answers, reduces "hallucinations", and avoids the high cost of model retraining. A typical RAG system includes three core components: a retriever, an augmenter, and a generator. The retriever is responsible for querying information from the knowledge base, the augmenter is responsible for processing, filtering, and integrating the retrieved information to match the query context, and the generator generates coherent and accurate answers based on the original query and the augmented context.

[0004] RAG technology has gone through several stages of development: the initial "Naive RAG" uses traditional sparse retrieval methods such as TF-IDF or BM25 to perform keyword matching on static text datasets. Although simple to implement, it lacks context awareness and the relevance of the search results is insufficient. The generated answers are scattered and generalized, and the performance on large-scale datasets is poor. "Advanced RAG" introduces dense vector retrieval, context rearrangement, and multi-hop retrieval capabilities, improving semantic understanding and retrieval accuracy. However, it has the problem of high computational overhead for dense vector calculation, and the fixed "retrieval-generation" pipeline is rigid when dealing with complex queries that require multiple steps and cross-domain knowledge. "Modular RAG" emphasizes flexibility, composability, and extensibility of the system, using hybrid retrieval strategies and integrating external tools. However, it does not achieve deep collaboration between "data-query-prompt", and in scenarios such as university enrollment counseling, it still has problems such as shallow query understanding, retrieval bias, and low answer accuracy when facing ambiguous queries such as "Can a student with 450 points in A province physics major in artificial intelligence?"

[0005] In specific fields such as university enrollment counseling, existing RAG methods have obvious shortcomings: Query understanding level: Ambiguous queries with colloquial language and incomplete information are processed by single vectorization, which cannot disassemble potential multi-dimensional information needs, resulting in lack of relevance of search results; Data processing level: The processing of raw data is rough, and there is a lack of collaboration with the prompt engineering. The "semantic sparsity" of raw data and the problem of similar vectors after structured table conversion lead to "semantic misalignment" between query intent and data representation; System robustness level: High-frequency words in specific domain knowledge bases can easily cause retrieval bias. When users query high-frequency words in the knowledge base, the system may mistakenly recall other unrelated entries containing the high-frequency word due to the proximity of the vector space, severely polluting the context environment of the generation model, and ultimately leading to incorrect answers; Reasoning guidance level: Existing fixed static prompt templates cannot be dynamically adjusted according to query types, complex structures of retrieval results, or specific states identified by the system's internal processes (such as potential retrieval bias), greatly limiting the quality and adaptive ability of the final generated answers.

[0006] Therefore, there is an urgent need for a retrieval-enhanced generation method that can solve the deep-seated problems of existing technologies in query understanding, data processing, system robustness, and reasoning guidance. SUMMARY

[0007] The application aims to solve the problems of the existing retrieval enhancement generation technology, such as insufficient query understanding, rough data processing, insufficient system robustness, and rigid reasoning guidance, and provides a retrieval enhancement generation method combining dynamic adaptive prompting engineering and semantic enhancement to improve the intelligent question and answer quality in specific fields such as college enrollment consultation.

[0008] The above object of the application is achieved by the following technical solutions: A retrieval enhancement generation method combining dynamic adaptive prompting engineering and semantic enhancement, comprising the following steps: Performing offline semantic enhancement preprocessing on original data in a target domain knowledge base; Performing offline vectorization processing on the data after semantic enhancement preprocessing, and constructing a corresponding vector knowledge base according to the data type and purpose, and storing the data after vectorization processing into the corresponding vector knowledge base; When receiving a natural language query input by a user, converting the natural language query into a structured subquery through a preset query rewriting node, wherein the natural language query is an unstructured query in a specific scenario of the target domain; Judging whether the structured subquery contains a preset high-frequency word to obtain a high-frequency word judgment result; Selecting a corresponding processing path according to the high-frequency word judgment result, constructing a context based on the structured subquery; Generating a target answer based on the context and the natural language query.

[0009] Preferably, the semantic enhancement preprocessing comprises: For unstructured text data, using a large language model to generate keywords based on the text data, and splicing the keywords with the original text to form enhanced data; For structured table data, first converting each row of data in the table data into an independent text row according to a preset text conversion template, then using a large language model to generate keywords for each text row, and splicing the keywords with the corresponding text row to form enhanced data.

[0010] Preferably, the preset text conversion template is a key-value pair splicing format of column name: cell value.

[0011] Preferably, the conversion of the natural language query into a structured subquery through a preset query rewriting node comprises: Obtaining the natural language query input by the user; Inputting the natural language query into a large language model configured with a preset prompting engineering, wherein the preset prompting engineering is a special prompting strategy designed for the specific scenario of the target domain; The preset prompt engineering guides the large language model to perform intent disassembly, element extraction, and format reconstruction on the natural language query, and outputs a structured subquery that is adapted to the data representation of the target domain knowledge base.

[0012] Preferably, the preset prompt engineering includes a format constraint module, an example training module, and an intent guide module, wherein, The format constraint module is configured to define the output format and field naming rules of the structured subquery. The example training module is configured to provide a labeled natural language query-structured subquery sample pair. The intent guide module is configured to instruct the large language model to identify the core requirements in the natural language query.

[0013] Preferably, the output format of the structured subquery is JSON format, and the JSON format structured subquery includes a retrieval matching field and an attribute marking field, wherein the retrieval matching field is used to accurately point to the data category in the target domain knowledge base, and the attribute marking field is used to identify whether the natural language query contains a preset high-frequency word.

[0014] Preferably, the selection of the corresponding processing path based on the high-frequency word judgment result includes: When the high-frequency word judgment result indicates that the structured subquery contains a preset high-frequency word, a high-frequency word adaptive processing path is selected to build a context based on the structured subquery. When the high-frequency word judgment result indicates that the structured subquery does not contain a preset high-frequency word, a regular processing path is selected to build a context based on the structured subquery.

[0015] Preferably, building a context based on the structured subquery through the high-frequency word adaptive processing path includes: Hard coding data corresponding to a preset high-frequency word into a prompt word to obtain hard coded data. Parsing the JSON format structured subquery to extract retrieval matching fields corresponding to each vector knowledge base. Based on the extracted retrieval matching fields, send retrieval requests to each vector knowledge base and obtain retrieval results. The hard coded data and the retrieval results are aggregated together as a context.

[0016] Preferably, building a context based on the structured subquery through the regular processing path includes: Parsing the JSON format structured subquery to extract retrieval matching fields corresponding to each vector knowledge base. sending a search request to each vector knowledge base based on the extracted search matching field and obtaining a search result; aggregating the search result into a context.

[0017] Preferably, the generating a target answer based on the context and the natural language query comprises: inputting the context and the natural language query into a preset generation model to output the target answer, wherein the generation model is a large language model fine-tuned by target domain data.

[0018] The retrieval enhancement generation method fusing dynamic adaptive prompt engineering and semantic enhancement of the application optimizes data representation through semantic enhancement preprocessing, combines the precise matching of structured subqueries, solves the "semantic misplacement" problem between query intent and data representation in the prior art, and significantly improves retrieval accuracy. The dynamic adaptive path can effectively identify and process high-frequency word queries, avoid retrieval bias through hard-coded authoritative data, reduce answer errors caused by uneven data distribution, and enhance system robustness. Multi-knowledge base parallel retrieval and context aggregation provide comprehensive basis for the generation model, combine dynamic prompt guidance, effectively optimize answer quality, make answers more accurate and logical, and meet specific professional field requirements. The core architecture of the method can be flexibly adapted to medical, financial and other professional fields. Only by replacing the corresponding knowledge base data and fine-tuning the generation model, cross-field applications can be realized. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the application. Those skilled in the art can obtain other drawings according to these drawings without creating any creative labor.

[0020] Figure 1 A flowchart of a retrieval enhancement generation method fusing dynamic adaptive prompt engineering and semantic enhancement in an embodiment of the application; Figure 2 A logic block diagram of a retrieval enhancement generation method fusing dynamic adaptive prompt engineering and semantic enhancement in an embodiment of the application. DETAILED DESCRIPTION

[0021] In order to make the skilled in the art better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of the present application.

[0022] In the embodiments provided in the present application, it should be understood that the disclosed methods and systems can be implemented in other ways. The system embodiments described below are only schematic. For example, the division of units and modules is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the various components shown or discussed can be indirect coupling or communication connection through some interfaces, devices or modules, and can be electrical, mechanical or other forms.

[0023] In addition, each functional unit in each embodiment of the present application can be integrated into one processor, or each unit can be a separate device, or two or more units can be integrated into one device. Each functional unit in each embodiment of the present application can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0024] Those skilled in the art can understand that all or part of the steps of the following method embodiments can be completed by program instructions and related hardware. The aforementioned program instructions can be stored in a computer readable storage medium, and the program instructions are executed to perform the steps of the method embodiments. The aforementioned storage medium includes mobile storage devices, read-only memory (ROM), magnetic or optical disks, and various media that can store program codes.

[0025] In addition, the terms "first", "second" are used only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "a plurality of" or "several" is two or more, unless otherwise explicitly specified.

[0026] In certain fields such as college enrollment counseling, the existing RAG method has the following obvious shortcomings: query understanding level: user's colloquial and incomplete fuzzy query is processed by single vectorization, which cannot disassemble potential multi-dimensional information demand, resulting in lack of pertinence of retrieval results; data processing level: the processing of original data is rough, and there is lack of coordination between the prompt engineering and the original data, the "semantic sparseness" of the original data and the problem of similar vectors after the conversion of structured tables, resulting in "semantic misplacement" between query intention and data representation; system robustness level: high-frequency words in specific domain knowledge base are easy to cause retrieval bias, that is, when the user queries the high-frequency words in the knowledge base, the system may mistakenly recall other irrelevant entries containing the high-frequency words due to the proximity of the vector space, thereby seriously polluting the context environment of the generation model, and finally leading to wrong answers; reasoning guidance level: the existing fixed static prompt word template cannot be dynamically adjusted according to the query type, the complex structure of the retrieval results, or the specific state (such as potential retrieval bias) recognized by the system internal process, thereby greatly limiting the quality and adaptive ability of the final generated answer.

[0027] To solve the deep-seated problems of the existing technology in the four core dimensions of query understanding, data processing, system robustness and reasoning guidance, the primary purpose of the present application is to provide an innovative retrieval enhancement generation method that combines dynamic adaptive prompt engineering and semantic enhancement. The ultimate goal of the present application is to build a new generation of RAG question and answer solution that can deeply understand the user's multi-dimensional complex intention, efficiently process heterogeneous data in a collaborative design way, actively identify and avoid internal bias in the retrieval process, and intelligently and adaptively reason. Through the present application, it is intended to significantly improve the accuracy, fact consistency and system robustness of the RAG system in complex professional fields such as college enrollment counseling, so that users can obtain an intelligent interactive experience comparable to or even surpassing that of a real expert.

[0028] To achieve the above object, the present application proposes a complete and innovative technical solution, the core of which is to deeply integrate and dynamically collaborate data processing and model inference through a multi-node collaborative workflow. To deal with the superficiality of query understanding, the present application designs a query rewriting and multi-dimensional expansion node to intelligently decompose the user's fuzzy query; to solve the roughness of data processing and its disconnection with the hinting engineering, the present application enhances the signal-to-noise ratio and readability of the knowledge source through semantic enhancement preprocessing. To overcome the problem of system robustness and break the static deadlock of inference guidance, the present application introduces a dynamically adaptive context fusion and hinting engineering mechanism. This mechanism can actively identify specific trigger conditions (such as potential bias keywords) in the query and accordingly reconstruct the context by forcibly injecting high-weight authoritative data, while strategically switching the hint words in the generation stage to guide the large language model to perform more reliable and targeted inference. Through this series of innovative designs that are closely linked and one-to-one corresponding, the present application aims to fundamentally solve the many shortcomings of the prior art.

[0029] The traditional RAG adopts a simple "retrieval-generation" mode, which often leads to "semantic misplacement" in the retrieval stage due to the fuzziness of query intent and the roughness of data representation, and the static hinting engineering lacks flexibility and scenario awareness, and cannot correct upstream errors. As shown in Figure 2 The present application optimizes the query end and data end through "query rewriting" and "semantic enhancement" respectively, and introduces a "dynamically adaptive hinting" node, which can make collaborative decisions based on the original query and retrieval content, improving the system robustness and inference flexibility, and improving the answer quality of the RAG system.

[0030] The present application improves the challenges faced by existing retrieval-enhanced generation methods in handling complex domain questions and answers through a multi-node workflow that deeply integrates and dynamically collaborates data processing and model inference. The present application first optimizes the data representation of the knowledge base through an innovative semantic enhancement method, and then uses a query rewriting node based on hinting engineering to convert the user's fuzzy query into structured subqueries for different data sources. In addition, the present application introduces a dynamically adaptive hinting engineering mechanism that can intelligently select the optimal execution path and context construction strategy according to the specific properties of the query content, thereby ensuring information retrieval accuracy while effectively avoiding robustness problems caused by data long-tail distribution. This series of collaborative design innovations significantly improves the accuracy, robustness and intelligence level of the RAG system in handling complex and multi-dimensional queries.

[0031] As shown in Figure 1 The embodiment of the present application provides a retrieval-enhanced generation method that integrates dynamically adaptive hinting engineering and semantic enhancement, which can include the following steps: S1, performing offline semantic enhancement preprocessing on original data in a target domain knowledge base; S2, performing offline vectorization processing on the data after semantic enhancement preprocessing, and constructing a corresponding vector knowledge base according to data types and uses, and storing the data after vectorization processing into the corresponding vector knowledge base; S3, when receiving a natural language query input by a user, converting the natural language query into a structured subquery through a preset query rewriting node, wherein the natural language query is an unstructured query under a specific scenario of the target domain; S4, determining whether the structured subquery contains a preset high-frequency word, to obtain a high-frequency word determination result; S5, selecting a corresponding processing path based on the high-frequency word determination result to construct a context based on the structured subquery; S6, generating a target answer based on the context and the natural language query.

[0032] The embodiment of the present application provides a retrieval enhancement generation method fusing a dynamic adaptive prompt engineering and semantic enhancement. First, original data in a target domain knowledge base is subjected to offline semantic enhancement preprocessing. Then, data after semantic enhancement preprocessing is subjected to offline vectorization processing, and a corresponding vector knowledge base is constructed according to data types and uses, and data after vectorization processing is stored into the corresponding vector knowledge base. Next, when a natural language query input by a user is received, the natural language query is converted into a structured subquery through a preset query rewriting node. Next, it is determined whether the structured subquery contains a preset high-frequency word, to obtain a high-frequency word determination result. Next, a corresponding processing path is selected based on the high-frequency word determination result to construct a context based on the structured subquery. Finally, a target answer is generated based on the context and the natural language query.

[0033] In summary, the embodiment of the present application optimizes data representation through semantic enhancement preprocessing, converts a natural language query into a structured subquery to achieve accurate matching, selects a dynamic processing path in combination with high-frequency word determination, and finally generates a target answer. The scheme of the embodiment of the present application constructs a complete technical framework of "data enhancement-query understanding-dynamic retrieval-intelligent generation", solves the defects of traditional RAG in four core dimensions of query understanding, data processing, system robustness and reasoning guidance, and overall improves the accuracy and robustness of an intelligent question answering system.

[0034] Specifically, the large language model in the embodiment is selected from any one of the GPT series, the LLaMA series or the ERNIE series.

[0035] It should be noted that steps S1 and S2 in the embodiments of the present application are preparation steps, which are completed in an offline manner to reduce network occupation in the query process. In the implementation process of processing user queries, only four steps S3-S6 need to be executed, which effectively improves the query efficiency.

[0036] In step S2, the specific processing process is as follows: all data after semantic enhancement processing is converted into vector form through a vector encoding model, and is stored in the corresponding vector knowledge base according to the source and type of the data. In the field of college enrollment consultation, a multi-source heterogeneous vector knowledge base system can be constructed, which at least includes a college enrollment question and answer knowledge base, a college entrance examination one-to-ten table knowledge base, a historical enrollment situation knowledge base, and a current enrollment plan knowledge base.

[0037] It should be noted that the target domain knowledge base in step S1 refers to a knowledge base in a specific field such as college enrollment consultation.

[0038] In one embodiment, the semantic enhancement preprocessing includes: For unstructured text data, a large language model is used to generate keywords based on the text data, and the keywords are spliced with the original text to form enhanced data; For structured table data, each row of data in the table data is first converted into an independent text row according to a preset text conversion template, and then a large language model is used to generate keywords for each text row, and the keywords are spliced with the corresponding text row to form enhanced data.

[0039] In this embodiment, different types of data are designed with different enhancement strategies. For unstructured text data (such as enrollment question and answer content), each original data is input into a large language model, which generates a set of keywords summarizing the core content according to a preset instruction, and the keywords are spliced with the original text to form unstructured enhanced data after semantic enhancement. For structured table data (such as college entrance examination one-to-ten table and historical enrollment data table), each row in the table is first converted into an independent text record according to a preset text conversion template; then, the same semantic enhancement operation as unstructured text is performed on each text record, that is, a large language model is used to generate keywords and splice them with the original text to form structured derived enhanced data after semantic enhancement, which strengthens the semantic features of the data, makes the semantic information of the original data more rich, solves the problem of "semantic sparseness", and effectively improves the matching accuracy of the query data.

[0040] In one embodiment, the preset text conversion template is a key-value pair splicing format of column name: cell value.

[0041] The embodiment adopts the key-value pair format of "column name: cell value" to texturize table data, retains data structure information, avoids information loss after structured data conversion, and effectively improves the semantic matching degree of table data and queries.

[0042] In one embodiment, converting the natural language query into a structured subquery by the preset query rewriting node includes: obtaining a natural language query input by a user; inputting the natural language query into a large language model configured with a preset prompt project, wherein the preset prompt project is a special prompt strategy designed for a target domain-specific scene; guiding the large language model to perform intent disassembly, element extraction, and format reconstruction on the natural language query through the preset prompt project, and outputting a structured subquery adapted to the data representation of the target domain knowledge base.

[0043] The embodiment utilizes the preset prompt project to guide the large language model to deeply analyze the natural language query and convert it into a structured subquery, structures the fuzzy query, realizes precise disassembly of multi-dimensional requirements, and effectively improves the accuracy of query intent recognition.

[0044] In one embodiment, the preset prompt project includes a format constraint module, an example training module, and an intent guiding module, wherein: the format constraint module is configured to limit the output format and field naming rules of the structured subquery; the example training module is configured to provide a labeled natural language query-structured subquery sample pair; the intent guiding module is configured to instruct the large language model to identify core requirements in the natural language query.

[0045] The embodiment greatly improves the format specification and core element extraction completeness of the structured subquery through the cooperative action of the format constraint module, the example training module, and the intent guiding module, and ensures the quality of query rewriting.

[0046] In one embodiment, the output format of the structured subquery is JSON format, and the structured subquery in JSON format includes a retrieval matching field and an attribute marking field, wherein the retrieval matching field is used to accurately point to a data category in the target domain knowledge base, and the attribute marking field is used to identify whether the natural language query contains a preset high-frequency word.

[0047] When the system receives a natural language query input by a user (such as "A province physics class test 450 points can report to artificial intelligence major"), the query is input into the query rewriting node. The core of the query rewriting node is a large language model configured with a specific prompt engineering, which contains not less than 50 groups of format-unified examples (each group of examples contains a user query example and a corresponding JSON format structured subquery example). Under the guidance of the prompt engineering, the large language model decomposes the user's fuzzy query into structured subqueries in JSON format, which contains retrieval matching fields corresponding to different vector knowledge bases (such as query2024AdmissionStatistics, queryEnrollmentPlan, etc.), and attribute marking fields that identify whether the query involves preset high-frequency words (such as query_is_A).

[0048] The present embodiment uses JSON format to store structured subqueries, and realizes accurate retrieval and high-frequency word recognition through retrieval matching fields and attribute marking fields, respectively, efficiently connects subqueries with knowledge bases, and effectively improves the accuracy of high-frequency word recognition.

[0049] In one embodiment, selecting a corresponding processing path according to the high-frequency word judgment result includes: When the high-frequency word judgment result indicates that the structured subquery contains the preset high-frequency word, the high-frequency word adaptive processing path is selected to build the context based on the structured subquery; When the high-frequency word judgment result indicates that the structured subquery does not contain the preset high-frequency word, the regular processing path is selected to build the context based on the structured subquery.

[0050] The system parses the structured subquery in JSON format output by the query rewriting node, and focuses on identifying the value of the attribute marking field that identifies high-frequency words. If the value is "yes", it indicates that the user query involves a preset high-frequency word (such as "A province"), and the workflow is guided to the high-frequency word adaptive processing path; if the value is "no", the workflow enters the regular processing path.

[0051] The present embodiment dynamically selects a processing path based on the high-frequency word judgment result, implements a differentiated retrieval strategy, and takes special processing for high-frequency word queries, greatly improving the robustness of the system.

[0052] In one embodiment, building a context based on a structured subquery through a high-frequency word adaptive processing path includes: Hard code data corresponding to the preset high-frequency word into the prompt word to obtain hard-coded data; Parse the structured subquery in JSON format and extract the retrieval matching fields corresponding to each vector knowledge base; sending a retrieval request to each vector knowledge base based on the extracted retrieval matching field and obtaining retrieval results; aggregating the hard-coded data and the retrieval results into the context.

[0053] In the high-frequency word adaptive processing path of the embodiment, the system hard-codes the authority data (such as the enrollment policy of the province, typical enrollment cases, etc.) corresponding to the preset high-frequency word and verified by human to the prompt word, simultaneously sends a retrieval request to the corresponding vector knowledge base based on the retrieval matching field in the structured subquery, and obtains retrieval results. Then, the hard-coded authority data and the retrieval results are integrated, filtered noise information, and effectively reduce the answer error rate of high-frequency word query, and the factual consistency is significantly improved.

[0054] In one embodiment, constructing the context based on the structured subquery through the conventional processing path includes: parsing the structured subquery in JSON format, and extracting the retrieval matching field corresponding to each vector knowledge base; sending a retrieval request to each vector knowledge base based on the extracted retrieval matching field and obtaining retrieval results; aggregating the retrieval results into the context.

[0055] In the conventional processing path of the embodiment, the retrieval fields of the structured subquery are directly used to retrieve multiple vector knowledge bases in parallel, the obtained retrieval results are de-duplicated and sorted, and then aggregated into the context, which simplifies the processing flow, effectively improves the processing efficiency of non-high-frequency word query, and ensures the system response speed.

[0056] In one embodiment, generating the target answer based on the context and the natural language query includes: inputting the context and the natural language query into a preset generation model to output the target answer, wherein the generation model is a large language model fine-tuned on target domain data.

[0057] In the embodiment, the aggregated context and the original query (natural language query) are input into the final generation model, which is a large language model fine-tuned on college enrollment counseling domain question and answer data, to generate a precise, fluent, and user-intention-compliant target answer.

[0058] The embodiment uses the generation model fine-tuned on the domain data to generate answers in combination with the context and the original query, effectively improving the professional relevance of the generated answers, and ensuring that the language fluency of the generated target answers approaches or reaches the level of professional consultants.

[0059] In order to more clearly illustrate the implementation principle of the present application, the specific field knowledge enhanced question and answer task solved by the present application is defined as follows: Given a natural language query Q inputted by a user, and a set of heterogeneous knowledge bases K = {K_1, K_2, …, K_n}, where each K_i represents a specific type of knowledge base (e.g. question-answer pairs, score tables, etc.), the goal of this application is to generate an answer A that is accurate in content, coherent in logic, and highly relevant to the user’s intent. This process can be formally represented as: A = G(Q, C) where G represents the final generation model, and C represents the context built for answering the query. The challenge of existing methods lies in how to effectively build C from K. This application optimizes the construction of C through a two-stage process, namely: Q’ = R_w(Q) C = R_a(Q’, K) where R_w represents the query rewriting function proposed in this application, which converts the original query Q into a structured query object Q’. R_a represents the adaptive retrieval and aggregation function, which intelligently retrieves and combines information from K based on the content of Q’, ultimately generating a high-quality context C.

[0060] To address the problem of “semantic sparsity” in the original data and lay the foundation for subsequent collaborative retrieval, this application first performs offline semantic enhancement preprocessing on all knowledge base data.

[0061] For unstructured text data, such as enrollment Q&A, each data entry D_original is first inputted into an auxiliary large language model. The model generates a set of keywords T = {t_1, t_2, …, t_m} that can summarize the core content of D_original according to pre-set instructions. The enhanced data entry D_enhanced is then formed by concatenating the original text with this set of tags.

[0062] For structured table data, text conversion is first performed. Each row in the table is converted into an independent text record, with the format being the concatenation of multiple “column name: cell value” pairs. Subsequently, for each converted text record, the same semantic enhancement process as for unstructured data is applied, i.e. a set of keyword tags is generated and attached using a large language model.

[0063] All semantically enhanced data is finally vectorized and stored in independent vector knowledge bases K_i according to its source and type.

[0064] The query rewriting and adaptive retrieval process is the core of the online processing stage of this application. It achieves deep collaboration between queries and data through sophisticated prompting engineering, and solves the robustness problem of the system through dynamic workflows.

[0065] The input of this flow is a query rewriting node, whose core is a large language model configured with a specific prompt. This prompt forces the model to perform two key tasks: 1) convert the user's natural language query Q into a JSON object Q' containing multiple predefined keys; 2) during the conversion process, determine whether the query involves a predefined high-frequency word (such as "A province") and mark it in the JSON object through a specific key-value pair. These predefined keys directly correspond to the various heterogeneous knowledge bases K_i built. The prompt ensures that the large language model can stably and accurately generate JSON objects that conform to the predetermined structure by providing a large number of highly uniform examples.

[0066] Output example of the query rewriting node: <example> <user_query> Hello, could you please tell me what your school's admission score was this year? I'm a physics student from province A, and I scored 450 points. I'd like to apply for the Artificial Intelligence major.

[0067] < / user_query> <expected_output> { "is_school_inquiry":"Yes", "query_is_A":"Yes", "query2024AdmissionStatistics":"<#school name#> 2024 Admission Statistics in Province A", "queryEnrollmentPlan":"<#school name#> 2025 Enrollment Plan in Province A", "queryExamScores": "450 points for Physics (Science) candidate in Province A, 2025 College Entrance Examination". "querySchoolProfile":"<#school name#>Introduction to the Artificial Intelligence Major and its Programs" } < / expected_output> < / example> After obtaining the structured query object Q', the system enters the adaptive retrieval phase. The system first checks the value of the query_is_A field in Q'. If the value is "yes", the workflow is directed to a processing node specially designed for "A province" queries. In the prompt of this node, data about A province is directly hardcoded to avoid errors caused by inaccurate vector retrieval. If the value is "no", the workflow enters the regular processing node.

[0068] Regardless of which processing path is entered, the system uses other query fields in Q' (such as query2024AdmissionStatistics, queryEnrollmentPlan, etc.) to retrieve information from the corresponding vector knowledge base K_i in parallel. Subsequently, all the recalled information fragments, together with the fixed data in the prompt (in the adaptive path), are aggregated to form the final context C. This context aggregation strategy ensures that the generation model can obtain comprehensive and sufficient judgment basis for complex problems that require comprehensive information.

[0069] The following three typical implementation examples in the context of university enrollment consultation will elaborate on the specific implementation process of this application, covering high-frequency word queries, regular queries, and complex multi-demand queries in different scenarios. The large language model used is GPT-4 (semantic enhancement and query rewriting), ERNIE4.0 (answer generation, fine-tuned on 5000 enrollment question and answer data), and the vector encoding model is Sentence-BERT.

[0070] Example 1: High-frequency word query scenario (involving "B province" high-frequency province) Scenario description: User input "B province history class 530 points, want to report Chinese literature major, can you go to your school?" ("B province" is a pre-defined high-frequency word with an 8% frequency in the domain knowledge base) Step 1: Semantic Enhancement Preprocessing Unstructured data: The original text "The minimum admission score for the Chinese Language and Literature major at a certain university in Province B for the history category in 2024 was 542 points, and the maximum score was 568 points" was processed by GPT-4 to generate keywords "2024-Chinese Language and Literature-Province B-History Category-542-568", which were concatenated into enhanced data: "The minimum admission score for the Chinese Language and Literature major at a certain university in Province B for the history category in 2024 was 542 points, and the maximum score was 568 points [2024-Chinese Language and Literature-Province B-History Category-542-568]", and stored in the knowledge base of past admission situations.

[0071] Structured data: The line "Score: 530; Cumulative number of people: 28000; Rank: 27501-28000" in the one-point-one-section table for the history category in Province B in 2025 was converted according to the template to "Score: 530; Cumulative number of people: 28000; Rank: 27501-28000", and after generating the keyword "2025-Province B-History Category-530-Rank 28000", it was concatenated and stored in the knowledge base of the college entrance examination one-point-one-section table.

[0072] Step 2: Vector Knowledge Base Construction The above enhanced data was vectorized by Sentence-BERT and stored in the knowledge base of past admission situations and the knowledge base of the college entrance examination one-point-one-section table respectively; at the same time, the enhanced data of "A certain university plans to recruit 25 students for the Chinese Language and Literature major in Province B in 2025" was stored in the knowledge base of the current year's enrollment plan.

[0073] Step 3: Query Rewriting Prompt engineering guides GPT-4 to output a JSON structured sub-query: { "is_school_inquiry": "Yes", "query_is_B": "Yes", / / High-frequency word judgment field "query2024Admission": "Admission data for the Chinese Language and Literature major in the history category in Province B at a certain university in 2024", "query2025Plan": "Enrollment plan for the Chinese Language and Literature major in Province B at a certain university in 2025", "queryScoreRank": "Rank corresponding to 530 points in the history category in Province B in 2025" } Step 4: Dynamic Path Selection Analyze "query_is_B = Yes", enter the high-frequency word adaptive processing path, and hard-code the authoritative data: "The undergraduate admission line for the history category in Province B in 2025 was 490 points".

[0074] Step 5: Parallel Retrieval and Context Aggregation Based on the structured subquery parallel retrieval of three knowledge bases—the historical admission data knowledge base, the college entrance examination score distribution table knowledge base, and the current year's enrollment plan knowledge base—the following information is obtained: Historical admissions data: "Lowest score in 2024 was 542". Enrollment plan database: "25 students planned for enrollment in 2025"; The "530 points corresponds to position 28000" database is used for segmented analysis. The context is as follows: "In 2025, the undergraduate admission score for history majors in Province B was 490; in 2024, the lowest admission score for the Chinese Language and Literature major in history majors at a certain university in Province B was 542; the planned enrollment for this major in 2025 was 25; in 2025, a score of 530 in history majors in Province B corresponded to a ranking of 28,000."

[0075] Step 6: Response Generation ERNIE4.0 outputs: "Hello! The undergraduate admission score for History in Province B in 2025 was 490, and your score of 530 meets the undergraduate admission score. In 2024, the lowest admission score for Chinese Language and Literature major in Province B for History majors was 542, and your score is below that score. In 2025, this major plans to enroll 25 students in Province B, the same as last year. Considering that 530 points corresponds to a provincial ranking of 28,000 and the admission rankings in previous years (the lowest admission ranking in 2024 was 25,000), the possibility of being admitted to this major is low. We suggest you consider our university's Chinese International Education major (the lowest score in 2024 was 525)." Example 2: Typical query scenario (not involving high-frequency words) Scenario Description: A user enters "How many years is the clinical medicine program at your school? What are the core courses?" (This does not involve frequently visited provinces / majors; the user will proceed through the standard path.) Steps 1-2: Data Preprocessing and Knowledge Base Construction The unstructured text "Clinical Medicine major has a 5-year program, and core courses include human anatomy, pathology, internal medicine, and surgery" is enhanced by generating the keywords "Clinical Medicine - 5-year program - core courses" and then stored in the admissions Q&A knowledge base.

[0076] Step 3: Query Rewrite Output JSON subquery: { "is_school_inquiry":"Yes", "query_is_highfreq":"No", / / Non-high-frequency query "queryMajorInfo": "Length of study and core courses of a clinical medicine major at a certain university" } Step 4: Dynamic Path Selection Parse "query_is_highfreq=no" and enter the regular processing path.

[0077] Step 5-6: Retrieval and Generation Retrieve the admissions Q&A knowledge base to obtain enhanced data. After aggregating the context, generate the answer: "The school's clinical medicine major has a 5-year schooling system. The core courses include human anatomy, histology and embryology, physiology, pathology, pharmacology, internal medicine, surgery, obstetrics and gynecology, pediatrics, etc., and it cultivates medical talents with clinical diagnosis, treatment and scientific research capabilities." Example 3: Complex multi-requirement query scenario (involving collaboration of multiple knowledge bases) Scenario description: The user inputs "In Province C, with a science score of 620, I want to apply for the computer science and technology major. May I ask what was the admission ranking last year? How many students will be recruited this year? And what is the tuition fee?" (involving requirements from 3 different knowledge bases) Step 3: Query Rewriting Output JSON sub-query: { "is_school_inquiry": "yes", "query_is_C": "yes", / / Province C is a high-frequency word (word frequency 6%) "query2024Rank": "The admission ranking of a certain university's computer major for science students in Province C in 2024", "query2025Plan": "The enrollment number of a certain university's computer major in Province C in 2025", "queryTuition": "The tuition standard of a certain university's computer major" } Step 4-5: Dynamic Retrieval and Aggregation Hard-code "The science undergraduate line in Province C in 2025 is 513 points"; Parallelly retrieve the historical admission database (the admission ranking in 2024 is 8,000), the enrollment plan database (40 students will be recruited in 2025), and the admissions Q&A database (tuition is 5,200 yuan per year); Aggregate the context to include all the above information.

[0078] Step 6: Answer Generation Output an answer that integrates multi-dimensional information, clearly respond to the questions of ranking, enrollment number and tuition fee, and give application suggestions. The generated answer can be as follows: "Hello! Combining your score, intended major and the retrieved authoritative information, the relevant questions are explained as follows: First, about the admission rank and enrollment plan: the computer science and technology major of our school in C province in 2024 has an admission rank of 8000; the enrollment plan of this major in C province in 2025 is 40, which is 5 more than last year, and the admission opportunity is slightly improved.

[0079] Second, about the tuition standard: the tuition of computer science and technology major in our school is 5200 yuan per student per year, which meets the relevant regulations of C province education charge, in addition, it also needs to pay the accommodation fee and other expenses according to the unified standard of the school (for details, please refer to the admission notice).

[0080] Combined with the examination suggestion: the control score line of C province science undergraduate in 2025 is 513 points, and your 620 points are far beyond the undergraduate line, which has strong competitiveness. But need to pay attention, computer science as a popular engineering major, the competition has always been fierce, it is suggested that you focus on the provincial rank after the announcement of the college entrance examination results: if your rank is around 8000 or earlier, the probability of successful application is higher; if the rank is slightly lower than 8000, you can consider filling in the computer-related majors such as big data technology and artificial intelligence in our school, which are highly related to the curriculum system of computer science and technology, and the admission rank has been relatively mild in the past two years, forming a reasonable volunteer gradient.

[0081] In addition, computer science requires high logical thinking and continuous learning ability, if you are interested in technology research and algorithm design, you can participate in the artificial intelligence laboratory project of the school after enrollment, relying on the discipline platform to improve the practical ability and lay the foundation for future further study or employment. The above embodiments of the present application propose a retrieval enhancement generation method for deep collaborative design of data processing and prompt engineering. The method generates a structured subquery highly matched with the semantic enhanced data representation format by querying and rewriting the prompt engineering, effectively solving the "semantic misplacement" problem between query intent and data representation, significantly improving the accuracy of retrieval; a dynamic adaptive workflow based on query content recognition is designed, which can actively identify and isolate high-frequency word queries that may cause retrieval bias, by switching to a specific node for processing high-frequency words, the problem of system robustness caused by uneven data distribution is alleviated; a complete solution supporting parallel retrieval and context aggregation is constructed, which can simultaneously obtain information from multiple different types of data sources and effectively integrate them for a single complex problem, enabling large language models to perform complex reasoning tasks that require multiple information sources, greatly expanding the application potential and answer depth of RAG system.

[0082] The various embodiments described in this specification are presented by way of example, and embodiments disclosed herein can be implemented in any number of different ways. Each embodiment is presented for the purpose of illustration only, and is not intended to limit the scope of the disclosure. Embodiments disclosed herein can be implemented in software, hardware, firmware, or any combination thereof. Embodiments disclosed herein can be implemented in one or more computer programs that are executable on a programmable computer or processing device. Embodiments disclosed herein can be implemented in a computer program product that can be executed on a programmable computer or processing device. Embodiments disclosed herein can be implemented in a computer program tangibly embodied in a computer readable storage medium.

[0083] Those skilled in the art will further appreciate that the units and algorithms described in connection with the examples disclosed herein can be implemented in electronic hardware, computer software, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various components will be described herein generally in terms of their functionality, without reference to the particular manner in which they are implemented. Skilled persons will appreciate that the described functionality can be implemented in one or more general purpose or specially designed components or routines.

[0084] The steps of a method or algorithm described in connection with the examples disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM, flash memory, ROM, electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, hard disk can be used as a storage medium.

[0085] The above description of disclosed embodiments is meant to be illustrative of the disclosure and not limiting. Many modifications of the embodiments disclosed herein will occur to persons of ordinary skill in the art. Those modifications will be within the scope of this disclosure as defined by the appended claims and their equivalents. Accordingly, it is not intended that the application be limited, except as by the appended claims and their equivalents.

Claims

1. A retrieval enhancement generation method integrating dynamic adaptive suggestion engineering and semantic enhancement, characterized in that, Includes the following steps: Offline semantic enhancement preprocessing is performed on the raw data in the target domain knowledge base; The data that has undergone semantic enhancement preprocessing is vectorized offline, and a corresponding vector knowledge base is built according to the data type and purpose. The vectorized data is then stored in the corresponding vector knowledge base. When a natural language query is received from a user, the natural language query is converted into a structured subquery through a preset query rewriting node. The natural language query is an unstructured query in a specific scenario of the target domain. Determine whether the structured subquery contains preset high-frequency words, and obtain the high-frequency word determination result; Based on the high-frequency word judgment results, the corresponding processing path is selected, and a context is constructed based on the structured subquery; The target answer is generated based on the context and the natural language query.

2. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to claim 1, characterized in that, The semantic enhancement preprocessing includes: For unstructured text data, a large language model is used to generate keywords based on the text data, and the keywords are concatenated with the original text to form enhanced data; For structured tabular data, each row of data in the table is first converted into an independent text line according to a preset text conversion template. Then, keywords are generated for each text line using a large language model, and the keywords are concatenated with the corresponding text lines to form enhanced data.

3. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to claim 2, characterized in that, The preset text conversion template is a key-value pair format consisting of column name and cell value.

4. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to claim 1, characterized in that, The step of converting the natural language query into a structured subquery through a preset query rewriting node includes: Obtain the natural language query input by the user; The natural language query is input into a large language model configured with a preset suggestion project, wherein the preset suggestion project is a dedicated suggestion strategy designed for a specific scenario in the target domain; The preset prompting process guides the large language model to deconstruct the intent, extract elements, and reconstruct the format of the natural language query, outputting a structured subquery that is adapted to the data representation of the target domain knowledge base.

5. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to claim 4, characterized in that, The preset prompt project includes a format constraint module, an example training module, and an intent guidance module, wherein... The format constraint module is used to limit the output format and field naming rules of structured subqueries; The example training module is used to provide labeled natural language query-structured subquery sample pairs; The intent guidance module is used to instruct the large language model to identify the core requirements in natural language queries.

6. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to claim 5, characterized in that, The output format of the structured subquery is JSON format. The JSON format structured subquery includes a search matching field and an attribute tag field. The search matching field is used to accurately point to the data category in the target domain knowledge base, and the attribute tag field is used to identify whether the natural language query contains preset high-frequency words.

7. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to claim 6, characterized in that, The step of selecting the corresponding processing path based on the high-frequency word judgment result and constructing the context based on the structured subquery includes: When the high-frequency word judgment result indicates that the structured subquery contains preset high-frequency words, the high-frequency word adaptive processing path is selected to construct a context based on the structured subquery; When the high-frequency word judgment result indicates that the structured subquery does not contain the preset high-frequency words, the conventional processing path is selected to construct a context based on the structured subquery.

8. The retrieval enhancement generation method integrating dynamic adaptive suggestion engineering and semantic enhancement according to claim 7, characterized in that, The context constructed based on the structured subquery through the high-frequency word adaptive processing path includes: Hard-coded data is obtained by hard-coding the data corresponding to preset high-frequency words into the prompt words. Parse the structured subquery in JSON format and extract the search matching fields corresponding to each vector knowledge base; Based on the extracted search matching fields, search requests are sent to each vector knowledge base simultaneously and search results are obtained. The hard-coded data and the search results are aggregated together to form a context.

9. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to claim 7, characterized in that, Constructing a context based on the structured subquery through a conventional processing path includes: Parse the structured subquery in JSON format and extract the search matching fields corresponding to each vector knowledge base; Based on the extracted search matching fields, search requests are sent to each vector knowledge base simultaneously and search results are obtained. The search results are aggregated into context.

10. The retrieval enhancement generation method integrating dynamic adaptive prompting engineering and semantic enhancement according to any one of claims 1-9, characterized in that, The generation of the target answer based on the context and the natural language query includes: The context and the natural language query are input into a preset generative model, and the target answer is output. The generative model is a large language model that has been fine-tuned with target domain data.

Citation Information

Patent Citations

  • Prompt word generation method and device, equipment, medium and program product

    CN118296119A

  • Model acceleration database retrieval optimization system and method based on knowledge graph

    CN119046315A

  • Dynamic knowledge retrieval enhancement method based on large language model

    CN120407570A

  • Low-altitude intelligent question and answer construction method and system based on dynamic parameters

    CN120632055A

  • Structured query statement generation method and system

    CN120780730A