Water conservancy knowledge question and answer method and system
By building a water conservancy knowledge base and combining large language models to build general agents and behavioral agents, the problem of insufficient accuracy and practicality of existing water conservancy knowledge Q&A methods in complex queries and numerical calculations is solved, and more efficient and accurate water conservancy knowledge Q&A effects are achieved.
Patent Information
- Application Number
- CN202510077764.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-23
AI Technical Summary
The existing water conservancy knowledge question and answer methods are insufficient in the accuracy and practicality of handling complex queries and numerical calculations, especially in the aspects of multi-subqueries and numerical calculations.
By building a water conservancy knowledge base, using text blocks for knowledge base construction and search, and combining large language models to build general agents and behavioral agents, including knowledge base search and numerical calculation behavioral agents, to improve the accuracy of the question-and-answer system and the ability to handle complex queries.
It significantly improves the accuracy and practicality of the water conservancy knowledge question and answer system, can more effectively respond to users' diverse water conservancy knowledge needs, and enhances the system's ability to handle complex queries and numerical calculations.
Smart Images

Figure CN120030119A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a water conservancy knowledge question-answering method and system. Background Art
[0002] When executing water conservancy knowledge question answering methods, large language models (LLMs) are often used. Large language models (LLMs) play an important role in improving human knowledge acquisition and work efficiency. However, due to the limitation of training data, these models are insufficient in professional domain knowledge and real-time knowledge, and are prone to "hallucination" phenomena. In order to solve these problems, the retrieval augmented generation (RAG) method came into being. It provides reference for LLM by retrieving water conservancy knowledge based on user queries, thus avoiding the above problems to a large extent.
[0003] However, the retrieval enhancement generation method also has limitations. First, the quality of its query response depends entirely on the accuracy and coverage of information retrieval. Secondly, due to the model's ability to understand long texts and the memory capacity of computing equipment, the length of information for model reference is usually limited, which makes it difficult for retrieval enhancement generation to effectively cope with some complex query tasks. For example, in the process of water conservancy knowledge question and answer, when a user's single query includes multiple sub-queries, the information retrieved needs to accurately hit the information corresponding to these sub-queries, which is usually difficult to achieve for the general retrieval enhancement generation process, and the combination of multiple sub-queries will also have a certain negative impact on the retrieval. In addition, if the query contains numerical calculations, it is difficult for LLM to accurately output the calculation results, and there will usually be some deviations in the actual numerical values, and the credibility is low. In summary, the general process of retrieval enhancement generation in the process of water conservancy knowledge question and answer has great limitations and is difficult to meet the complex query needs of users. Therefore, in the process of water conservancy knowledge question and answer, there is an urgent need for a method to improve the accuracy of the water conservancy question and answer process, and a more effective retrieval scheme is needed to break through the limitations of retrieval enhancement generation. Summary of the invention
[0004] In order to solve the above technical problems, the present invention proposes a water conservancy knowledge question and answer method and system, which can more effectively respond to users' diverse water conservancy knowledge needs and improve the accuracy and practicality of the question and answer system.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A water conservancy knowledge question-answering method comprises the following steps:
[0007] Acquire water conservancy field data, pre-process the water conservancy field data into text data, divide the text data into text blocks according to length, and use the text blocks to build a water conservancy knowledge base;
[0008] Constructing a behavior space; the behavior space includes knowledge base retrieval and numerical calculation; wherein the behavior description of knowledge base retrieval is to search the water conservancy knowledge base and obtain useful information; the behavior description of numerical calculation is to be helpful for problems related to calculation;
[0009] A general agent and a behavior agent are constructed based on a large language model; the general agent is used to receive user questions, execute the general agent process according to the user questions in the behavior space, decide whether to call the behavior agent or which behavior agent to call; and decide whether to perform behavior iteration; the behavior agent includes a first behavior agent for knowledge base retrieval and a second behavior agent for numerical calculation; after the general agent is executed, the answer corresponding to the user question is completed to respond to the user question.
[0010] Furthermore, the process of preprocessing data in the water conservancy field includes: defining different loading methods for files of different formats, and cleaning the loaded text data.
[0011] Furthermore, different loading methods are defined for files of different formats:
[0012] The Document function in the python-docx package is used to load Word documents with the suffix ".docx". The soffice command of LibreOffice is used to automatically convert Word documents with the suffix ".doc" to ".docx". The Document function is used to extract text and tables from the file.
[0013] The text in the image is extracted based on the text detection model, text recognition model and text direction classification model in PaddleOCR; the image formats include ".jpg", ".jpeg" and ".png";
[0014] The table format is loaded using the openpyxl library, and the merged cells are filled and converted to Markdown format after loading;
[0015] The presentation format is read using the Presentation function in the python-pptx package;
[0016] Files with the suffix ".pdf" are loaded based on the fitz and PaddleOCR in the pyMuPDF library.
[0017] Furthermore, dividing the text data into text blocks according to length specifically includes: dividing the text data into text blocks of preset length, and setting overlapping areas between adjacent text blocks.
[0018] Furthermore, the process of constructing a water conservancy knowledge base using text blocks includes: storing the text blocks in two forms: text storage and vector storage; wherein the text storage specifically includes: storing the text blocks in a binary pickle file for subsequent text similarity retrieval; and the vector storage specifically includes: using a text embedding model M E The text blocks are embedded into d-dimensional vectors one by one and stored in the Faiss vector library for subsequent semantic similarity retrieval; therefore, the water conservancy knowledge base consists of the original file, the text block binary pickle file and the Faiss vector library file.
[0019] Furthermore, the general agent includes thoughts, behaviors, behavior inputs and behavior results; thoughts are: how to solve user problems and what behaviors to take; behaviors represent: behaviors to be performed and belong to the behavior space; behavior inputs are: input information carried when the behavior starts to be performed; behavior results are: the final results after the behavior is performed;
[0020] During the execution of the general agent: a first prompt word template of the general agent execution example is constructed, and the user question, behavior and behavior description are embedded into the first prompt word template to obtain a first prompt word; the first prompt word is input into a large language model, and the behavior result is used as a first stop mark.
[0021] Furthermore, the first behavior agent includes: search idea, search knowledge base, search input and search result; the search idea is: to obtain the information corresponding to the original query object, consider whether it is necessary to continue to perform the search and what method to take for the search; the search knowledge base is: the name of the knowledge base to be searched; the search input is: the original query object or the query object after the search is performed; the search result is: the final result after the search is performed;
[0022] During the execution of the first behavior agent: construct a second prompt word template for the first behavior agent execution example; embed the water conservancy knowledge base name, knowledge base description and behavior input in the general agent into the second prompt word template to obtain a second prompt word; input the second prompt word into the large language model, and use the retrieval result as the second stop mark to obtain the final result after the retrieval, and feed the final result back to the general agent.
[0023] Furthermore, the method also includes searching the knowledge base using a no-replacement search method during the execution of the first behavior agent.
[0024] Furthermore, the second behavior proxy includes a mathematical expression and a calculation result; the mathematical expression is: a single-line mathematical expression for solving a mathematical problem; the operation result is: the execution result of the mathematical expression;
[0025] During the execution of the second behavior agent: a third prompt word template of the second behavior agent execution example is constructed; the behavior in the general agent is input into the third prompt word template to obtain a third prompt word; the third prompt word is input into the large language model to obtain an execution result, and the execution result is fed back to the general agent.
[0026] The present invention also proposes a water conservancy knowledge question-answering system, comprising: a first building module, a second building module and an agent execution module;
[0027] The first construction module is used to obtain water conservancy field data, pre-process the water conservancy field data into text data, divide the text data into text blocks according to length, and use the text blocks to construct a water conservancy knowledge base;
[0028] The second construction module is used to construct a behavior space; the behavior space includes knowledge base retrieval and numerical calculation; wherein the behavior description of knowledge base retrieval is to search the water conservancy knowledge base and obtain useful information; the behavior description of numerical calculation is to be helpful for problems related to calculation;
[0029] The agent execution module is used to build a general agent and a behavior agent based on a large language model; the general agent is used to receive user questions, execute the general agent process according to the user questions in the behavior space, decide whether to call the behavior agent or which behavior agent to call; and decide whether to perform behavior iteration; the behavior agent includes a first behavior agent for knowledge base retrieval and a second behavior agent for numerical calculation; after the general agent is executed, the answer corresponding to the user question is completed to respond to the user question.
[0030] The effects provided in the content of the invention are only the effects of the embodiments, not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects:
[0031] The present invention proposes a water conservancy knowledge question and answer method and system, which includes the following steps: obtaining water conservancy field data, preprocessing the water conservancy field data into text data, dividing the text data into text blocks according to length, and using the text blocks to build a water conservancy knowledge base; constructing a behavior space; the behavior space includes knowledge base retrieval and numerical calculation; wherein the behavior description of the knowledge base retrieval is to search the water conservancy knowledge base and obtain useful information; the behavior description of the numerical calculation is to help the problem of the calculation; building a general agent and a behavior agent based on a large language model; the general agent is used to receive user questions, execute the general agent process according to the user questions in the behavior space, decide whether to call the behavior agent or which behavior agent to call; and decide whether to perform behavior iteration; the behavior agent includes a first behavior agent for knowledge base retrieval and a second behavior agent for numerical calculation; after the general agent is executed, the answer corresponding to the user question is terminated to respond to the user question. Based on a water conservancy knowledge question and answer method, a water conservancy knowledge question and answer system is also proposed. The present invention can more effectively respond to the user's diverse water conservancy knowledge needs and improve the accuracy and practicality of the question and answer system.
[0032] The present invention aims to make full use of the autonomy advantage of LLM and inject it with initiative by providing callable tools to enhance its ability to handle complex queries. In addition, the present invention provides a knowledge base without replacement retrieval, which significantly improves the retrieval recall rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A flow chart of a water conservancy knowledge question-answering method proposed in Embodiment 1 of the present invention;
[0034] Figure 2 This is an architecture diagram for implementing a water conservancy knowledge question-answering method proposed in Example 1 of the present invention;
[0035] Figure 3 This is a general agent execution flow chart proposed in Example 1 of the present invention;
[0036] Figure 4 This is a flow chart of the knowledge base retrieval agent execution proposed in Example 1 of the present invention;
[0037] Figure 5 This is a schematic diagram of the hybrid search proposed in Example 1 of the present invention;
[0038] Figure 6 This is an example diagram comparing the standard RAG process proposed in Example 1 of the present invention and the answer results enhanced by proxy retrieval;
[0039] Figure 7 This is a schematic diagram of a water conservancy knowledge question and answer system proposed in Example 2 of the present invention. DETAILED DESCRIPTION
[0040] In order to clearly illustrate the technical features of the present solution, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings. The disclosure below provides many different embodiments or examples for realizing different structures of the present invention. In order to simplify the disclosure of the present invention, the components and settings of specific examples are described below. In addition, the present invention may repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed. It should be noted that the components illustrated in the accompanying drawings are not necessarily drawn to scale. The present invention omits the description of known components and processing techniques and processes to avoid unnecessary limitations on the present invention.
[0041] Example 1
[0042] Embodiment 1 of the present invention proposes a water conservancy knowledge question-answering method, which is used to solve the technical problems existing in the water conservancy knowledge question-answering process in the prior art. Figure 1 This is a flow chart of a water conservancy knowledge question and answer method proposed in Example 1 of the present invention.
[0043] Figure 1 A flow chart of a water conservancy knowledge question-answering method proposed in Embodiment 1 of the present invention;
[0044] In step S1, water conservancy field data is obtained, and after the water conservancy field data is preprocessed into text data, the text data is divided into text blocks according to length, and a water conservancy knowledge base is constructed using the text blocks.
[0045] Acquire water conservancy data. Acquire water conservancy data through automated crawlers or manual collection, including books or documents on hydrology and water resources, river basin flood control, irrigation and drainage, etc. It may also include the user's private data, such as basic information of a reservoir in the jurisdiction, emergency plans, standardized management manuals, and maintenance records.
[0046] The process of preprocessing data in the water conservancy field includes: defining different loading methods for files of different formats, and cleaning the loaded text data.
[0047] Different loading methods are defined for files of different formats:
[0048] The Document function in the python-docx package is used to load Word documents with the suffix ".docx". The soffice command of LibreOffice is used to automatically convert Word documents with the suffix ".doc" to ".docx". The Document function is used to extract text and tables in the file. After reading, the table is converted into Markdown format for easy understanding of the model, such as "\n\n|Dam characteristics|Eigenvalue|\n|-|-|\n|Dam type|Mash masonry gravity dam|\n|Designed dam crest elevation|203.5m|\n\n".
[0049] The text in the image is extracted based on the text detection model, text recognition model and text direction classification model in PaddleOCR; the image formats include ".jpg", ".jpeg" and ".png".
[0050] The table format is loaded using the openpyxl library, and the merged cells are filled and converted to Markdown format after loading; the table format suffixes include "xlsx", "xlsm", "xls", etc.
[0051] The presentation format is read using the Presentation function in the python-pptx package; the presentation suffix is ".pptx".
[0052] Files with the suffix ".pdf" are loaded based on fitz and PaddleOCR in the pyMuPDF library. First, fitz is used to try to load the text in the current page. If there is no text, the current page is regarded as an image, and the text in the page is extracted based on the text detection model, text recognition model and text direction classification model in PaddleOCR.
[0053] Secondly, we customize the cleaning rules for the loaded text data, such as removing special characters and redundant line breaks, spaces, and tabs. For PDF files, since optical character recognition (OCR) is used, there is a line break after each line of text. Inappropriate line breaks in the text will have a certain negative impact on the semantic understanding of LLM. Therefore, in the PDF file, only the line breaks in the "."\n" pattern are retained, and the rest are replaced with spaces.
[0054] Finally, the cleaned text data is segmented into text blocks with a length less than l. The segmentation process follows a strategy based on punctuation priority (['\n', '!', '.', ';', ',']) rather than simple hard truncation. Under this strategy, a certain overlap area is set between adjacent text blocks to enhance the semantic continuity and splicing ability of LLM when the reference information is adjacent text blocks.
[0055] The process of constructing a water conservancy knowledge base using text blocks includes: storing the text blocks in two forms: text storage and vector storage; the text storage is specifically: storing the text blocks in binary pickle files for subsequent text similarity retrieval; the vector storage is specifically: using the text embedding model M E The text blocks are embedded into d-dimensional vectors one by one and stored in the Faiss vector library for subsequent semantic similarity retrieval; therefore, the water conservancy knowledge base consists of the original file, the text block binary pickle file and the Faiss vector library file.
[0056] Figure 2 This is an architecture diagram for implementing a water conservancy knowledge question and answer method proposed in Example 1 of the present invention.
[0057] In step S2, a behavior space is constructed; the behavior space includes knowledge base retrieval and numerical calculation; wherein the behavior description of knowledge base retrieval is to search the water conservancy knowledge base and obtain useful information; the behavior description of numerical calculation is to be helpful for problems related to calculation.
[0058] The behavior space B includes knowledge base retrieval (search_knowledgebase) and numerical calculation (calculate), that is, B = {search_knowledgebase, calculate}. Defining behavior descriptions allows LLM to fully understand the behavior function and make more accurate behavior decisions. Among them, the behavior description of knowledge base retrieval is "this behavior can search the local knowledge base and obtain useful information", and the behavior description of numerical calculation is "this behavior is very helpful for you to answer questions about calculations."
[0059] The above behaviors are only necessary behaviors in the intranet environment. If the device can connect to the Internet, the behavior space can also be expanded, such as weather query, search engine search, and academic literature search.
[0060] In step S3, a general agent and a behavior agent are constructed based on a large language model; the general agent is used to receive user questions, execute the general agent process according to the user questions in the behavior space, decide whether to call the behavior agent or which behavior agent to call; and decide whether to perform behavior iteration; the behavior agent includes a first behavior agent for knowledge base retrieval and a second behavior agent for numerical calculation; after the general agent is executed, the answer corresponding to the user question is completed to respond to the user question.
[0061] The general agent includes thoughts, behaviors, behavior inputs and behavior results; thoughts are: how to solve user problems and what behaviors to take; behaviors represent: behaviors to be performed and belong to the described behavior space; behavior inputs are: input information carried when the behavior starts to execute; behavior results are: the final result after the behavior is executed.
[0062] During the execution of the general agent: a first prompt word template of the general agent execution example is constructed, and the user question, behavior and behavior description are embedded into the first prompt word template to obtain a first prompt word; the first prompt word is input into a large language model, and the behavior result is used as a first stop mark.
[0063] Figure 3 This is a general agent execution flow chart proposed in Example 1 of the present invention;
[0064] The general agent includes an iterative formulaic process of thought → behavior → behavior input → behavior result.
[0065] Idea: Use the autonomy of LLM to make it think about how to do it and what actions to take to solve user problem q. In this embodiment, LLM uses Qwen2-72B and uses the vllm library for reasoning acceleration.
[0066] Behavior: The behavior to be performed B i , should belong to the predefined behavior space B, that is: B i ∈B={B 1 ,B 2}={search_knowledgebase,calculate}, where i∈{1,2}.
[0067] Behavior input: the input information carried when the behavior starts to execute. Behavior result: the final result after the behavior is executed. If an exception occurs during the execution of the behavior, the exception information will be used as the behavior result. If you think that the problem cannot be solved at present, you can repeat the above process, and the maximum number of iterations is n times, n = 6. The specific implementation is as follows:
[0068] Construct the first prompt word template P containing the general agent execution process example, embed the user question q, behavior and behavior description into the first prompt word template P, and obtain the prompt word Will Input to LLM, and use "behavior output:" as the stop mark, that is, when LLM generates "behavior output:", it stops generating and obtains the generated content C 1 .
[0069] If C 1 If the end identifier defined in the first prompt word template P is included, the top-level agent execution ends, that is, the dialogue ends, where the end identifier is "final answer:".
[0070] If C 1 contains the string identifiers "behavior:" and "behavior input:", then from C 1 Match the actual behavior B i and its input And execute the behavior proxy function F i (·), get behavior output F i (·) represents a series of execution processes of the behavior agent; Input to LLM again to get output C 2 Repeat the above steps until C m There is an end marker defined in the first prompt word template P, where cat(·) represents string concatenation, t = "thought:", m∈{1,2,...,n}.
[0071] The first behavior agent includes: retrieval idea, retrieval knowledge base, retrieval input and retrieval result; the retrieval idea is: to obtain the information corresponding to the original query object, consider whether it is necessary to continue the retrieval and what method to take for the retrieval; the retrieval knowledge base is: the name of the knowledge base to be retrieved; the retrieval input is: the original query object or the revised query object; the retrieval result is: the final result after executing the retrieval.
[0072] During the execution of the first behavior agent: construct a second prompt word template for the first behavior agent execution example; embed the water conservancy knowledge base name, knowledge base description and behavior input in the general agent into the second prompt word template to obtain a second prompt word; input the second prompt word into the large language model, and use the retrieval result as the second stop mark to obtain the final result after the retrieval, and feed the final result back to the general agent.
[0073] In the process of the first line agent execution, the knowledge base is searched using the no-replacement retrieval method. Figure 4 This is a flow chart of the knowledge base retrieval agent execution proposed in Example 1 of the present invention;
[0074] Retrieval idea: To obtain the query The corresponding information enables LLM to independently consider whether it is necessary to continue the search and what method to adopt for the search, such as search without replacement or search after query rewriting.
[0075] Search knowledge base: The name of the knowledge base to be searched u is the total number of knowledge bases.
[0076] Search input: original query or the rewritten query
[0077] Search results: The final result after executing the search. The result here is the comprehensive answer of LLM based on the retrieved text blocks. The specific implementation is as follows:
[0078] Construct a second prompt word template P containing an example of the behavior agent execution process R , enter the names of all water conservancy knowledge bases, knowledge base descriptions, and top-level agent behaviors Embedded into the second prompt word template P R , get the prompt word Will Input to LLM and use "Search results:" as a stop sign to get the output In order to enable the behavior agent to actively perform no-replacement retrieval sampling, the second prompt word template P R It should include a description like "Each search retrieves only a small amount of text, which may not include the information you want. When searching the same knowledge base again with the same search input, a search without replacement will be used. If a search does not retrieve useful information, it is recommended to perform a search without replacement no more than three times."
[0079] like The second prompt word template P is included R The behavior agent ends execution and feeds back the final result to the general agent, where the end identifier is, for example, a string of "final answer:", "final answer:", "final result:", etc.
[0080] The second line agent includes a mathematical expression and a calculation result; the mathematical expression is: a single-line mathematical expression for solving a mathematical problem; the operation result is: the execution result of the mathematical expression;
[0081] During the execution of the second behavior agent: a third prompt word template of the second behavior agent execution example is constructed; the behavior in the general agent is input into the third prompt word template to obtain a third prompt word; the third prompt word is input into the large language model to obtain an execution result, and the execution result is fed back to the general agent.
[0082] The third prompt word template PC , input the behavior Embedded in P C To the third prompt word Will Input to LLM, get output C C , according to P C Defined expression output mode, matching C C The mathematical expression in the code is executed by using the evaluate function in the numexpr library, and the execution result is fed back to the general agent.
[0083] Figure 5 This is a schematic diagram of the hybrid search proposed in Example 1 of the present invention; if contains the string identifiers "Search knowledge base:" and "Search input:", then Match the actual knowledge base to be searched and search input And execute the knowledge base retrieval function f i (·), get the search results Will Input to LLM again to get output Repeat the above steps until There is a prompt word template P R The end marker defined in , where cat(·) represents string concatenation, t = "Thought:", m∈{1,2,...,n}.
[0084] Retrieve function f i (·) is the standard RAG process, including hybrid retrieval → LLM answer extraction. Specifically, hybrid retrieval consists of vector retrieval based on semantic similarity and BM25 retrieval based on text similarity. Vector retrieval: using text embedding model M E Enter the search Embed it into a d-dimensional vector and perform similarity calculation with all semantic vectors in the Faiss vector library, taking the text block corresponding to the semantic vector of the top 4k Top4k, i.e. the top 4k with the highest similarity, can be calculated by inner product, Euclidean distance, etc. BM25 search: BM25 algorithm is used to calculate the search input Text similarity with all text blocks, for easy display, search input Temporarily denote it as q. Taking text block c as an example, the similarity S(q,c) between q and text block c is calculated as follows:
[0085]
[0086] Among them, l is the vocabulary size of user query q, q Z is the Zth word of q; IDF(qZ ) is the inverse document frequency,
[0087]
[0088] Where N is the total number of text blocks; n(q Z ) that is, the word q appears Z The number of text blocks; TF(q z ,c) is q Z frequency of occurrence in c; is the average length of all text blocks; k 1 and b are hyperparameters that adjust the effect of word frequency and text block length on the similarity score. Here, k 1 =1.5, b=0.75. After calculating the similarity of each text block, take the top 4k text blocks The text block R retrieved by the vector V The text block R retrieved by BM25 T Alternately mixed into a collection of 8k text blocks Remove R A After the repeated text blocks in R, keep the first 4k text blocks F ={c 1 ,c 2 ,...c 4k} for subsequent sampling without replacement; k is the number of text blocks for LLM reference in a single search. The specific steps of LLM answer extraction are: F The first k text blocks in are used as known information and the retrieval input Embedded into the fourth prompt word template P D In the prompt word Will Input to LLM, get output o j , which is the search function Feedback j Behavioral agent for knowledge base retrieval.
[0089] No-replacement sampling is specifically to use no-replacement retrieval when the retrieval knowledge base and retrieval input are exactly the same as the previous iteration to avoid the problem that the actual relevant information is ranked at the bottom. The specific implementation mechanism of no-replacement retrieval is to cache the retrieval knowledge base name of the previous iteration Retrieve Input and text block set R F The 3k text blocks R after the k+1th C ={c k+1 ,c k+2 ,...c 4k}; If the name of the search knowledge base and the search input of the current iteration are exactly the same as those of the previous iteration, no search is performed and the cache RC The first k text blocks of are used as reference information for LLM, and R is removed C Update cache R with the first k text blocks in C ={c 2k+1 ,c 2k+2 ,...c 4k}; If the cache is empty or the search knowledge base name and search input of the current iteration are inconsistent with those of the previous iteration, search again based on the query and update the cache R B , R I and R C , where R C The text block set R retrieved for the current iteration F The 3k related text blocks after the k+1th one.
[0090] In step S4, according to the user question q, the general agent first relies on the thought link to conduct independent thinking, including analyzing the nature of the problem, breaking down the problem components, and forming the steps of the solution.
[0091] Then, we make decisions and execute behaviors based on behaviors and behavior input. If the "behavior:" tag is followed by "search_knowledgebase", the behavior agent for knowledge base retrieval is executed; if the "behavior:" tag is followed by "calculate", the behavior agent for numerical calculation is executed. Finally, we return to the thought link to think about the behavior results and decide whether to iterate the behavior.
[0092] In step S5, specifically, after the general agent iteration is completed, the text after the string end string identifier is the answer corresponding to the user question q, which is used to respond to the user question. Figure 6 This is an example diagram comparing the standard RAG process provided by the present invention and the answer results enhanced by proxy retrieval.
[0093] A water conservancy knowledge question-answering method proposed in Embodiment 1 of the present invention uses execution units including: a human-computer interaction unit, a general agent unit, a behavior agent unit, an LLM service unit, and a knowledge base unit;
[0094] Human-computer interaction unit: a front-end interface that allows users to input questions and view answers to questions.
[0095] General agent unit: receives user questions, executes predefined general agent processes based on the questions, and independently decides whether to call the action agent or which action agent to call.
[0096] Behavior agent unit: includes behavior agents such as knowledge base retrieval and numerical calculation, which are used to process and respond to behavior calls from the general agent unit according to the predefined behavior agent execution process.
[0097] LLM service unit: deploys LLM services to provide cognitive capabilities for the general agent unit and the behavior agent unit, that is, generates text based on the input of calling the service and uses the generated text to respond to the service call.
[0098] Knowledge base unit: includes the original files, text block storage library and vector library of each water conservancy knowledge base, which are used by the RAG engine in the knowledge base retrieval behavior agent to perform retrieval and feedback the retrieval results.
[0099] A water conservancy knowledge question and answer method proposed in Example 1 of the present invention can more effectively respond to users' diverse water conservancy knowledge needs and improve the accuracy and practicality of the question and answer system.
[0100] The water conservancy knowledge question-answering method proposed in Example 1 of the present invention aims to make full use of the autonomy advantage of LLM and inject initiative into it by providing callable tools to enhance its ability to handle complex queries. In addition, the present invention provides a knowledge base without replacement retrieval, which significantly improves the retrieval recall rate.
[0101] Example 2
[0102] Based on the water conservancy knowledge question-answering method proposed in Example 1 of the present invention, Example 2 of the present invention further proposes a water conservancy knowledge question-answering system, which modularizes the process of executing the water conservancy knowledge question-answering method. Figure 7 This is a schematic diagram of a water conservancy knowledge question-answering system proposed in Embodiment 2 of the present invention; the system comprises: a first building module, a second building module and an agent execution module;
[0103] The first construction module is used to obtain water conservancy field data, pre-process the water conservancy field data into text data, divide the text data into text blocks according to length, and use the text blocks to construct a water conservancy knowledge base;
[0104] The second building block is used to build a behavior space; the behavior space includes knowledge base retrieval and numerical calculation; wherein the behavior description of knowledge base retrieval is to search the water conservancy knowledge base and obtain useful information; the behavior description of numerical calculation is to be helpful for problems related to calculation;
[0105] The agent execution module is used to build a general agent and a behavior agent based on a large language model; the general agent is used to receive user questions, execute the general agent process according to the user questions in the behavior space, decide whether to call the behavior agent or which behavior agent to call; and decide whether to perform behavior iteration; the behavior agent includes a first behavior agent for knowledge base retrieval and a second behavior agent for numerical calculation; after the general agent is executed, the answer corresponding to the user question is completed to respond to the user question.
[0106] During the execution of the first building module, the process of preprocessing the water conservancy field data includes: defining different loading methods for files of different formats, and cleaning the loaded text data.
[0107] Different loading methods are defined for files of different formats: the Document function in the python-docx package is used to load Word documents with the suffix ".docx"; the soffice command of LibreOffice is used to automatically convert Word documents with the suffix ".doc" to ".docx"; the Document function is used to extract text and tables from the file;
[0108] The text in the image is extracted based on the text detection model, text recognition model and text direction classification model in PaddleOCR; the image formats include ".jpg", ".jpeg" and ".png";
[0109] The table format is loaded using the openpyxl library, and the merged cells are filled and converted to Markdown format after loading;
[0110] The presentation format is read using the Presentation function in the python-pptx package;
[0111] Files with the suffix ".pdf" are loaded based on the fitz and PaddleOCR in the pyMuPDF library.
[0112] Dividing the text data into text blocks according to length specifically includes: dividing the text data into text blocks of preset length, and setting overlapping areas between adjacent text blocks.
[0113] The process of constructing a water conservancy knowledge base using text blocks includes: storing the text blocks in two forms: text storage and vector storage; the text storage is specifically: storing the text blocks in binary pickle files for subsequent text similarity retrieval; the vector storage is specifically: using the text embedding model M E The text blocks are embedded into d-dimensional vectors one by one and stored in the Faiss vector library for subsequent semantic similarity retrieval; therefore, the water conservancy knowledge base consists of the original file, the text block binary pickle file and the Faiss vector library file.
[0114] In the agent execution module, the general agent includes thoughts, behaviors, behavior inputs and behavior results; thoughts are: how to solve user problems and what behaviors to take; behaviors represent: behaviors to be executed and belong to the behavior space; behavior inputs are: input information carried when the behavior starts to execute; behavior results are: the final result after the behavior is executed.
[0115] During the execution of the general agent: a first prompt word template of the general agent execution example is constructed, and the user question, behavior and behavior description are embedded into the first prompt word template to obtain a first prompt word; the first prompt word is input into a large language model, and the behavior result is used as a first stop mark.
[0116] The first behavior agent includes: retrieval idea, retrieval knowledge base, retrieval input and retrieval result; the retrieval idea is: to obtain the information corresponding to the original query object, consider whether it is necessary to continue the retrieval and what method to take for the retrieval; the retrieval knowledge base is: the name of the knowledge base to be retrieved; the retrieval input is: the original query object or the revised query object; the retrieval result is: the final result after executing the retrieval.
[0117] During the execution of the first behavior agent: construct a second prompt word template for the first behavior agent execution example; embed the water conservancy knowledge base name, knowledge base description and behavior input in the general agent into the second prompt word template to obtain a second prompt word; input the second prompt word into the large language model, and use the retrieval result as the second stop mark to obtain the final result after the retrieval, and feed the final result back to the general agent.
[0118] In the process of the first line agent execution, the knowledge base is searched using the no-replacement retrieval method.
[0119] The second behavior agent includes a mathematical expression and a calculation result; the mathematical expression is: a single-line mathematical expression that solves a mathematical problem; the operation result is: the execution result of the mathematical expression.
[0120] During the execution of the second behavior agent: a third prompt word template of the second behavior agent execution example is constructed; the behavior in the general agent is input into the third prompt word template to obtain a third prompt word; the third prompt word is input into the large language model to obtain an execution result, and the execution result is fed back to the general agent.
[0121] The description of the relevant parts of a water conservancy knowledge question and answer system provided in Example 2 of the present application can be found in the detailed description of the corresponding parts of a water conservancy knowledge question and answer method provided in Example 1 of the present application, and will not be repeated here.
[0122] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or equipment that includes a series of elements are inherent to the elements. In the absence of more restrictions, the elements limited by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or equipment that includes the elements. In addition, the above-mentioned technical solution provided in the embodiment of the present application is consistent with the corresponding technical solution in the prior art in principle, and the part is not described in detail, so as not to repeat too much.
[0123] Although the above describes the specific implementation of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. For those skilled in the art, other different forms of modifications or deformations can be made on the basis of the above description. It is not necessary and impossible to list all the implementation methods here. On the basis of the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.
Claims
1. A water conservancy knowledge question-answering method, characterized in that: The following steps are involved: Acquire water conservancy field data, pre-process the water conservancy field data into text data, divide the text data into text blocks according to length, and use the text blocks to build a water conservancy knowledge base; Constructing a behavior space; the behavior space includes knowledge base retrieval and numerical calculation; wherein the behavior description of knowledge base retrieval is to search the water conservancy knowledge base and obtain useful information; the behavior description of numerical calculation is to be helpful for problems related to calculation; A general agent and a behavior agent are constructed based on a large language model; the general agent is used to receive user questions, execute the general agent process according to the user questions in the behavior space, decide whether to call the behavior agent or which behavior agent to call; and decide whether to perform behavior iteration; the behavior agent includes a first behavior agent for knowledge base retrieval and a second behavior agent for numerical calculation; after the general agent is executed, the answer corresponding to the user question is completed to respond to the user question.
2. A water conservancy knowledge question-answering method according to claim 1, characterized in that: The process of preprocessing data in the water conservancy field includes: defining different loading methods for files of different formats, and cleaning the loaded text data.
3. A water conservancy knowledge question-answering method according to claim 2, characterized in that: Different loading methods are defined for files of different formats: The Document function in the python-docx package is used to load Word documents with the suffix ".docx". The soffice command of LibreOffice is used to automatically convert Word documents with the suffix ".doc" to ".docx". The Document function is used to extract text and tables from the file. The text in the image is extracted based on the text detection model, text recognition model and text direction classification model in PaddleOCR; the image formats include ".jpg", ".jpeg" and ".png"; The table format is loaded using the openpyxl library, and the merged cells are filled and converted to Markdown format after loading; The presentation format is read using the Presentation function in the python-pptx package; Files with the suffix ".pdf" are loaded based on the fitz and PaddleOCR in the pyMuPDF library.
4. A water conservancy knowledge question-answering method according to claim 1, characterized in that: Dividing the text data into text blocks according to length specifically includes: dividing the text data into text blocks of preset length, and setting overlapping areas between adjacent text blocks.
5. A water conservancy knowledge question-answering method according to claim 4, characterized in that: The process of constructing a water conservancy knowledge base using text blocks includes: storing the text blocks in two forms: text storage and vector storage; the text storage specifically includes: storing the text blocks in a binary pickle file for subsequent text similarity retrieval; the vector storage specifically includes: using a text embedding model M E The text blocks are embedded into d-dimensional vectors one by one and stored in the Faiss vector library for subsequent semantic similarity retrieval; therefore, the water conservancy knowledge base consists of the original file, the text block binary pickle file and the Faiss vector library file.
6. A water conservancy knowledge question-answering method according to claim 1, characterized in that: The general agent includes thoughts, behaviors, behavior inputs and behavior results; thoughts are: how to solve user problems and what behaviors to take; behaviors represent: behaviors to be performed and belong to the behavior space; behavior inputs are: input information carried when the behavior starts to be performed; behavior results are: the final results after the behavior is completed; During the execution of the general agent: constructing a first prompt word template for the general agent execution example, embedding the user's question, behavior and behavior description into the first prompt word template to obtain a first prompt word; The first prompt word is input into the large language model, and the behavior result is used as the first stop mark.
7. A water conservancy knowledge question-answering method according to claim 1, characterized in that: The first behavior agent includes: search idea, search knowledge base, search input and search result; the search idea is: to obtain the information corresponding to the original query object, consider whether it is necessary to continue to perform the search and what method to take for the search; the search knowledge base is: the name of the knowledge base to be searched; the search input is: the original query object or the query object after the search is performed; the search result is: the final result after the search is performed; During the execution of the first behavior agent: construct a second prompt word template for the first behavior agent execution example; embed the water conservancy knowledge base name, knowledge base description and behavior input in the general agent into the second prompt word template to obtain a second prompt word; input the second prompt word into the large language model, and use the retrieval result as the second stop mark to obtain the final result after the retrieval, and feed the final result back to the general agent.
8. A water conservancy knowledge question-answering method according to claim 7, characterized in that: The method further includes searching the knowledge base by using a no-replacement search method during the execution of the first behavior agent.
9. A water conservancy knowledge question-answering method according to claim 1, characterized in that: The second behavior proxy includes a mathematical expression and a calculation result; the mathematical expression is: a single-line mathematical expression that solves a mathematical problem; the operation result is: the execution result of the mathematical expression; During the execution of the second behavior agent: constructing a third prompt word template of the second behavior agent execution example; Inputting the behavior in the general agent into the third prompt word template to obtain the third prompt word; The third prompt word is input into the large language model to obtain an execution result, and the execution result is fed back to the general agent.
10. A water conservancy knowledge question-answering system, characterized in that: include: A first building module, a second building module and an agent execution module; The first construction module is used to obtain water conservancy field data, pre-process the water conservancy field data into text data, divide the text data into text blocks according to length, and use the text blocks to construct a water conservancy knowledge base; The second construction module is used to construct a behavior space; the behavior space includes knowledge base retrieval and numerical calculation; wherein the behavior description of knowledge base retrieval is to search the water conservancy knowledge base and obtain useful information; the behavior description of numerical calculation is to be helpful for problems related to calculation; The agent execution module is used to build a general agent and a behavior agent based on a large language model; the general agent is used to receive user questions, execute the general agent process according to the user questions in the behavior space, decide whether to call the behavior agent or which behavior agent to call; and decide whether to perform behavior iteration; the behavior agent includes a first behavior agent for knowledge base retrieval and a second behavior agent for numerical calculation; after the general agent is executed, the answer corresponding to the user question is completed to respond to the user question.
Citation Information
Cited By
Water conservancy intelligent question-answering system and method based on knowledge enhancement and data driving
CN121860062A