Large model retrieval enhancement generation method based on vector database and SQL (Structured Query Language) database
By combining vector database and SQL database methods, a large model search system is built, which solves the problems of insufficient semantic understanding and complex operations in traditional database search methods, and achieves more efficient and accurate knowledge base retrieval.
Patent Information
- Application Number
- CN202510417504.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional database search methods rely on keyword matching and SQL queries and cannot process complex semantic queries, resulting in inaccurate search results and complex operations, and fail to make full use of database structured information.
Combining vector databases and SQL databases, through text vectorization and model inference, a large model retrieval system is built, and a domain knowledge file is used to build a vector database and SQL database, conduct comprehensive analysis and reasoning, and improve retrieval efficiency and accuracy.
It realizes more accurate semantic understanding, reduces user manual operations, improves retrieval efficiency and comprehensiveness of results, and does not require user manual viewing to process appendix information, simplifying the operation process.
Smart Images

Figure CN120296032A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information retrieval, and particularly to a large model retrieval enhanced generation method and system based on a vector database and an SQL database, which are used to improve the efficiency and accuracy of knowledge base retrieval. Background Art
[0002] Traditional database retrieval methods are usually based on keyword matching or Structured Query Language (SQL). Users need to manually input query conditions, and the retrieval results depend on exact keyword matching. Its disadvantages are: keyword matching cannot handle complex semantic queries; for the SQL query statement method, users need to have certain database query knowledge, and the retrieval results may not be accurate enough.
[0003] The retrieval methods in the prior art fail to make full use of the structured information in the database, such as the relationship between the main text and the appendix, which may lead to incomplete retrieval results; or only visualize the appendix file name or table name, and users need to manually view further and judge the required information by themselves. This not only has a long operation process, but also users may have misjudgments and missed judgments. Summary of the Invention
[0004] The present invention aims to avoid the deficiencies of the above-mentioned prior art, and provides a large model retrieval enhanced generation method and system based on a vector database and an SQL database. By analyzing the file structure, fully combining the advantages of the vector database and the SQL database, the efficiency and accuracy of knowledge base retrieval are improved, and problems such as low retrieval efficiency, insufficient semantic understanding, and imperfect appendix information processing in the prior art are solved.
[0005] The present invention adopts the following technical solutions to achieve the invention purpose:
[0006] The characteristics of the large model retrieval enhanced generation method of the present invention based on a vector database and an SQL database are: constructing a vector database and an SQL database based on domain knowledge files, vectorizing the user problem text and the model inferring SQL statements, executing the SQL statements to obtain corresponding structured information, and then comprehensively analyzing and inferring by combining the content of the vector database and the SQL database, effectively improving the efficiency and accuracy of knowledge base retrieval.
[0007] The characteristics of the large model retrieval enhanced generation method of the present invention based on a vector database and an SQL database are carried out according to the following steps:
[0008] Step 1: Construct a vector database and an SQL database based on domain knowledge files:
[0009] Divide the domain knowledge file into two parts according to text A and appendix table B, and the intersection point of the two parts is the appendix name;
[0010] The domain knowledge file contains one or more files; the appendix table B is a structured table;
[0011] Expand the original appendix information C in text A to form expanded appendix information E, so as to establish a corresponding relationship with appendix table B;
[0012] Slice text A based on corresponding rules, including delimiters and sliding windows;
[0013] The sliced text fragments are converted into vectors of a fixed length through a text vectorization method, and all the vectors converted from the text fragments constitute a vector database;
[0014] The text vectorization method utilizes a text embedding model;
[0015] Extract and transform appendix table B in the domain knowledge file to constitute an SQL database;
[0016] Step 2: Vectorize the user's question and infer SQL statements:
[0017] Vectorize the user's question text M in the same way as in step 1 to obtain a question vector, and perform a similarity match between the question vector and the vectors in the vector database;
[0018] The method of the similarity match is: adopt the cosine similarity index Dis(u, v), and return the first k vectors corresponding to text fragments N with higher similarity rankings. k is a set value or dynamically adjusted according to the similarity score. The cosine similarity index Dis(u, v) is:
[0019]
[0020] u represents the question vector, and v represents the vector in the vector database;
[0021] "·" represents the dot product of vectors, "∥u∥" represents the norm of vector u, and ∥v∥ represents the norm of vector v;
[0022] Synthesize the question text M and the k vectors corresponding to text fragments N into the first prompt word L;
[0023] Input the first prompt word L into the first inference model M1, and the first inference model M1 outputs a prompt on whether to use the appendix table according to the content of the first prompt word L;
[0024] If the first inference model M1 prompts that it is not needed, there is no need to retrieve the appendix table, and then output the answer text for the question text M to the user; if the user continues to ask questions, re-enter step 2, otherwise end the entire event;
[0025] If the first inference model M1 indicates a need, output an SQL statement for retrieving relevant appendix tables and proceed to step 3;
[0026] Step 3, execute the SQL statement to obtain the retrieval result:
[0027] When the first inference model M1 indicates a need to use appendix tables, output an SQL statement for retrieving relevant appendix tables. The retrieval module executes the corresponding SQL statement in the SQL database to obtain the SQL retrieval result and proceed to step 4;
[0028] Step 4, perform comprehensive analysis and reasoning by combining the content of the vector database and the SQL database:
[0029] Organize and synthesize the question text M, the text fragments N corresponding to k vectors, and the SQL retrieval result in step 3 into the second prompt word H; input the second prompt word H into the second inference model M2, and the second inference model M2 outputs the final inference result as the answer text to the user. If the user continues to ask questions, re-enter step 2; otherwise, end the entire event.
[0030] The feature of the large model retrieval enhancement generation method based on the vector database and the SQL database in the present invention also lies in that the expanded appendix information E is obtained by the following method: adding the table names and header information in the appendix table B to the original appendix information C in the form of text description.
[0031] The feature of the large model retrieval enhancement generation method based on the vector database and the SQL database in the present invention also lies in that the first inference model M1 and the second inference model M2 are pre-trained language models with natural language generation and SQL statement parsing capabilities, including the open-source model deepseek-r1 or the closed-source model GPT-4.
[0032] The present invention can significantly improve the efficiency and accuracy of knowledge base retrieval, reduce the time for users to perform manual operations, and enhance the user experience by vectorizing the text information in the knowledge base and combining the collaboration of multiple inference models for reasoning and induction. Compared with the prior art, the advantages of the present invention are:
[0033] 1. More accurate semantic understanding of user questions: Through the vectorization technology, it is possible to better understand the semantics of user questions and avoid the limitations of keyword matching.
[0034] 2. Appendix table information processing: The present invention not only processes the main text information but also processes the appendix information, eliminating the need for users to manually view the appendix and improving the retrieval efficiency.
[0035] 3. Structured Information Utilization: Make full use of the structured information in the database (such as the relationship between the main text and the appendix) to ensure the comprehensiveness of the retrieval results and improve the problem that the existing retrieval enhanced generation method based on the vector database has weak query capabilities for tabular data. Brief Description of the Drawings
[0036] Figure 1 It is a schematic diagram of the principle of the method of the present invention.
[0037] Figure 2 It is a schematic diagram of the construction of the vector database and the SQL database in the method of the present invention.
[0038] Figure 3 It is a flowchart of the method of the present invention from the user's question to obtaining the result. Detailed Embodiment
[0039] Refer to Figure 1 and Figure 2 In this embodiment, the large model retrieval enhanced generation method based on the vector database and the SQL database is carried out according to the following steps:
[0040] Step 1. Construct a vector database and an SQL database based on the domain knowledge file:
[0041] Divide the domain knowledge file into two parts according to text A and appendix table B. The intersection point of the two parts is the appendix name. The domain knowledge file contains one or more files, and appendix table B are all structured tables.
[0042] Expand the original appendix information C in text A to form expanded appendix information E, so as to establish a corresponding relationship with appendix table B; the expanded appendix information E is obtained by the following method: add the table name and header information in appendix table B to the original appendix information C in text description.
[0043] Slice text A based on corresponding rules, including delimiters and sliding windows.
[0044] One slicing method is: use the paragraph mark "\n\n" as the delimiter, and further slice the long paragraph by a 512-character sliding window with a step size of 256 characters to slice text A into multiple text fragments.
[0045] The sliced text fragments are converted into fixed-length vectors through the text vectorization method. All the vectors converted from the text fragments constitute the vector database; the text vectorization method is to convert text into vectors for machine learning and data analysis. By converting text data into vectors, fast matching between the user's question text and the vector database can be achieved, so as to extract relevant text content; the text vectorization method generally uses text embedding models, such as BGE M3-Embedding.
[0046] Extract and transform Appendix Table B in the domain knowledge file to form an SQL database.
[0047] The extraction of tables in the domain knowledge file can be achieved by using scripts. For docx format files, VBA (Visual Basic for Applications) technology can be used for extraction. By traversing the document object model, the target table can be accurately located. After generating a standardized CSV intermediate file, it can be stored in the SQL database through the Open Database Connectivity (ODBC) interface. For pdf format files, the python language can be used for extraction (using the Tabula library), and pre-compiled SQL statements can be used to batch write to the database.
[0048] The SQL mentioned above, namely Structured Query Language, is a database query and programming language used to access data and query, update, and manage relational database systems.
[0049] Step 2: Vectorize the user's question and infer SQL statements in the model:
[0050] Vectorize the user's question text M in the same way as in Step 1 to obtain a question vector, and match the similarity between the question vector and the vectors in the vector database.
[0051] The method of similarity matching is as follows: Use the cosine similarity index Dis(u, v) to return the top k text fragments N corresponding to the vectors with higher similarity rankings. k is a set value or dynamically adjusted according to the similarity score. One method of dynamically adjusting the k value is: Set a similarity threshold of 0.7. If more than 3 scores among the top 5 results are ≥ the threshold, then k = 5; otherwise, k = 10. The cosine similarity index Dis(u, v) is:
[0052]
[0053] u represents the question vector, and v represents the vector in the vector database;
[0054] “·” represents the dot product of vectors, “∥u∥” represents the norm of vector u, and ∥v∥ represents the norm of vector v;
[0055] Synthesize the question text M and the k text fragments N corresponding to the vectors into the first piece of prompt L.
[0056] Input the first piece of prompt L into the first inference model M1, and the first inference model M1 outputs a prompt indicating whether the appendix table needs to be used according to the content of the first piece of prompt L.
[0057] If the first inference model M1 indicates that it is not needed, there is no need to retrieve the appendix table, and then the answer text for the question text M is output to the user; if the user continues to ask questions, go back to step 2, otherwise end the whole event;
[0058] If the first inference model M1 indicates that it is needed, output the SQL statement for retrieving the relevant appendix table and enter step 3;
[0059] Since the appendix information C already contains the complete table structure information in the first step, the first inference model M1 can infer the SQL statement for retrieving the relevant appendix table based on the complete table structure information.
[0060] Example of the first prompt word L:
[0061] "You are an expert in the relevant field. Now you need to answer the user's question based on the reference materials.
[0062] The user's question is: {user's question}
[0063] The reference materials are: {text corresponding to k vectors}
[0064] If the text corresponding to the k vectors contains key appendix table information and the relevant table data is needed to answer the user's question, please reply "SQL database query is required" in the first line, and then write a relevant SQL statement without returning other content. If the text corresponding to the k vectors does not contain key appendix table information, or the relevant table data is not needed to answer the user's question, please reply "SQL database query is not required" in the first line, and then directly answer the user's question."
[0065] Step 3: Execute the SQL statement to obtain the retrieval result:
[0066] When the first inference model M1 indicates that the appendix table is needed, output the SQL statement for retrieving the relevant appendix table. The retrieval module executes the corresponding SQL statement in the SQL database to obtain the SQL retrieval result and enter step 4;
[0067] Step 4: Conduct comprehensive analysis and inference by combining the contents of the vector database and the SQL database:
[0068] Collate and synthesize the question text M, the text fragments N corresponding to the k vectors, and the SQL retrieval result in step 3 into the second prompt word H; input the second prompt word H into the second inference model M2, and the second inference model M2 outputs the final inference result as the answer text to the user. If the user continues to ask questions, go back to step 2, otherwise end the whole event.
[0069] Examples of the second prompt word H are as follows:
[0070] "You are an expert in the relevant field and now need to answer the user's question based on the reference materials."
[0071] The user's question is: {user's question}
[0072] The reference materials are: {texts corresponding to k vectors}{retrieval results in step 3}"
[0073] In specific implementation, the first inference model M1 and the second inference model M2 are pre-trained language models with natural language generation and SQL statement parsing capabilities, including the open-source model deepseek-r1 or the closed-source model GPT-4. For users, the inference model is used to semantically understand the user's question and output accurate SQL statements. Users do not need to further learn SQL management technology. The vector database and the SQL database are equivalent to black boxes for users. Users only need to input text information to directly obtain the answer.
[0074] From the user's perspective, the flowchart from the question to the result obtained by the user is as Figure 3 shown.
[0075] Through the above specific implementation manners, those skilled in the art of the present technology can easily implement the present invention. However, it should be understood that the present invention is not limited to the above specific implementation manners. Based on the disclosed implementation manners, those skilled in the art can arbitrarily combine different technical features to implement different technical solutions.
Claims
1. A large model retrieval enhanced generation method based on a vector database and an SQL database, characterized in that: Construct a vector database and an SQL database based on the domain knowledge file, vectorize the user's question text and model the SQL statement, execute the SQL statement to obtain the corresponding structured information, and then comprehensively analyze and reason based on the content of the vector database and the SQL database to effectively improve the efficiency and accuracy of knowledge base retrieval.
2. The large model retrieval enhanced generation method based on a vector database and an SQL database according to claim 1, characterized in that The steps are as follows: Step 1: Construct a vector database and an SQL database based on the domain knowledge file: Divide the domain knowledge file into two parts according to Text A and Appendix Table B. The intersection of the two parts is the appendix name; The domain knowledge file contains one or more files; the Appendix Table B is a structured table; Expand the original appendix information C in Text A to form expanded appendix information E, so as to establish a corresponding relationship with Appendix Table B; Slice Text A based on corresponding rules, including delimiters and sliding windows; The sliced text fragments are converted into fixed-length vectors through a text vectorization method, and all the vectors converted from the text fragments constitute the vector database; The text vectorization method utilizes a text embedding model; Extract and transform Appendix Table B in the domain knowledge file to constitute the SQL database; Step 2: Vectorize the user's question and model the SQL statement: Vectorize the user's question text M in the same way as in Step 1 to obtain a question vector, and perform a similarity match between the question vector and the vectors in the vector database; The method of the similarity match is: adopt the cosine similarity index Dis(u, v), and return the top k vectors corresponding to the text fragments N with higher similarity rankings. k is a set value or dynamically adjusted according to the similarity score. The cosine similarity index Dis(u, v) is: u represents the question vector, and v represents the vector in the vector database; "·" represents the dot product of vectors, "∥u∥" represents the norm of vector u, and ∥v∥ represents the norm of vector v; Synthesize the question text M and the k vectors corresponding to the text fragments N into the first prompt word L; Input the first prompt word L into the first inference model M1, and the first inference model M1 outputs a prompt on whether to use the appendix table according to the content of the first prompt word L; If the first inference model M1 prompts that it is not needed, there is no need to retrieve the appendix table, and then output the answer text for the question text M to the user; if the user continues to ask questions, re-enter Step 2, otherwise end the entire event; If the first inference model M1 prompts that it is needed, output the SQL statement for retrieving the relevant appendix table and enter Step 3; Step 3: Execute the SQL statement to obtain the retrieval result: When the first inference model M1 prompts that it is necessary to use the appendix table, output the SQL statement for retrieving the relevant appendix table. The retrieval module executes the corresponding SQL statement in the SQL database to obtain the SQL retrieval result and enter Step 4; Step 4: Conduct comprehensive analysis and reasoning by combining the content of the vector database and the SQL database: Collate the problem text M, the text fragments N corresponding to k vectors, and the SQL retrieval results in step 3 into the second prompt H; input the second prompt H into the second inference model M2, and the second inference model M2 outputs the final inference result as the answer text to the user. If the user continues to ask questions, re-enter step 2, otherwise end the entire event.
3. The method for retrieval-enhanced generation of large models based on a vector database and an SQL database according to claim 1, characterized in that: The extended appendix information E is obtained as follows: Add the table names and header information in the appendix table B to the original appendix information C in the form of a text description.
4. The large model retrieval enhanced generation method based on a vector database and an SQL database according to claim 1, characterized in that: The first inference model M1 and the second inference model M2 are pre-trained language models with natural language generation and SQL statement parsing capabilities, including the open-source model deepseek-r1 or the closed-source model GPT-4.