Database query method, device and equipment based on large model retrieval enhancement generation
By introducing vector knowledge database and large-model retrieval enhancement generation technology into database queries, the problems of inaccuracy and efficiency in traditional query solutions are solved, and a more efficient and accurate database query process is achieved.
Patent Information
- Application Number
- CN202510578192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The query accuracy and query efficiency of traditional database query schemes are low, and multiple Q&A processes are required to obtain satisfactory results.
The pre-constructed vector knowledge database searches the query problem, obtains relevant table structure data, domain knowledge and reference examples, and then enhances the generation of the query problem, obtains multiple alternative SQL statements, and generates more accurate query results by extracting the optimal SQL statement, bias correction and execution, and combining the language model to generate more accurate query results.
It significantly improves the accuracy and efficiency of database queries, can obtain satisfactory query results more conveniently, and reduces the complexity of the query link.
Smart Images

Figure CN120104762A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a database query method, device and equipment based on large model retrieval enhanced generation. Background Art
[0002] In the field of language processing, large models mainly refer to large language models, or large language models (LLM), which are deep learning models trained based on massive text data. They can not only generate natural language text, but also deeply understand the meaning of text and handle various natural language tasks, such as text summarization, question answering, and translation.
[0003] In the related art, traditional query solutions are usually implemented only based on large models. Specifically, the query results of the questions to be queried can be directly obtained based on the large models. However, in the above query solutions, since the questions to be queried are usually not perfect and specific enough, and the performance of the model itself is difficult to achieve the best, it usually takes multiple question-and-answer processes to obtain relatively satisfactory query results.
[0004] It can be seen that traditional query solutions have technical problems of low query accuracy and query efficiency. Summary of the invention
[0005] The present invention provides a database query method, device and equipment based on large model retrieval enhanced generation, which are used to solve the defects of low query accuracy and query efficiency of traditional query solutions.
[0006] In one aspect, the present invention provides a database query method based on large model retrieval enhanced generation, comprising: Searching the query question based on the pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question; According to the table structure data, domain knowledge and reference examples, the query question is enhanced and generated to obtain multiple candidate SQL statements; Extracting an optimal SQL statement from the multiple candidate SQL statements, correcting deviations of the optimal SQL statement and executing it to obtain an SQL statement execution result; Extracting the summary instruction in the question to be queried, and constructing a query prompt word according to the question to be queried, the summary instruction, the optimal SQL statement and the execution result of the SQL statement; The query prompt word is input into the language model to obtain the query result output by the language model.
[0007] According to the database query method based on large model retrieval enhancement generation provided by the present invention, a query question is searched based on a pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question, including: Vectorize the query question to obtain a question vector; Retrieving table structure data, domain knowledge and a first SQL example related to the vector semantics of the question vector from a pre-built vector knowledge database; Fusion screening is performed based on the first SQL sample to obtain a reference sample.
[0008] According to the database query method based on large model retrieval enhancement generation provided by the present invention, fusion screening is performed according to the first SQL sample to obtain a reference sample, including: Directly generate a corresponding initial SQL statement according to the query question, and retrieve a second SQL example most similar to the initial SQL statement from a vector knowledge database; Replace the table name, column name and example value in the table structure data by mask marks to obtain general table data; Establishing a query vector according to the general table data, and retrieving a third SQL example most similar to the query vector from a vector knowledge database; The first SQL sample, the second SQL sample, and the third SQL sample are reordered and fused and screened to obtain a reference sample.
[0009] According to the database query method based on large model retrieval enhanced generation provided by the present invention, the query question is enhanced and generated according to the table structure data, domain knowledge and reference examples to obtain multiple candidate SQL statements, including: Determine key instructions and constraints in the query question, and randomly sort the query question, the table structure data, the domain knowledge, the reference examples, the key instructions and the constraints to generate multiple prompt word templates; Generate a corresponding SQL statement according to each prompt word template to obtain multiple candidate SQL statements.
[0010] According to the database query method based on large model retrieval enhanced generation provided by the present invention, extracting the best SQL statement from the multiple candidate SQL statements includes: Performing syntax verification on the multiple candidate SQL statements to obtain pre-selected SQL statements that pass the syntax verification; If there are multiple pre-selected SQL statements, perform performance evaluation on the multiple pre-selected SQL statements to obtain a performance evaluation value of each pre-selected SQL statement; The pre-selected SQL statement with the highest performance evaluation value is used as the optimal SQL statement.
[0011] According to the database query method based on large model retrieval enhanced generation provided by the present invention, a plurality of pre-selected SQL statements are subjected to performance evaluation to obtain a performance evaluation value of each pre-selected SQL statement, including: Determine the execution time, resource consumption value, and number of scanned rows for each pre-selected SQL statement; Normalizing the execution time, the resource consumption value, and the number of scanned lines respectively to obtain a normalized execution time value, a normalized resource consumption value, and a normalized scanned line number value; The performance evaluation value of each pre-selected SQL statement is calculated based on the normalized value of the execution time, the normalized value of the resource consumption, and the normalized value of the number of scanned rows.
[0012] According to the database query method based on large model retrieval enhanced generation provided by the present invention, extracting the best SQL statement from the multiple candidate SQL statements includes: Execute each candidate SQL statement separately to obtain the execution result corresponding to each candidate SQL statement; Determine the accuracy and completeness of the execution results corresponding to each candidate SQL statement respectively; Voting on the execution results corresponding to each of all candidate SQL statements according to the accuracy and completeness; The alternative SQL statement corresponding to the execution result with the most votes after the voting is completed is regarded as the optimal SQL statement.
[0013] According to the database query method based on large model retrieval enhanced generation provided by the present invention, deviation correction is performed on the optimal SQL statement, including: The optimal SQL statement is respectively corrected for spelling errors, missing keywords, incorrect bracket matching, data type, and logic errors.
[0014] On the other hand, the present invention also provides a database query device based on large model retrieval enhancement generation, comprising: A retrieval module is used to search the query question based on the pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question; An enhanced generation module is used to enhance and generate the query question based on the table structure data, domain knowledge and reference examples to obtain multiple candidate SQL statements; An execution module, used for extracting an optimal SQL statement from the multiple candidate SQL statements, correcting deviations of the optimal SQL statement and executing it, to obtain an SQL statement execution result; A prompt word construction module, used to extract the summary instruction in the question to be queried, and construct a query prompt word according to the question to be queried, the summary instruction, the optimal SQL statement and the execution result of the SQL statement; The query module is used to input the query prompt word into the language model to obtain the query result output by the language model.
[0015] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements a database query method based on large model retrieval enhanced generation as described in any one of the above.
[0016] The database query method, device and equipment based on large model retrieval enhanced generation provided by the present invention can search the query problem based on a pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query problem; based on the table structure data, domain knowledge and reference examples, enhance the generation of the query problem to obtain multiple candidate SQL statements; extract the optimal SQL statement from the multiple candidate SQL statements, correct the deviation of the optimal SQL statement and execute it to obtain the SQL statement execution result; construct a query prompt word based on the query problem, the summary instruction in the query problem, the optimal SQL statement and the SQL statement execution result; input the query prompt word into the language large model to obtain the query result output by the language large model. Since the retrieval enhancement generation technology is introduced in the query stage, a series of data related to the query question can be retrieved through the vector knowledge database, and then multiple alternative SQL statements can be obtained through enhanced generation. Subsequently, the optimal SQL statement is extracted, and query prompts are constructed based on the query question, the summary instructions in the query question, the optimal SQL statement, and the SQL statement execution results. More accurate query prompts can be obtained, and more satisfactory query results can be obtained conveniently, thereby improving the accuracy and efficiency of the query stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flowchart of a database query method based on large model retrieval enhanced generation provided by an embodiment of the present invention; Figure 2is a schematic diagram of the determination principle of the reference sample in the embodiment of the present invention; Figure 3 1 is a schematic diagram of the implementation principle of the present invention of using a random sorting method to perform different combinations of six elements; Figure 4 It is a structural schematic diagram of a database query device based on large model retrieval enhanced generation provided by an embodiment of the present invention; Figure 5 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] Combine the following Figures 1 to 5 Describe the detailed scheme of the database query method, device and equipment based on large model retrieval enhanced generation provided by the embodiments of the present invention.
[0021] Figure 1 It is a flowchart of a database query method based on large model retrieval enhanced generation provided by an embodiment of the present invention.
[0022] like Figure 1 As shown, the database query method based on large model retrieval enhanced generation provided in the embodiment of the present invention can be executed by a computer or server with data receiving and sending and data processing capabilities. The method mainly includes the following steps: Step 110: Search the query question based on the pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question.
[0023] In this embodiment, the vector knowledge database is constructed mainly based on relevant database table creation statements, SQL query records accumulated during actual use, and relevant domain knowledge.
[0024] Among them, the database table creation statement provides database table structure information, SQL query records serve as reference examples, and relevant domain knowledge serves as prior knowledge to guide the generation of SQL statements that meet the real intention.
[0025] It should be noted that this embodiment uses L-Schema to present the hierarchical structure between databases, tables and columns, as shown in Table 1 below, which respectively shows the use of DDL Schema (data definition language model) and L-Schema to describe a school-related database structure.
[0026] Table 1 Description forms of DDL Schema and L-Schema
[0027] As shown in Table 1 above, in the description statement corresponding to the DDL Schema, the student table is created using the CREATE TABLE statement, which contains the fields id (integer, auto-increment primary key), name (a string of up to 50 characters, not allowed to be empty), and age (integer).
[0028] The score table is also created using CREATE TABLE, and contains the fields id (integer, auto-increment primary key), student_id (integer, foreign key associated with the student table), subject (a string of up to 50 characters, not allowed to be empty), and score (a decimal number with a total length of 5 and an accuracy of two decimal places). At the same time, the student_id field is specified through FOREIGN KEY to reference the student_id of the student table.
[0029] In the description statement corresponding to L-Schema, the database ID is marked as school. The table structure description of the student table includes id (integer, primary key, example student IDs such as 23011111, 240134), name (string of up to 50 characters, example name "Li Ming"), age (integer, example age such as 13, 12), and the score table contains id (integer, primary key, example score unique code such as 2105031144), student_id (string, example corresponding student ID such as 23011111, 240134), score (decimal number with two decimal places and a total length of 5, example score such as 96.00). In the foreign key relationship, it is clearly stated that student.id=score.student_id, indicating the association relationship between the two tables.
[0030] It is not difficult to find that the L-Schema proposed in this embodiment presents the hierarchical structure between databases, tables, and columns in a semi-structured form, and uses specific tags to identify them. Specifically, "[DB]" represents the database,
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046] Figure 2 Figure 2
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053] Figure 3
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066] norm norm norm
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089] Figure 4
[0090]
[0091]
[0092] Figure 5
[0093] Figure 5
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] " indicates a table, and "[Foreign keys]" indicates a foreign key. For each table, the field name and description are provided. The information in the table is converted into a list, where each item is a tuple, representing the detailed information of the column. Each column includes the column name, data type, column description, primary key identifier, and example value. In addition, due to the importance of foreign keys, foreign keys need to be listed. Subsequently, the embedding model bge-m3-large can be used to vectorize the L-Schema description content of the database, SQL samples, and related domain knowledge and store them in the vector knowledge database. The hierarchical structure between the database, tables, and columns is presented in a semi-structured form through L-Schema, and is identified by specific tags, making the database representation more compact and clear, and can It can better present the hierarchical structure between databases, tables and columns. Step 120: Based on the table structure data, domain knowledge and reference examples, the query question is enhanced and generated to obtain multiple alternative SQL statements. It can be understood that multiple alternative SQL statements can be obtained based on their corresponding prompt word templates. Step 130: Extract the optimal SQL statement from multiple alternative SQL statements, correct the deviation of the optimal SQL statement and execute it to obtain the SQL statement execution result. In this embodiment, the SQL statement execution result obtained by extracting the optimal SQL statement, correcting the deviation and executing it can provide a more effective data basis for the construction of subsequent query prompt words. Step 140: Extract the query question Summarize the instructions, and construct query prompts based on the question to be queried, the summary instructions, the optimal SQL statement, and the SQL statement execution results. It can be understood that since the query prompts are obtained based on a variety of data during the construction process, they can express the purpose of the query more comprehensively and specifically. Step 150: Input the query prompts into the large language model to obtain the query results output by the large language model. The method provided in this embodiment can overcome the problems of limited storage capacity of large language models, difficulty in obtaining the latest information in real time, and insufficient knowledge in specific fields through retrieval enhancement generation. It assists the model in generating more accurate, detailed and targeted answers through an integrated retrieval mechanism. It combines the advantages of large language models and retrieval systems to improve the large language model. The accuracy, relevance and timeliness of the content generated by the model can be improved. Compared with the generation that only relies on the large language model, information can be retrieved from the external knowledge base, avoiding the model's hallucination problem, and improving the processing ability of problems with high real-time requirements. It greatly simplifies the database query process, allowing non-technical users to easily interact with the database, improving the efficiency of database operations, and promoting the popularization of database technology, and improving the convenience and accuracy of the query link. In one embodiment, the query question is retrieved based on the pre-constructed vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question, specifically including: first, the query question is vectorized to obtain a question vector.In practical applications, the question to be queried can be input into the embedding model bge-m3-large for vectorization processing, thereby providing an effective data basis for subsequent retrieval. Then, the table structure data, domain knowledge and the first SQL sample related to the vector semantics of the question vector are retrieved from the pre-built vector knowledge database. In this embodiment, the similarity between the question vector and the vector in the vector knowledge database can be calculated. Specifically, the cosine similarity or Euclidean distance between the question vector and the vector in the vector knowledge database can be calculated, and then the similarity between the question vector and the vector in the vector knowledge database can be judged. Finally, fusion screening is performed based on the first SQL sample to obtain a reference sample. The principle of determining the reference sample in this embodiment can be referred to, as shown, fusion screening is performed based on the first SQL sample to obtain a reference sample, specifically including: on the one hand, directly generating a corresponding initial SQL statement based on the question to be queried, and retrieving the second SQL sample that is most similar to the initial SQL statement from the vector knowledge database. In practical applications, the initial SQL statement and the SOL examples in the vector knowledge database can be converted into their respective corresponding statement vectors, and then the similarity between the initial SQL statement and the SQL examples in the vector knowledge database can be determined by calculating the cosine similarity or Euclidean distance between the two vectors, and then the second SQL example most similar to the initial SQL statement can be determined. On the other hand, the table name, column name and example value in the table structure data are replaced by mask marks to obtain general table data. Then, a query vector is established based on the general table data, and the third SQL example most similar to the query vector is retrieved from the vector knowledge database. It can be understood that using mask marks to replace all table names, column names and example values in the table structure data of the query problem can eliminate specific information in the field. Then, the nearest neighbor algorithm is used to calculate the similarity between the query vector and the corresponding vector of the data in the vector knowledge database, and then the third SQL example most similar to the query vector is screened. Finally, the first SQL example, the second SQL example and the third SQL example are reordered and fused to obtain a reference example. In practical applications, the reordering model can be used to merge and reorder the first SQL example, the second SQL example, and the third SQL example, and finally select one or more reference examples with the highest similarity. In one embodiment, based on the table structure data, domain knowledge, and reference examples, the query problem is enhanced and generated to obtain multiple alternative SQL statements, specifically including: first, determining the key instructions and constraints in the query problem, and randomly sorting the query problem, table structure data, domain knowledge, reference examples, key instructions, and constraints to generate multiple prompt word templates. It can be understood that this embodiment divides the main content of the prompt word template into six elements, including: key instructions, table structure data, reference examples, constraints, domain knowledge, and the query problem.Considering that the six elements are arranged in different orders, different element formats and contents will affect the effect of the prompt word. Therefore, as shown, the present invention uses a random sorting method to combine the six elements in different ways to generate n prompt word templates. Then, a corresponding SQL statement is generated according to each prompt word template to obtain multiple candidate SQL statements. Subsequently, n candidate SQL statements can be obtained according to the n prompt word templates, and then the optimal SQL statement is obtained by screening the candidate SQL statements. In one embodiment, the optimal SQL statement among the multiple candidate SQL statements is extracted, specifically including: first, performing syntax verification on the multiple candidate SQL statements to obtain a pre-selected SQL statement that passes the syntax verification. It can be understood that when extracting the optimal SQL statement, it is necessary to first ensure that the statement syntax is correct, and the syntax verification tool provided by the database or the third-party library can be used to perform syntax verification on the candidate SQL statements, and the statements with syntax errors are eliminated, so as to obtain the pre-selected SQL statements that pass the syntax verification. Then, if there are multiple pre-selected SQL statements, the performance evaluation of the multiple pre-selected SQL statements is performed to obtain the performance evaluation value of each pre-selected SQL statement. In actual applications, if there is only one pre-selected SQL statement, the pre-selected SQL statement that has passed the language verification can be directly used as the optimal SQL statement. If there are multiple pre-selected SQL statements, since performance is a key factor in measuring the quality of SQL statements, it is necessary to further screen through performance evaluation to determine the optimal SQL statement. Finally, the pre-selected SQL statement with the highest performance evaluation value is used as the optimal SQL statement. In a specific implementation, a performance evaluation is performed on multiple pre-selected SQL statements to obtain a performance evaluation value of each pre-selected SQL statement, specifically including: first, determining the execution time, resource consumption value, and number of scanned rows of each pre-selected SQL statement. In actual applications, each pre-selected SQL statement can be actually executed and the execution time can be recorded. The shorter the execution time, the better the statement performance. The resource consumption value can be determined by indicators such as CPU usage and memory occupancy, and relevant information can be obtained with the help of the database performance monitoring tool. The fewer the number of scanned rows, the better the statement performance is. The number of scanned rows can be viewed through the database execution plan. Then, the execution time, resource consumption value, and number of scanned rows are normalized to obtain the normalized execution time value, the normalized resource consumption value, and the normalized scanned row value. It is understandable that the numerical ranges of the execution time, resource consumption value, and number of scanned rows may vary greatly. In order to make them comparable in performance evaluation, these indicators need to be normalized and their values mapped to the numerical interval [0,1]. In practical applications, normalization can be implemented using the maximum and minimum value normalization function. Finally, the performance evaluation value of each pre-selected SQL statement is calculated based on the normalized execution time value, resource consumption value, and scanned row value.In this embodiment, the performance evaluation value of the pre-selected SQL statement can be specifically calculated as follows: wherein P represents the performance evaluation value of the pre-selected SQL statement, E represents the normalized value of the execution time, R represents the normalized value of the resource consumption, and S represents the normalized value of the number of scanned rows. Since the performance evaluation value can comprehensively evaluate the performance of each pre-selected SQL statement, the pre-selected SQL statement with the largest performance evaluation value can be used as the pre-selected SQL statement. It is not difficult to find that this embodiment can accurately and effectively implement the screening of the optimal SQL statement by combining syntax verification with performance evaluation. In another embodiment, the optimal SQL statement from multiple candidate SQL statements is extracted, specifically including: the first step, executing each candidate SQL statement separately, and obtaining the execution result corresponding to each candidate SQL statement. This embodiment adopts a voting method to implement the screening of the optimal SQL statement. In this process, it is necessary to execute each candidate SQL statement separately to obtain the execution result corresponding to each candidate SQL statement. The second step is to determine the accuracy and completeness of the execution result corresponding to each candidate SQL statement. First, the expected number of rows and expected fields of the execution result are set respectively, and then the actual number of rows and actual fields of the execution result corresponding to each candidate SQL statement are determined respectively, and the actual number of rows is subtracted from the expected number of rows to obtain the row number deviation value, which is used to measure the accuracy of the execution result. All the actual fields in each execution result are compared with the expected fields to determine whether the execution result contains all the expected fields and each expected field has a corresponding data value, and the judgment result is used to measure the completeness of the execution result. In the third step, according to the accuracy and completeness, the execution results corresponding to all the candidate SQL statements are voted. During the voting process, one vote is cast for the execution result whose row number deviation value is positive and whose value is less than the corresponding preset threshold, and one vote is cast for the execution result whose number of expected fields is greater than the corresponding preset threshold and each expected field has a corresponding data value. In the fourth step, the alternative SQL statement corresponding to the execution result with the most votes after the voting is completed is used as the optimal SQL statement. It is not difficult to find that this embodiment can conveniently and efficiently realize the screening of the optimal SQL statement by voting. In one embodiment, the optimal SQL statement is corrected for deviations, including: correcting spelling errors, missing keywords, incorrect bracket matching, data type and logic errors in the optimal SQL statement. In this embodiment, in order to ensure the accuracy of the optimal SQL statement, a series of correction operations are subsequently performed to correct spelling errors, missing keywords, incorrect bracket matching, incorrect data types and logic errors in the optimal SQL statement, thereby further improving data accuracy. In some embodiments, query prompts are constructed based on the question to be queried, summary instructions, optimal SQL statements and SQL statement execution results, specifically including: on the one hand, based on the question to be queried, an introduction statement is generated at the beginning to clearly and clearly explain the question to be queried originally raised by the user, so that the big model understands the core and background of the problem.On the other hand, an instruction description statement is generated based on the summary instruction, and the specific requirements for summarizing and processing the query results are described in detail, so that the big model can clearly understand the content format that needs to be presented in the end. It should be noted that the summary instruction is further extracted based on the key instructions that have been determined before. On the other hand, based on the optimal SQL statement and the SQL execution result, the statement display content and the result presentation form content are generated. In practical applications, appropriate comments can be added to the statement display content to explain its role to help the big model understand how to obtain data. At the same time, the result presentation form content displays the SQL statement execution results in a clear format (such as a table), which is convenient for the big model to summarize based on this form, and clearly informs the big model to process the results according to the summary instruction and output the summary content. Finally, the beginning introduction statement, instruction description statement, statement display content and result presentation form content are integrated to obtain the query prompt word. Subsequently, the query prompt word is input into the language big model to make it summarize and reply to obtain the query result. In summary, the database query method based on large model retrieval enhanced generation provided by the embodiment of the present invention has at least the following advantages: First, it can not only understand the user's natural language query needs, but also retrieve key information from existing knowledge to generate more accurate SQL statements, significantly improving the accuracy and intelligence of SQL generation. Second, a new model representation method of L-Schema is proposed, which presents the hierarchical structure between databases, tables and columns in a semi-structured form, enhancing the understanding of the database structure by the large language model. Third, by integrating multiple example selection strategies to screen similar SQL statements, it helps LLM to learn by analogy and improve the accuracy and rationality of SQL statement generation. Fourth, the six elements that make up the prompt words are arranged and combined in multiple sequences to generate multiple prompt word templates, and SQL statements are generated according to different prompt word templates. Through reasonable SQL statement screening strategies and deviation correction methods, the accuracy of SQL statement generation is significantly improved. Based on the same general inventive concept, the present invention also protects a database query device based on large model retrieval enhanced generation. The database query device based on large model retrieval enhanced generation provided by the present invention is described below. The database query device based on large model retrieval enhanced generation described below and the database query method based on large model retrieval enhanced generation described above can be referenced to each other.As shown, the database query device based on large model retrieval and enhanced generation provided by the embodiment of the present invention specifically includes: a retrieval module 210, which is used to search the query question based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question; an enhanced generation module 220, which is used to enhance the query question based on the table structure data, domain knowledge and reference examples to obtain multiple alternative SQL statements; an execution module 230, which is used to extract the optimal SQL statement from multiple alternative SQL statements, correct the deviation of the optimal SQL statement and execute it to obtain the SQL statement execution result; a prompt word construction module 240, which is used to extract the summary instruction in the query question, and construct a query prompt word based on the query question, the summary instruction, the optimal SQL statement and the SQL statement execution result; a query module 250, which is used to input the query prompt word into the language large model to obtain the query result output by the language large model. It can be seen that the database query device based on large model retrieval enhancement generation provided by the embodiment of the present invention, due to the introduction of retrieval enhancement generation technology in the query link, can obtain a series of data related to the question to be queried through the vector knowledge database retrieval, and then obtain multiple candidate SQL statements through enhanced generation, and then extract the optimal SQL statement, and construct the query prompt word according to the question to be queried, the summary instruction in the question to be queried, the optimal SQL statement and the SQL statement execution result, so as to obtain more accurate query prompt words, and then more conveniently obtain more satisfactory query results, thereby improving the accuracy and query efficiency of the query link. Regarding the device in the above embodiment, the specific way in which each module performs the operation has been described in detail in the embodiment of the method, and will not be described in detail here. It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. As shown, the electronic device may include: a processor (processor) 310, a communication interface (Communications Interface) 320, a memory (memory) 330 and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 complete mutual communication through the communication bus 340.The processor 310 can call the logic instructions in the memory 330 to execute the database query method based on the large model retrieval enhancement generation provided by the above embodiments, the method comprising: searching the query problem based on the pre-built vector knowledge database to obtain the table structure data, domain knowledge and reference examples related to the query problem; based on the table structure data, domain knowledge and reference examples, enhancing the query problem to obtain multiple candidate SQL statements; extracting the optimal SQL statement from the multiple candidate SQL statements, correcting the deviation of the optimal SQL statement and executing it to obtain the SQL statement execution result; extracting the summary instruction in the query problem, and constructing the query prompt word based on the query problem, the summary instruction, the optimal SQL statement and the SQL statement execution result; inputting the query prompt word into the language large model to obtain the query result output by the language large model. In addition, the logic instructions in the above memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., various media that can store program codes. On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the database query method based on large model retrieval enhanced generation provided in the above embodiments, the method including: searching the query question based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question; enhancing and generating the query question based on the table structure data, domain knowledge and reference examples to obtain multiple alternative SQL statements; extracting the optimal SQL statement from the multiple alternative SQL statements, correcting the deviation of the optimal SQL statement and executing it to obtain the SQL statement execution result; extracting the summary instruction in the query question, and constructing a query prompt word based on the query question, the summary instruction, the optimal SQL statement and the SQL statement execution result; inputting the query prompt word into the language large model to obtain the query result output by the language large model.On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the database query method based on the large model retrieval enhancement generation provided in the above embodiments is implemented, and the method includes: searching the query problem based on the pre-constructed vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query problem; enhancing the query problem based on the table structure data, domain knowledge and reference examples to obtain multiple candidate SQL statements; extracting the optimal SQL statement from the multiple candidate SQL statements, correcting the deviation of the optimal SQL statement and executing it to obtain the SQL statement execution result; extracting the summary instruction in the query problem, and constructing the query prompt word based on the query problem, the summary instruction, the optimal SQL statement and the SQL statement execution result; inputting the query prompt word into the language large model to obtain the query result output by the language large model. The above-described device embodiments are merely illustrative, wherein the unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative labor. Through the description of the above implementation methods, technicians in this field can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment. Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention is described in detail with reference to the above embodiments, ordinary technicians in this field should understand that it can still modify the technical solutions recorded in the above embodiments, or replace some of the technical features therein by equivalent; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A database query method based on large model retrieval enhancement generation, characterized in that: include: Searching the query question based on the pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question; According to the table structure data, domain knowledge and reference examples, the query question is enhanced and generated to obtain multiple candidate SQL statements; Extracting an optimal SQL statement from the multiple candidate SQL statements, correcting deviations of the optimal SQL statement and executing the statement to obtain an SQL statement execution result; Extracting the summary instruction in the question to be queried, and constructing a query prompt word according to the question to be queried, the summary instruction, the optimal SQL statement and the execution result of the SQL statement; The query prompt word is input into the language model to obtain the query result output by the language model.
2. The database query method based on large model retrieval enhancement generation according to claim 1 is characterized in that: Search the query question based on the pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question, including: Vectorize the query question to obtain a question vector; Retrieving table structure data, domain knowledge and a first SQL example related to the vector semantics of the question vector from a pre-built vector knowledge database; Fusion screening is performed based on the first SQL sample to obtain a reference sample.
3. The database query method based on large model retrieval enhancement generation according to claim 2 is characterized in that: Based on the first SQL sample, fusion screening is performed to obtain reference samples, including: Directly generate a corresponding initial SQL statement according to the query question, and retrieve a second SQL example most similar to the initial SQL statement from a vector knowledge database; Replacing the table name, column name and example value in the table structure data by mask marks to obtain general table data; Establishing a query vector according to the general table data, and retrieving a third SQL example most similar to the query vector from a vector knowledge database; The first SQL sample, the second SQL sample, and the third SQL sample are reordered and fused and screened to obtain a reference sample.
4. The database query method based on large model retrieval enhancement generation according to claim 1 is characterized in that: Based on the table structure data, domain knowledge and reference examples, the query question is enhanced and generated to obtain multiple candidate SQL statements, including: Determine key instructions and constraints in the query question, and randomly sort the query question, the table structure data, the domain knowledge, the reference examples, the key instructions and the constraints to generate multiple prompt word templates; Generate a corresponding SQL statement according to each prompt word template to obtain multiple candidate SQL statements.
5. The database query method based on large model retrieval enhancement generation according to claim 1 is characterized in that: Extracting the optimal SQL statement from the multiple candidate SQL statements includes: Performing syntax verification on the multiple candidate SQL statements to obtain pre-selected SQL statements that pass the syntax verification; If there are multiple pre-selected SQL statements, perform performance evaluation on the multiple pre-selected SQL statements to obtain a performance evaluation value of each pre-selected SQL statement; The pre-selected SQL statement with the highest performance evaluation value is used as the optimal SQL statement.
6. The database query method based on large model retrieval enhancement generation according to claim 5 is characterized in that: Perform performance evaluation on multiple pre-selected SQL statements to obtain the performance evaluation value of each pre-selected SQL statement, including: Determine the execution time, resource consumption value, and number of scanned rows for each pre-selected SQL statement; Normalizing the execution time, the resource consumption value, and the number of scanned lines respectively to obtain a normalized execution time value, a normalized resource consumption value, and a normalized scanned line number value; The performance evaluation value of each pre-selected SQL statement is calculated based on the normalized value of the execution time, the normalized value of the resource consumption, and the normalized value of the number of scanned rows.
7. The database query method based on large model retrieval enhancement generation according to claim 1 is characterized in that: Extracting the optimal SQL statement from the multiple candidate SQL statements includes: Execute each candidate SQL statement separately to obtain the execution result corresponding to each candidate SQL statement; Determine the accuracy and completeness of the execution results corresponding to each candidate SQL statement respectively; Voting on the execution results corresponding to each of all candidate SQL statements according to the accuracy and completeness; The alternative SQL statement corresponding to the execution result with the most votes after the voting is completed is regarded as the optimal SQL statement.
8. The database query method based on large model retrieval enhancement generation according to claim 1 is characterized in that: Correcting deviations of the optimal SQL statement includes: The optimal SQL statement is respectively corrected for spelling errors, missing keywords, incorrect bracket matching, data type, and logic errors.
9. A database query device based on large model retrieval enhancement generation, characterized in that: include: A retrieval module is used to search the query question based on the pre-built vector knowledge database to obtain table structure data, domain knowledge and reference examples related to the query question; An enhanced generation module is used to enhance and generate the query question based on the table structure data, domain knowledge and reference examples to obtain multiple candidate SQL statements; An execution module, used for extracting an optimal SQL statement from the multiple candidate SQL statements, correcting deviations of the optimal SQL statement and executing it, to obtain an SQL statement execution result; A prompt word construction module, used to extract the summary instruction in the question to be queried, and construct a query prompt word according to the question to be queried, the summary instruction, the optimal SQL statement and the execution result of the SQL statement; The query module is used to input the query prompt word into the language model to obtain the query result output by the language model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the database query method based on large model retrieval enhanced generation as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Knowledge retrieval enhancement generation method and system based on large language model
CN118394890A
Database query method and device, electronic equipment and storage medium
CN118779315A
Data query and visualization method and system based on large language model
CN118820315A
Method, system and equipment for generating SQL (Structured Query Language) statement based on large model
CN119127913A
Text-to-structured query language conversion method based on large language model
CN119415546A
Cited By
Information query method, terminal equipment and computer readable storage medium
CN121009106A