Database Query Method, Device and Equipment Based on Retrieval-Augmented Generation of Large Language Models

By building a method that combines vector knowledge database and large language model, SQL statements are generated and optimized, and the problems of low accuracy and efficiency in traditional query solutions are solved, and more efficient and accurate database queries are achieved.

CN120104762BActive Publication Date: 2025-07-25INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510578192.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-25
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In traditional database query solutions, query accuracy and query efficiency are low, making it difficult to directly obtain satisfactory query results through large language models.

Method used

By constructing a vector knowledge database for search, table structure data, domain knowledge and reference samples are obtained, multiple alternative SQL statements are generated, and the optimal SQL statement is selected through syntax verification, performance evaluation and bias correction, and finally, the query prompt word input is input to a large language model to obtain query results.

Benefits of technology

It improves the accuracy and efficiency of database queries, ensures the accuracy and timeliness of query results, and simplifies the database operation process of non-technical users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104762B_ABST
    Figure CN120104762B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of data processing and relates to a database query method, device and equipment based on retrieval augmented generation of large models. The method includes: retrieving the problem to be queried based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge and reference examples; performing augmented generation on the problem to be queried based on the table structure data, domain knowledge and reference examples to obtain multiple alternative SQL statements; extracting the optimal SQL statement from the multiple alternative SQL statements, correcting the deviation of the optimal SQL statement and executing it to obtain the execution result of the SQL statement; extracting the summary instruction in the problem to be queried, and constructing a query prompt word based on the problem to be queried, the summary instruction, the optimal SQL statement and the execution result of the SQL statement; inputting the query prompt word into a language large model to obtain the query result output by the language large model. It can conveniently obtain a more satisfactory query result, improving the accuracy and query efficiency of the query process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a database query method, device, and equipment based on retrieval-enhanced generation of a large model. Background Art

[0002] In the field of language processing, a large model mainly refers to a language large model, or a large language model (LLM). This model is a deep learning model trained based on a vast amount of text data. It can not only generate natural language text but also deeply understand the meaning of the text and process various natural language tasks, such as text summarization, question answering, and translation.

[0003] In related technologies, traditional query solutions are usually implemented only based on a large model. Specifically, the query result of the problem to be queried can be directly obtained based on the large model. However, in the above query solution, since the problem to be queried is usually not perfect and specific enough, and the performance of the model itself is difficult to reach the optimal, it usually takes multiple question-and-answer processes to obtain a relatively satisfactory query result.

[0004] It can be seen that the traditional query solution has the technical problems of low query accuracy and low query efficiency. Summary of the Invention

[0005] The present invention provides a database query method, device, and equipment based on retrieval-enhanced generation of a large model to solve the defects of low query accuracy and low query efficiency of the traditional query solution.

[0006] On the one hand, the present invention provides a database query method based on retrieval-enhanced generation of a large model, including:

[0007] Retrieving the problem to be queried based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the problem to be queried;

[0008] Enhancing and generating the problem to be queried based on the table structure data, domain knowledge, and reference examples to obtain multiple alternative SQL statements;

[0009] Extracting the optimal SQL statement from the multiple alternative SQL statements, correcting the deviation of the optimal SQL statement and executing it to obtain the execution result of the SQL statement;

[0010] Extracting the summary instruction in the problem to be queried, and constructing a query prompt word based on the problem to be queried, the summary instruction, the optimal SQL statement, and the execution result of the SQL statement;

[0011] Inputting the query prompt word into the language large model to obtain the query result output by the language large model.

[0012] According to the database query method based on large model retrieval enhanced generation provided by the present invention, retrieve the query problem according to the pre-constructed vector knowledge database, and obtain table structure data, domain knowledge, and reference examples related to the query problem, including:

[0013] Vectorize the query problem to obtain a problem vector;

[0014] Retrieve table structure data, domain knowledge, and the first SQL example related to the vector semantics of the problem vector from the pre-constructed vector knowledge database;

[0015] Perform fusion screening according to the first SQL example to obtain a reference example.

[0016] According to the database query method based on large model retrieval enhanced generation provided by the present invention, perform fusion screening according to the first SQL example to obtain a reference example, including:

[0017] Directly generate a corresponding initial SQL statement according to the query problem, and retrieve the second SQL example most similar to the initial SQL statement from the vector knowledge database;

[0018] Replace the table name, column name, and example value in the table structure data with mask tags to obtain general table data;

[0019] Establish a query vector according to the general table data, and retrieve the third SQL example most similar to the query vector from the vector knowledge database;

[0020] Reorder and perform fusion screening on the first SQL example, the second SQL example, and the third SQL example to obtain a reference example.

[0021] According to the database query method based on large model retrieval enhanced generation provided by the present invention, perform enhanced generation on the query problem according to the table structure data, domain knowledge, and reference example, and obtain multiple alternative SQL statements, including:

[0022] Determine the key instructions and constraint conditions in the query problem, and randomly sort the query problem, the table structure data, the domain knowledge, the reference example, the key instructions, and the constraint conditions to generate multiple prompt word templates;

[0023] Generate respective corresponding SQL statements according to each prompt word template to obtain multiple alternative SQL statements.

[0024] According to the database query method based on retrieval-enhanced generation of large models provided by the present invention, extracting the optimal SQL statement from the multiple alternative SQL statements includes:

[0025] Perform syntax verification on the multiple alternative SQL statements to obtain preselected SQL statements that pass the syntax verification;

[0026] If there are multiple preselected SQL statements, perform performance evaluation on the multiple preselected SQL statements to obtain the performance evaluation values of each preselected SQL statement;

[0027] Take the preselected SQL statement with the highest performance evaluation value as the optimal SQL statement.

[0028] According to the database query method based on retrieval-enhanced generation of large models provided by the present invention, performing performance evaluation on the multiple preselected SQL statements to obtain the performance evaluation values of each preselected SQL statement includes:

[0029] Determine the execution duration, resource consumption value, and number of scanned rows of each preselected SQL statement;

[0030] Perform normalization processing on the execution duration, the resource consumption value, and the number of scanned rows respectively to obtain the normalized execution duration value, the normalized resource consumption value, and the normalized number of scanned rows value;

[0031] Calculate the performance evaluation value of each preselected SQL statement based on the normalized execution duration value, the normalized resource consumption value, and the normalized number of scanned rows value.

[0032] According to the database query method based on retrieval-enhanced generation of large models provided by the present invention, extracting the optimal SQL statement from the multiple alternative SQL statements includes:

[0033] Execute each alternative SQL statement separately to obtain the execution result corresponding to each alternative SQL statement;

[0034] Determine the accuracy and completeness of the execution result corresponding to each alternative SQL statement respectively;

[0035] Vote on the execution results corresponding to all alternative SQL statements based on the accuracy and completeness;

[0036] Take the alternative SQL statement corresponding to the execution result with the most votes after the voting as the optimal SQL statement.

[0037] According to the database query method based on retrieval-enhanced generation of large models provided by the present invention, performing deviation correction on the optimal SQL statement includes:

[0038] Perform spelling error correction, missing keyword correction, incorrect parenthesis matching correction, data type correction, and logical error correction on the optimal SQL statement respectively.

[0039] On the other hand, the present invention also provides a database query device based on retrieval augmented generation of a large model, including:

[0040] A retrieval module, configured to retrieve the problem to be queried based on a pre-constructed vector knowledge database, and obtain table structure data, domain knowledge, and reference examples related to the problem to be queried;

[0041] An enhancement generation module, configured to perform enhancement generation on the problem to be queried based on the table structure data, domain knowledge, and reference examples, and obtain multiple alternative SQL statements;

[0042] An execution module, configured to extract the optimal SQL statement from the multiple alternative SQL statements, perform deviation correction on the optimal SQL statement and execute it, and obtain the SQL statement execution result;

[0043] A prompt word construction module, configured to extract the summary instruction in the problem to be queried, and construct a query prompt word based on the problem to be queried, the summary instruction, the optimal SQL statement, and the SQL statement execution result;

[0044] A query module, configured to input the query prompt word into a language large model, and obtain the query result output by the language large model.

[0045] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the database query method based on retrieval augmented generation of a large model as described in any one of the above.

[0046] The database query method, device, and equipment based on retrieval augmented generation of large models provided by the present invention can retrieve the problem to be queried based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the problem to be queried; enhance and generate the problem to be queried based on the table structure data, domain knowledge, and reference examples to obtain multiple alternative SQL statements; extract the optimal SQL statement from the multiple alternative SQL statements, correct the deviation of the optimal SQL statement and execute it to obtain the execution result of the SQL statement; construct a query prompt word based on the problem to be queried, the summary instruction in the problem to be queried, the optimal SQL statement, and the execution result of the SQL statement; input the query prompt word into the language large model to obtain the query result output by the language large model. Since the retrieval augmented generation technology is introduced in the query link, a series of data related to the problem to be queried can be retrieved through the vector knowledge database, and then multiple alternative SQL statements can be obtained through enhanced generation. Subsequently, by extracting the optimal SQL statement and constructing a query prompt word based on the problem to be queried, the summary instruction in the problem to be queried, the optimal SQL statement, and the execution result of the SQL statement, a more accurate query prompt word can be obtained, and then a more satisfactory query result can be conveniently obtained, improving the accuracy and query efficiency of the query link. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 is a schematic flowchart of the database query method based on retrieval augmented generation of large models provided by the embodiments of the present invention;

[0049] Figure 2 is a schematic diagram of the determination principle of the reference example in the embodiments of the present invention;

[0050] Figure 3 is a schematic diagram of the implementation principle of different combinations of six elements using random sorting in the embodiments of the present invention;

[0051] Figure 4 is a schematic structural diagram of the database query device based on retrieval augmented generation of large models provided by the embodiments of the present invention;

[0052] Figure 5 is a schematic structural diagram of the electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the protection scope of the present invention.

[0054] The following will describe in detail the solutions of the database query method, device, and equipment based on retrieval-enhanced generation with a large model provided by the embodiments of the present invention in conjunction with Figures 1 to 5 to describe the details of the database query method, device, and equipment based on retrieval-enhanced generation with a large model provided by the embodiments of the present invention.

[0055] Figure 1 is a schematic flowchart of the database query method based on retrieval-enhanced generation with a large model provided by the embodiments of the present invention.

[0056] As Figure 1 shown, for the database query method based on retrieval-enhanced generation with a large model provided by the embodiments of the present invention, the execution subject can be a computer or server with data transceiver and data processing capabilities. The above method mainly includes the following steps:

[0057] Step 110: Retrieve the problem to be queried based on the pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the problem to be queried.

[0058] In this embodiment, the vector knowledge database is mainly constructed based on relevant database table creation statements, SQL query records accumulated during actual use, and relevant domain knowledge.

[0059] Among them, the database table creation statements provide database table structure information, the SQL query records serve as reference examples, and the relevant domain knowledge serves as prior knowledge to guide the generation of SQL statements that conform to the true intention.

[0060] It should be noted that in this embodiment, L-Schema is used to present the hierarchical structure between databases, tables, and columns. As shown in Table 1 below, it respectively shows the description of a school-related database structure using DDL Schema (Data Definition Language Schema) and L-Schema.

[0061] Table 1 Description forms of DDL Schema and L-Schema

[0062]

[0063] As shown in Table 1 above, in the description statement corresponding to the DDL Schema, the student table is created using the CREATE TABLE statement, which contains fields such as id (integer, auto-incrementing primary key), name (string with a maximum of 50 characters, not allowing null), and age (integer).

[0064] The score table is also created using CREATE TABLE, containing fields such as id (integer, auto-incrementing primary key), student_id (integer, foreign key associated with the student table), subject (string with a maximum of 50 characters, not allowing null), and score (decimal number with a precision of two decimal places and a total length of 5). At the same time, the FOREIGN KEY is used to specify that the student_id field references the student_id of the student table.

[0065] In the description statement corresponding to the L-Schema, the database ID is marked as school. The table structure description of the student table includes id (integer, primary key, example student number such as 23011111, 240134), name (string with a maximum of 50 characters, example name "Li Ming"), age (integer, example age such as 13, 12), and the score table contains id (integer, primary key, example unique score code such as 2105031144), student_id (string, example corresponding student number such as 23011111, 240134), score (decimal number with a precision of two decimal places and a total length of 5, example score such as 96.00). In the foreign key relationship, it is clearly stated that student.id = score.student_id, indicating the association relationship between the two tables.

[0066] It is not difficult to find that the L-Schema proposed in this embodiment presents the hierarchical structure between the database, tables, and columns in a semi-structured form and uses specific tags for identification. Specifically, "[DB]" represents the database, "

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083] Figure 2 Figure 2

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092] Figure 3

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109] norm norm norm

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136] Figure 4

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144] Figure 5

[0145] Figure 5

[0146]

[0147]

[0148]

[0149]

[0150]

[0151] "Table" represents a table, and "[Foreign keys]" represents foreign keys. For each table, the field names and descriptions are provided. The information in the table will be converted into a list, where each item is a tuple representing the detailed information of a column. Each column includes the column name, data type, column description, primary key identifier, and example value. Additionally, due to the importance of foreign keys, the foreign keys need to be listed. Subsequently, the L-Schema description content of the database, SQL examples, and their related domain knowledge can be vector-encoded using the embedded model bge-m3-large and stored in a vector knowledge database. The L-Schema presents the hierarchical structure between the database, tables, and columns in a semi-structured form and uses specific tags for identification, making the database representation more compact and clear, and better presenting the hierarchical structure between the database, tables, and columns. Step 120: Based on the table structure data, domain knowledge, and reference examples, perform enhanced generation on the query problem to obtain multiple alternative SQL statements. It can be understood that the multiple alternative SQL statements can be obtained according to their corresponding prompt templates. Step 130: Extract the optimal SQL statement from the multiple alternative SQL statements, correct the deviation of the optimal SQL statement and execute it to obtain the SQL statement execution result. In this embodiment, the SQL statement execution result obtained through the extraction of the optimal SQL statement, deviation correction, and execution can provide a more effective data basis for the construction of subsequent query prompts. Step 140: Extract the summary instruction in the query problem and construct a query prompt based on the query problem, summary instruction, optimal SQL statement, and SQL statement execution result. It can be understood that since the query prompt is obtained based on multiple data in the construction process, it can express the query purpose more comprehensively and specifically. Step 150: Input the query prompt into the language large model to obtain the query result output by the language large model. The method provided in this embodiment can overcome the problems of limited storage capacity of the large language model, difficulty in obtaining the latest information immediately, and insufficient knowledge in specific fields through retrieval-enhanced generation, assist the model in generating more accurate, detailed, and targeted answers by integrating the retrieval mechanism, improve the accuracy, relevance, and timeliness of the content generated by the large language model by combining the advantages of the large language model and the retrieval system. Compared with relying solely on the generation of the large language model, it can retrieve information from an external knowledge base, avoid the hallucination problem of the model, and improve the processing ability of problems with high real-time requirements, greatly simplifying the database query process, enabling non-technical users to easily interact with the database, improving the efficiency of database operations, promoting the popularization of database technology, and enhancing the convenience and accuracy of the query link. In one embodiment, retrieve the query problem based on the pre-constructed vector knowledge database to obtain the table structure data, domain knowledge, and reference examples related to the query problem, specifically including: First, perform vectorization processing on the query problem to obtain a problem vector.In practical applications, the problem to be queried can be input into the embedding model bge-m3-large for vectorization processing, so as to provide an effective data basis for subsequent retrieval. Then, table structure data, domain knowledge, and the first SQL example related to the vector semantics of the problem vector are retrieved from the pre-constructed vector knowledge database. In this embodiment, the similarity between the problem vector and the vectors in the vector knowledge database can be calculated. Specifically, the cosine similarity or Euclidean distance between the problem vector and the vectors in the vector knowledge database can be calculated to evaluate the similarity between the problem vector and the vectors in the vector knowledge database. Finally, the reference example is obtained through fusion screening based on the first SQL example. The determination principle of the reference example in this embodiment can be referred to. As shown, the reference example is obtained through fusion screening based on the first SQL example, which specifically includes: on the one hand, an initial SQL statement corresponding to the problem to be queried is directly generated, and the second SQL example most similar to the initial SQL statement is retrieved from the vector knowledge database. In practical applications, the initial SQL statement and the SQL examples in the vector knowledge database can be converted into their respective corresponding statement vectors, and then the similarity between the initial SQL statement and the SQL examples in the vector knowledge database is determined by calculating the cosine similarity or Euclidean distance between the two vectors, and then the second SQL example most similar to the initial SQL statement is determined. On the other hand, the table name, column name, and example value in the table structure data are replaced with mask tokens to obtain general table data. Then, a query vector is established based on the general table data, and the third SQL example most similar to the query vector is retrieved from the vector knowledge database. It can be understood that replacing all the table names, column names, and example values in the table structure data of the problem to be queried with mask tokens can eliminate the specific information of the domain. Then, the similarity between the query vector and the data corresponding vectors in the vector knowledge database is calculated using the nearest neighbor algorithm, and then the third SQL example most similar to the query vector is screened out. Finally, the first SQL example, the second SQL example, and the third SQL example are re-ranked and fusion-screened to obtain the reference example. In practical applications, a re-ranking model can be used to fuse and re-rank the first SQL example, the second SQL example, and the third SQL example, and finally one or more reference examples with higher similarity are screened out. In one embodiment, based on the table structure data, domain knowledge, and reference example, the problem to be queried is enhanced and generated to obtain multiple alternative SQL statements, which specifically includes: first, the key instructions and constraint conditions in the problem to be queried are determined, and the problem to be queried, table structure data, domain knowledge, reference example, key instructions, and constraint conditions are randomly sorted to generate multiple prompt templates. It can be understood that the main content in the prompt template in this embodiment is divided into six elements, including: key instructions, table structure data, reference example, constraint conditions, domain knowledge, and the problem to be queried.Considering that the effects of prompt words are affected by the six elements arranged in different orders, different element formats and contents. Therefore, as shown, the present invention uses a random sorting method to combine the six elements in different ways to generate n prompt word templates. Then, respective corresponding SQL statements are generated according to each prompt word template to obtain multiple alternative SQL statements. Subsequently, n alternative SQL statements can be obtained according to the n prompt word templates, and then the optimal SQL statement can be obtained by screening the alternative SQL statements. In one embodiment, extracting the optimal SQL statement from multiple alternative SQL statements specifically includes: First, perform syntax verification on multiple alternative SQL statements to obtain preselected SQL statements that pass the syntax verification. It can be understood that when extracting the optimal SQL statement, it is necessary to ensure that the statement syntax is correct first. The syntax verification tool provided by the database or a third-party library can be used to perform syntax verification on the alternative SQL statements, and the statements with syntax errors are excluded, so as to obtain preselected SQL statements that pass the syntax verification. Then, if there are multiple preselected SQL statements, perform performance evaluation on the multiple preselected SQL statements to obtain the performance evaluation value of each preselected SQL statement. In practical applications, if there is only one preselected SQL statement, the preselected SQL statement that passes the language verification can be directly used as the optimal SQL statement. If there are multiple preselected SQL statements, since performance is the key factor to measure the quality of SQL statements, further screening is required through performance evaluation to determine the optimal SQL statement. Finally, the preselected SQL statement with the highest performance evaluation value is used as the optimal SQL statement. In a specific implementation, performing performance evaluation on multiple preselected SQL statements to obtain the performance evaluation value of each preselected SQL statement specifically includes: First, determine the execution duration, resource consumption value, and number of scanned rows of each preselected SQL statement. In practical applications, each preselected SQL statement can be actually executed and the execution duration can be recorded. The shorter the execution duration, the better the statement performance. The resource consumption value can be determined by indicators such as CPU usage rate and memory occupancy rate, and relevant information can be obtained with the help of the performance monitoring tool of the database. The fewer the number of scanned rows, the better the statement performance usually is, and the number of scanned rows can be viewed through the execution plan of the database. Then, perform normalization processing on the execution duration, resource consumption value, and number of scanned rows respectively to obtain the normalized execution duration value, normalized resource consumption value, and normalized number of scanned rows value. It can be understood that the numerical ranges of the execution duration, resource consumption value, and number of scanned rows may vary greatly. In order to make them comparable in performance evaluation, these indicators need to be normalized, and their values are mapped to the numerical interval [0,1]. In practical applications, the normalization processing can be implemented using the maximum-minimum normalization function. Finally, according to the normalized execution duration value, normalized resource consumption value, and normalized number of scanned rows value, calculate the performance evaluation value of each preselected SQL statement.In this embodiment, the performance evaluation value of the preselected SQL statement can be specifically calculated as follows: where P represents the performance evaluation value of the preselected SQL statement, E represents the normalized execution duration value, R represents the normalized resource consumption value, and S represents the normalized number of scanned rows value. Since the performance evaluation value can comprehensively evaluate the performance of each preselected SQL statement, finally, the preselected SQL statement with the largest performance evaluation value can be used as the preselected SQL statement. It is not difficult to find that in this embodiment, through the method of syntax verification combined with performance evaluation, the screening of the optimal SQL statement can be accurately and effectively achieved. In another embodiment, extracting the optimal SQL statement from multiple alternative SQL statements specifically includes: First step, execute each alternative SQL statement respectively to obtain the execution result corresponding to each alternative SQL statement. In this embodiment, the screening of the optimal SQL statement is implemented by voting. In this process, it is necessary to execute each alternative SQL statement respectively to obtain the corresponding execution result. Second step, determine the accuracy and integrity of the execution result corresponding to each alternative SQL statement respectively. First, set the expected number of rows and expected fields of the execution result respectively, and then determine the actual number of rows and actual fields of the execution result corresponding to each alternative SQL statement respectively. Subtract the actual number of rows from the expected number of rows to obtain the row deviation value, which is used to measure the accuracy of the execution result. Compare all the actual fields in each execution result with the expected fields to determine whether the execution result contains all the expected fields and there are corresponding data values under each expected field. Use this judgment result to measure the integrity of the execution result. Third step, vote on the execution results corresponding to all alternative SQL statements according to the accuracy and integrity. During the voting process, vote once for the execution result whose row deviation value is positive and the value is less than the corresponding preset threshold, and vote once for the execution result whose number of included expected fields is greater than the corresponding preset threshold and there are corresponding data values under each expected field. Fourth step, use the alternative SQL statement corresponding to the execution result with the most votes after the voting as the optimal SQL statement. It is not difficult to find that in this embodiment, through the voting method, the screening of the optimal SQL statement can be conveniently and efficiently achieved. In one embodiment, correcting the deviation of the optimal SQL statement includes: correcting spelling errors, missing keywords, incorrect bracket matching, data type correction, and logical error correction for the optimal SQL statement respectively. In this embodiment, in order to ensure the accuracy of the optimal SQL statement, subsequent correction operations are performed on the spelling errors, missing keywords, incorrect bracket matching, incorrect data types, and logical errors in the optimal SQL statement, thereby further improving the data accuracy. In some embodiments, according to the query problem, summary instruction, optimal SQL statement, and SQL statement execution result, a query prompt is constructed, specifically including: On the one hand, according to the query problem, a beginning introduction statement is generated to clearly and explicitly elaborate on the query problem initially proposed by the user, so that the large model can understand the core and background of the problem.On the other hand, according to the summary instruction, generate an instruction description statement to describe in detail the specific requirements for summarizing and processing the query results, so that the large model can clearly understand the content form that needs to be presented finally. It should be noted that the summary instruction is further extracted based on the key instructions determined previously. On the other hand, according to the optimal SQL statement and the SQL execution result, generate the statement display content and the result presentation form content. In practical applications, appropriate comments can be added to the statement display content to explain its function, helping the large model understand the data acquisition method. At the same time, the result presentation form content displays the SQL statement execution result in a clear format (such as a table), facilitating the large model to summarize based on this form, and clearly informing the large model to process the result according to the summary instruction and output the summary content. Finally, integrate the opening introduction statement, the instruction description statement, and the statement display content and the result presentation form content to obtain the query prompt. Subsequently, input the query prompt into the language large model to make it summarize and reply to obtain the query result. In summary, the database query method based on large model retrieval enhanced generation provided by the embodiments of the present invention has at least the following advantages: First, it can not only understand the natural language query requirements of users, but also retrieve key information from existing knowledge to generate more accurate SQL statements, significantly improving the accuracy and intelligence of SQL generation. Second, a new mode representation method of L-Schema is proposed to present the hierarchical structure between the database, tables, and columns in a semi-structured form, enhancing the understanding of the large language model for the database structure. Third, by integrating multiple example selection strategies to screen similar SQL statements, it helps the LLM learn by analogy and improve the accuracy and rationality of SQL statement generation. Fourth, various permutations and combinations of the six elements constituting the prompt are carried out to generate multiple prompt templates, and SQL statements are generated respectively according to different prompt templates. Through reasonable SQL statement screening strategies and deviation correction methods, the accuracy of SQL statement generation is significantly improved. Based on the same general inventive concept, the present invention also protects a database query device based on large model retrieval enhanced generation. The database query device based on large model retrieval enhanced generation provided below will be described, and the database query device based on large model retrieval enhanced generation described below can be correspondingly referred to the database query method based on large model retrieval enhanced generation described above.As shown, the database query device based on retrieval-augmented generation of large models provided by the embodiments of the present invention specifically includes: a retrieval module 210, configured to retrieve the problem to be queried based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the problem to be queried; an augmentation generation module 220, configured to perform augmentation generation on the problem to be queried based on the table structure data, domain knowledge, and reference examples to obtain multiple alternative SQL statements; an execution module 230, configured to extract the optimal SQL statement from the multiple alternative SQL statements, correct the deviation of the optimal SQL statement and execute it to obtain the execution result of the SQL statement; a prompt word construction module 240, configured to extract the summary instruction in the problem to be queried and construct a query prompt word based on the problem to be queried, the summary instruction, the optimal SQL statement, and the execution result of the SQL statement; a query module 250, configured to input the query prompt word into a language large model to obtain the query result output by the language large model. Thus, it can be seen that for the database query device based on retrieval-augmented generation of large models provided by the embodiments of the present invention, since the retrieval-augmented generation technology is introduced in the query link, a series of data related to the problem to be queried can be retrieved through the vector knowledge database, and then multiple alternative SQL statements can be obtained through augmentation generation. Subsequently, by extracting the optimal SQL statement and constructing a query prompt word based on the problem to be queried, the summary instruction in the problem to be queried, the optimal SQL statement, and the execution result of the SQL statement, a more accurate query prompt word can be obtained, and thus a more satisfactory query result can be obtained more conveniently, improving the accuracy and query efficiency of the query link. Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here. It is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. As shown, the electronic device may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communication interface 320, and the memory 330 complete communication with each other through the communication bus 340.The processor 310 may invoke the logic instructions in the memory 330 to execute the database query method based on large model retrieval enhanced generation provided in the above embodiments. The method includes: retrieving the problem to be queried based on the pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the problem to be queried; enhancing and generating the problem to be queried based on the table structure data, domain knowledge, and reference examples to obtain multiple alternative SQL statements; extracting the optimal SQL statement from the multiple alternative SQL statements, correcting and executing the deviation of the optimal SQL statement to obtain the SQL statement execution result; extracting the summary instruction in the problem to be queried, and constructing a query prompt word based on the problem to be queried, the summary instruction, the optimal SQL statement, and the SQL statement execution result; inputting the query prompt word into the language large model to obtain the query result output by the language large model. In addition, when the logic instructions in the above-mentioned memory 330 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the database query method based on large model retrieval enhanced generation provided in the above embodiments. The method includes: retrieving the problem to be queried based on the pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the problem to be queried; enhancing and generating the problem to be queried based on the table structure data, domain knowledge, and reference examples to obtain multiple alternative SQL statements; extracting the optimal SQL statement from the multiple alternative SQL statements, correcting and executing the deviation of the optimal SQL statement to obtain the SQL statement execution result; extracting the summary instruction in the problem to be queried, and constructing a query prompt word based on the problem to be queried, the summary instruction, the optimal SQL statement, and the SQL statement execution result; inputting the query prompt word into the language large model to obtain the query result output by the language large model.In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the database query method based on large model retrieval enhanced generation provided in the above embodiments. The method includes: retrieving a problem to be queried according to a pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the problem to be queried; enhancing and generating the problem to be queried according to the table structure data, domain knowledge, and reference examples to obtain multiple alternative SQL statements; extracting the optimal SQL statement from the multiple alternative SQL statements, correcting and executing the deviation of the optimal SQL statement to obtain the execution result of the SQL statement; extracting the summary instruction in the problem to be queried, and constructing a query prompt word according to the problem to be queried, the summary instruction, the optimal SQL statement, and the execution result of the SQL statement; inputting the query prompt word into a language large model to obtain the query result output by the language large model. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor. Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments. Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A database query method based on retrieval-augmented generation of large models, characterized in that, Including: Retrieve the query problem based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the query problem; among them, L-Schema is used to present the hierarchical structure between the database, tables, and columns in a semi-structured form. Each table provides a field name and description, and each column includes a column name, data type, column description, primary key identifier, and example value; Determine the key instructions and constraints in the query problem, and randomly sort the query problem, table structure data, domain knowledge, reference examples, key instructions, and constraints to generate multiple prompt templates; generate respective corresponding SQL statements according to each prompt template to obtain multiple alternative SQL statements; Extract the optimal SQL statement from the multiple alternative SQL statements, correct the deviation of the optimal SQL statement and execute it to obtain the execution result of the SQL statement; Extract the summary instruction in the query problem, and construct a query prompt according to the query problem, the summary instruction, the optimal SQL statement, and the execution result of the SQL statement; Input the query prompt into the language large model to obtain the query result output by the language large model.

2. The database query method based on retrieval-augmented generation of large models according to claim 1, wherein Retrieve the query problem based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the query problem, including: Vectorize the query problem to obtain a problem vector; Retrieve table structure data, domain knowledge, and the first SQL example with vector semantics related to the problem vector from the pre-constructed vector knowledge database; Perform fusion screening based on the first SQL example to obtain a reference example.

3. The database query method based on retrieval-enhanced generation of a large model according to claim 2, characterized in that Perform fusion screening based on the first SQL example to obtain a reference example, including: Directly generate a corresponding initial SQL statement according to the query problem, and retrieve the second SQL example most similar to the initial SQL statement from the vector knowledge database; Replace the table name, column name, and example value in the table structure data with mask tags to obtain general table data; Establish a query vector based on the general table data, and retrieve the third SQL example most similar to the query vector from the vector knowledge database; Reorder and perform fusion screening on the first SQL example, the second SQL example, and the third SQL example to obtain a reference example.

4. The database query method based on retrieval-augmented generation of large models according to claim 1, wherein Extract the optimal SQL statement from the multiple alternative SQL statements, including: Perform syntax verification on the multiple alternative SQL statements to obtain preselected SQL statements that pass the syntax verification; If there are multiple preselected SQL statements, perform performance evaluation on the multiple preselected SQL statements to obtain the performance evaluation value of each preselected SQL statement; Use the preselected SQL statement with the highest performance evaluation value as the optimal SQL statement.

5. The database query method based on retrieval-augmented generation of a large model according to claim 4, wherein Perform performance evaluation on the multiple preselected SQL statements to obtain the performance evaluation value of each preselected SQL statement, including: Determine the execution duration, resource consumption value, and number of scanned rows of each preselected SQL statement; Normalize the execution duration, the resource consumption value, and the number of scanned rows respectively to obtain a normalized execution duration value, a normalized resource consumption value, and a normalized number of scanned rows; Calculate the performance evaluation value of each preselected SQL statement based on the normalized execution duration value, the normalized resource consumption value, and the normalized number of scanned rows.

6. The database query method based on retrieval-enhanced generation of large models according to claim 1, wherein Extract the optimal SQL statement from the multiple alternative SQL statements, including: Execute each alternative SQL statement respectively to obtain the execution result corresponding to each alternative SQL statement; Determine the accuracy and integrity of the execution result corresponding to each alternative SQL statement respectively; Vote on the execution results corresponding to all alternative SQL statements based on the accuracy and integrity; Take the alternative SQL statement corresponding to the execution result with the most votes after the voting as the optimal SQL statement.

7. The database query method based on retrieval-augmented generation of large models according to claim 1, wherein, Perform deviation correction on the optimal SQL statement, including: Perform spelling error correction, missing keyword correction, incorrect parenthesis matching correction, data type correction, and logical error correction on the optimal SQL statement respectively.

8. A database query device based on retrieval-augmented generation of large models, characterized in that, Including: A retrieval module for retrieving the query problem based on a pre-constructed vector knowledge database to obtain table structure data, domain knowledge, and reference examples related to the query problem; wherein, L-Schema is used to present the hierarchical structure between the database, tables, and columns in a semi-structured form, each table provides a field name and a description, and each column includes a column name, a data type, a column description, a primary key identifier, and an example value; An enhancement generation module for determining the key instructions and constraint conditions in the query problem, randomly sorting the query problem, table structure data, domain knowledge, reference examples, key instructions, and constraint conditions to generate multiple prompt templates; generating respective corresponding SQL statements according to each prompt template to obtain multiple alternative SQL statements; An execution module for extracting the optimal SQL statement from the multiple alternative SQL statements, performing deviation correction on the optimal SQL statement and executing it to obtain an SQL statement execution result; A prompt construction module for extracting the summary instructions in the query problem and constructing a query prompt according to the query problem, the summary instructions, the optimal SQL statement, and the SQL statement execution result; A query module for inputting the query prompt into a language large model to obtain a query result output by the language large model.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the database query method based on large model retrieval enhancement generation according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method, system and equipment for generating SQL (Structured Query Language) statement based on large model

    CN119127913A