Intelligent number asking method for large model index query

By building an auxiliary information database and a consistent alignment mechanism, the low accuracy rate and insufficient self-improvement of the text-to-SQL method in small sample scenarios are solved, and the efficient generation of accurate SQL queries and self-learning capabilities are achieved, which is suitable for a variety of database environments.

CN120336360AActive Publication Date: 2025-07-18INSPUR SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202510830537.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing text-to-SQL methods have low accuracy in small sample scenarios and lack the ability to check the result and self-improve the results, resulting in the generated SQL statements that may not conform to the database structure or have deviations from user intentions, making it difficult to accurately obtain the required information.

Method used

Build an auxiliary information library, including database structure and historical Q&A examples, generate accurate SQL queries through a low sample prompt and consistency alignment mechanism, and achieve self-learning enhancement through a feedback mechanism to ensure that the generated queries are consistent with user questions and database schema.

Benefits of technology

It significantly improves the accuracy and reliability of text-to-SQL generation of large models, and the system has the ability to learn by itself, adapts to different databases and business scenarios, and reduces the dependence on large-scale training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336360A_ABST
    Figure CN120336360A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent number asking method for large model index query, and belongs to the field of natural language processing and database query. An auxiliary information base is constructed to store a database structure and historical question and answer examples, and key entities and fields are extracted by conducting structured format processing on natural language questions; a large language model guided by a small number of examples is utilized to generate an intermediate expression close to an SQL grammar and a plurality of SQL query candidates, and a consistency alignment mechanism is designed to verify candidate SQL in the aspects of field matching, aggregation functions, data tables and the like so as to screen out an optimal SQL query. Finally, the generated correct SQL and the corresponding question and answer pairs are fed back to the auxiliary information base, and self-learning enhancement is achieved. According to the method, under the condition that only a small number of training examples are needed, the accuracy and expandability of large language model text generation SQL query are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of natural language processing and database query, and particularly to an intelligent question-asking method for querying large model metrics. Background Art

[0002] With the rapid development of database technology and artificial intelligence, users hope to directly query the database through everyday language to obtain the required information. The Text-to-SQL technology has emerged, enabling the computer to convert natural language questions into Structured Query Language (SQL) statements for database retrieval. Existing Text-to-SQL methods usually rely on model training with large-scale labeled data or require model fine-tuning for specific database schemas. However, when large pre-trained language models (LLMs) directly generate SQL queries in the absence of sufficient training data or context information, some problems often occur. For example, the generated SQL statements may not conform to the actual structure of the database, resulting in incorrect query execution; or the generated queries may deviate from the user's intention and fail to accurately obtain the required "metric" information. In addition, many current solutions lack a verification mechanism for the generated SQL results. It is difficult to correct in a timely manner if there are field errors or semantic inconsistencies in the query statements output by the model, reducing the reliability of the system. To solve the above problems, some improvement measures have been proposed in the industry, such as introducing database schema information for schema alignment when generating SQL, or adopting few-shot learning to improve the pertinence of model generation by providing examples in the prompt. However, these improvements are usually used independently, and there is still a lack of a general method that combines few-shot examples, advanced language models, and result consistency verification. At the same time, existing technologies rarely consider enabling the system to self-learn from newly generated correct queries to continuously improve the performance of subsequent query generation. Therefore, it is necessary to propose a new technical solution to comprehensively utilize a small amount of training examples, formatted prompts, and consistency alignment mechanisms, thereby significantly improving the accuracy of large model Text-to-SQL generation and enabling the system to have the ability of continuous self-improvement. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides an intelligent question answering method for large model metric queries, which overcomes the deficiencies of existing text-to-SQL generation methods in small sample scenarios, such as low accuracy, lack of result verification, and self-improving ability. By using a small amount of example data to guide the large language model to generate accurate SQL queries and ensuring that the generated queries match the user's questions and database schema through consistency alignment checks, the accuracy of query generation is significantly improved. At the same time, the present invention also enables the system to use each successfully generated query as a new learning sample through a feedback mechanism, gradually enhancing the adaptability and scalability of the model under different databases and problem domains; in the case of scarce training data, enabling the pre-trained large language model to accurately understand the user's intent and generate effective SQL queries, and continuously improving performance through an automatic verification and continuous learning mechanism.

[0004] The technical solution of the present invention is as follows: An intelligent question answering method for large model metric queries, comprising the following steps: constructing an auxiliary information library containing target database structure information and historical question and answer examples; performing structured formatting processing on the obtained natural language questions to extract the entities, fields, and query condition information involved therein; using the example data in the auxiliary information library to perform few-shot prompting on the large language model to guide it to generate a SQL-like intermediate expression corresponding to the natural language question and multiple SQL query candidate statements; performing consistency alignment verification on the SQL query candidates, and screening out the optimal SQL query that is consistent with the semantics of the natural language question and the database structure; outputting the optimal SQL query and storing it together with the corresponding natural language question in the auxiliary information library as new example data for self-learning enhancement.

[0005] Further, The auxiliary information library contains the schema metadata of the target database, including the names of one or more tables, the field names and types of each table, the relationships between the tables, and example pairs of pre-collected natural language questions and corresponding SQL queries, which are used to provide domain context and reference examples for the large language model.

[0006] Further, The step of performing structured formatting processing on the natural language question includes: using predefined parsing rules or named entity recognition techniques to identify the entity names, constraint values, and target fields involved in the question text input by the user, and determining the corresponding database tables and their relationships to generate a structured query intent representation.

[0007] Further, The steps of using examples in the auxiliary information library to perform few-shot prompting on the large language model to guide it to generate queries include: selecting several natural language questions semantically related to the current problem and their corresponding SQL examples from the auxiliary information library as prompt samples, and inputting them into the context of the large language model, so that the model generates a SQL-like expression corresponding to the natural language question on the basis of referring to the prompt samples, and then generates an abstract syntax tree, and outputs multiple SQL query candidate solutions according to the target database features.

[0008] Furthermore, The consistency alignment check includes performing a matching degree check on the constituent elements of the candidate SQL query: verifying whether the data tables and fields in the candidate SQL correspond to the entities and metrics mentioned in the natural language question, whether the aggregation functions or filtering conditions used in the candidate SQL meet the intention requirements of the user's question, performing a syntax check, and whether the data source targeted by the candidate SQL belongs to the target database scope in the auxiliary information library, so as to screen out query candidates with semantic inconsistencies or non-executability, select the SQL statement that meets all alignment conditions as the final output result, and record the error information in the screened query candidates as counterexamples and store them in the augmented example dataset.

[0009] Furthermore, After outputting the optimal SQL query, store the SQL query and the corresponding natural language question pair in the auxiliary information library to augment the example dataset; the large language model can use the newly added examples for reference during subsequent query generation, so as to continuously optimize the accuracy of query generation and achieve self-learning enhancement of the system.

[0010] Furthermore, The consistency alignment check also includes field alignment: checking whether the tables and fields used in each candidate SQL match the entity and metric requirements mentioned in the natural language question, and whether there are any missing or misused fields Calculate the matching degree of the column set F(Si) used in the candidate query and the question-expected column set F ref The matching degree. The field alignment score F field (Si) can be defined as: , Its value range is [0,1], indicating what proportion of the target columns the candidate SQL correctly contains. If the candidate references columns (wrong columns) not mentioned or non-existent in the question, then F field decreases. At the same time, check whether any key columns are missing.

[0011] Operation alignment: Check whether the operations such as aggregation functions, sorting, and filtering conditions used in the candidate SQL match the user's question meaning Figure 1For example, if the user's question requires statistical summarization, the correct aggregation function should be used in SQL; if the question contains filtering conditions, the WHERE clause of SQL should correctly reflect the conditions. Define the operation alignment score F op (Si) is the similarity between the set of operations to be used as candidates and the set of operations expected by the question. For example, it can be calculated by the ratio of the number of matches to the number of expected operations. If the question requires "statistical sum", but the candidate uses the wrong aggregation function or has no aggregation, then F op will be low.

[0012] Dataset alignment: Check whether the database table (or view) targeted by the SQL query is correct, and ensure that the selected data source is consistent with the data domain involved in the user's question. For example, when the auxiliary information library provides schema information of multiple tables, it is necessary to confirm that SQL extracts data from the correct table or joins the corresponding tables. Define the table alignment score F table (Si), which takes the value of 1 when the table selected by the candidate is consistent with the question domain, and 0 otherwise. For the case involving multi-table joins, it can also be further verified whether the correct JOIN relationship is included.

[0013] The consistency alignment mechanism can perform consistency scoring on each candidate SQL through rule checking and scoring. The above alignment scores can be combined with weights to obtain a comprehensive consistency score, and the SQL with the highest score is selected as the optimal result according to the matching degree. Among them, the weights $w_1, w_2, w_3$ can be set according to actual needs; 。

[0014] The beneficial effects of the present invention are as follows: Few-shot efficient learning: The present invention uses the few-shot learning mechanism to guide the large language model to generate SQL queries. Only a small number of example question-and-answer pairs are required to significantly improve the model's understanding and generation ability for new natural language questions, reducing the dependence on large-scale training data. Such a design enables the system to quickly adapt to new domains or databases and can be put into use with only extremely limited examples.

[0015] Structured prompts enhance understanding: By performing structured formatting on natural language questions and extracting the entities, fields, and condition information involved, the model can obtain a clearer representation of the question than the original unprocessed text. Such formatted prompts help eliminate ambiguity, enabling the large model to more accurately understand the user's intention and thus generate SQL queries that better match the requirements.

[0016] Multi-candidate Generation and Optimization: The method of the present invention does not rely on a single-path output. Instead, the model generates multiple query candidates and selects the optimal ones through a consistency alignment mechanism. Such a diverse generation strategy increases the probability of finding the correct query. At the same time, by means of automatic verification, candidates that do not meet the semantic requirements are screened out, ensuring that the finally output query statement is correct.

[0017] Consistency Alignment Ensures Accuracy: The designed consistency alignment mechanism strictly verifies the generated SQL candidates from multiple aspects including fields, functions, and data sources, and compares them one by one with the user's question and the database schema. This mechanism effectively prevents possible knowledge misuse or omission by the model, ensuring that the output SQL is consistent with the expectation both semantically and grammatically. Therefore, the SQL queries generated by the present invention have a higher accuracy rate and can be directly used to query the database to obtain the required metric data, with a significant improvement in reliability.

[0018] Self-learning for Continuous Improvement: By storing each successfully generated pair of natural language questions and corresponding SQLs in the auxiliary information repository, the knowledge base of the system will be continuously enriched. As time goes by, the number of examples that the model can refer to increases, the coverage of various queries expands, and thus the confidence and accuracy in generating SQLs for new questions are improved. This self-learning enhancement mechanism makes the system become more intelligent with use, gradually reducing the error rate, and having a strong ability of continuous evolution without frequent manual intervention or retraining of the model.

[0019] Generality and Scalability: The method of the present invention has general applicability and does not rely on a specific database or data in a fixed domain. Through the guidance and adaptive expansion of the auxiliary information repository, the system can be flexibly applied to different database environments and business scenarios. Whether it is used for metric analysis queries in business intelligence or database retrieval with a natural language interface, this method can be quickly deployed at a low cost and provide accurate query generation services.

[0020] In summary, by innovatively combining few-shot examples, format guidance, and result alignment verification, the present invention significantly improves the performance of large models in the text-to-SQL task, ensuring both the correctness of query results and the self-evolution of the system, and having great practical value and broad application prospects. Brief Description of the Drawings

[0021] Figure 1 is a schematic diagram of the working process of the present invention. Detailed Embodiments

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0023] The present invention provides an intelligent question-asking method for querying large model metrics. As Figure 1 shown, it demonstrates the complete process from natural language input to SQL query output and self-learning feedback. First, the user's natural language question is input into the system. The parsing module extracts structured information such as entity names, fields, and conditions from it, and then sends it together with relevant examples stored in the auxiliary information library to the large language model for processing. Under the guidance of few-shot examples, the large language model generates multiple candidate SQL-like query statements. Then, through a consistency alignment mechanism, these candidate queries are verified and compared, and the SQL statement that is most consistent with the user's intention and the database schema is selected as the final output result. Finally, the system saves the optimal SQL and its corresponding natural language question pair back to the auxiliary information library as new learning samples to strengthen the subsequent query generation process. The closed-loop design ensures that the system can continuously optimize itself to improve the accuracy of query generation.

[0024] Specifically, it includes the following steps: Construct an auxiliary information library: Establish an auxiliary information library for storing prior knowledge and example data related to the target database. This information library includes, but is not limited to, the structured metadata of the database (such as table names, field names, field types, relationships, etc.) and several historical Q&A examples and their corresponding SQL queries. The auxiliary information library provides the necessary context and reference for the large language model to understand the database schema and typical query methods.

[0025] Structured processing of natural language questions: After obtaining the user's natural language query question, first preprocess and parse the question to convert it into a structured representation. Specifically, using predefined parsing rules or named entity recognition and other technologies, extract key information such as potential entity names, constraint values, target metric fields, and column names in the database from the question text, and identify the operation types involved in the user's question (such as filtering conditions, aggregation types, etc.). Through this formatting process, unstructured natural language is converted into structured information containing the corresponding relationship between entities and fields, which is convenient for subsequent generation processes to utilize.

[0026] Few-shot Example-guided Query Generation: Combine the above-structured information with relevant content in the auxiliary information repository and input it into a large language model to generate a preliminary query expression. Specifically, adopt the few-shot learning method to guide the large model to generate SQL. For example, select several natural language questions and their SQL examples similar to the current question from the auxiliary information repository, add them as prompt examples to the model input to form a Few-shot prompt. Under the guidance of the prompt examples and structured information, the large language model first generates a SQL-like expression (an intermediate representation form) close to the SQL syntax, then generates an abstract syntax tree, and verifies the abstract syntax tree. Then traverse the abstract syntax tree and generate multiple possible SQL query statements as candidate outputs according to the target database features. This step utilizes the powerful natural language generation ability of the large model and enables it to produce syntactically reasonable query candidates even in a small data environment by providing format examples.

[0027] Consistency Alignment Mechanism Screening: For the multiple SQL query candidate results generated above, design and execute a consistency alignment mechanism to select the final query that best matches the user's intention and the actual situation of the database. This mechanism verifies and compares the candidate SQLs from the following aspects: Field Alignment: Check whether the tables and fields used in each candidate SQL match the entities and metric requirements mentioned in the natural language question, whether there are any missing or misused relevant fields, and calculate the matching degree between the set of columns F(Si) used in the candidate query and the set of columns F expected by the question. ref The matching degree can be defined as the field alignment score F field (Si) as: ,

[0028] Its value range is (-1, 1], indicating what proportion of the target columns the candidate SQL correctly contains. And if the candidate references columns not mentioned or non-existent in the question (wrong columns), then F field (Si) decreases. Also, check whether any key columns are missing.

[0029] Operation Alignment: Check whether the operations such as aggregate functions, sorting, and filtering conditions used in the candidate SQL are consistent with the user's question intention. Figure 1 For example, if the user's question requires statistical aggregation, the correct aggregate function should be used in the SQL; if the question contains filtering conditions, the WHERE clause of the SQL should correctly reflect the conditions. Define the operation alignment score F op (Si) as the similarity between the set of operations used by the candidate and the set of operations expected by the question. For example, it can be calculated by the ratio of the number of matches to the number of expected operations. If the question requires "statistical sum", but the candidate uses the wrong aggregate function or no aggregation, then F op (Si) will be lower.

[0030] Dataset Alignment: Check whether the database table (or view) targeted by the SQL query is correct to ensure that the selected data source is consistent with the data domain involved in the user's question. For example, when the auxiliary information repository provides schema information for multiple tables, it is necessary to confirm that the SQL extracts data from the correct table or the joined corresponding tables. Define the table alignment score F table (Si), which takes the value of 1 when the candidate selected table is consistent with the problem domain, and 0 otherwise. For the case involving multi-table joins, it can also be further verified whether the correct JOIN relationship is included.

[0031] The consistency alignment mechanism can perform a consistency score for each candidate SQL through rule verification and scoring. The above alignment scores can be combined with weights to obtain a comprehensive consistency score, and the SQL with the highest score is selected as the optimal result according to the matching degree. Among them, the weights $w_1, w_2, w_3$ can be set according to actual needs; ,

[0032] If there are obvious candidates that do not conform to the consistency, such as referring to non-existent fields or missing key conditions, they will be excluded. Through this alignment screening process, the semantic or syntactic deviations that may be introduced by the large language model are significantly reduced, ensuring that the output SQL query highly matches the input question semantically and can be correctly executed on the actual database.

[0033] Result Feedback and Self-Learning: Store the finally screened correct SQL query and the original natural language question corresponding to this SQL back into the auxiliary information repository as new example data for accumulation. Through this result feedback mechanism, the method of the present invention can continuously enrich the Q&A pair content of the auxiliary information repository. When the system processes more questions, the auxiliary information repository will accumulate more and more high-quality Q&A examples. When the large language model faces similar new questions, it can refer to more instances, thereby realizing self-learning enhancement. This incremental learning does not require retraining the large language model, but improves the overall performance of the system by expanding the example library, ensuring that the method has good scalability and continuous improvement ability.

[0034] In summary, the method of the present invention combines technical means such as few-shot prompting, structured parsing, and result consistency verification, and can effectively generate accurate SQL queries under the condition of limited training data. Its core lies in providing context example guidance for the large model through the auxiliary information repository to generate, and adopting a consistency alignment mechanism to automatically verify the model output, thereby significantly improving the reliability and accuracy of text-to-SQL conversion. The final feedback learning link further ensures that the system will become more intelligent, accumulate experience over time, and continuously optimize its performance.

[0035] The above are only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. An intelligent question-asking method for querying large model metrics, characterized in that: It includes the following steps: Construct an auxiliary information library containing target database structure information and historical Q&A examples; Perform structured formatting on the obtained natural language questions, and extract the entities, fields, and query condition information involved therein; Use the example data in the auxiliary information library to perform few-shot prompting on the large language model, guiding it to generate a SQL-like intermediate expression corresponding to the natural language question and several SQL query candidate statements; Perform consistency alignment verification on the SQL query candidates, and screen out the optimal SQL query that is consistent with the semantics of the natural language question and the database structure; Output the optimal SQL query and store it together with the corresponding natural language question in the auxiliary information library as new example data for self-learning enhancement.

2. The method according to claim 1, characterized in that: The auxiliary information library contains the schema metadata of the target database, including the names of one or more tables, the field names and types of each table, the relationships between tables, and pre-collected examples of natural language questions and corresponding SQL queries, which are used to provide domain context and reference examples for the large language model.

3. The method according to claim 1, characterized in that: When performing structured formatting on the natural language question, after obtaining the natural language query question input by the user, first preprocess and parse the question to convert it into a structured representation form; specifically: use predefined parsing rules or named entity recognition technology to identify the entity names, constraint values, and target fields involved from the question text input by the user, and determine the corresponding database tables and their relationships to generate a structured query intent representation.

4. The method according to claim 1, characterized in that: The steps of using the examples in the auxiliary information library to perform few-shot prompting on the large language model to guide its generation of queries include: Select several natural language questions and their corresponding SQL examples related to the semantics of the current question from the auxiliary information library as prompt samples, and input them into the context of the large language model, so that the model generates a SQL-like expression corresponding to the natural language question on the basis of referring to the prompt samples, and then generates an abstract syntax tree, and outputs several SQL query candidate solutions according to the characteristics of the target database.

5. The method according to claim 4, characterized in that: The consistency alignment verification includes checking the matching degree of the constituent elements of the candidate SQL query, specifically including: verifying whether the data tables and fields in the candidate SQL correspond to the entities and metrics mentioned in the natural language question, whether the aggregation functions or filtering conditions used in the candidate SQL meet the intent requirements of the user's question, performing syntax checking, and whether the data source targeted by the candidate SQL belongs to the scope of the target database in the auxiliary information library, so as to screen out query candidates with inconsistent semantics or non-executable ones, select the SQL statement that meets all alignment conditions as the final output result, and record the error information in the screened query candidates as counterexamples and store them in the extended example dataset.

6. The method according to claim 5, wherein The consistency alignment check also includes field alignment: checking whether the tables and fields used in each candidate SQL match the entity and metric requirements mentioned in the natural language question, and whether relevant fields are omitted or misused to calculate the matching degree between the column set F(Si) used in the candidate query and the expected column set F of the question. ref of the match.

7. The method according to claim 6, wherein Define the field alignment score F field where (Si) is: , Its value range is [0, 1], indicating what proportion of the target columns the candidate SQL correctly contains; if the candidate references columns not mentioned or non-existent in the question, i.e., incorrect columns, then F field decreases; at the same time, check whether any key columns are missing.

8. The method according to claim 5 or 7, wherein The consistency alignment mechanism can perform a consistency score on each candidate SQL by means of rule verification and scoring. The above alignment scores can be combined with weights to obtain a comprehensive consistency score, and the SQL with the highest score is selected as the optimal result according to the matching degree; among them, the weights $w_1, w_2, w_3$ can be set according to actual needs; ; If there are candidates that are significantly inconsistent, they will be eliminated.

9. The method according to claim 1, wherein After outputting the optimal SQL query, the SQL query and the corresponding natural language question pair record are stored in the auxiliary information library to expand the example dataset; the large language model uses the newly added examples for reference during the subsequent query generation process, so as to continuously optimize the accuracy of query generation and achieve self-learning enhancement of the system.

Citation Information

Patent Citations

  • Retrieval method and device based on large language model and medium

    CN118277442A

  • Large model SQL (Structured Query Language) generation method integrating few-sample prompt and multi-choice mechanism

    CN119088818A

  • Database query generation method and system based on multi-selection optimization and storage medium

    CN119690987A

  • Task processing method, model training method and code generating method

    CN119692464A

  • ICL large language model data query generation method and system based on diversity SQL reinforcement

    CN119692466A

Cited By

  • Traffic question and answer method, device and equipment based on large model

    CN120929575A

  • A large model-based traffic question and answer method, device and equipment

    CN120929575B

  • Method, system and device for generating reinsurance contract text based on large language model and medium

    CN121032687A

  • Vehicle structured data query and visualization method based on large model

    CN121412303A

  • A small parameter model intelligent question answering method based on multi-agent cooperation and double-path processing strategy

    CN122654143A