Large model SQL (structured query language) generation method, system and equipment based on business domain semantic enhancement and multi-round voting
Through business domain semantic enhancement and multi-round voting methods, the accuracy and robustness issues of NL2SQL technology in enterprise-level diversified scenarios are solved, and high-precision SQL generation under complex query conditions is achieved.
Patent Information
- Application Number
- CN202510705636.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-03
AI Technical Summary
Existing NL2SQL technology has difficulty accurately identifying the business sub-domains of user queries in diverse enterprise-level scenarios, resulting in field matching errors and low SQL generation accuracy. It also lacks effective a posteriori verification and error correction mechanisms, and is particularly lacking in robustness under multi-table joins and complex conditions.
A large-model SQL generation method based on business domain semantic enhancement and multi-round voting is adopted. Through business domain division, data structure semantic annotation, multi-round candidate generation and voting screening, the business domain of user queries is identified, semantic annotation and expansion are performed, multiple candidate SQL statements are generated, and the final output is selected through consistent voting.
It significantly improves the accuracy of SQL generation and the robustness of the system, and can generate correct SQL statements under complex conditions and multi-table associations, reducing field matching errors and hallucination risks, and improving the stability and availability of the system.
Smart Images

Figure CN120743983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to a method, system and device for generating large-model SQL based on business domain semantic enhancement and multi-round voting. Background Art
[0002] In enterprise data analytics, natural language query-to-sql (NL2SQL) technology significantly lowers the barrier to entry for non-technical users to access relational databases by converting natural language queries into executable SQL statements. This technology has important applications in business intelligence (BI) and data science, enabling enterprises to rapidly analyze and make decisions. However, in real-world enterprise environments, NL2SQL faces numerous challenges. First, traditional rule-based or fixed template-based approaches can only handle queries within fixed domains or table structures, and struggle to account for the multi-ambiguity and diverse descriptions in natural language. Modern enterprise data is often distributed across multiple business domains (such as customer management, loan management, and account management), and databases may be distributed across different platforms, with complex and variable table structures. Second, user query intent varies widely, and wording is flexible and ambiguous, making it difficult for systems to accurately understand their needs. Furthermore, while large language models (such as GPT-like models) demonstrate powerful generation capabilities, end-to-end SQL generation carries the risk of "hallucination" (generating syntactically correct but semantically incorrect SQL), high computational costs and performance bottlenecks, and a lack of transparency in the generation process. The above technical difficulties make it difficult for traditional NL2SQL systems to guarantee query accuracy and availability in diverse enterprise scenarios.
[0003] Currently, the demand for natural language queries in enterprise-level data analysis scenarios (such as banking) is growing. However, existing NL2SQL technology has significant shortcomings in handling such queries: First, it cannot accurately identify the specific business sub-domains (such as loans, account opening, and transactions) associated with user questions, resulting in unclear table and field selection and a lack of focus on the correct business domain. Second, its understanding of domain terms and synonyms in natural language is imperfect, leading to discrepancies in field matching. For example, it is difficult for the system to accurately map "account opening customer" to the account opening date field in the customer table, or "loan amount" to the loan amount field. As a result, the generated SQL may omit key fields or introduce erroneous fields. Finally, traditional NL2SQL methods typically use a single-pass generation strategy, whose output relies on a one-time generation result and lacks effective a posteriori verification and error correction mechanisms. This results in low generation accuracy and insufficient robustness when dealing with complex conditions and multi-table joins. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to design a large-model SQL generation method, system and device based on business domain semantic enhancement and multi-round voting. Through business domain division, data structure semantic annotation, multi-round candidate generation and voting screening and other technologies, the SQL generation accuracy and system robustness of natural language queries are significantly improved. The method and system of the present invention can accurately identify query intentions and domain constraints in business scenarios of enterprises such as banks, avoid field matching errors, and ensure the correctness and stability of the output SQL statements through multi-round iterative generation and consistency voting screening, thereby solving existing technical problems.
[0005] To solve the above technical problems, the present invention provides a large model SQL generation method based on business domain semantic enhancement and multi-round voting, comprising the following steps: Step S1: Obtain the natural language input by the user, identify and determine the business field involved in the user query.
[0006] Step S2: semantically annotate and expand the input query, and output an enhanced query intent representation and related table field mapping results.
[0007] Step S3: The user's natural language question, the domain information, and the matched key table field information are organized into a prompt input large language model to trigger it to generate the corresponding SQL statement.
[0008] Step S4: For the same natural language question, generate multiple candidate SQL statements and select the SQL statement with the highest frequency as the output.
[0009] Step S5: Submit the final output SQL statement to the actual database system for execution, return the query results to the user, and display them through the front end.
[0010] Furthermore, in the present invention, in step S1, a set of related tables is selected from the table structure of the enterprise database in combination with business rules or classification models to limit the query scope and provide domain context for subsequent query.
[0011] Furthermore, in the present invention, in step S2, a semantic enhancement engine is called to semantically annotate and expand the input query using a pre-established business domain knowledge base or ontology.
[0012] Furthermore, in the present invention, in step S2, the semantic annotation and expansion include synonym replacement, entity standardization and concept extraction of business domain terms.
[0013] Furthermore, in step S3 of the present invention, the large language model includes GPT-4, fine-tuned BERT and Deepseek.
[0014] Furthermore, in the present invention, in step S3, a domain-focused prompt design is adopted, which only includes table structures and examples related to the current domain, thereby avoiding interference from information in other domains.
[0015] Furthermore, in step S4 of the present invention, multiple candidate SQL statements are obtained by configuring multiple different large language models or repeating reasoning, and then the frequency of the corresponding query results being consistent after the execution of these SQL statements is counted.
[0016] The present invention also provides a large model SQL generation system based on business domain semantic enhancement and multi-round voting, which specifically includes the following modules: The business domain identification module is used to automatically determine the business domain involved in the user query.
[0017] The semantic enhancement module is used to semantically annotate and expand the input query, and the output includes the enhanced query intent representation and related table field mapping results.
[0018] The NL2SQL conversion module is used to organize the user's natural language questions, domain information, and matched key table field information into a Prompt input large language model, triggering it to generate the corresponding SQL statement.
[0019] The voting module is used to call the NL2SQL conversion module multiple times for the same natural language question to generate candidate SQL statements to obtain multiple SQL outputs, and then count the frequency of consistent corresponding query results after these SQLs are executed. Finally, the SQL with the most occurrences is regarded as the best output for the question.
[0020] The SQL execution module is used to submit the final SQL statement output by the voting module to the actual database system for execution and return the query results to the user.
[0021] Furthermore, in the present invention, the SQL execution module utilizes pre-configured database connection information (including SQL dialect, link address, etc.) to ensure that different database sources can be connected and portable queries can be achieved in a multi-data source environment.
[0022] The present invention further provides an electronic device, comprising: at least one processor; and at least one memory in communication with the processor; The memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the electronic device to execute the aforementioned large model SQL generation method based on business domain semantic enhancement and multi-round voting.
[0023] Compared with the prior art, the method and system for generating large model SQL based on business domain semantic enhancement and multi-round voting of the present invention has the following beneficial effects: (1) Strong cross-table structure generalization capability: The SQL generation method of the present invention utilizes business domain semantic annotation to automatically identify the relevant tables and fields required for user queries. This allows for flexible association of data with different table structures in a multi-table, multi-business domain environment. Therefore, even if the database table structure is complex or variable, the system can accurately match the required fields, thereby enhancing the generalization capability for multi-table association queries.
[0024] (2) High reusability and low maintenance costs: The banking business knowledge base and mapping rules built into the semantic enhancement engine can be reused in a variety of query scenarios. When introducing new business domains, only the domain dictionary and mapping rules need to be expanded, without the need to redesign or train the entire model, significantly reducing system expansion and maintenance costs. Compared with traditional methods that require a lot of manual writing and updating of rules, this invention is easier to maintain.
[0025] (3) More stable output: The use of a multi-round generation and voting control mechanism greatly improves the consistency of SQL output. By cross-validating and voting multiple candidate statements, the risk of accidental errors or logical omissions in a single generation is reduced, thereby ensuring the semantic correctness and stability of the final SQL. This is superior to traditional methods that only use a single generation, which is easily affected by differences in input representation and produces unstable output.
[0026] (4) Strong adaptability to complex queries: Through multiple rounds of iterative generation and consensus voting, the present invention can handle more complex query requirements, including multi-condition filtering, complex aggregation, and nested queries. These functions are usually difficult to implement or have low accuracy in traditional template or single-round generation systems, but the present invention can stably output correct SQL.
[0027] (5) Significantly improved accuracy and robustness: By combining domain recognition and semantic enhancement, the method of the present invention can more accurately understand query intent, avoiding field matching errors caused by improper handling of synonyms and professional terms in traditional methods. In complex query scenarios with multi-table associations, the system of the present invention relies on additional semantic information and verification mechanisms to maintain high accuracy, and has greater tolerance for natural language errors and robustness to diverse expressions. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The specific embodiments of the present invention will be further explained below with reference to the accompanying drawings.
[0029] Figure 1 This is a flowchart of the large model SQL generation method based on business domain semantic enhancement and multi-round voting of the present invention.
[0030] Figure 2This is a block diagram of the large model SQL generation system based on business domain semantic enhancement and multi-round voting of the present invention. DETAILED DESCRIPTION Example 1
[0031] Combine Figure 1 The large model SQL generation method based on business domain semantic enhancement and multi-round voting in this embodiment specifically includes the following steps: Step S1: Obtain the natural language input by the user, identify and determine the business areas involved in the user's query, such as customer management, loan management, account management, etc. in banking business.
[0032] In this embodiment, preferably, in step S1, a set of related tables is selected from the table structure of the enterprise database in combination with business rules or classification models to limit the query scope and provide domain context.
[0033] Step S2: semantically annotate and expand the input query, and output an enhanced query intent representation and related table field mapping results.
[0034] In this embodiment, preferably, in step S2, a semantic enhancement engine is called to semantically annotate and expand the input query using a pre-established business domain knowledge base or ontology.
[0035] In this embodiment, preferably, in step S2, the semantic annotation and expansion includes synonym replacement, entity standardization, and concept extraction of business domain terms. For example, industry terms in the user's language are mapped to corresponding field names in the database, and time expressions are converted to specific time ranges, thereby enriching the input information and reducing inconsistencies with the database structure.
[0036] Step S3: The user's natural language question, the domain information, and the matched key table field information are organized into a prompt input large language model to trigger it to generate the corresponding SQL statement.
[0037] Preferably, in this embodiment, in step S3, the large language model includes GPT-4, fine-tuned BERT, and Deepseek.
[0038] Preferably, in this embodiment, in order to improve the accuracy in various scenarios, in step S3, a domain-focused prompt design is adopted, which only includes table structures and examples related to the current domain, thereby avoiding interference from information in other domains.
[0039] Step S4: For the same natural language question, generate multiple candidate SQL statements and select the SQL statement with the highest frequency as the output.
[0040] In this embodiment, in step S4, multiple candidate SQL statements are obtained by configuring multiple large language models or repeating inference. The frequency with which these SQL statements produce consistent query results after execution is then counted. The SQL statement with the most occurrences is ultimately selected as the optimal output for the problem. This "self-consistency" voting method can improve the consistency and accuracy of model output and further reduce the hallucination rate of large models.
[0041] Step S5: Submit the final output SQL statement to the actual database system for execution, return the query results to the user, and display them through the front end.
[0042] Specifically, in this embodiment, loan management in banking business is taken as an example for demonstration. The process in this embodiment is only illustrative, and the steps and modules used can be equivalently replaced according to specific needs.
[0043] Step S1: Natural language input and domain identification. The user enters a question through the system interface: "What was the total loan amount for customer Zhang San last year?" The system calls the business domain identification module to determine that the question involves the "loan management" domain and selects relevant tables such as customer_info and loan_account as candidate tables.
[0044] Step S2: Semantic enhancement and structural annotation. The semantic enhancement engine is called and the system identifies: Customer name: Zhang San (converted into a unique identifier customer_id: C1234 through the customer information system); Time range: Last year → 2024-01-01 to 2024-12-31; Field mapping: "Loan amount": SUM(loan_amount).
[0045] Finally, the following structured semantic fragments are formed: { "domain": "Loan Management"; "entities": {"customer_id": "C1234"}; "fields": ["loan_amount"]; "conditions": ["loan_date BETWEEN '2024-01-01' AND '2024-12-31'"]; } Step S3: Prompt construction and SQL generation. The system converts the above structured content into a prompt template, combines it with the table structure information of the target database, and inputs it into a large language model (such as GPT-4). The output candidate SQL is as follows: SELECT SUM(loan_amount) FROM loan_account WHERE customer_id = 'C1234' AND loan_date BETWEEN '2024-01-01' AND '2024-12-31'.
[0046] Step S4: A multi-round generation and voting mechanism uses different models or repeated reasoning to generate 10 candidate SQL statements. After executing them, the system obtains the return value. If seven of them return the same result, "2,000,000 yuan," the system identifies this as the final consistent result and selects the SQL statement with the highest frequency as the output.
[0047] Step S5: SQL execution and result return. The final SQL statement is submitted to the back-end database system. The system returns the query result "¥2,000,000" and displays it through the front-end.
[0048] The method of the embodiment of the present invention improves the accuracy and fault tolerance of natural language questions through structured semantic enhancement and a self-consistent voting mechanism, and is particularly suitable for industry scenarios such as banking and insurance with rich structured data and complex business semantics. Example 2
[0049] Combine Figure 2 The large model SQL generation system based on business domain semantic enhancement and multi-round voting in this embodiment specifically includes the following modules: The business domain identification module is used to automatically determine the business domain involved in the user query.
[0050] Specifically, in this embodiment, the business domain identification module automatically determines the business domain (e.g., customer management, loan management, account management, etc. in banking) associated with the user's query based on the natural language input. The business domain identification module, combined with business rules or classification models, selects a set of relevant tables from the enterprise database's table structure to subsequently limit the query scope and provide domain context.
[0051] The business domain identification module adopts a query context limitation method based on business domain identification. By identifying the business domain information implicit in natural language queries (such as customer management, loan management, etc.), and combining it with the enterprise database table structure to automatically screen the relevant table set, it is used to provide context constraints and table field range limitations for subsequent NL2SQL conversion.
[0052] The semantic enhancement module is used to semantically annotate and expand the input query, and the output includes the enhanced query intent representation and related table field mapping results.
[0053] Specifically, in this embodiment, the semantic enhancement module uses a pre-established business domain knowledge base or ontology to semantically annotate and expand the input query. This includes synonym replacement, entity standardization, concept extraction, and the like for business domain terms. For example, industry terms in the user's language are mapped to corresponding field names in the database, and time expressions are converted into specific time ranges, thereby enriching the input information and reducing inconsistencies with the database structure. The output of the semantic enhancement engine includes the enhanced query intent representation and the mapping results of related table fields.
[0054] The semantic enhancement module adopts a semantic enhancement and entity standardization mechanism that integrates domain knowledge, and uses a predefined business domain knowledge base to perform semantic enhancement on natural language queries, including term synonym replacement, concept extraction, time expression parsing, entity normalization, etc., to improve the mapping accuracy and consistency of natural language to database fields.
[0055] The NL2SQL conversion module is used to organize the user's natural language questions, domain information, and matched key table field information into a Prompt input large language model, triggering it to generate the corresponding SQL statement.
[0056] Specifically, in this embodiment, SQL queries are generated based on large language models (such as GPT-4, fine-tuned BERT, Deepseek, etc.). The system organizes the user's natural language question, domain information, and matched key table field information into a prompt input to the large model, triggering the generation of the corresponding SQL statement. To improve accuracy in various scenarios, a domain-focused prompt design can be adopted, for example, only including table structures and examples related to the current domain, thereby avoiding interference from information in other domains.
[0057] The voting module is used to call the NL2SQL conversion module multiple times for the same natural language question to generate candidate SQL statements to obtain multiple SQL outputs, and then count the frequency of consistent corresponding query results after these SQLs are executed. Finally, the SQL with the most occurrences is regarded as the best output for the question.
[0058] Specifically, in this embodiment, the NL2SQL conversion model is invoked multiple times to generate candidate SQL statements for the same natural language question. The system configures multiple large language models for parallel execution, generating multiple SQL outputs. The system then calculates the frequency with which these SQL statements produce consistent results for the query. Ultimately, the SQL statement with the most occurrences is selected as the optimal output for the question. This "self-consistent" voting method improves the consistency and accuracy of model outputs, further reducing the likelihood of hallucinations from large models.
[0059] The voting module uses a multi-round SQL generation and voting mechanism based on consistent execution results. It uses a large language model to generate candidate SQL statements multiple times and verify the self-consistency of their execution results. It then votes and screens through statistical result consistency to select the SQL output with the most accurate semantics and the most stable execution, which is used to improve the robustness and accuracy of complex query tasks.
[0060] The SQL execution module is used to submit the final SQL statement output by the voting module to the actual database system for execution and return the query results to the user.
[0061] In this embodiment, preferably, the SQL execution module uses pre-configured database connection information, which includes SQL dialect, link address, etc., to ensure that different database sources can be connected and portable queries can be achieved in a multi-data source environment. Example 3
[0062] This embodiment provides an electronic device, including: at least one processor; and at least one memory in communication with the processor; The memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the electronic device to execute the large model SQL generation method based on business domain semantic enhancement and multi-round voting in Example 1.
[0063] The present invention integrates business domain semantic annotation, semantic enhancement engine, multi-round generation and voting mechanism to build a high-performance NL2SQL generation system. Business domain semantic annotation is used to extract and mark domain information and terms in user queries, providing context constraints for subsequent processing; the semantic enhancement engine further uses the domain knowledge base to expand and standardize the extracted information (such as synonym mapping, entity recognition, time expression parsing, etc.), improving the model's understanding of domain terms; the multi-round generation and voting mechanism allows the large model to generate multiple candidate SQLs under different prompts, and then screens the multiple candidate results through execution result consistency verification and feedback, and finally outputs the SQL statement that best meets the query intent. The close cooperation of the above modules enables the system to accurately understand complex business queries and generate accurate and reliable SQL results.
[0064] In the above description, many specific details are set forth in order to fully understand the present invention. However, the above description is only a preferred embodiment of the present invention. The present invention can be implemented in many other ways different from those described herein, so the present invention is not limited to the specific implementation disclosed above. At the same time, any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention using the methods and technical contents disclosed above without departing from the scope of the technical solution of the present invention, or modify it into an equivalent embodiment of equivalent changes. Any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still falls within the scope of protection of the technical solution of the present invention.
Claims
1. A large model SQL generation method based on business domain semantic enhancement and multi-round voting, characterized by: The steps include: Step S1: Obtain the natural language input by the user, identify and determine the business field involved in the user query; Step S2: semantically annotate and expand the input query, and output an enhanced query intent representation and related table field mapping results; Step S3: The user's natural language question, the domain information, and the matched key table field information are organized into a prompt input large language model to trigger it to generate the corresponding SQL statement; Step S4: For the same natural language question, generate multiple candidate SQL statements and select the SQL statement with the highest frequency as output; Step S5: Submit the final output SQL statement to the actual database system for execution, return the query results to the user, and display them through the front end.
2. The method for generating large-model SQL based on business domain semantic enhancement and multi-round voting according to claim 1 is characterized by: In step S1, a set of related tables is selected from the table structure of the enterprise database in combination with business rules or classification models to limit the query scope and provide domain context.
3. The method for generating large-model SQL based on business domain semantic enhancement and multi-round voting according to claim 1 is characterized by: In step S2, a semantic enhancement engine is called to semantically annotate and expand the input query using a pre-established business domain knowledge base or ontology.
4. The method for generating large model SQL based on business domain semantic enhancement and multi-round voting according to claim 3 is characterized by: In step S2, the semantic annotation and expansion include synonym replacement, entity standardization and concept extraction of business domain terms.
5. The method for generating large-model SQL based on business domain semantic enhancement and multi-round voting according to claim 1 is characterized by: In step S3, the large language model includes GPT-4, fine-tuned BERT and Deepseek.
6. The method for generating large model SQL based on business domain semantic enhancement and multi-round voting according to claim 1 is characterized by: In step S3, a domain-focused prompt design is adopted, which only includes table structures and examples related to the current domain, thereby avoiding interference from information in other domains.
7. The method for generating large-model SQL based on business domain semantic enhancement and multi-round voting according to claim 1 is characterized by: In step S4, multiple candidate SQL statements are obtained by configuring multiple different large language models or repeated reasoning, and then the frequency of consistent corresponding query results after these SQL statements are executed is counted.
8. A large model SQL generation system based on business domain semantic enhancement and multi-round voting, characterized by: Includes the following modules: Business domain identification module, used to automatically determine the business domain involved in the user query; The semantic enhancement module is used to semantically annotate and expand the input query, and the output includes the enhanced query intent representation and related table field mapping results; The NL2SQL conversion module is used to organize the user's natural language questions, domain information, and matched key table field information into a prompt input large language model, triggering it to generate the corresponding SQL statement; The voting module is used to call the NL2SQL conversion module multiple times for the same natural language question to generate candidate SQL statements to obtain multiple SQL outputs. It then counts the frequency of consistent query results after these SQL statements are executed, and ultimately selects the SQL statement with the most occurrences as the optimal output for the question. The SQL execution module is used to submit the final SQL statement output by the voting module to the actual database system for execution and return the query results to the user.
9. The large model SQL generation system based on business domain semantic enhancement and multi-round voting according to claim 8 is characterized by: The SQL execution module uses pre-configured database connection information to ensure that it can connect to different database sources and realize portable queries in a multi-data source environment.
10. An electronic device, characterized in that: include: at least one processor; as well as at least one memory in communication with the processor; The memory stores instructions that can be executed by the processor, and the instructions are executed by the processor to enable the electronic device to execute the large model SQL generation method based on business domain semantic enhancement and multi-round voting according to any one of claims 1 to 7.
Citation Information
Cited By
NL2SQL generation method based on large language model
CN120910089A
Database table retrieval method and device, storage medium, product and electronic equipment
CN121166696A
Multi-round dialogue system and method based on conversion from natural language to SQL
CN121455988A
Multi-modal system fusing dynamic semantic arrangement and cross-platform agent collaborative reasoning
CN121683858A