SQL generation method and system combining GraphRAG and large model

By combining GraphRAG and large model technologies, the problems of understanding user intent, integrating business knowledge, handling redundant fields, and optimizing SQL fragments in NL2SQL technology have been solved, achieving efficient and accurate SQL generation, lowering the technical threshold, and improving the real-time performance and security of the system.

CN120849446APending Publication Date: 2025-10-28FUJIAN NEWLAND SOFTWARE ENGINEERING CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510743386.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing NL2SQL technology has shortcomings in understanding user query intent, integrating business knowledge, handling redundant fields, optimizing SQL fragments, and improving real-time performance, resulting in low accuracy and efficiency of the generated SQL statements.

Method used

By combining GraphRAG and a large model, the system identifies user intent using TextCNN, matches similarity metrics using a text slider similarity algorithm, queries related table fields using GraphRAG technology, generates query suggestions using preset prompt word templates, and finally inputs the data into the large model to generate SQL statements.

Benefits of technology

It improves the accuracy and efficiency of SQL generation, lowers the technical threshold for data analysis, supports real-time requirements, improves the accuracy and generation efficiency of complex queries, and enhances the security and maintainability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849446A_ABST
    Figure CN120849446A_ABST
Patent Text Reader

Abstract

The invention provides an SQL generation method and system combining GraphRAG and a large model in the technical field of natural language processing and artificial intelligence crossing, and the method comprises the steps: S1, obtaining an input natural language query statement, and recognizing the user intention of the natural language query statement through a TextCNN text classification model; s2, on the basis of the user intention, matching similar indexes of the natural language query statement through a text sliding block similarity algorithm; s3, querying associated table fields from a preset data knowledge vector library on the basis of the similar indexes through a GraphRAG technology, and filtering each table field; s4, through a preset cue word template, generating a query cue word based on the natural language query statement and the filtered table field; and S5, inputting the query prompt word into a large model to obtain an SQL statement corresponding to the natural language query statement. The method has the advantage that the SQL generation accuracy and efficiency are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing and artificial intelligence, and in particular to a method and system for generating SQL that combines GraphRAG and large models. Background Technology

[0002] With the rapid development of information technology and AI big data models, enterprises and organizations have accumulated massive amounts of structured data, making relational databases the core of data storage and management. However, SQL, as the primary means of database querying, presents a high barrier to entry for non-technical users, especially the extremely long SQL statements that incorporate complex business logic, which are highly challenging for them to understand and write. Non-technical users often struggle to master SQL syntax and cannot intuitively understand the relationships between tables in a database, leading to inefficient data analysis and information extraction, heavy reliance on technical personnel, and impacting the real-time nature and flexibility of data-driven decision-making.

[0003] To address this issue, Natural Language to SQL (NL2SQL) technology emerged, aiming to allow users to directly query databases using natural language, significantly lowering the barrier to data access and improving data utilization efficiency. Traditional NL2SQL technologies mainly include rule-based methods, template-based methods, and deep learning-based end-to-end methods. However, traditional NL2SQL technologies face the following challenges:

[0004] 1. Insufficient semantic understanding: In practical applications, users' query intentions are often extremely diverse and complex; natural language expression has a high degree of flexibility and ambiguity, and the same query target can be expressed in many different ways; for example, "find the product with the highest sales last year" and "find the best-selling product in 2022" are essentially the same requirement, but the expressions are completely different; while the traditional NL2SQL model has obvious shortcomings in understanding the user's true intention, identifying the query target, and handling ambiguity, especially when it involves complex scenarios such as business terminology, contextual dependencies, and compound conditions, it is easy to have misunderstandings, resulting in the generated SQL statement not matching the user's needs.

[0005] 2. Disconnect between table structure and business knowledge: SQL generation relies solely on database table structure information (such as table names, field names, and field types), neglecting the integration of business knowledge, indicator definitions, and data definitions. In actual business scenarios, user queries are often based on business indicators, management definitions, and industry terminology, which are not directly reflected in the database table structure. For example, indicators such as "number of active users" and "gross profit margin" may involve multiple tables, complex calculation logic, and business rules. NL2SQL systems lacking business knowledge struggle to accurately understand and map user business needs, resulting in generated SQL statements that often fail to meet actual business scenarios, severely impacting practical value and user experience.

[0006] 3. Redundant field interference: In large databases or multi-table join query scenarios, a single data table may contain dozens or even hundreds of fields. If all field information is input into the large model indiscriminately, it will greatly increase the input noise of the large model, making it difficult for the large model to focus on the core fields related to the user's problem. Redundant fields not only increase the consumption of computing resources, but also interfere with the judgment of the large model and reduce the accuracy of SQL generation.

[0007] 4. Insufficient SQL Fragment Reuse and Optimization: In complex business scenarios, SQL statements are often composed of multiple business fragments (such as subqueries, aggregations, filtering, sorting, etc.). Traditional NL2SQL technologies mostly adopt an end-to-end generation approach, lacking structured management and reuse mechanisms for SQL fragments. This results in generated SQL statements that are difficult to reuse existing business logic fragments, leading to low efficiency. At the same time, there is a lack of SQL fragment-level optimization capabilities, making it difficult to automatically adjust the SQL structure to improve query performance. For example, common business queries such as "group statistics," "year-on-year and month-on-month comparisons," and "multi-table joins" could benefit from fragment-level vectorized retrieval and automatic concatenation, which would not only improve generation efficiency but also ensure the correctness and maintainability of the SQL.

[0008] 5. Poor real-time performance: In big data and real-time analysis scenarios, users have increasingly higher requirements for the real-time performance of query results. Traditional NL2SQL systems often suffer from response latency and performance bottlenecks when processing large-scale data and complex queries. Improving real-time response capabilities is key to ensuring user experience.

[0009] Therefore, how to provide a SQL generation method and system that combines GraphRAG and large models to improve the accuracy and efficiency of SQL generation has become an urgent technical problem to be solved. Summary of the Invention

[0010] The technical problem to be solved by this invention is to provide a SQL generation method and system that combines GraphRAG and large models to improve the accuracy and efficiency of SQL generation.

[0011] In a first aspect, the present invention provides a SQL generation method combining GraphRAG and large models, comprising the following steps:

[0012] Step S1: Obtain the input natural language query statement and identify the user intent of the natural language query statement through the TextCNN text classification model;

[0013] Step S2: Based on the user intent, match the similarity index of the natural language query statement using a text slider similarity algorithm;

[0014] Step S3: Using GraphRAG technology, query the associated table fields from the preset data knowledge vector base based on the similarity index, and filter each of the table fields;

[0015] Step S4: Generate query suggestions based on the natural language query statement and the filtered table fields using a preset suggestion template;

[0016] Step S5: Input the query suggestion words into the large model to obtain the SQL statement corresponding to the natural language query statement.

[0017] Furthermore, step S1 specifically includes:

[0018] The input natural language query statement is obtained through a visual interface. The natural language query statement is preprocessed, including at least irrelevant character deletion, format standardization, word segmentation, and stop word removal. The user intent of the natural language query statement is identified through a pre-trained TextCNN text classification model. The user intent includes at least SQL generation, SQL annotation, and data table lookup.

[0019] Furthermore, step S2 specifically includes:

[0020] Based on the user intent, a text slider similarity algorithm is used to match similarity indicators of the natural language query from a preset indicator library. The matching process is as follows:

[0021] Define a first text slider and a second text slider, wherein the length of the first text slider is longer than that of the second text slider;

[0022] The first text slider iterates through the characters of the natural language query statement to obtain a long string. The long string is then iterated through using a similarity algorithm to calculate the first similarity score between the long string and the indicators in the preset indicator library. Based on the first similarity score, candidate indicators are selected from the indicator library for first-level matching.

[0023] The second text slider iterates through the characters of the natural language query statement to obtain a short string. A similarity algorithm is used to iterate through the short string and calculate the second similarity score between it and each candidate indicator. Based on the second similarity score, similar indicators are selected from each candidate indicator for secondary matching.

[0024] Furthermore, step S3 specifically includes:

[0025] Using GraphRAG technology, entities associated with the similarity index are retrieved from a preset business knowledge graph. Based on the entities, the corresponding data tables are queried. Using RAG technology, SQL fragments similar to the data tables are queried. Based on the SQL fragments, the associated table fields are queried from a preset data knowledge vector library. The table fields are then filtered based on preset filtering rules or filtering conditions.

[0026] Furthermore, step S4 specifically includes:

[0027] Using a preset prompt word template that includes Question, Business Knowledge, Examples, Requirements, and Table Information, query prompt words are generated based on the natural language query statement and the filtered table fields.

[0028] The Question field stores natural language query statements; the Business Knowledge field stores business terminology knowledge; the Examples field stores examples of SQL fragments; the Requirements field stores SQL format requirements; and the Table Information field stores table fields.

[0029] Step S5 specifically involves:

[0030] The query suggestions are input into the pre-trained Qwen large model to obtain the SQL statement corresponding to the natural language query. Based on the input adjustment instructions, the SQL statement is optimized by at least adjusting the JOIN order, removing redundant conditions, and merging subqueries.

[0031] Secondly, this invention provides a SQL generation system that combines GraphRAG and large models, comprising the following modules:

[0032] The user intent recognition module is used to acquire the input natural language query statement and identify the user intent of the natural language query statement through the TextCNN text classification model.

[0033] The similarity index matching module is used to match the similarity index of the natural language query statement based on the user intent using a text slider similarity algorithm;

[0034] The table field query and filtering module is used to query related table fields from a preset data knowledge vector base based on the similarity index using GraphRAG technology, and to filter each of the table fields.

[0035] The query suggestion generation module is used to generate query suggestion words based on the natural language query statement and the filtered table fields using a preset suggestion word template;

[0036] The SQL statement generation module is used to input the query suggestion words into the large model and obtain the SQL statement corresponding to the natural language query statement.

[0037] Furthermore, the user intent recognition module is specifically used for:

[0038] The input natural language query statement is obtained through a visual interface. The natural language query statement is preprocessed, including at least irrelevant character deletion, format standardization, word segmentation, and stop word removal. The user intent of the natural language query statement is identified through a pre-trained TextCNN text classification model. The user intent includes at least SQL generation, SQL annotation, and data table lookup.

[0039] Furthermore, the similarity index matching module is specifically used for:

[0040] Based on the user intent, a text slider similarity algorithm is used to match similarity indicators of the natural language query from a preset indicator library. The matching process is as follows:

[0041] Define a first text slider and a second text slider, wherein the length of the first text slider is longer than that of the second text slider;

[0042] The first text slider iterates through the characters of the natural language query statement to obtain a long string. The long string is then iterated through using a similarity algorithm to calculate the first similarity score between the long string and the indicators in the preset indicator library. Based on the first similarity score, candidate indicators are selected from the indicator library for first-level matching.

[0043] The second text slider iterates through the characters of the natural language query statement to obtain a short string. A similarity algorithm is used to iterate through the short string and calculate the second similarity score between it and each candidate indicator. Based on the second similarity score, similar indicators are selected from each candidate indicator for secondary matching.

[0044] Furthermore, the table field query filtering module is specifically used for:

[0045] Using GraphRAG technology, entities associated with the similarity index are retrieved from a preset business knowledge graph. Based on the entities, the corresponding data tables are queried. Using RAG technology, SQL fragments similar to the data tables are queried. Based on the SQL fragments, the associated table fields are queried from a preset data knowledge vector library. The table fields are then filtered based on preset filtering rules or filtering conditions.

[0046] Furthermore, the query suggestion term generation module is specifically used for:

[0047] Using a preset prompt word template that includes Question, Business Knowledge, Examples, Requirements, and Table Information, query prompt words are generated based on the natural language query statement and the filtered table fields.

[0048] The Question field stores natural language query statements; the Business Knowledge field stores business terminology knowledge; the Examples field stores examples of SQL fragments; the Requirements field stores SQL format requirements; and the Table Information field stores table fields.

[0049] The SQL statement generation module is specifically used for:

[0050] The query suggestions are input into the pre-trained Qwen large model to obtain the SQL statement corresponding to the natural language query. Based on the input adjustment instructions, the SQL statement is optimized by at least adjusting the JOIN order, removing redundant conditions, and merging subqueries.

[0051] The advantages of this invention are:

[0052] 1. By acquiring the input natural language query, the TextCNN text classification model is used to identify the user intent of the natural language query. Then, based on the user intent, a text slider similarity algorithm is used to match the similarity index of the natural language query. Using GraphRAG technology, related table fields are queried from a pre-defined data knowledge vector library based on the similarity index, and each table field is filtered. Next, using a pre-defined prompt word template, query prompt words are generated based on the natural language query and the filtered table fields. These prompt words are then input into a large model to obtain the corresponding SQL statement. In other words, the Qwen large model effectively improves semantic understanding capabilities. By combining GraphRAG technology and business knowledge graphs, the disconnect between table structure and business knowledge is avoided. Filtering each table field overcomes the problem of redundant field interference. By incorporating SQL fragments during the table field query process, SQL fragments are managed and reused in a structured manner, overcoming the traditional problem of insufficient SQL fragment reuse and optimization. Through identifying user intent, matching similarity indexes, filtering table fields, and using prompt word templates, response capabilities are effectively improved, ultimately greatly enhancing the accuracy and efficiency of SQL generation.

[0053] 2. By integrating the semantic understanding capabilities of the Qwen large model, the utilization of structured business knowledge graphs by GraphRAG, and the accurate modeling of user intent using prompt word templates, it can more accurately transform natural language problems into corresponding SQL statements. Compared with traditional rule-based or shallow learning methods, it exhibits higher accuracy when facing scenarios such as complex syntax, nested queries, and cross-table joins.

[0054] 3. Through natural language interaction, even non-technical personnel without SQL writing skills can easily complete data query operations. Users only need to ask questions in everyday language (natural language query statements), and the system can automatically parse their user intent and generate corresponding SQL statements, greatly reducing the technical threshold for data analysis.

[0055] 4. A dynamic Prompt generation mechanism is adopted, which combines predefined prompt templates and automatically generates the most suitable prompts for the large model based on the user's natural language query, historical dialogue records, and current task objectives. This not only improves the accuracy of generated SQL statements but also enhances generalization ability.

[0056] 5. By automatically identifying the database tables and their fields that may be involved when parsing user queries (natural language query statements), the search space is reduced, invalid calculations are decreased, and the speed of SQL statement generation is accelerated. At the same time, it supports the caching and reuse strategy of SQL fragments. For common query patterns or frequently used subquery structures, they are abstracted into reusable SQL fragments and directly called in subsequent queries to avoid repeated parsing and generation, thereby improving overall performance.

[0057] 6. Through the built-in SQL optimization engine, after generating the original SQL statement, further optimizations can be made such as adjusting the JOIN order, removing redundant conditions, merging subqueries, and suggesting the use of indexes, thereby improving execution efficiency and resource utilization.

[0058] 7. The TextCNN text classification model is used to initially screen user intent (such as SQL generation, annotation, table lookup), which reduces the complexity of subsequent processing. Compared with traditional single-model processing, staged intent recognition can improve processing efficiency and reduce interference from irrelevant information. The preprocessing process (removing irrelevant characters, word segmentation, and stop word removal) improves the standardization of model input and reduces the interference of noise on classification results.

[0059] 8. The system uses a first text slider (long text) for coarse-grained filtering and a second text slider (short text) for fine-grained matching, which balances the comprehensiveness and accuracy of the matching. Compared with the traditional single sliding window, it can handle fuzzy query needs in complex natural language more efficiently. Through the two-level matching mechanism, the similarity score threshold can be adaptively adjusted to avoid mismatch problems caused by a single threshold.

[0060] 9. GraphRAG technology is used to retrieve entities associated with similarity metrics from the knowledge graph, overcoming the limitations of traditional RAG which relies solely on vector similarity for retrieval; through relational reasoning in the graph structure, implicitly related table fields (such as foreign key relationships and business logic dependencies) can be mined, improving the accuracy of field retrieval; by retrieving similar SQL fragments using RAG technology and combining them with entity associations in the knowledge graph, candidate fields that conform to business logic can be generated, reducing the risk of generating irrelevant fields in large models.

[0061] 10. Filter candidate fields using preset rules / conditions (such as access control and data timeliness) to avoid generating redundant or invalid SQL; for example, exclude deprecated table fields or sensitive data fields to enhance system security.

[0062] 11. By setting prompt word templates to include multi-dimensional information such as Question, Business Knowledge, Examples, Requirements, and Table Information, context injection guides the large model to generate SQL statements that conform to business specifications.

[0063] 12. By combining TextCNN (convolutional neural network) with GraphRAG (graph retrieval augmentation generation), the complementarity of text semantic understanding and structured knowledge reasoning is achieved; for example, TextCNN processes fuzzy expressions in natural language, while GraphRAG supplements business logic constraints, forming an end-to-end semantic-to-SQL mapping.

[0064] 13. By constraining the generation range of large models through external knowledge bases (data knowledge vector library, business knowledge graph), the "illusion" problem caused by the lack of domain knowledge in large models is solved; at the same time, the knowledge base can be updated independently without retraining large models, which improves the maintainability of the system.

[0065] 14. By using TextCNN to quickly filter irrelevant intents (such as SQL comment generation) during the preprocessing stage, the entire computational burden is avoided in the subsequent large model inference stage, significantly reducing the overall system resource consumption; by reducing the number of candidate indicators through the first text slider (coarse screening) and the second text slider (fine screening) to be calculated only within a limited candidate set, the computational efficiency is effectively improved compared to traversing the entire indicator library.

[0066] 15. By deeply integrating TextCNN intent recognition, two-level text slider similarity matching, GraphRAG knowledge graph retrieval, and large model generation capabilities, it achieves efficient and accurate conversion of natural language to SQL statements. Its multi-stage collaborative mechanism (intent filtering → dynamic slider matching → graph relationship reasoning → prompt word constraints) significantly improves the accuracy of semantic parsing, solving pain points such as difficulty in adapting fuzzy queries and lack of cross-table association logic in traditional solutions. Combining field filtering and SQL optimization rules from the business knowledge graph, it ensures that the generated results meet data specifications and security requirements. At the same time, it supports rapid cross-industry migration and incremental updates through modular design, reducing development costs. With the help of the generalization capabilities of large models, it covers long-tail complex query scenarios, providing an efficient, secure, and scalable intelligent data analysis solution. Attached Figure Description

[0067] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0068] Figure 1 This is a flowchart of a SQL generation method combining GraphRAG and large models according to the present invention.

[0069] Figure 2 This is a schematic diagram of the structure of an SQL generation system that combines GraphRAG and large models according to the present invention.

[0070] Figure 3 This is a flowchart illustrating the present invention. Detailed Implementation

[0071] The overall approach of the technical solution in this application is as follows: It effectively enhances semantic understanding capabilities through the Qwen big data model; it avoids the disconnect between table structure and business knowledge by combining GraphRAG technology and business knowledge graphs; it overcomes the problem of redundant field interference by filtering each table field; it overcomes the traditional problem of insufficient SQL fragment reuse and optimization by combining SQL fragments during table field queries and performing structured management and reuse of SQL fragments; and it effectively improves response capabilities by recognizing user intent, matching similar indicators, filtering table fields, and using prompt word templates, thereby improving the accuracy and efficiency of SQL generation.

[0072] Please refer to Figures 1 to 3 As shown, a preferred embodiment of the present invention, which combines GraphRAG and large models for SQL generation, includes the following steps:

[0073] Step S1: Obtain the input natural language query statement and identify the user intent of the natural language query statement through the TextCNN text classification model;

[0074] By using natural language interaction, even non-technical personnel who do not have SQL writing skills can easily complete data query operations. Users only need to ask questions in everyday language (natural language query statements), and the system can automatically parse their user intent and generate corresponding SQL statements, which greatly reduces the technical threshold for data analysis.

[0075] Step S2: Based on the user intent, match the similarity index of the natural language query statement using a text slider similarity algorithm;

[0076] Step S3: Using GraphRAG technology, query the associated table fields from the preset data knowledge vector base based on the similarity index, and filter each of the table fields;

[0077] GraphRAG combines knowledge graphs with vector retrieval technology, enabling intelligent association between users' natural language business descriptions and structured information such as tables, fields, and metrics in the database, thus solving the problem of the disconnect between table structure and business knowledge in the traditional NL2SQL method.

[0078] Step S4: Generate query prompts based on the natural language query statement and the filtered table fields using a preset prompt template;

[0079] Step S5: Input the query suggestion words into the large model to obtain the SQL statement corresponding to the natural language query statement.

[0080] Step S1 specifically involves:

[0081] The system obtains input natural language query statements (e.g., "the number of broadband users among male users in Quanzhou at the end of January by county and city") through a visual interface. The system performs preprocessing on the natural language query statements, including at least the removal of irrelevant characters, format standardization, word segmentation, and removal of stop words. The system then identifies the user intent of the natural language query statements through a pre-trained TextCNN text classification model. The user intent includes at least SQL generation, SQL annotation, and data table lookup.

[0082] By using the TextCNN text classification model to initially filter user intent (such as SQL generation, annotation, and table lookup), the complexity of subsequent processing is reduced. Compared with traditional single-model processing, staged intent recognition can improve processing efficiency and reduce interference from irrelevant information. The preprocessing process (removing irrelevant characters, word segmentation, and stop word removal) improves the standardization of model input and reduces noise interference on classification results.

[0083] SQL generation refers to users' desire for SQL statements to be automatically generated based on natural language; SQL commenting refers to users' input of an SQL statement and their expectation for explanation, annotation, or optimization suggestions; data table lookup refers to users' desire to know which data table a certain business indicator or field belongs to, or to understand the structure information of a data table.

[0084] Step S2 specifically involves:

[0085] Based on the user intent, a text slider similarity algorithm is used to match similar indicators (such as "user broadband usage") from a preset indicator library to the natural language query statement. The matching process is as follows:

[0086] Define a first text slider (sliding window) and a second text slider, wherein the length of the first text slider is longer than that of the second text slider;

[0087] The first text slider iterates through the characters of the natural language query statement to obtain a long string. The long string is then iterated through using a similarity algorithm to calculate the first similarity score between the long string and the indicators in the preset indicator library. Based on the first similarity score, candidate indicators are selected from the indicator library for first-level matching.

[0088] The text slider algorithm iterates through the characters of a natural language query to obtain a short string. A similarity algorithm is then used to calculate the second similarity score between this short string and each candidate indicator. Based on this second similarity score, similar indicators are selected from the candidate indicators for secondary matching. This text slider similarity algorithm is suitable for handling string matching problems with partial overlap, spelling errors, or local variations, and can effectively find the optimal matching substring and its similarity score.

[0089] By using a first text slider (long text) for coarse-grained filtering and a second text slider (short text) for fine-grained matching, the system balances comprehensiveness and accuracy in matching. Compared to a traditional single sliding window, it can more efficiently handle fuzzy query needs in complex natural language. Through the two-level matching mechanism, the similarity score threshold can be adaptively adjusted to avoid mismatches caused by a single threshold.

[0090] By using TextCNN to quickly filter irrelevant intents (such as SQL comment generation) during the preprocessing stage, the computational burden is avoided in the subsequent large model inference stage, significantly reducing the overall system resource consumption. The first text slider (coarse screening) reduces the number of candidate indicators, and the second text slider (fine screening) is calculated only within a limited candidate set, which effectively improves computational efficiency compared to traversing the entire indicator library.

[0091] Step S3 specifically involves:

[0092] Using GraphRAG technology, entities associated with the similarity indicators are retrieved from a pre-defined business knowledge graph. Based on these entities, corresponding data tables are queried. RAG technology is then used to query SQL fragments similar to these data tables. These SQL fragments are then used to query associated table fields from a pre-defined data knowledge vector library. Finally, each table field is filtered based on pre-defined filtering rules or conditions. Specifically, semantic similarity algorithms can be used to filter the table fields. The SQL fragments are business-fragmented SQL statements written by developers. A large model automatically extracts SQL fragments with independent business meaning and stores them in vector form for easy RAG retrieval and reuse. The business knowledge graph abstracts tables, fields, indicators, and business logic from the database into nodes and edges in the graph.

[0093] By automatically identifying the database tables and their fields that may be involved when parsing user queries (natural language queries), the search space is narrowed, unnecessary calculations are reduced, and the speed of SQL statement generation is accelerated. At the same time, it supports caching and reuse strategies for SQL fragments. For common query patterns or frequently used subquery structures, they are abstracted into reusable SQL fragments that can be directly called in subsequent queries, avoiding repeated parsing and generation, thereby improving overall performance.

[0094] By leveraging GraphRAG technology to retrieve entities associated with similarity metrics from a knowledge graph, the limitations of traditional RAG retrieval, which relies solely on vector similarity, are overcome. Through relational reasoning within the graph structure, implicitly related table fields (such as foreign key relationships and business logic dependencies) can be mined, improving the accuracy of field retrieval. By retrieving similar SQL fragments using RAG technology and combining them with entity associations in the knowledge graph, candidate fields that conform to business logic can be generated, reducing the risk of generating irrelevant fields in large models.

[0095] Candidate fields can be filtered by pre-defined rules / conditions (such as access control and data timeliness) to avoid generating redundant or invalid SQL; for example, deprecated table fields or sensitive data fields can be excluded to enhance system security.

[0096] By combining TextCNN (convolutional neural network) with GraphRAG (graph retrieval augmented generation), a complementary relationship between text semantic understanding and structured knowledge reasoning is achieved. For example, TextCNN handles fuzzy expressions in natural language, while GraphRAG supplements business logic constraints, forming an end-to-end semantic-to-SQL mapping.

[0097] By constraining the generation scope of large models through external knowledge bases (data knowledge vector library, business knowledge graph), the "illusion" problem caused by the lack of domain knowledge in large models is solved; at the same time, the knowledge base can be updated independently without retraining large models, thus improving system maintainability.

[0098] Step S4 specifically involves:

[0099] Using a preset prompt word template that includes Question, Business Knowledge, Examples, Requirements, and Table Information, query prompt words are generated based on the natural language query statement and the filtered table fields; that is, the corresponding query prompt words are dynamically assembled based on the actual data according to the prompt word template.

[0100] The Question field stores natural language query statements; the Business Knowledge field stores business terminology knowledge; the Examples field stores examples of SQL fragments; the Requirements field stores SQL format requirements; and the Table Information field stores table fields.

[0101] By adopting a dynamic Prompt generation mechanism, combined with predefined prompt templates, the system automatically generates the most suitable prompts for the large model based on the user's natural language query, historical dialogue records, and current task objectives. This not only improves the accuracy of generated SQL statements but also enhances generalization capabilities.

[0102] By setting prompt word templates that include multi-dimensional information such as Question, Business Knowledge, Examples, Requirements, and Table Information, context injection guides the large model to generate SQL statements that conform to business specifications.

[0103] Step S5 specifically involves:

[0104] The query suggestions are input into the pre-trained Qwen large model to obtain the SQL statement corresponding to the natural language query. Based on the input adjustment instructions, the built-in SQL optimization engine optimizes the SQL statement by at least adjusting the JOIN order, removing redundant conditions, and merging subqueries.

[0105] By integrating the semantic understanding capabilities of the Qwen large model, the utilization of structured business knowledge graphs by GraphRAG, and the accurate modeling of user intent using prompt word templates, it can more accurately transform natural language problems into corresponding SQL statements. Compared with traditional rule-based or shallow learning methods, it exhibits higher accuracy when facing scenarios such as complex syntax, nested queries, and cross-table joins.

[0106] With the built-in SQL optimization engine, after generating the original SQL statement, further optimizations can be made, such as adjusting the JOIN order, removing redundant conditions, merging subqueries, and suggesting the use of indexes, thereby improving execution efficiency and resource utilization.

[0107] By deeply integrating TextCNN intent recognition, two-level text slider similarity matching, GraphRAG knowledge graph retrieval, and large model generation capabilities, it achieves efficient and accurate conversion of natural language to SQL statements. Its multi-stage collaborative mechanism (intent filtering → dynamic slider matching → graph relationship reasoning → prompt word constraints) significantly improves the accuracy of semantic parsing, solving pain points such as difficulty in adapting fuzzy queries and lack of cross-table association logic in traditional solutions. Combining field filtering and SQL optimization rules from the business knowledge graph, it ensures that the generated results meet data specifications and security requirements. At the same time, its modular design supports rapid cross-industry migration and incremental updates, reducing development costs. Furthermore, it leverages the generalization capabilities of large models to cover long-tail complex query scenarios, providing an efficient, secure, and scalable intelligent data analysis solution.

[0108] A preferred embodiment of the present invention, which combines GraphRAG and large models, includes the following modules:

[0109] The user intent recognition module is used to acquire the input natural language query statement and identify the user intent of the natural language query statement through the TextCNN text classification model.

[0110] By using natural language interaction, even non-technical personnel who do not have SQL writing skills can easily complete data query operations. Users only need to ask questions in everyday language (natural language query statements), and the system can automatically parse their user intent and generate corresponding SQL statements, which greatly reduces the technical threshold for data analysis.

[0111] The similarity index matching module is used to match the similarity index of the natural language query statement based on the user intent using a text slider similarity algorithm;

[0112] The table field query and filtering module is used to query related table fields from a preset data knowledge vector base based on the similarity index using GraphRAG technology, and to filter each of the table fields.

[0113] GraphRAG combines knowledge graphs with vector retrieval technology, enabling intelligent association between users' natural language business descriptions and structured information such as tables, fields, and metrics in the database, thus solving the problem of the disconnect between table structure and business knowledge in the traditional NL2SQL method.

[0114] The query suggestion generation module is used to generate query suggestions based on the natural language query statement and the filtered table fields using a preset suggestion template.

[0115] The SQL statement generation module is used to input the query suggestion words into the large model and obtain the SQL statement corresponding to the natural language query statement.

[0116] The user intent recognition module is specifically used for:

[0117] The system obtains input natural language query statements (e.g., "the number of broadband users among male users in Quanzhou at the end of January by county and city") through a visual interface. The system performs preprocessing on the natural language query statements, including at least the removal of irrelevant characters, format standardization, word segmentation, and removal of stop words. The system then identifies the user intent of the natural language query statements through a pre-trained TextCNN text classification model. The user intent includes at least SQL generation, SQL annotation, and data table lookup.

[0118] By using the TextCNN text classification model to initially filter user intent (such as SQL generation, annotation, and table lookup), the complexity of subsequent processing is reduced. Compared with traditional single-model processing, staged intent recognition can improve processing efficiency and reduce interference from irrelevant information. The preprocessing process (removing irrelevant characters, word segmentation, and stop word removal) improves the standardization of model input and reduces noise interference on classification results.

[0119] SQL generation refers to users' desire for SQL statements to be automatically generated based on natural language; SQL commenting refers to users' input of an SQL statement and their expectation for explanation, annotation, or optimization suggestions; data table lookup refers to users' desire to know which data table a certain business indicator or field belongs to, or to understand the structure information of a data table.

[0120] The similarity index matching module is specifically used for:

[0121] Based on the user intent, a text slider similarity algorithm is used to match similar indicators (such as "user broadband usage") from a preset indicator library to the natural language query statement. The matching process is as follows:

[0122] Define a first text slider (sliding window) and a second text slider, wherein the length of the first text slider is longer than that of the second text slider;

[0123] The first text slider iterates through the characters of the natural language query statement to obtain a long string. The long string is then iterated through using a similarity algorithm to calculate the first similarity score between the long string and the indicators in the preset indicator library. Based on the first similarity score, candidate indicators are selected from the indicator library for first-level matching.

[0124] The text slider algorithm iterates through the characters of a natural language query to obtain a short string. A similarity algorithm is then used to calculate the second similarity score between this short string and each candidate indicator. Based on this second similarity score, similar indicators are selected from the candidate indicators for secondary matching. This text slider similarity algorithm is suitable for handling string matching problems with partial overlap, spelling errors, or local variations, and can effectively find the optimal matching substring and its similarity score.

[0125] By using a first text slider (long text) for coarse-grained filtering and a second text slider (short text) for fine-grained matching, the system balances comprehensiveness and accuracy in matching. Compared to a traditional single sliding window, it can more efficiently handle fuzzy query needs in complex natural language. Through the two-level matching mechanism, the similarity score threshold can be adaptively adjusted to avoid mismatches caused by a single threshold.

[0126] By using TextCNN to quickly filter irrelevant intents (such as SQL comment generation) during the preprocessing stage, the computational burden is avoided in the subsequent large model inference stage, significantly reducing the overall system resource consumption. The first text slider (coarse screening) reduces the number of candidate indicators, and the second text slider (fine screening) is calculated only within a limited candidate set, which effectively improves computational efficiency compared to traversing the entire indicator library.

[0127] The table field query and filtering module is specifically used for:

[0128] Using GraphRAG technology, entities associated with the similarity indicators are retrieved from a pre-defined business knowledge graph. Based on these entities, corresponding data tables are queried. RAG technology is then used to query SQL fragments similar to these data tables. These SQL fragments are then used to query associated table fields from a pre-defined data knowledge vector library. Finally, each table field is filtered based on pre-defined filtering rules or conditions. Specifically, semantic similarity algorithms can be used to filter the table fields. The SQL fragments are business-fragmented SQL statements written by developers. A large model automatically extracts SQL fragments with independent business meaning and stores them in vector form for easy RAG retrieval and reuse. The business knowledge graph abstracts tables, fields, indicators, and business logic from the database into nodes and edges in the graph.

[0129] By automatically identifying the database tables and their fields that may be involved when parsing user queries (natural language queries), the search space is narrowed, unnecessary calculations are reduced, and the speed of SQL statement generation is accelerated. At the same time, it supports caching and reuse strategies for SQL fragments. For common query patterns or frequently used subquery structures, they are abstracted into reusable SQL fragments that can be directly called in subsequent queries, avoiding repeated parsing and generation, thereby improving overall performance.

[0130] By leveraging GraphRAG technology to retrieve entities associated with similarity metrics from a knowledge graph, the limitations of traditional RAG retrieval, which relies solely on vector similarity, are overcome. Through relational reasoning within the graph structure, implicitly related table fields (such as foreign key relationships and business logic dependencies) can be mined, improving the accuracy of field retrieval. By retrieving similar SQL fragments using RAG technology and combining them with entity associations in the knowledge graph, candidate fields that conform to business logic can be generated, reducing the risk of generating irrelevant fields in large models.

[0131] Candidate fields can be filtered by pre-defined rules / conditions (such as access control and data timeliness) to avoid generating redundant or invalid SQL; for example, deprecated table fields or sensitive data fields can be excluded to enhance system security.

[0132] By combining TextCNN (convolutional neural network) with GraphRAG (graph retrieval augmented generation), a complementary relationship between text semantic understanding and structured knowledge reasoning is achieved. For example, TextCNN handles fuzzy expressions in natural language, while GraphRAG supplements business logic constraints, forming an end-to-end semantic-to-SQL mapping.

[0133] By constraining the generation scope of large models through external knowledge bases (data knowledge vector library, business knowledge graph), the "illusion" problem caused by the lack of domain knowledge in large models is solved; at the same time, the knowledge base can be updated independently without retraining large models, thus improving system maintainability.

[0134] The query suggestion term generation module is specifically used for:

[0135] Using a preset prompt word template that includes Question, Business Knowledge, Examples, Requirements, and Table Information, query prompt words are generated based on the natural language query statement and the filtered table fields; that is, the corresponding query prompt words are dynamically assembled based on the actual data according to the prompt word template.

[0136] The Question field stores natural language query statements; the Business Knowledge field stores business terminology knowledge; the Examples field stores examples of SQL fragments; the Requirements field stores SQL format requirements; and the Table Information field stores table fields.

[0137] By adopting a dynamic Prompt generation mechanism, combined with predefined prompt templates, the system automatically generates the most suitable prompts for the large model based on the user's natural language query, historical dialogue records, and current task objectives. This not only improves the accuracy of generated SQL statements but also enhances generalization capabilities.

[0138] By setting prompt word templates that include multi-dimensional information such as Question, Business Knowledge, Examples, Requirements, and Table Information, context injection guides the large model to generate SQL statements that conform to business specifications.

[0139] The SQL statement generation module is specifically used for:

[0140] The query suggestions are input into the pre-trained Qwen large model to obtain the SQL statement corresponding to the natural language query. Based on the input adjustment instructions, the built-in SQL optimization engine optimizes the SQL statement by at least adjusting the JOIN order, removing redundant conditions, and merging subqueries.

[0141] By integrating the semantic understanding capabilities of the Qwen large model, the utilization of structured business knowledge graphs by GraphRAG, and the accurate modeling of user intent using prompt word templates, it can more accurately transform natural language problems into corresponding SQL statements. Compared with traditional rule-based or shallow learning methods, it exhibits higher accuracy when facing scenarios such as complex syntax, nested queries, and cross-table joins.

[0142] With the built-in SQL optimization engine, after generating the original SQL statement, further optimizations can be made, such as adjusting the JOIN order, removing redundant conditions, merging subqueries, and suggesting the use of indexes, thereby improving execution efficiency and resource utilization.

[0143] By deeply integrating TextCNN intent recognition, two-level text slider similarity matching, GraphRAG knowledge graph retrieval, and large model generation capabilities, it achieves efficient and accurate conversion of natural language to SQL statements. Its multi-stage collaborative mechanism (intent filtering → dynamic slider matching → graph relationship reasoning → prompt word constraints) significantly improves the accuracy of semantic parsing, solving pain points such as difficulty in adapting fuzzy queries and lack of cross-table association logic in traditional solutions. Combining field filtering and SQL optimization rules from the business knowledge graph, it ensures that the generated results meet data specifications and security requirements. At the same time, its modular design supports rapid cross-industry migration and incremental updates, reducing development costs. Furthermore, it leverages the generalization capabilities of large models to cover long-tail complex query scenarios, providing an efficient, secure, and scalable intelligent data analysis solution.

[0144] In summary, the advantages of this invention are:

[0145] 1. By acquiring the input natural language query, the TextCNN text classification model is used to identify the user intent of the natural language query. Then, based on the user intent, a text slider similarity algorithm is used to match the similarity index of the natural language query. Using GraphRAG technology, related table fields are queried from a pre-defined data knowledge vector library based on the similarity index, and each table field is filtered. Next, using a pre-defined prompt word template, query prompt words are generated based on the natural language query and the filtered table fields. These prompt words are then input into a large model to obtain the corresponding SQL statement. In other words, the Qwen large model effectively improves semantic understanding capabilities. By combining GraphRAG technology and business knowledge graphs, the disconnect between table structure and business knowledge is avoided. Filtering each table field overcomes the problem of redundant field interference. By incorporating SQL fragments during the table field query process, SQL fragments are managed and reused in a structured manner, overcoming the traditional problem of insufficient SQL fragment reuse and optimization. Through identifying user intent, matching similarity indexes, filtering table fields, and using prompt word templates, response capabilities are effectively improved, ultimately greatly enhancing the accuracy and efficiency of SQL generation.

[0146] 2. By integrating the semantic understanding capabilities of the Qwen large model, the utilization of structured business knowledge graphs by GraphRAG, and the accurate modeling of user intent using prompt word templates, it can more accurately transform natural language problems into corresponding SQL statements. Compared with traditional rule-based or shallow learning methods, it exhibits higher accuracy when facing scenarios such as complex syntax, nested queries, and cross-table joins.

[0147] 3. Through natural language interaction, even non-technical personnel without SQL writing skills can easily complete data query operations. Users only need to ask questions in everyday language (natural language query statements), and the system can automatically parse their user intent and generate corresponding SQL statements, greatly reducing the technical threshold for data analysis.

[0148] 4. A dynamic Prompt generation mechanism is adopted, which combines predefined prompt templates and automatically generates the most suitable prompts for the large model based on the user's natural language query, historical dialogue records, and current task objectives. This not only improves the accuracy of generated SQL statements but also enhances generalization ability.

[0149] 5. By automatically identifying the database tables and their fields that may be involved when parsing user queries (natural language query statements), the search space is reduced, invalid calculations are decreased, and the speed of SQL statement generation is accelerated. At the same time, it supports the caching and reuse strategy of SQL fragments. For common query patterns or frequently used subquery structures, they are abstracted into reusable SQL fragments and directly called in subsequent queries to avoid repeated parsing and generation, thereby improving overall performance.

[0150] 6. Through the built-in SQL optimization engine, after generating the original SQL statement, further optimizations can be made such as adjusting the JOIN order, removing redundant conditions, merging subqueries, and suggesting the use of indexes, thereby improving execution efficiency and resource utilization.

[0151] 7. The TextCNN text classification model is used to initially screen user intent (such as SQL generation, annotation, table lookup), which reduces the complexity of subsequent processing. Compared with traditional single-model processing, staged intent recognition can improve processing efficiency and reduce interference from irrelevant information. The preprocessing process (removing irrelevant characters, word segmentation, and stop word removal) improves the standardization of model input and reduces the interference of noise on classification results.

[0152] 8. The system uses a first text slider (long text) for coarse-grained filtering and a second text slider (short text) for fine-grained matching, which balances the comprehensiveness and accuracy of the matching. Compared with the traditional single sliding window, it can handle fuzzy query needs in complex natural language more efficiently. Through the two-level matching mechanism, the similarity score threshold can be adaptively adjusted to avoid mismatch problems caused by a single threshold.

[0153] 9. GraphRAG technology is used to retrieve entities associated with similarity metrics from the knowledge graph, overcoming the limitations of traditional RAG which relies solely on vector similarity for retrieval; through relational reasoning in the graph structure, implicitly related table fields (such as foreign key relationships and business logic dependencies) can be mined, improving the accuracy of field retrieval; by retrieving similar SQL fragments using RAG technology and combining them with entity associations in the knowledge graph, candidate fields that conform to business logic can be generated, reducing the risk of generating irrelevant fields in large models.

[0154] 10. Filter candidate fields using preset rules / conditions (such as access control and data timeliness) to avoid generating redundant or invalid SQL; for example, exclude deprecated table fields or sensitive data fields to enhance system security.

[0155] 11. By setting prompt word templates to include multi-dimensional information such as Question, Business Knowledge, Examples, Requirements, and Table Information, context injection guides the large model to generate SQL statements that conform to business specifications.

[0156] 12. By combining TextCNN (convolutional neural network) with GraphRAG (graph retrieval augmentation generation), the complementarity of text semantic understanding and structured knowledge reasoning is achieved; for example, TextCNN processes fuzzy expressions in natural language, while GraphRAG supplements business logic constraints, forming an end-to-end semantic-to-SQL mapping.

[0157] 13. By constraining the generation range of large models through external knowledge bases (data knowledge vector library, business knowledge graph), the "illusion" problem caused by the lack of domain knowledge in large models is solved; at the same time, the knowledge base can be updated independently without retraining large models, which improves the maintainability of the system.

[0158] 14. By using TextCNN to quickly filter irrelevant intents (such as SQL comment generation) during the preprocessing stage, the entire computational burden is avoided in the subsequent large model inference stage, significantly reducing the overall system resource consumption; by reducing the number of candidate indicators through the first text slider (coarse screening) and the second text slider (fine screening) to be calculated only within a limited candidate set, the computational efficiency is effectively improved compared to traversing the entire indicator library.

[0159] 15. By deeply integrating TextCNN intent recognition, two-level text slider similarity matching, GraphRAG knowledge graph retrieval, and large model generation capabilities, it achieves efficient and accurate conversion of natural language to SQL statements. Its multi-stage collaborative mechanism (intent filtering → dynamic slider matching → graph relationship reasoning → prompt word constraints) significantly improves the accuracy of semantic parsing, solving pain points such as difficulty in adapting fuzzy queries and lack of cross-table association logic in traditional solutions. Combining field filtering and SQL optimization rules from the business knowledge graph, it ensures that the generated results meet data specifications and security requirements. At the same time, it supports rapid cross-industry migration and incremental updates through modular design, reducing development costs. With the help of the generalization capabilities of large models, it covers long-tail complex query scenarios, providing an efficient, secure, and scalable intelligent data analysis solution.

[0160] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A SQL generation method combining GraphRAG and large models, characterized in that: Includes the following steps: Step S1: Obtain the input natural language query statement and identify the user intent of the natural language query statement through the TextCNN text classification model; Step S2: Based on the user intent, match the similarity index of the natural language query statement using a text slider similarity algorithm; Step S3: Using GraphRAG technology, query the associated table fields from the preset data knowledge vector base based on the similarity index, and filter each of the table fields; Step S4: Generate query suggestions based on the natural language query statement and the filtered table fields using a preset suggestion template; Step S5: Input the query suggestion words into the large model to obtain the SQL statement corresponding to the natural language query statement.

2. The SQL generation method combining GraphRAG and large models as described in claim 1, characterized in that: Step S1 specifically involves: The input natural language query statement is obtained through a visual interface. The natural language query statement is preprocessed, including at least irrelevant character deletion, format unification, word segmentation and stop word removal. The user intent of the natural language query statement is identified through a pre-trained TextCNN text classification model. The user intent includes at least SQL generation, SQL commenting, and data table lookup.

3. The SQL generation method combining GraphRAG and large models as described in claim 1, characterized in that: Step S2 specifically involves: Based on the user intent, a text slider similarity algorithm is used to match similarity indicators of the natural language query from a preset indicator library. The matching process is as follows: Define a first text slider and a second text slider, wherein the length of the first text slider is longer than that of the second text slider; The first text slider iterates through the characters of the natural language query statement to obtain a long string. The long string is then iterated through using a similarity algorithm to calculate the first similarity score between the long string and the indicators in the preset indicator library. Based on the first similarity score, candidate indicators are selected from the indicator library for first-level matching. The second text slider iterates through the characters of the natural language query statement to obtain a short string. A similarity algorithm is used to iterate through the short string and calculate the second similarity score between it and each candidate indicator. Based on the second similarity score, similar indicators are selected from each candidate indicator for secondary matching.

4. The SQL generation method combining GraphRAG and large models as described in claim 1, characterized in that: Step S3 specifically involves: Using GraphRAG technology, entities associated with the similarity index are retrieved from a preset business knowledge graph. Based on the entities, the corresponding data tables are queried. Using RAG technology, SQL fragments similar to the data tables are queried. Based on the SQL fragments, the associated table fields are queried from a preset data knowledge vector library. The table fields are then filtered based on preset filtering rules or filtering conditions.

5. The SQL generation method combining GraphRAG and large models as described in claim 1, characterized in that: Step S4 specifically involves: Using a preset prompt word template that includes Question, Business Knowledge, Examples, Requirements, and TableInformation, query prompt words are generated based on the natural language query statement and the filtered table fields. The Question field stores natural language query statements; the Business Knowledge field stores business terminology knowledge; the Examples field stores examples of SQL fragments; the Requirements field stores SQL format requirements; and the Table Information field stores table fields. Step S5 specifically involves: The query suggestions are input into the pre-trained Qwen large model to obtain the SQL statement corresponding to the natural language query. Based on the input adjustment instructions, the SQL statement is optimized by at least adjusting the JOIN order, removing redundant conditions, and merging subqueries.

6. A SQL generation system combining GraphRAG and large models, characterized in that: Includes the following modules: The user intent recognition module is used to acquire the input natural language query statement and identify the user intent of the natural language query statement through the TextCNN text classification model. The similarity index matching module is used to match the similarity index of the natural language query statement based on the user intent using a text slider similarity algorithm; The table field query and filtering module is used to query related table fields from a preset data knowledge vector base based on the similarity index using GraphRAG technology, and to filter each of the table fields. The query suggestion generation module is used to generate query suggestion words based on the natural language query statement and the filtered table fields using a preset suggestion word template; The SQL statement generation module is used to input the query suggestion words into the large model and obtain the SQL statement corresponding to the natural language query statement.

7. The SQL generation system combining GraphRAG and large models as described in claim 6, characterized in that: The user intent recognition module is specifically used for: The input natural language query statement is obtained through a visual interface. The natural language query statement is preprocessed, including at least irrelevant character deletion, format unification, word segmentation and stop word removal. The user intent of the natural language query statement is identified through a pre-trained TextCNN text classification model. The user intent includes at least SQL generation, SQL commenting, and data table lookup.

8. The SQL generation system combining GraphRAG and large models as described in claim 6, characterized in that: The similarity index matching module is specifically used for: Based on the user intent, a text slider similarity algorithm is used to match similarity indicators of the natural language query from a preset indicator library. The matching process is as follows: Define a first text slider and a second text slider, wherein the length of the first text slider is longer than that of the second text slider; The first text slider iterates through the characters of the natural language query statement to obtain a long string. The long string is then iterated through using a similarity algorithm to calculate the first similarity score between the long string and the indicators in the preset indicator library. Based on the first similarity score, candidate indicators are selected from the indicator library for first-level matching. The second text slider iterates through the characters of the natural language query statement to obtain a short string. A similarity algorithm is used to iterate through the short string and calculate the second similarity score between it and each candidate indicator. Based on the second similarity score, similar indicators are selected from each candidate indicator for secondary matching.

9. The SQL generation system combining GraphRAG and large models as described in claim 6, characterized in that: The table field query and filtering module is specifically used for: Using GraphRAG technology, entities associated with the similarity index are retrieved from a preset business knowledge graph. Based on the entities, the corresponding data tables are queried. Using RAG technology, SQL fragments similar to the data tables are queried. Based on the SQL fragments, the associated table fields are queried from a preset data knowledge vector library. The table fields are then filtered based on preset filtering rules or filtering conditions.

10. The SQL generation system combining GraphRAG and large models as described in claim 6, characterized in that: The query suggestion term generation module is specifically used for: Using a preset prompt word template that includes Question, Business Knowledge, Examples, Requirements, and TableInformation, query prompt words are generated based on the natural language query statement and the filtered table fields. The Question field stores natural language query statements; the Business Knowledge field stores business terminology knowledge; the Examples field stores examples of SQL fragments; the Requirements field stores SQL format requirements; and the Table Information field stores table fields. The SQL statement generation module is specifically used for: The query suggestions are input into the pre-trained Qwen large model to obtain the SQL statement corresponding to the natural language query. Based on the input adjustment instructions, the SQL statement is optimized by at least adjusting the JOIN order, removing redundant conditions, and merging subqueries.

Citation Information

Cited By

  • Heat supply industry advanced report generation method and system based on large language model

    CN121052227A

  • Query statement generation method and device, electronic equipment and medium

    CN121412257A

  • Table query method and device

    CN121579471A

  • Query method, query device, medium, electronic equipment and product

    CN121764957A