SQL query interpretation generation method and system based on heterogeneous graph
By constructing heterogeneous graphs of SQL queries and using improved heterogeneous graph neural networks, the problem of insufficient accuracy and readability of SQL interpretation in the prior art is solved, and more efficient SQL query interpretation generation is achieved.
Patent Information
- Application Number
- CN202510214090.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art fails to fully utilize heterogeneous characteristics when processing SQL queries, resulting in a lack of uniqueness and accuracy in AST structure expression, affecting the effectiveness of graph neural networks in learning node relationships, and thus affecting the accuracy and readability of SQL interpretation.
Using a heterogeneous graph-based method, by constructing heterogeneous graphs of SQL queries, six different types of edges are introduced, the heterogeneous features of SQL queries are extracted using the improved heterogeneous graph neural network, and natural language interpretations are generated in the encoder-decoder framework.
It significantly improves the accuracy and readability of SQL query interpretation, can effectively handle common and multi-table complex queries, and has important application value for database query optimization and code maintenance.
Smart Images

Figure CN120296024A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of databases, and particularly to a method and system for generating SQL query explanations based on heterogeneous graphs. Background Art
[0002] In the digital and data-driven era, various applications continuously generate a large amount of data, and databases have become the core tools for storing and managing data. With the increasing complexity and diversity of modern database management systems (DBMS), the optimization and maintenance of databases have become increasingly difficult. SQL (Structured Query Language), as an important tool for database access and management, is crucial for database operations and data analysis. However, due to the complexity of the SQL query language, especially for non-technical users, SQL queries are often difficult to understand. Even for experts or the original creators of the queries, complex SQL queries often pose challenges in reading and maintenance.
[0003] Currently, many existing methods convert SQL queries into equivalent Abstract Syntax Trees (ASTs) and extract the structural features of the ASTs through encoders for generating natural language explanations of SQL. However, these methods usually treat the syntax structure of the AST as a homogeneous graph, ignoring the different relationships between nodes in the SQL query, such as parent-child relationships and sibling relationships. This processing method fails to fully utilize the inherent heterogeneous characteristics of SQL queries, resulting in a lack of uniqueness and accuracy in the expression of the AST structure, which in turn affects the effectiveness of graph neural networks (GNNs) in learning the relationships between AST nodes. These limitations will all lead to deficiencies in the accuracy and readability of the SQL explanations generated by the model. Summary of the Invention
[0004] To overcome the above problems, the purpose of the present invention is to provide a method and system for generating SQL query explanations based on heterogeneous graphs to solve the technical problem of how to convert SQL queries into natural language explanations.
[0005] The present invention is implemented as follows: A method for generating SQL query explanations based on heterogeneous graphs, the method comprising the following steps:
[0006] S1: After cleaning the SQL query, extract the source tokens;
[0007] S2: Based on the source tokens, construct an SQL abstract syntax tree, where the abstract syntax tree includes keyword nodes, table name nodes, column name nodes, and query condition nodes of the SQL query;
[0008] S3: Based on the abstract syntax tree, introduce six different types of edges between parent and child nodes, sibling nodes, and leaf nodes to construct a heterogeneous graph of the SQL query;
[0009] S4: Input the heterogeneous graph into an improved heterogeneous graph neural network, and extract heterogeneous features of the SQL query through a two-stage attention mechanism-based aggregation process;
[0010] S5: Incorporate the learned SQL features into an encoder-decoder framework to generate a natural language explanation of the SQL query.
[0011] Preferably, the operations of data cleaning and extraction in step S1 include: reading each word in the SQL statement, lowercasing it, splitting it by underscores, and performing splitting and lemmatization operations on tokens in CamelCase and Snake_case forms to obtain tokens of the code.
[0012] Preferably, the method for constructing the SQL abstract syntax tree in step S2 is:
[0013] Create corresponding abstract syntax tree nodes for each token, including SQL keyword nodes such as Select, From, and Where, table name nodes, column name nodes, and query condition nodes;
[0014] Create a new SQL node as the root node of the abstract syntax tree, and create new abstract keyword nodes such as Table, Column, and Condition;
[0015] The keyword nodes such as Select, From, Where, and the abstract keyword nodes are connected to the root node, the Table node is connected to the table name node, the Column node is connected to the column name node, and the Condition node is connected to the query condition node.
[0016] Preferably, the method for constructing the SQL heterogeneous graph in step S3 is:
[0017] Introduce bidirectional directed edges between the parent-child nodes, sibling nodes, and leaf nodes of the abstract syntax tree to construct a heterogeneous graph of the SQL query;
[0018] The parent-child edges represent the basic hierarchical syntax structure of the SQL sequence; the edges between sibling nodes represent the parallel relationship in the heterogeneous graph; the edges between leaf nodes represent the inherent order between SQL tokens sequences, reflecting the syntax sequence of the SQL query.
[0019] Preferably, the two-stage aggregation of the heterogeneous graph neural network in step S4 is:
[0020] In the first stage, calculate the weights of the adjacent neighbors of each node through a node-level attention mechanism, and aggregate them into a neighbor group to obtain a preliminary semantic representation of the node;
[0021] In the second stage, a semantic-level attention mechanism is used to distinguish the differences between different neighbor groups, so as to obtain the optimal weighted combination and generate the final embedding representation of each node.
[0022] Preferably, the step S5 is further specifically as follows:
[0023] Stack six layers of Transformer encoders to process the node embedding representation generated by the heterogeneous graph neural network, and use residual connection and graph normalization operations in each layer;
[0024] Stack eight layers of Transformer decoders to generate the natural language explanation of the SQL query. Each decoder layer includes three sub-layers: masked multi-head attention, standard multi-head attention, and positional feed-forward network (FFN), and each sub-layer is optimized using layer normalization and residual connection.
[0025] The present invention also provides a SQL query explanation generation system based on a heterogeneous graph, including: a preprocessing module, an abstract syntax tree construction module, a heterogeneous graph construction module, a heterogeneous graph neural network module, and an encoder-decoder module;
[0026] The preprocessing module is used to clean the data of the SQL query and extract the source tokens;
[0027] The abstract syntax tree construction module is used to construct the abstract syntax tree of the SQL query according to the source tokens;
[0028] The heterogeneous graph construction module is used to construct the heterogeneous graph of the SQL query according to the abstract syntax tree;
[0029] The heterogeneous graph neural network module is used to extract the heterogeneous features of the SQL query and generate the node embedding representation;
[0030] The encoder-decoder module is used to encode the SQL features and decode to generate the natural language explanation of the SQL query.
[0031] The beneficial effects of the present invention are as follows: 1. By introducing a heterogeneous graph neural network to parse SQL queries, the present invention can more comprehensively understand the semantics and structure of SQL queries. 2. Compared with the traditional AST-based processing method, the present invention can effectively overcome the limitation that the AST model fails to fully consider the diverse relationships between SQL query nodes, and significantly improve the accuracy and readability of SQL query explanations. 3. This method can not only handle common simple SQL queries, but also effectively cope with multi-table complex queries, and has important application value especially in database query optimization, code maintenance, and data analysis. Description of the Drawings
[0032] Figure 1 is the schematic diagram of the method flow of the present invention.
[0033] Figure 2 It is a method model architecture diagram of the present invention. DETAILED DESCRIPTION
[0034] The present invention will be further described below in conjunction with the accompanying drawings.
[0035] See also Figure 1 , Figure 1 A flowchart of a method for generating SQL query explanations based on heterogeneous graphs.
[0036] The above SQL query interpretation generation method comprises the following steps:
[0037] Step S1: Extract source tokens after data cleaning of SQL query.
[0038] In the application, Python toolkits and specific cleaning methods can be used to clean the SQL queries to be processed and extract their tokens. Examples of data cleaning include lowercase, underscore splitting, camel case splitting, and word form restoration.
[0039] During the training process, the SQL queries used include the WikiSQL dataset and the Spider dataset; the WikiSQL dataset contains a total of 80,654 instances, 56,355 instances are used for training, and 8421 instances are used for verification. The Spider dataset contains a total of 9692 data instances, 8658 instances are used for training, and 1034 instances are used for verification.
[0040] In one embodiment, the following operations may be performed when data cleaning is performed on the corpus to be processed:
[0041] Tokenize the SQL query to be processed and convert each token to lowercase;
[0042] Use the NLTK library to split each identifier based on camel case and underscores; for example, the column name "sale_price" will be broken down into "sale" and "price";
[0043] Each word was lemmatized using spaCy2, standardizing the form to improve uniformity.
[0044] Step S2: Build a SQL abstract syntax tree based on the source tokens.
[0045] With the help of the tree-sitter-sql toolkit, the SQL query is constructed as an abstract syntax tree. The steps are as follows:
[0046] Create corresponding abstract syntax tree nodes for each token, including SQL keyword nodes such as Select, From, and Where, table name nodes, column name nodes, and query condition nodes;
[0047] Create a new SQL node as the root node of the abstract syntax tree, and create abstract keyword nodes such as Table, Column, and Condition;
[0048] The keyword nodes such as Select, From, Where, and the abstract keyword nodes are connected to the root node. The Table node is connected to the table name node, the Column node is connected to the column name node, and the Condition node is connected to the query condition node.
[0049] Compared with the traditional method, this method of constructing the abstract syntax tree can avoid the increase in the depth of the SQL tree as the complexity of the SQL code increases, which may make it more difficult for the nodes at the bottom of the AST to communicate with the nodes at the top, and this may affect the extraction of structural features during the coding process.
[0050] Step S3: Construct a heterogeneous graph for the SQL query based on the abstract syntax tree.
[0051] Introduce bidirectional directed edges between the parent and child nodes, sibling nodes, and leaf nodes of the abstract syntax tree to construct a heterogeneous graph for the SQL query;
[0052] The parent-child edges represent the basic hierarchical syntax structure of the SQL sequence; the edges between sibling nodes represent the parallel relationship in the heterogeneous graph; the edges between leaf nodes represent the inherent order between SQL tokens sequences, reflecting the syntax sequence of the SQL query.
[0053] Since the leaf nodes correspond to SQL query tokens, the directed edges between the leaves can represent the code sequence of SQL, so the sequential features of the SQL query can be learned from the heterogeneous graph.
[0054] Step S4: Input the heterogeneous graph into an improved heterogeneous graph neural network to extract the heterogeneous features of the SQL query. The heterogeneous graph neural network includes a two-stage attention-based aggregation process, as Figure 2 shown.
[0055] Due to the heterogeneity of the nodes, different nodes have different feature spaces. First, project the features of all nodes into the same feature space, and the projection process is as follows:
[0056] h i ′ = M i h i ;
[0057] where hi and h i ′ are the original feature and projected feature of the node respectively, and M i is a specific projection matrix.
[0058] In the first stage, a node-level attention aggregation process. Neighbors with different edge types are aggregated into different neighbor groups respectively, and self-attention is used to learn the weights between nodes. Given a pair of nodes (i, j) connected by edge type g, the node-level attention can learn the importance of node j to node i. The expression of the importance of node pair (i, j) based on edge type g is as follows:
[0059]
[0060] Inject the structural information into the model through masked attention. After obtaining the importance between node pairs based on edge type g, normalize them, and the weight coefficients obtained through the softmax function are:
[0061]
[0062] where represents the set of neighbors of node i with edge type g, σ is the activation function, ∥ is the concatenation operation, and a g is the node-level attention vector of edge type g. Then, the node embedding based on edge type g can be aggregated into the neighbor group through the projected features of the neighbors, and the corresponding coefficients are as follows:
[0063]
[0064] where represents the neighbor group embedding representation of node i with edge type g. In summary, for node i, given the set of edge types After feeding the node features into the node-level attention, the neighbor group of the semantic-specific node embedding can be obtained, denoted as
[0065] In the second stage, a semantic-level attention aggregation process. Take the P neighbor group embeddings learned from the node-level attention as input, and the weight representation learned for each neighbor group is:
[0066]
[0067] Here att semA deep neural network that performs semantic-level attention. To understand the importance of each neighbor group, first transform the semantics-specific embeddings through a non-linear transformation (e.g., a single-layer MLP). Then, measure the importance of the semantics-specific embeddings using the similarity between the transformed embeddings and the semantic-level attention vector q. Further, average the importance of all semantics-specific node embeddings. The importance of each neighbor group is expressed as as follows:
[0068]
[0069] where W is the weight matrix, b is the bias vector, and q is the semantic-level attention vector. After obtaining the importance of each neighbor group, normalize them using the softmax function:
[0070]
[0071] The higher, the more important the neighbor group g P is. Taking the learned weights as coefficients, these semantics-specific embeddings can be fused to obtain the final embedding Z i , as follows:
[0072]
[0073] After all node states are updated, the state vectors are concatenated and sent to the ReLU activation function for non-linear transformation:
[0074] Z n = ReLU(Z1, Z2, …, Z i , …)
[0075] Step S5: Incorporate the learned SQL features into the encoder-decoder framework to generate a natural language explanation of the SQL query.
[0076] For the Kth encoder layer, represent the embedding as As more encoder layers are stacked, the nodes collect their neighboring information from farther distances and extract more heterogeneous features. To mitigate the vanishing gradients and excessive vector offsets in multi-layer computations, residual connections and graph normalization are adopted in each layer, which are formalized as follows:
[0077]
[0078] where represents the heterogeneous graph node state vector output by the (K-1)th encoder layer. GraphNorm represents the graph normalization operation.
[0079] In the decoder, eight Transformer decoding layers are stacked to generate SQL interpretations. Each module contains three sub-layers, including masked multi-head attention for self-attention encoding of tokens of the existing query interpretation, standard multi-head attention for decoding heterogeneous graph nodes of the representation, and a fully-connected positional feed-forward network (FFN). Each sub-layer performs a residual connection with layer normalization.
[0080] Given the query interpretation tokens from the (K-1)th layer and the extracted heterogeneous graph representation The decoding process of the Kth Transformer layer is as follows:
[0081]
[0082] where MaskAtt and Att represent masked multi-head attention and standard multi-head attention respectively, both taking query, key, and value vectors as inputs to explore the relationships between them. LayerNorm represents the layer normalization operation. Finally, the decoded vector is fed into the FFN for non-linear transformation.
[0083] The present invention also proposes a SQL query interpretation generation system based on a heterogeneous graph, which can execute the above-mentioned SQL query interpretation generation method based on a heterogeneous graph, including: a preprocessing module, an abstract syntax tree construction module, a heterogeneous graph construction module, a heterogeneous graph neural network module, and an encoder-decoder module;
[0084] The preprocessing module is used to clean the data of the SQL query and extract source tokens;
[0085] The abstract syntax tree construction module is used to construct an abstract syntax tree of the SQL query according to the source tokens;
[0086] The heterogeneous graph construction module is used to construct a heterogeneous graph of the SQL query according to the abstract syntax tree;
[0087] The heterogeneous graph neural network module is used to extract heterogeneous features of the SQL query and generate node embedding representations;
[0088] The encoder-decoder module is used to encode SQL features and decode to generate a natural language interpretation of the SQL query.
[0089] The above-mentioned embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for generating SQL query explanations based on heterogeneous graphs, characterized in that, It includes the following steps: S1: Extract source tokens after cleaning the SQL query data; S2: Build an SQL abstract syntax tree based on the source tokens, where the abstract syntax tree includes keyword nodes, table name nodes, column name nodes, and query condition nodes of the SQL query; S3: Based on the abstract syntax tree, introduce six different types of edges between parent and child nodes, sibling nodes, and leaf nodes to build a heterogeneous graph of the SQL query; S4: Input the heterogeneous features into an improved heterogeneous graph neural network, and extract the heterogeneous features of the SQL query through a two-stage attention mechanism-based aggregation process; S5: Incorporate the learned SQL features into an encoder-decoder framework to generate a natural language explanation of the SQL query.
2. The method for generating an SQL query interpretation based on a heterogeneous graph according to claim 1, wherein The data cleaning in step S1 includes standardization processes such as lowercasing, underscore splitting, camel case splitting, and lemmatization.
3. A method for generating SQL query explanations based on heterogeneous graphs according to claim 1, characterized in that, Step S2 is further specifically: Create corresponding abstract syntax tree nodes for each token, including SQL keyword nodes such as Select, From, and Where, table name nodes, column name nodes, and query condition nodes; Create a new SQL node as the root node of the abstract syntax tree, and create new abstract keyword nodes such as Table, Column, and Condition; The keyword nodes such as Select, From, Where, etc., and the abstract keyword nodes are connected to the root node, the Table node is connected to the table name node, the Column node is connected to the column name node, and the Condition node is connected to the query condition node.
4. A method for generating SQL query explanations based on heterogeneous graphs according to claim 1, characterized in that, Step S3 is further specifically: Introduce bidirectional directed edges between parent and child nodes, sibling nodes, and leaf nodes of the abstract syntax tree to build a heterogeneous graph of the SQL query; The parent-child edges represent the basic hierarchical syntax structure of the SQL sequence; the edges between sibling nodes represent the parallel relationship in the heterogeneous graph; the edges between leaf nodes represent the inherent order between SQL token sequences, reflecting the syntax sequence of the SQL query.
5. A method for generating SQL query explanations based on heterogeneous graphs according to claim 1, characterized in that, The two-stage aggregation of the heterogeneous graph neural network in step S4 is: In the first stage, calculate the weights of the adjacent neighbors of each node through a node-level attention mechanism and aggregate them into a neighbor group to obtain a preliminary semantic representation of the node; In the second stage, distinguish the differences between different neighbor groups through a semantic-level attention mechanism to obtain an optimal weighted combination and generate the final embedding representation of each node.
6. The method for generating an SQL query interpretation based on a heterogeneous graph according to claim 1, wherein Step S5 is further specifically: Stack six layers of Transformer encoders to process the node embedding representations generated by the heterogeneous graph neural network, and use residual connections and graph normalization operations in each layer; Stack eight layers of Transformer decoders to generate a natural language explanation of the SQL query. Each decoder layer includes three sub-layers: masked multi-head attention, standard multi-head attention, and positional feed-forward network (FFN), and each sub-layer is optimized using layer normalization and residual connections.
7. An SQL query interpretation generation system based on a heterogeneous graph, which is used to execute any one of the SQL query interpretation generation methods based on a heterogeneous graph described in claims 1 to 6, characterized in that, It includes: A preprocessing module, an abstract syntax tree construction module, a heterogeneous graph construction module, a heterogeneous graph neural network module, and an encoder-decoder module; The preprocessing module is used to perform data cleaning on the SQL query and extract source tokens; The abstract syntax tree construction module is used to construct an abstract syntax tree of the SQL query based on the source tokens; The heterogeneous graph construction module is used to construct a heterogeneous graph of the SQL query based on the abstract syntax tree; The heterogeneous graph neural network module is used to extract heterogeneous features of the SQL query and generate node embedding representations; The encoder-decoder module is used to encode the SQL features and decode to generate a natural language interpretation of the SQL query.
Citation Information
Cited By
Structured query language generation method for semantic proper subset decomposition based on LLM
CN122064708A
Structured query language generation method based on semantic subset decomposition of llm
CN122064708B