Method, system, terminal and medium for constructing reasoning path based on decomposing SQL

By parsing and conditional splitting the original SQL, generating subSQL collections and building inference paths, the problems of incomplete SQL resolution and single paths in the existing technology are solved, and the accuracy and optimization effect of SQL debugging are achieved.

CN119293074BActive Publication Date: 2025-05-09INT DIGITAL ECONOMY ACAD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411814412.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-09
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Existing SQL decomposition methods are prone to miss potential subSQL when processing complex queries, with a single inference path, and cannot accurately locate error conditions during debugging SQL.

Method used

By parsing the original SQL, we identify splitable conditions, build a condition set, and generate a subSQL set based on preset deletion rules, and finally build an inference path based on the subSQL set.

Benefits of technology

It realizes comprehensive analysis of complex SQL, covers all syntax structures, and generates multiple inference paths, which facilitates precise positioning and optimization in SQL debugging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293074B_ABST
    Figure CN119293074B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, terminal and medium for constructing reasoning paths based on decomposing SQL. The method comprises: parsing the original SQL, identifying the conditions that can be split in the original SQL, and obtaining a condition set, wherein the condition is a clause in the original SQL; deleting several conditions in the condition set based on a preset deletion rule, obtaining a subSQL each time a condition is deleted, and obtaining a subSQL set corresponding to the original SQL; and constructing reasoning paths based on the conditions contained in each subSQL in the subSQL set. The present invention can parse complex SQL into SQL clauses and multiple subSQL sets containing partial conditions, and construct reasoning paths by analyzing the conditions contained in the subSQL set, so as to increase the understandability and execution efficiency of SQL.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of SQL decomposition technology, and in particular to a method, system, terminal and medium for constructing an inference path based on decomposing SQL. Background Art

[0002] In database management systems, SQL is widely used to query and manipulate data. However, with the increase in data volume and query complexity, traditional SQL parsing methods face many challenges in processing complex queries.

[0003] The existing SQL decomposition methods do not fully support SQL syntax, only support part of SQL syntax, and the generated subSQL is not comprehensive enough, and there is a problem of missing some potential subSQL. In addition, the existing SQL decomposition methods only infer one path, which cannot achieve flexible SQL debugging. In addition, the existing technology usually directly parses the complete SQL statement, fails to fully utilize the relationship between query conditions in SQL, and cannot accurately locate error conditions and intuitively display the SQL intermediate execution process during SQL debugging. Summary of the invention

[0004] The technical problem to be solved by the present invention is that, in view of the above-mentioned defects of the prior art, a method, system, terminal and medium for constructing an inference path based on decomposing SQL are provided, aiming to solve the problems that the SQL parsing method in the prior art is prone to miss some potential subSQL when processing complex queries, the inferred path is single, and the error conditions cannot be accurately located during the SQL debugging process.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0006] In a first aspect, the present invention provides a method for constructing an inference path based on decomposing SQL, wherein the method comprises:

[0007] Parse the original SQL and identify the separable conditions in the original SQL to obtain a condition set, wherein the conditions are clauses in the original SQL;

[0008] Deleting several conditions in the condition set based on a preset deletion rule, obtaining a subSQL each time a condition is deleted, and obtaining a subSQL set corresponding to the original SQL;

[0009] The reasoning paths are constructed based on the conditions contained in each subSQL in the subSQL set.

[0010] In one implementation, the original SQL is parsed, and the separable conditions in the original SQL are identified to obtain a condition set, including:

[0011] Parse the original SQL to obtain the abstract syntax tree structure;

[0012] Identify the type of each node in the abstract syntax tree structure and determine the conditions for splitting the abstract syntax tree structure;

[0013] The determined separable conditions are post-processed to obtain the condition set.

[0014] In one implementation, identifying the type of each node in the abstract syntax tree structure and determining the conditions for splitting the abstract syntax tree structure include:

[0015] If the type of a node in the abstract syntax tree structure is an operation type, then identifying the types of all child nodes of the node;

[0016] If the type of all child nodes is a non-operation type, the node is determined to be a conditional node, and the node and all corresponding child nodes are used as splittable conditions.

[0017] In one implementation, identifying the type of each node in the abstract syntax tree structure and determining the conditions for splitting the abstract syntax tree structure further includes:

[0018] If a node in the abstract syntax tree structure is a nested query node, the nested query node is used as a splittable condition.

[0019] In one implementation, identifying the type of each node in the abstract syntax tree structure further includes:

[0020] Add attribute information to each separable conditional node identified from the abstract syntax tree structure, the attribute information including: an identifier of the conditional node, a condition identifier, and a transition identifier of the conditional node. In one implementation, post-processing is performed on the determined separable conditions to obtain the condition set, including:

[0021] Identify the FROM type conditions among the separable conditions, and determine whether to delete the FROM type conditions;

[0022] Identify the SELECT conditions in the separable conditions and merge the conditions based on the SELECT conditions;

[0023] If a conditional node in the separable conditions does not have a corresponding sibling node, or the sibling nodes corresponding to the conditional node do not include the conditional node, the conditional node is used as the corresponding parent node.

[0024] In one implementation, several conditions in the condition set are deleted based on a preset deletion rule, and a subSQL is obtained by deleting each condition, and a subSQL set corresponding to the original SQL is obtained, including:

[0025] Determining a subset of conditions in the condition set that can be deleted based on dependencies between conditions in the condition set;

[0026] Based on the deletion rule, several conditions in the condition set are deleted according to the subset of conditions that can be deleted, and a subSQL is obtained each time a condition is deleted, and the subSQL set is obtained, wherein the deletion rule is that if the dependent condition is deleted, all conditions that depend on the condition are deleted.

[0027] In one implementation, determining a subset of conditions in the condition set that can be deleted based on the dependency relationship between the conditions in the condition set includes:

[0028] Determine the dependent condition and the dependent condition in the condition set based on the dependency relationship between the conditions in the condition set;

[0029] All of the dependent conditions and all of the dependent conditions are taken as a condition subset that can be deleted; wherein, only when all of the dependent conditions are included in the condition subset that can be deleted, the dependent conditions are included in the condition subset that can be deleted.

[0030] In one implementation, based on the deletion rule, several conditions in the condition set are deleted according to the condition subset that can be deleted, and a subSQL is obtained each time a condition is deleted, and the subSQL set is obtained, including:

[0031] Obtain each condition in the condition subset that can be deleted in turn, and delete the conditions corresponding to the condition set in turn, and obtain a subSQL each time a condition is deleted;

[0032] Based on all subSQLs, the subSQL set is obtained.

[0033] In one implementation, obtaining each condition in the condition subset that can be deleted in turn, and deleting the conditions corresponding to the condition set in turn, includes:

[0034] If the type of the condition obtained from the condition subset that can be deleted is an arithmetic operation type, then two condition nodes corresponding to the condition of the arithmetic operation type are determined, wherein the two condition nodes are sibling nodes of each other;

[0035] When deleting one of the conditional nodes, the deletion is achieved by replacing the parent node with the corresponding sibling node.

[0036] In one implementation, a reasoning path is constructed based on the conditions contained in each subSQL in the subSQL set, including:

[0037] Based on the subSQL set, determine the subSQL with the least conditions and the subSQL with the most conditions, where the subSQL with the most conditions is the original SQL;

[0038] An inference path is constructed based on the subSQL containing the least conditions and the subSQL containing the most conditions.

[0039] In one implementation, a reasoning path is constructed based on the subSQL containing the least conditions and the subSQL containing the most conditions, including:

[0040] The subSQL with the least conditions is used as the starting point of the reasoning path, and the subSQL with the most conditions is used as the end point of the reasoning path;

[0041] All subSQLs are traversed one by one from the starting point, and if the current subSQL traversed has one more condition than the previous subSQL, the current subSQL is used as the next starting point, and so on, until the end point is reached, and all reasoning paths are obtained.

[0042] In a second aspect, an embodiment of the present invention further provides a system for constructing an inference path based on decomposing SQL, wherein the system includes:

[0043] A condition set determination module, used to parse the original SQL and identify the separable conditions in the original SQL to obtain a condition set, wherein the conditions are clauses in the original SQL;

[0044] A partial condition deletion module is used to delete several conditions in the condition set based on a preset deletion rule, and obtain a subSQL each time a condition is deleted, thereby obtaining a subSQL set corresponding to the original SQL;

[0045] The reasoning path building module is used to build reasoning paths based on the conditions contained in each subSQL in the subSQL set.

[0046] In a third aspect, an embodiment of the present invention further provides a terminal, wherein the terminal includes a memory, a processor, and a program for constructing an inference path based on decomposed SQL, which is stored in the memory and can be run on the processor. When the processor executes the program for constructing an inference path based on decomposed SQL, the steps of the method for constructing an inference path based on decomposed SQL of any one of the above-mentioned schemes are implemented.

[0047] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein a program for constructing an inference path based on decomposing SQL is stored on the computer-readable storage medium, and when the program for constructing an inference path based on decomposing SQL is executed by a processor, the steps of the method for constructing an inference path based on decomposing SQL described in any one of the above-mentioned schemes are implemented.

[0048] Beneficial effects: Compared with the prior art, the present invention provides a method for constructing an inference path based on decomposing SQL. The present invention first parses the original SQL and identifies the separable conditions in the original SQL to obtain a condition set, wherein the conditions are clauses in the original SQL. Then, based on the preset deletion rules, several conditions in the condition set are deleted respectively, and a subSQL is obtained each time a condition is deleted, and a subSQL set corresponding to the original SQL is obtained. Finally, inference paths are constructed based on the conditions contained in each subSQL in the subSQL set. The present invention can parse complex SQL into SQL clauses and multiple subSQL sets containing partial conditions, which is conducive to covering all grammatical structures in SQL and realizing a wider range of applications. In addition, the present invention constructs an inference path by analyzing the conditions contained in the subSQL set, which is convenient for increasing the comprehensibility and execution efficiency of SQL, and is conducive to accurately locating potential error conditions in SQL debugging. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 A flowchart of a preferred embodiment of a method for constructing an inference path based on decomposing SQL provided in an embodiment of the present invention.

[0050] Figure 2 A diagram showing a specific application implementation of the method for constructing an inference path based on decomposing SQL provided in an embodiment of the present invention.

[0051] Figure 3 A schematic diagram of an AST obtained by parsing the original SQL in the method for constructing an inference path based on decomposing SQL provided in an embodiment of the present invention.

[0052] Figure 4 This is a schematic diagram of an AST after identifying conditions in the method for constructing an inference path based on decomposing SQL provided by an embodiment of the present invention.

[0053] Figure 5 A schematic diagram of a reasoning path in a method for constructing a reasoning path based on decomposing SQL provided in an embodiment of the present invention.

[0054] Figure 6 A schematic diagram of the architecture of a system for constructing an inference path based on decomposing SQL provided in an embodiment of the present invention.

[0055] Figure 7 A functional block diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0056] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0057] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations or steps, nor must they be executed in the order described. For example, some operations or steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0058] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0059] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, in the embodiments of the present invention, words such as "first" and "second" are used to distinguish between identical or similar items with substantially identical functions and effects. For example, the first control information and the second control information are only used to distinguish different control information, and their order is not limited.

[0060] Those skilled in the art can understand that the words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not necessarily limit the differences.

[0061] It should be further understood that the term “and / or” used in the present specification and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0062] To solve the problem of difficulty in debugging due to complex query conditions or unclear logical structure in the SQL debugging process in the prior art. This embodiment provides a method for constructing an inference path based on decomposing SQL, which can parse complex SQL into SQL clauses and multiple subSQL sets containing partial conditions, which is conducive to covering all grammatical structures in SQL and realizing wider application. In specific application, this embodiment first parses the original SQL, and identifies the conditions that can be split in the original SQL to obtain a condition set, wherein the condition is a clause in the original SQL. Then, based on the preset deletion rules, several conditions in the condition set are deleted respectively, and a subSQL is obtained for each condition deleted, and the subSQL set corresponding to the original SQL is obtained. Finally, the inference path is constructed based on the conditions contained in each subSQL in the subSQL set. It can be seen that this embodiment constructs the inference path by analyzing the conditions contained in the subSQL set, which is convenient for increasing the comprehensibility and execution efficiency of SQL, and is conducive to accurately locating potential error conditions in SQL debugging.

[0063] The method of constructing an inference path based on decomposing SQL in this embodiment can be applied to a terminal, which is an intelligent product terminal such as a computer, a smart TV, a mobile phone, etc. Figure 1 As shown in , the method for constructing an inference path based on decomposing SQL in this embodiment includes the following steps:

[0064] Step S100: parse the original SQL, identify the separable conditions in the original SQL, and obtain a condition set, wherein the conditions are clauses in the original SQL.

[0065] Existing technologies are all based on the abstract syntax tree structure (AST) of SQL (Structured Query Language) for semantic understanding, query optimization and analysis. When a database system (such as Spark SQL database and MySQL database, etc.) receives a SQL query statement, it will first use the parser to perform lexical analysis and syntax analysis to convert the query statement into AST. AST is an abstract representation of the SQL syntax structure, and each node represents a grammatical construct, such as an expression, statement, or declaration. In this process, each keyword, table name, column name, condition, etc. in the SQL statement will be represented as a node in AST, and these nodes are connected to each other through edges to form a tree that reflects the structure of the query statement. This representation method helps the database system better understand the intent of the query and provides a basis for subsequent optimization and execution.

[0066] Based on the AST form of SQL, the query optimizer in the database system (such as the optimizer of the PostgreSQL database) optimizes the execution efficiency of the query by analyzing the execution plan of SQL. Although these optimizers do not have the function of generating subSQL, they usually provide some tools and functions for debugging SQL queries to help users diagnose and optimize query performance, detect errors, and understand query execution plans. For example, the EXPLAIN command in MySQL can view the query plan of SQL, and the OPTIMIZER_TRACE command can view the decision process and record every step of the SQL query optimization process in detail. The traditional debugging method in SQL is trial and error debugging, which requires repeated execution of queries to manually obtain intermediate data in SQL queries. Users usually manually decompose SQL into SQL containing partial clauses, and then run these SQL containing partial clauses separately to find out the problems, which is an extremely time-consuming process.

[0067] To this end, combined Figure 2 As shown, in this embodiment, a third-party SQL parsing tool SQLGlot can be used to parse the original SQL into an abstract syntax tree structure (AST). For example:

[0068] SQL: SELECT COUNT(movies.movie_title), ratings.critic FROM ratingsINNER JOIN movies ON ratings.movie_id = movies.movie_id WHEREmovies.director_name = 'Francis Ford Coppola' AND movies.movie_popularity>1000

[0069] The parsed AST is as follows Figure 3 shown. Figure 3 It contains the SQL form of each node and the type of the node, including operation type and non-operation type.

[0070] Furthermore, the present embodiment identifies the type of each node in the above-mentioned abstract syntax tree structure, and determines the conditions that can be split in the abstract syntax tree structure. Among them, the condition is a specific node of the AST and its corresponding child nodes, that is, the clauses in the original SQL. When determining the splittable conditions, the present embodiment analyzes the nodes of the abstract syntax tree structure (AST), and if the type of a node in the abstract syntax tree structure is an operation type, the types of all child nodes of the node are identified. If the type of all child nodes is a non-operation type, it is determined that the node is defined as a conditional node, and at this time the node and all corresponding child nodes can be used as splittable conditions.

[0071] In this embodiment, the operation type node can be divided into three categories according to the type, including: main type, arithmetic operation type and built-in function type. Among them, the main type is the main structure and core logic of SQL; the arithmetic operation type is the logical operation and basic arithmetic operation in SQL; the built-in function type is the commonly used built-in function in SQL, as shown in Table 1 below.

[0072] Table 1 Operation type nodes and non-operation type node collection

[0073]

[0074] The SUBQUERY node (nested query node, another SQL query embedded in an SQL query) in the AST usually represents an independent SQL query. For this reason, in this embodiment, if a node in the abstract syntax tree structure is a nested query node, it is used as a separate condition, so that the nested query node can be used as a separable condition.

[0075] In this embodiment, three attribute information are added to each separable conditional node identified from the abstract syntax tree structure (AST): conID, label, and transferID, and all separable conditional nodes constitute a condition set.

[0076] 1) conID: The identifier of the conditional node, used to quickly locate and delete the conditional node.

[0077] 2) label: conditional identifier. If label is True, it means that the current node is a conditional node. The default value is False.

[0078] 3) transferID: The transfer identifier of the condition node. The default value is False. If transferID is the identifier of another node, it means that the current condition is transferred from the node corresponding to conID.

[0079] Furthermore, after the identification of the separable conditions is completed, the present embodiment performs post-processing on the determined separable conditions to obtain the condition set. Specifically, the present embodiment first identifies the FROM type conditions in the separable conditions, and determines whether to delete the FROM type conditions. The most indispensable thing in SQL is the FROM type condition, and a SQL cannot be constructed without the FROM type condition. Therefore, if there is only one FROM type condition in the separable conditions parsed by the AST, the FROM type condition will not be deleted. If there are more than one FROM type conditions, only the FROM type conditions will remain, and then the remaining FROM type conditions will be deleted.

[0080] Next, this embodiment identifies the SELECT conditions in the separable conditions, and merges the conditions based on the SELECT conditions. Specifically, if the separable conditions include the SELECT conditions, the conditions can be merged. If there are non-operation type nodes in the SELECT conditions, generally speaking, the object of SELECT is the column name of the table, then the SELECT conditions of the same table can merge all the SELECT conditions of the table into the same SELECT condition. For example, in "SELECT name" and "SELECT year", the name and year columns both belong to the person table, so the two SELECT conditions can be merged into one "SELECT name, year". However, arithmetic expressions appearing in the SELECT conditions and non-operation type nodes belonging to different tables cannot be merged.

[0081] In addition, the present embodiment can also transfer the divisible conditions, using transferID: as the transfer identifier of the condition node. The default value of transferID is False. The process will perform the same operation on each node, traversing from the bottom to the top. The specific requirement for the condition node transfer is that if a condition node in the divisible condition does not have a corresponding sibling node, or the sibling node corresponding to the condition node does not include the condition node, then the node identifier of the condition node will be migrated to the corresponding parent node. Specifically, if the current node is identified as a condition node, if the current node has no sibling node or the sibling node of the current node does not contain the condition node, the condition identifier of the current node will be transferred to the parent node, and the process will continue until the transfer requirements are not met or the parent node of the condition identifier is traversed upward to a SELECT node. The purpose of this is to meet the maximum divisible condition and ensure that the generated subSQL syntax is correct. For example Figure 4 The JOIN node with conID 19 has its condition label transferred from the node with conID 18.

[0082] It is understandable that if the conditional node transfer occurs multiple times, the algorithm will use recursion based on the post-order traversal of the binary tree nodes, and only pass the conID of the initial condition. For example, the condition node with conID 19 in the condition identification is transferred from conID: 19 to conID: 20, and conID: 20 is transferred to conID: 21. Then the transferID of conID: 20 and conID: 21 are both 19.

[0083] like Figure 4 As shown in , the condition set is {3,6,9,19,24,29}, and the transferID of the node conID19 is 18, that is, it is transferred from 18, then the condition of conID18 replaces the condition of conID19, that is, the final condition set is {3,6,9,18,24,29}.

[0084] Step S200: Delete several conditions in the condition set based on a preset deletion rule, obtain a subSQL each time a condition is deleted, and obtain a subSQL set corresponding to the original SQL.

[0085] Specifically, this embodiment can summarize all the condition subsets that can be deleted based on the condition set obtained above and the dependency relationship between the conditions in the condition set, and then delete several conditions in the condition set according to the condition subset that can be deleted based on the preset deletion rule, and obtain a subSQL each time a condition is deleted, so that several subSQLs can be obtained, and the subSQL is a subset of the original SQL, thereby forming the subSQL set. In actual applications, when summarizing the condition subset that can be deleted, the dependency relationship between the conditions needs to be considered. When all the dependent conditions are not included in the deleted condition subset, the dependent condition cannot be included in the deleted condition subset. The situation with dependency relationship is shown in Table 2 below.

[0086] Table 2 Two cases of dependency

[0087]

[0088] If there is a dependency between JOIN type conditions, it is necessary to determine whether to delete them when deleting. For example, in SQL involving JOIN type conditions, if other conditions in the SQL contain columns in the JOIN table, then the current condition and the JOIN type condition have a dependency relationship; the nodes with a dependency relationship also include GROUP BY type conditions and HAVING type conditions, and the HAVING type condition can only exist when the GROUP BY type condition exists. Figure 4As can be seen from the AST diagram, conID 3, conID 24, and conID 29 all depend on the node of the JOIN type condition with transferID 18. Because these nodes all contain the tables connected to the nodes of the JOIN type condition, the dependent nodes cannot be deleted if the dependent nodes exist. In order to ensure that the dependent nodes are deleted in the final result, the node with transferID 18 has the opportunity to be added to the deleted condition subset only when conID 3, 24, and 29 are included in the condition subset that can be deleted.

[0089] In addition, when deleting the deletable condition subset, this embodiment sequentially obtains each condition in the deletable condition subset, and sequentially deletes the conditions corresponding to the condition set, and obtains a subSQL for each condition deleted. Then, based on all subSQLs, the subSQL set is obtained. Specifically, in order to ensure that the SQL syntax structure is not destroyed, if the type of the condition obtained from the deletable condition subset is an arithmetic operation type, the node of the arithmetic operation type will involve two nodes, and these two nodes may be condition nodes, so the two condition nodes corresponding to the condition of the arithmetic operation type can be determined, wherein the two condition nodes are brother nodes to each other; when deleting one of the condition nodes, the deletion is achieved by replacing the parent node with the corresponding brother node. If it is deleted directly, a syntax error will occur, for example: "SELECT * FROM person WHERE name = 'jack' and year = '18'", where the condition nodes are "name = 'jack'" and "year = '18'", the two condition nodes are brother nodes to each other, and the parent node is an AND type node. If you delete one of the conditional nodes directly, the resulting SQL will have syntax errors, such as "SELECT * FROM personWHERE and year = '18'". This is the parsed result after deleting "name = 'jack'". It can be found that there is only one condition in the child nodes of the AND type node, so the parent node AND type node needs to be deleted and only year = '18' is retained. In this way, the parsed SQL syntax is correct. Use the remaining sibling nodes to replace the parent node, that is, delete the AND type node, and get the correct SQL "SELECT * FROM person WHERE year = '18'".

[0090] Combined with the contents of step 100 and step 200, an example is given. For example, after the original SQL is parsed and the conditions are identified in the present embodiment, the condition set {3, 6, 9, 18, 24, 29} is obtained, and the FROM type condition with conID 9 is removed from the condition set, indicating that the node with conID 9 is not considered. Then, the condition subsets that can be deleted are summarized, and based on the deletion rule, several conditions in the condition set are deleted according to the condition subsets that can be deleted, and the subSQL set is finally obtained, which includes all possible subSQLs, as shown in Table 3 below.

[0091] Table 3 All possible subSQL

[0092]

[0093] It can be seen that this embodiment can automatically identify and decompose conditions in complex SQL queries, parse out several separable conditions in SQL, and obtain several subSQLs containing some conditions after deleting some conditions, so as to show the dependency relationship and inclusion order between conditions in subsequent reasoning paths.

[0094] Step S300: construct reasoning paths based on the conditions contained in each subSQL in the subSQL set.

[0095] This embodiment can construct a relationship graph based on the relationship between the conditions contained in the subSQLIDs, and then obtain the reasoning path based on the relationship graph. Specifically, Figure 5 As shown in the figure, since each subSQL in the subSQL set is a subset of SQL, first determine the subSQL with the least conditions and the subSQL with the most conditions based on the subSQL set, where the subSQL with the most conditions is the original SQL. Set an ID for each subSQL and record it as subSQLID, then use the subSQL with the least conditions as the starting point of the reasoning path, that is, Figure 5 The subSQL represented by the node with subSQLID 1 has the least conID conditions and is used as the starting point of the reasoning path; the subSQLID with the most conditions is used as the end point of the reasoning path, that is, Figure 5The node with subSQLID 18 in the traversal is taken as the end point. Then, all subSQLs are traversed one by one from the starting point, and if the current subSQL traversed has one more condition than the previous subSQL, the current subSQL is taken as the next starting point, and so on, until the end point is reached. A subSQLID sequence passed during the traversal process is recorded as an inference path to obtain all inference paths.

[0096] Refer to Table 3 and Figure 5 , Reasoning path: gradually add a conID condition based on the conID condition contained in subSQLID No. 1 until the subSQLID sequence that is sequentially experienced when the conID condition contained in the original SQL (subSQLID No. 18) is met. Figure 5 As shown in the figure, the first reasoning path is: the condition contained in subSQLID No. 1 is conID:9, and on this basis, the condition of conID:18 is added to become subSQLID No. 2; the condition of conID:29 is added to become subSQLID No. 4; the condition of conID:24 is added to become subSQLID No. 8; the condition of conID:6 is added to become subSQLID No. 14; the condition of conID:3 is added to become subSQLID No. 18, which is the original SQL. The final reasoning path can be expressed as the sequence of subSQLIDs passed through during the traversal process: [1,2,4,8,14,18].

[0097] By analogy, combined with Table 4 and Figure 6 As shown, 30 reasoning paths can be obtained, and the results are shown in Table 4.

[0098] Table 4 Reasoning path set

[0099]

[0100] It can be seen that, according to the abstract syntax tree (AST) structure of SQL, this embodiment can parse the complex conditions in the original SQL query into several separable conditions (SQL clauses) and several subSQLs (sub-SQLs), and further construct a reasoning path to show the inclusion relationship and dependency order of these conditions in the SQL structure. Specifically, this embodiment can automatically parse the original SQL into several SQL clauses and identify the separable conditions in SQL. The SQL clause set (except the From type clause) obtained after deleting part of the conditions from the original SQL each time is disassembled to obtain the corresponding subSQL set. Then, based on the subSQLs obtained by disassembly, a reasoning path from the minimum condition SQL to the complete SQL is constructed to show the inclusion relationship and dependency order between the conditions. This reasoning path not only provides a clear execution process for SQL debugging, but also provides new ideas for query optimization. In addition, in the condition identification step, the SQL keywords involved are very sufficient, basically covering all SQL syntax, which is more applicable than traditional methods and is suitable for a variety of database query scenarios.

[0101] Based on the above embodiments, the present invention also provides a system for constructing an inference path based on decomposing SQL, such as Figure 6 As shown in, the system includes: a condition set determination module 10, a partial condition deletion module 20 and an inference path construction module 30. The condition set determination module 10 is used to parse the original SQL and identify the separable conditions in the original SQL to obtain a condition set, wherein the conditions are clauses in the original SQL. The partial condition deletion module 20 is used to delete several conditions in the condition set based on preset deletion rules, obtain a subSQL each time a condition is deleted, and obtain the subSQL set corresponding to the original SQL. The inference path construction module 30 is used to construct inference paths based on the conditions contained in each subSQL in the subSQL set.

[0102] The working principles of each module in the system for constructing an inference path based on decomposing SQL in this embodiment are the same as the principles of each step in the above method embodiment, and will not be repeated here.

[0103] Each module in the above-mentioned system for constructing an inference path based on decomposing SQL can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a terminal in the form of hardware, or can be stored in a memory in the terminal in the form of software, so that the processor can call and execute operations corresponding to each of the above modules.

[0104] Based on the above embodiment, the present invention further provides a terminal, the principle block diagram of the terminal can be as follows: Figure 7The terminal may include one or more processors 100 ( Figure 7 Only one is shown in the figure), a memory 101 and a computer program 102 stored in the memory 101 and executable on one or more processors 100, for example, a sleep analysis program based on multi-sensor data. When one or more processors 100 execute the computer program 102, the various steps in the embodiment of the sleep analysis method based on multi-sensor data can be implemented. Alternatively, when one or more processors 100 execute the computer program 102, the functions of various modules / units in the embodiment of the sleep analysis system based on multi-sensor data can be implemented, which is not limited here.

[0105] In one embodiment, the processor 100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0106] In one embodiment, the memory 101 may be an internal storage unit of an electronic device, such as a hard disk or memory of the electronic device. The memory 101 may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 101 may also include both an internal storage unit of the electronic device and an external storage device. The memory 101 is used to store computer programs and other programs and data required by the terminal. The memory 101 may also be used to temporarily store data that has been output or is to be output.

[0107] Those skilled in the art will understand that Figure 7 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the scheme of the present invention, and does not constitute a limitation on the terminal to which the scheme of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0108] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, operating database or other media used in the embodiments provided by the present invention may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double operational data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing an inference path based on decomposing SQL, characterized in that: The method comprises: Parse the original SQL and identify the separable conditions in the original SQL to obtain a condition set, wherein the conditions are clauses in the original SQL; Deleting several conditions in the condition set based on a preset deletion rule, obtaining a subSQL each time a condition is deleted, and obtaining a subSQL set corresponding to the original SQL; Construct reasoning paths based on the conditions contained in each subSQL in the subSQL set; Based on the preset deletion rules, several conditions in the condition set are deleted respectively, and a subSQL is obtained each time a condition is deleted, and the subSQL set corresponding to the original SQL is obtained, including: Based on the dependency relationship between the conditions in the condition set, a condition subset in the condition set that can be deleted is determined.

2. The method for constructing an inference path based on decomposing SQL according to claim 1, characterized in that: Parse the original SQL and identify the separable conditions in the original SQL to obtain a condition set, including: Parse the original SQL to obtain the abstract syntax tree structure; Identify the type of each node in the abstract syntax tree structure and determine the conditions for splitting the abstract syntax tree structure; The determined separable conditions are post-processed to obtain the condition set.

3. The method for constructing an inference path based on decomposing SQL according to claim 2, characterized in that: Identifying the type of each node in the abstract syntax tree structure and determining the conditions for splitting the abstract syntax tree structure include: If the type of a node in the abstract syntax tree structure is an operation type, then identifying the types of all child nodes of the node; If the type of all child nodes is a non-operation type, the node is determined to be a conditional node, and the node and all corresponding child nodes are used as splittable conditions.

4. The method for constructing an inference path based on decomposing SQL according to claim 3, characterized in that: Identifying the type of each node in the abstract syntax tree structure and determining the conditions for splitting the abstract syntax tree structure also includes: If a node in the abstract syntax tree structure is a nested query node, the nested query node is used as a splittable condition.

5. The method for constructing an inference path based on decomposing SQL according to claim 2, characterized in that: Identifying the type of each node in the abstract syntax tree structure also includes: Attribute information is added to each separable conditional node identified from the abstract syntax tree structure, wherein the attribute information includes: an identifier of the conditional node, a condition identifier, and a transfer identifier of the conditional node.

6. The method for constructing an inference path based on decomposing SQL according to claim 2, characterized in that: Post-processing is performed on the determined separable conditions to obtain the condition set, including: Identify the FROM type conditions among the separable conditions, and determine whether to delete the FROM type conditions; Identify the SELECT conditions in the separable conditions and merge the conditions based on the SELECT conditions; If a conditional node in the separable conditions does not have a corresponding sibling node, or the sibling nodes corresponding to the conditional node do not include the conditional node, the conditional node is used as the corresponding parent node.

7. The method for constructing an inference path based on decomposing SQL according to claim 1, characterized in that: Based on the preset deletion rule, several conditions in the condition set are deleted respectively, and a subSQL is obtained by deleting each condition, so as to obtain the subSQL set corresponding to the original SQL, and further comprising: Based on the deletion rule, several conditions in the condition set are deleted according to the subset of conditions that can be deleted, and a subSQL is obtained each time a condition is deleted, and the subSQL set is obtained, wherein the deletion rule is that if the dependent condition is deleted, all conditions that depend on the condition are deleted.

8. The method for constructing an inference path based on decomposing SQL according to claim 1, characterized in that: Determining a subset of conditions in the condition set that can be deleted based on dependencies between conditions in the condition set includes: Determine the dependent condition and the dependent condition in the condition set based on the dependency relationship between the conditions in the condition set; All of the dependent conditions and all of the dependent conditions are taken as a condition subset that can be deleted; wherein, only when all of the dependent conditions are included in the condition subset that can be deleted, the dependent conditions are included in the condition subset that can be deleted.

9. The method for constructing an inference path based on decomposing SQL according to claim 7, characterized in that: Based on the deletion rule, several conditions in the condition set are deleted according to the condition subset that can be deleted, and a subSQL is obtained each time a condition is deleted, and the subSQL set is obtained, including: Obtain each condition in the condition subset that can be deleted in turn, and delete the conditions corresponding to the condition set in turn, and obtain a subSQL each time a condition is deleted; Based on all subSQLs, the subSQL set is obtained.

10. The method for constructing an inference path based on decomposing SQL according to claim 9, characterized in that: Obtaining each condition in the condition subset that can be deleted in turn, and deleting the conditions corresponding to the condition set in turn, including: If the type of the condition obtained from the condition subset that can be deleted is an arithmetic operation type, then two condition nodes corresponding to the condition of the arithmetic operation type are determined, wherein the two condition nodes are sibling nodes of each other; When deleting one of the conditional nodes, the deletion is achieved by replacing the parent node with the corresponding sibling node.

11. The method for constructing an inference path based on decomposing SQL according to claim 1, characterized in that: The reasoning paths are constructed based on the conditions contained in each subSQL in the subSQL set, including: Based on the subSQL set, determine the subSQL with the least conditions and the subSQL with the most conditions, where the subSQL with the most conditions is the original SQL; An inference path is constructed based on the subSQL containing the least conditions and the subSQL containing the most conditions.

12. The method for constructing an inference path based on decomposing SQL according to claim 11, characterized in that: Based on the subSQL with the least conditions and the subSQL with the most conditions, an inference path is constructed, including: The subSQL with the least conditions is used as the starting point of the reasoning path, and the subSQL with the most conditions is used as the end point of the reasoning path; All subSQLs are traversed one by one from the starting point, and if the current subSQL traversed has one more condition than the previous subSQL, the current subSQL is used as the next starting point, and so on, until the end point is reached, and all reasoning paths are obtained.

13. A system for constructing reasoning paths based on decomposing SQL, characterized in that: The system is used to implement the steps of the method for constructing an inference path based on decomposing SQL according to any one of claims 1 to 12, and the system includes: A condition set determination module, used to parse the original SQL and identify the separable conditions in the original SQL to obtain a condition set, wherein the conditions are clauses in the original SQL; A partial condition deletion module is used to delete several conditions in the condition set based on a preset deletion rule, and obtain a subSQL each time a condition is deleted, thereby obtaining a subSQL set corresponding to the original SQL; The reasoning path building module is used to build reasoning paths based on the conditions contained in each subSQL in the subSQL set.

14. A terminal, characterized in that: The terminal includes a memory, a processor, and a program for constructing a reasoning path based on decomposed SQL, which is stored in the memory and can be run on the processor. When the processor executes the program for constructing a reasoning path based on decomposed SQL, the steps of the method for constructing a reasoning path based on decomposed SQL as described in any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program for constructing a reasoning path based on decomposed SQL. When the program for constructing a reasoning path based on decomposed SQL is executed by a processor, the steps of the method for constructing a reasoning path based on decomposed SQL as described in any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Text-to-SQL method and system

    CN114282497A