Query statement detection method and device, equipment, storage medium and product

By converting the abstract syntax tree, execution plan tree, and context information of database query statements into vectors, constructing feature vectors, and using deep learning models to automatically detect the causes of slow queries, the problem of insufficient detection accuracy and adaptability in existing technologies is solved, and efficient slow query diagnosis is achieved.

CN120950531APending Publication Date: 2025-11-14WUHAN DAMENG DATABASE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511055520.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, database administrators rely on experience-based rules to detect slow query statements, which suffers from poor subjectivity, lack of objectivity and universality, resulting in insufficient diagnostic accuracy and consistency, and difficulty in adapting to changes and developments in database systems.

Method used

The abstract syntax tree, execution plan tree, and statement context information of the target query statement are converted into vectors respectively to construct the target feature vector, which is then input into a preset detection model to determine the type of slow cause. The objectivity and universality of the deep learning model are used for automatic detection.

Benefits of technology

It improves the accuracy of slow query detection, reduces subjective bias, quickly adapts to changes and developments in the database system, and lowers the cost of formulating, maintaining, and updating rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950531A_ABST
    Figure CN120950531A_ABST
Patent Text Reader

Abstract

The invention discloses a query statement detection method and device, equipment, a storage medium and a product. The method comprises the steps that an abstract syntax tree corresponding to a target query statement is converted into a first vector, the target query statement is a slow query statement, and the slow query statement is a query statement with the execution time larger than a preset threshold value; converting a target execution plan tree corresponding to the target query statement into a second vector; converting statement context information of the target query statement into a third vector; constructing a target feature vector of the target query statement according to the first vector, the second vector and the third vector; and inputting the target feature vector into a preset detection model, and determining the slow reason type of the target query statement according to the output of the preset detection model. By adopting the technical scheme, the detection accuracy of the slow query statement of the model can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of database technology, and in particular to query statement detection methods, apparatus, devices, storage media, and products. Background Technology

[0002] With the widespread application of databases across various industries, users are paying increasing attention to the stability of database operations. However, due to various internal or external factors, databases may experience performance anomalies during actual operation. Typically, database administrators (DBAs) diagnose slow queries by observing monitoring metrics, but there are hundreds of metrics related to database monitoring, making it difficult for database users or administrators to extract valuable information in a short period of time.

[0003] Currently, to improve efficiency, one detection scheme involves providing query interfaces at both the table and function levels within the database. The table query interface retrieves performance metrics for slow SQL queries, such as execution time and network overhead. The function query interface uses rules developed based on DBA experience to analyze the underlying causes of these performance metrics. The technical approach utilizes Boolean search to logically combine and analyze the diagnostic results from multiple dimensions, thereby pinpointing the root cause of slow SQL queries. However, DBA experience is influenced by subjective factors, potentially leading to rules that are not objective or universally applicable. Different DBAs may have different experiences and preferences, affecting the accuracy and consistency of the diagnosis. Developing, maintaining, and updating diagnostic rules based on DBA experience also requires significant human and time investment. Furthermore, with the development and changes in database systems, diagnostic methods based on fixed rules cannot keep pace with and adapt to new database architectures, application models, and query types. Summary of the Invention

[0004] This invention provides a query statement detection method, apparatus, device, storage medium, and product, which can improve the detection results for query statements.

[0005] According to one aspect of the present invention, a query statement detection method is provided, comprising:

[0006] The abstract syntax tree corresponding to the target query statement is converted into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold.

[0007] Convert the target execution plan tree corresponding to the target query statement into a second vector;

[0008] The statement context information of the target query statement is converted into a third vector;

[0009] Construct the target feature vector of the target query statement based on the first vector, the second vector, and the third vector;

[0010] The target feature vector is input into a preset detection model, and the slow cause type of the target query statement is determined based on the output of the preset detection model.

[0011] According to another aspect of the present invention, a query statement detection device is provided, comprising:

[0012] The first vector conversion module is used to convert the abstract syntax tree corresponding to the target query statement into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold.

[0013] The second vector conversion module is used to convert the target execution plan tree corresponding to the target query statement into a second vector.

[0014] The third vector conversion module is used to convert the statement context information of the target query statement into a third vector;

[0015] A vector construction module is used to construct a target feature vector of the target query statement based on the first vector, the second vector, and the third vector.

[0016] The cause type determination module is used to input the target feature vector into a preset detection model and determine the slow cause type of the target query statement based on the output of the preset detection model.

[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0018] At least one processor; and

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the query statement detection method according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the query statement detection method according to any embodiment of the present invention.

[0022] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the query statement detection method described in any embodiment of the present invention.

[0023] The technical solution of this invention converts the abstract syntax tree corresponding to the target query statement into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold; converts the target execution plan tree corresponding to the target query statement into a second vector; converts the statement context information of the target query statement into a third vector; constructs a target feature vector of the target query statement based on the first vector, the second vector, and the third vector; inputs the target feature vector into a preset detection model, and determines the slow cause type of the target query statement based on the output of the preset detection model. By adopting the above technical solution, the abstract syntax tree, target execution plan tree, and statement context information corresponding to the target query statement are converted into corresponding vectors, and then a target feature vector is constructed based on these vectors to generate a high-quality vector representation. This allows the target feature vector input into the preset detection model to more comprehensively represent the relevant information of the slow query statement, which can effectively improve the detection accuracy of the slow query statement of the model.

[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart of a query statement detection method provided by an embodiment of the present invention;

[0027] Figure 2 This is a flowchart of another query statement detection method provided by an embodiment of the present invention;

[0028] Figure 3 This is a schematic diagram of the structure of a query statement detection device according to an embodiment of the present invention;

[0029] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the query statement detection method of this invention. Detailed Implementation

[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first," "second," "original," and "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0032] Figure 1 This is a flowchart of a query statement detection method according to an embodiment of the present invention. This embodiment is applicable to diagnosing the causes of slow queries. The method can be executed by a query statement detection device, which can be implemented in hardware and / or software. The query statement detection device can be configured in an electronic device, such as a personal computer (PC) or a server. Figure 1 As shown, the method includes:

[0033] Step 101: Convert the abstract syntax tree corresponding to the target query statement into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold.

[0034] For example, the query statement in this embodiment of the invention can be a Structured Query Language (SQL) statement. A slow query statement is one whose execution time exceeds a preset threshold. The preset threshold can be set according to actual conditions; for example, in some systems, the preset threshold is 1 second, while for some high-concurrency systems, it may be 100 milliseconds. It can be set according to the purpose of the database and business requirements. Slow SQL statements generally consume a large amount of database resources, causing other queries to slow down. Long query response times also affect user experience. In high-concurrency scenarios, slow SQL statements may cause database overload or even system crashes. Therefore, when slow SQL statements occur, it is necessary to detect and analyze the reasons for their slowness in order to take targeted improvement measures. The target query statement can be any slow query statement to be detected.

[0035] An Abstract Syntax Tree (AST) is a tree-like structure used to represent the syntactic structure of source code. In SQL queries, the AST parses the query statement into its components, such as the SELECT clause, FROM clause, and WHERE clause. Each component serves as a node in the tree, and the hierarchical relationship between nodes reflects the syntactic structure. Through the AST, a deeper understanding of the SQL query structure can be achieved, which is helpful for semantic analysis, optimization, and transformation. In this step, the Abstract Syntax Tree is converted into vector form, denoted as the first vector (or the AST node vector). The specific conversion method is not limited.

[0036] Step 102: Convert the target execution plan tree corresponding to the target query statement into a second vector.

[0037] For example, an execution plan tree can be understood as an abstract syntax tree of the execution plan. The execution plan tree represents the detailed steps of how a database system executes an SQL query. Operators exist within the execution plan; to easily distinguish them from the nodes in the abstract syntax tree corresponding to the target query statement, these nodes are referred to as tree nodes in this paper. Tree nodes can be associated with operators, which represent the operation type of the database operation, such as SELECT STATEMENT and TABLE ACCESS. The target execution plan tree is the execution plan tree corresponding to the target query statement. In this step, the execution plan tree is converted into a vector form, denoted as the second vector (or AST structure vector). The specific conversion method is not limited.

[0038] Step 103: Convert the statement context information of the target query statement into a third vector.

[0039] For example, even if two SQL statements are identical, they may behave differently in different environments. For instance, database pressure and related statements may affect the execution speed of SQL statements. Therefore, information related to the execution environment of the target query statement can be used as the statement context information of the target query statement, and the statement context information can be converted into a vector form, denoted as the third vector (or context vector). The specific conversion method is not limited.

[0040] Step 104: Construct the target feature vector of the target query statement based on the first vector, the second vector, and the third vector.

[0041] For example, the first vector, the second vector, and the third vector can be concatenated in a certain way, such as sequential concatenation, and the resulting vector can be determined as the target feature vector of the constructed target query statement, that is, SQL vector = AST node vector + AST structure vector + context vector.

[0042] Step 105: Input the target feature vector into the preset detection model, and determine the slow cause type of the target query statement based on the output of the preset detection model.

[0043] For example, the preset detection model can be an artificial intelligence model, such as a deep learning model, specifically a model containing multi-layer neural networks (such as fully connected neural networks). The specific training process for the preset detection model is not limited. The pre-trained preset detection model can be sequentially deployed to the database system, enabling the database system to integrate slow SQL diagnostic capabilities and achieve real-time monitoring and diagnosis of slow SQL queries occurring in the database system.

[0044] For example, after the target feature vector is input into the preset detection model, the preset detection model can output the probability value of the target query statement as various preset slow cause types. Based on the comparison result of the probability value of various preset slow cause types with the corresponding threshold, the slow cause type of the target query statement is determined. For example, the slow cause type with a probability value greater than the corresponding threshold is determined as the slow cause type of the target query statement.

[0045] The query statement detection method of this invention converts the abstract syntax tree corresponding to the target query statement into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold; converts the target execution plan tree corresponding to the target query statement into a second vector; converts the statement context information of the target query statement into a third vector; constructs a target feature vector of the target query statement based on the first vector, the second vector, and the third vector; inputs the target feature vector into a preset detection model, and determines the slow cause type of the target query statement based on the output of the preset detection model. By adopting the above technical solution, the abstract syntax tree, target execution plan tree, and statement context information corresponding to the target query statement are converted into corresponding vectors, and then a target feature vector is constructed based on these vectors to generate a high-quality vector representation. This makes the target feature vector input into the preset detection model more comprehensively represent the relevant information of the slow query statement, which can effectively improve the detection accuracy of the model for slow query statements. In addition, automatically detecting the slow cause of slow query statements through a preset detection model can reduce subjective bias, possess the objectivity and universality of deep learning models, quickly adapt to changes and developments in database systems, and reduce the cost of formulating, maintaining, and updating rules.

[0046] In some embodiments, converting the abstract syntax tree corresponding to the target query statement into a first vector includes: parsing the target query statement into an abstract syntax tree; for each node in the abstract syntax tree, using a preset word embedding model to convert the current node into a corresponding node vector; and concatenating the node vectors corresponding to all nodes in the abstract syntax tree into a first vector. Thus, the word embedding model can effectively represent the semantic information of the query statement, enabling the preset detection model to quickly adapt to changes and developments in the database system.

[0047] For example, by writing SQL syntax rules corresponding to the currently used database, SQL queries can be parsed into an abstract syntax tree (AST) to extract the structural information of the query.

[0048] Optionally, before parsing the target query statement into an abstract syntax tree, sensitive information such as table names and field names can be anonymized to protect data privacy, such as by hashing the latter to map it to the anonymized value.

[0049] For example, suppose the target query is `SELECT xx FROM TABLE employees xx WHERE department xx = xx`, where `xx` represents sensitive information, representing name, employee, department ID, and value in that order. After hash mapping, we get `SELECT name_hash FROM TABLE employees_hash WHERE department_id hash = value_hash`. After parsing, we can obtain the following abstract syntax tree example:

[0050]

[0051] For each node in the abstract syntax tree, such as SELECT_LIST, a pre-defined word embedding model is used to convert the current node into a corresponding node vector. Currently, programs can be created to perform certain tasks and complete communication with machines. Natural Language Processing (NLP) is the machine processing of human language, aiming to teach machines how to process and understand human language, thereby establishing a simple communication channel between humans and machines. The main logic of this technology is as follows: segmenting, stemming, and restoring word forms to basic forms of text data; using deep learning technology, using neural network models to learn the language patterns of text, analyzing the dependencies between words, the phrase structure of sentences, and the sentiment tendency in the text; and using neural network-based machine translation to generate natural and intelligent dialogues using deep learning models. However, NLP has a necessity: it is based on artificial intelligence, and machine learning models and deep learning models are most effective for numerical data. Since numerical data is difficult to generate naturally for natural language, NLP converts text data into numerical data, thus enabling deep learning models to be applied to text data. Existing technologies typically use the bag-of-words model to count the frequency of occurrence of a natural language word and then convert it into numerical form for use in deep learning training. When applied to database query statements, the aforementioned existing technologies have the following drawbacks: 1. Lack of semantic analysis capabilities: SQL is a standardized language used to manage database systems. Unlike natural language, SQL cannot be used for natural language communication between humans and machines. Therefore, in deep learning tasks, SQL does not involve semantic analysis of natural language. 2. Information overload: In SQL, a table name may not only represent the name of the table but may also contain additional information about the table's structure, content, or performance. This increases the complexity of processing and understanding SQL statements, causing deep learning models to encounter difficulties when learning SQL statements. 3. Semantic loss: Slow SQL is analogous to "verbose or ambiguous language expressions" in natural language. Natural language processing can help simplify text and remove word variations through techniques such as stemming. However, for slow SQL, the model cannot fully capture the rich information contained in the query statement, leading to information omission or loss.

[0052] In this embodiment of the invention, word embeddings are used to vectorize the nodes of the abstract syntax tree so that they can be input into the model for processing. This vectorization process is to convert text data into a numerical form that can be processed by a computer so that it can be input into the deep learning model for further processing and analysis.

[0053] For example, a large number of SQL queries can be collected to build a corpus for training. The corpus can then be preprocessed, including removing comments, standardizing table and field names, and word segmentation. Based on the preprocessed corpus, a word embedding model can be trained using algorithms such as Word2Vec. This allows the model to map each word in the SQL query to a vector, resulting in a predefined word embedding model. By mapping words to a continuous vector space, the semantic and syntactic relationships between SQL terms can be captured, providing richer and more meaningful representations.

[0054] For example, there are pointing relationships between the nodes in an abstract syntax tree. After converting each node into a corresponding node vector, the node vectors can be concatenated according to the pointing relationships to obtain a first vector, so that the first vector can express the structural information of the abstract syntax tree.

[0055] In some embodiments, converting the target execution plan tree corresponding to the target query statement into a second vector includes: obtaining the target execution plan tree corresponding to the target query statement; for each tree node in the target execution plan tree, determining the information to be encoded based on the current tree node, the depth information of the current tree node from the root node, and the information of the parent tree node pointed to by the current tree node, and converting the information to be encoded into a tree node vector corresponding to the current tree node; and concatenating the tree node vectors corresponding to all tree nodes in the target execution plan tree into a second vector. This allows for a more accurate representation of the structural information of the execution plan tree.

[0056] For example, the execution plan tree can be determined based on the execution plan table, where tree nodes in the execution plan tree correspond to nodes in the execution plan table, such as one row in the table being one node. The execution plan table can be generated by obtaining the execution plan information of the target query statement. Optionally, the execution plan table may include the abstract syntax tree (STMT) corresponding to the target query statement, the unique identifier of the execution plan (PLAN_ID), the timestamp of generating the execution plan (TIMESTAMP), the type of operation (OPERATION) (such as SELECT STATEMENT, TABLE ACCESS, etc.), the options of the operation (OPTIONS), the name of the object involved in the operation (OBJECT_NAME) (such as table name), the type of object (OBJECT_TYPE) (such as TABLE, INDEX, etc.), the parent step ID of the current step (PARENT_ID), the operation cost estimated by the optimizer (COST), the estimated number of rows returned by the operation (CARDINALITY), the estimated amount of data processed by the operation (BYTES), the estimated CPU cost of the operation (CPU_COST), the estimated I / O cost of the operation (IO_COST), the estimated execution time of the operation (TIME), and remarks (REMARKS), etc. It should be noted that the actual structure of the execution plan table may vary depending on the database version and configuration; the above is only an example.

[0057] For example, when encoding structural information, the structural information of the execution plan tree can be considered, such as the relationships between tree nodes and the depth of tree nodes. An appropriate encoding method (such as positional encoding) can be used to integrate the structural information into the vector representation. For instance, if tree node B is a subtree node of tree node A, and the depth of tree node B is 2, tree node B can be encoded as "node type: tree node B → n-dimensional vector of vector word embedding" + "depth from root node: 2 → mapping to the above vector (positional encoding)" + "parent node type: tree node A → n-dimensional vector word embedding," resulting in a structural vector containing the meaning, structural position, and semantics of the parent tree node B. The above is merely an example; 2 represents the depth information, and the parent tree node information is tree node A. The actual encoding method is not limited; different tree nodes have different depth information, and the depth information of the parent tree node can also be used as the parent tree node information for encoding.

[0058] For example, after determining the tree node vectors corresponding to all tree nodes in the target execution plan tree, the tree node vectors can be concatenated to obtain a second vector. The specific concatenation method is not limited. Optionally, the concatenation can be performed in order from shallow to deep according to the depth information.

[0059] In some embodiments, converting the statement context information of the target query statement into a third vector includes: obtaining the previous query statement of the target query statement, the database load information during the execution of the target query statement, and the user identifier of the user who input the target query statement, to obtain the statement context information of the target query statement; and converting the statement context information into a third vector. This allows the third vector to contain rich context information, enabling the construction of context-aware feature vectors. This allows the model to better perceive the execution context of the target query statement, further improving the model's accuracy in detecting slow query statements.

[0060] The statement context information can be obtained from the execution plan table mentioned above. Database load information may include, for example, CPU utilization, memory usage, and the number of concurrent connections, which can be used to assess the system resource status during query execution. By including user identifiers in the third vector, the model can analyze the user's query history based on the user identifiers, identify common query patterns and behaviors, and assist in personalized optimization.

[0061] For example, the format of statement context information can be:

[0062] | Information Item | Value | Description |

[0063] Previous SQL statement | SELECT * FROM logs | Continuous query of log entries |

[0064] CPU utilization: 90% | High database pressure |

[0065] | Historical average query time| 0.4 seconds| This type of query is inherently slow|

[0066] |User ID |user_1234 |Some users habitually write slow queries|

[0067] The statement context information can be edited as: "vector of the previous SQL statement" + "normalized encoded vector of numerical information" + "user ID vector" to obtain a third vector, which can become the statement context vector. Numerical information includes things like CPU utilization.

[0068] Optionally, the third vector can be a new node pointing to any node in the abstract syntax tree, and thus concatenated with a node vector in the first vector according to the pointing relationship.

[0069] Figure 2 This is a flowchart of another query statement detection method provided by an embodiment of the present invention. This embodiment is an optimization based on the above-mentioned optional embodiments. Figure 2 As shown, the method includes:

[0070] Step 201: Parse the target query statement into an abstract syntax tree.

[0071] Step 202: For each node in the abstract syntax tree, use a pre-defined word embedding model to convert the current node into a corresponding node vector.

[0072] Step 203: Concatenate the node vectors corresponding to all nodes in the abstract syntax tree into the first vector.

[0073] Step 204: Obtain the target execution plan tree corresponding to the target query statement.

[0074] Optionally, this step may further include: obtaining the original execution plan tree corresponding to the target query statement; obtaining the execution plan information corresponding to the target query statement, wherein the execution plan information includes the node context information of the original tree nodes in the original execution plan tree; for each original tree node in the original execution plan tree, creating a subtree node of the current original tree node according to the node context information of the current original tree node to obtain the target execution plan tree.

[0075] Therefore, leaf nodes are added to the original tree nodes in the original execution plan tree, so that the second vector constructed subsequently based on the target execution plan tree can contain the node context information of the original tree nodes, thereby more comprehensively representing the structure and execution environment of the target query statement.

[0076] For example, node context information may include the execution frequency and historical execution time of the original tree node, and may also be obtained from the execution plan table, such as OPTIONS, OBJECT_NAME, COST, CARDINALITY, BYTES, CPU_COST, IO_COST, and TIME mentioned above. Subtree nodes point to their corresponding original tree nodes.

[0077] Step 205: For each tree node in the target execution plan tree, determine the information to be encoded based on the current tree node, the depth information of the current tree node from the root node, and the information of the parent tree node pointed to by the current tree node, and convert the information to be encoded into the tree node vector corresponding to the current tree node.

[0078] Optionally, a pre-defined word embedding model can be used to convert the information to be encoded into a tree node vector corresponding to the current tree node.

[0079] Step 206: Concatenate the tree node vectors corresponding to all tree nodes in the target execution plan tree into a second vector.

[0080] Step 207: Obtain the previous query statement of the target query statement, the database load information during the execution of the target query statement, and the user identifier of the user who entered the target query statement to obtain the statement context information of the target query statement.

[0081] Step 208: Convert the statement context information into a third vector.

[0082] Step 209: Construct the target feature vector of the target query statement based on the first vector, the second vector, and the third vector.

[0083] Optionally, this may also include converting the query features and / or performance metrics of the target query statement into a fourth vector. Accordingly, this step can specifically involve constructing a target feature vector for the target query statement based on the first vector, the second vector, the third vector, and the fourth vector. Thus, by helping the model additionally collect query feature and / or performance metric data related to slow SQL queries, the query statement features can be further enriched, improving the model's understanding of the semantics of slow SQL queries and increasing detection accuracy.

[0084] For example, query characteristics may include the number of conditional statements, the number of subqueries, index usage, sorting, and grouping. Performance metrics may include, for example, query response time, table information (such as table size), index information, lock usage (such as lock overhead), network information (such as network time or network overhead), and wait events.

[0085] For example, an initial feature vector of the target query statement is constructed based on the first vector, the second vector, the third vector, and the fourth vector. The initial feature vector is padded or truncated to obtain a target feature vector of a preset length, so as to ensure that the input vectors have the same length, which is convenient for subsequent neural network processing.

[0086] Step 210: Input the target feature vector into the preset detection model, and determine the slow cause type of the target query statement based on the output of the preset detection model.

[0087] For example, relevant data of historical slow query statements can be collected to build a training dataset for the detection model. The training dataset includes multiple training samples. Each training sample can be classified according to the slow cause that leads to the performance bottleneck, and the sample label corresponding to the training sample can be labeled according to the category. The detection model can be trained using the training dataset to obtain a preset detection model.

[0088] Optionally, optimization suggestions and measures corresponding to each type of slow cause can be pre-written for use by the preset detection model in diagnosis and prediction. The content can be shown in Table 1 below:

[0089] Table 1. Correspondence between slow query reasons and suggestions

[0090]

[0091]

[0092] For example, after determining the type of slowness, it can be matched with the table above and output corresponding optimization suggestions to help users quickly improve the performance of slow SQL queries.

[0093] The query statement detection method provided in this invention employs a pre-defined word embedding model to convert each node in the abstract syntax tree corresponding to the target query statement into a corresponding node vector and constructs an AST node vector. It also converts the tree nodes, depth information, and parent tree node information in the target execution plan tree into tree node vectors and constructs an AST structure vector. Furthermore, it constructs a context vector by including the previous query statement, database load information, and user identifier. By converting the target query statement into a vector representation containing semantic, structural, and contextual information, and using this vector representation as input to the subsequent pre-defined detection model, the SQL query vectorization method based on structured syntax trees and context-aware mechanisms generates a high-quality, rich vector representation containing multi-dimensional information, further effectively improving the detection accuracy of slow query statements.

[0094] In some embodiments, the preset detection model can be continuously optimized and updated, and the detection model can be periodically evaluated and optimized to adapt to changes in the core query patterns of the database system. For example, historical data and real-world cases can be used to verify the model's effectiveness, and the model can be continuously updated based on actual application scenarios and feedback to improve the accuracy and efficiency of diagnosis. Based on the model's performance and business needs, model parameters (such as learning rate, regularization parameters, and sample weights), new features can be added, or the model architecture can be improved to further optimize the model's performance.

[0095] In some embodiments, the processes of collecting training dataset-related data, algorithm integration, model training, and data management can all be centralized on a database server, enabling centralized management, simplifying implementation and maintenance processes, and improving management efficiency. Compared to traditional deep learning methods, database systems can also periodically monitor model performance and metrics, offering advantages such as real-time performance monitoring, dynamic adjustment, and business-driven optimization.

[0096] Figure 3 This is a schematic diagram of a query statement detection device according to an embodiment of the present invention. Figure 3 As shown, the device includes:

[0097] The first vector conversion module 301 is used to convert the abstract syntax tree corresponding to the target query statement into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold.

[0098] The second vector conversion module 302 is used to convert the target execution plan tree corresponding to the target query statement into a second vector.

[0099] The third vector conversion module 303 is used to convert the statement context information of the target query statement into a third vector;

[0100] Vector construction module 304 is used to construct the target feature vector of the target query statement based on the first vector, the second vector and the third vector;

[0101] The cause type determination module 305 is used to input the target feature vector into a preset detection model and determine the slow cause type of the target query statement based on the output of the preset detection model.

[0102] The query statement detection device of this invention converts the abstract syntax tree, target execution plan tree, and statement context information corresponding to the target query statement into corresponding vectors. Then, it constructs a target feature vector based on these vectors to generate a high-quality vector representation. This allows the target feature vector input into the preset detection model to more comprehensively represent the relevant information of the slow query statement, effectively improving the model's accuracy in detecting slow queries. Furthermore, automatically detecting the causes of slow queries using a preset detection model reduces subjective bias, possesses the objectivity and universality of deep learning models, quickly adapts to changes and developments in database systems, and reduces the cost of formulating, maintaining, and updating rules.

[0103] Optionally, the first vector transformation module includes:

[0104] The statement parsing unit is used to parse the target query statement into an abstract syntax tree;

[0105] The node vector transformation unit is used to transform the current node into a corresponding node vector for each node in the abstract syntax tree using a preset word embedding model;

[0106] The first vector concatenation unit is used to concatenate the node vectors corresponding to all nodes in the abstract syntax tree into a first vector.

[0107] Optionally, the first vector transformation module includes:

[0108] The plan tree acquisition unit is used to acquire the target execution plan tree corresponding to the target query statement;

[0109] The tree node vector conversion unit is used to determine the information to be encoded for each tree node in the target execution plan tree based on the current tree node, the depth information of the current tree node from the root node, and the information of the parent tree node pointed to by the current tree node, and convert the information to be encoded into the tree node vector corresponding to the current tree node.

[0110] The second vector concatenation unit is used to concatenate the tree node vectors corresponding to all tree nodes in the target execution plan tree into a second vector.

[0111] Optionally, the plan tree acquisition unit includes:

[0112] The original tree acquisition subunit is used to obtain the original execution plan tree corresponding to the target query statement;

[0113] The plan information acquisition subunit is used to acquire the execution plan information corresponding to the target query statement, wherein the execution plan information includes the node context information of the original tree node in the original execution plan tree;

[0114] The target tree determination subunit is used to create a subtree node of the current original tree node for each original tree node in the original execution plan tree, based on the node context information of the current original tree node, to obtain the target execution plan tree.

[0115] Optionally, the third vector transformation module includes:

[0116] The statement context information acquisition unit is used to acquire the previous query statement of the target query statement, the database load information during the execution of the target query statement, and the user identifier of the user who inputs the target query statement, so as to obtain the statement context information of the target query statement.

[0117] The third vector conversion unit is used to convert the statement context information into a third vector.

[0118] Optionally, the device may also include:

[0119] The fourth vector conversion module is used to convert the query features and / or performance indicator features of the target query statement into a fourth vector;

[0120] The vector construction module is used to construct the target feature vector of the target query statement based on the first vector, the second vector, the third vector, and the fourth vector.

[0121] The query statement detection device provided in this embodiment of the invention can execute the query statement detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0122] Figure 4 A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0123] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0124] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0125] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as query statement detection methods.

[0126] In some embodiments, the query statement detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the query statement detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the query statement detection method by any other suitable means (e.g., by means of firmware).

[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0128] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0129] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0131] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0132] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0133] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the query statement detection method provided in the above embodiments.

[0134] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A query statement detection method, characterized in that, include: The abstract syntax tree corresponding to the target query statement is converted into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold. Convert the target execution plan tree corresponding to the target query statement into a second vector; The statement context information of the target query statement is converted into a third vector; Construct the target feature vector of the target query statement based on the first vector, the second vector, and the third vector; The target feature vector is input into a preset detection model, and the slow cause type of the target query statement is determined based on the output of the preset detection model.

2. The method according to claim 1, characterized in that, The step of converting the abstract syntax tree corresponding to the target query statement into a first vector includes: Parse the target query statement into an abstract syntax tree; For each node in the abstract syntax tree, a preset word embedding model is used to convert the current node into a corresponding node vector; The node vectors corresponding to all nodes in the abstract syntax tree are concatenated into a first vector.

3. The method according to claim 1, characterized in that, The step of converting the target execution plan tree corresponding to the target query statement into a second vector includes: Obtain the target execution plan tree corresponding to the target query statement; For each tree node in the target execution plan tree, the information to be encoded is determined based on the current tree node, the depth information of the current tree node from the root node, and the information of the parent tree node pointed to by the current tree node. The information to be encoded is then converted into a tree node vector corresponding to the current tree node. The tree node vectors corresponding to all tree nodes in the target execution plan tree are concatenated into a second vector.

4. The method according to claim 3, characterized in that, Obtaining the target execution plan tree corresponding to the target query statement includes: Obtain the original execution plan tree corresponding to the target query statement; Obtain the execution plan information corresponding to the target query statement, wherein the execution plan information includes the node context information of the original tree node in the original execution plan tree; For each original tree node in the original execution plan tree, a subtree node of the current original tree node is created based on the node context information of the current original tree node to obtain the target execution plan tree.

5. The method according to claim 1, characterized in that, The step of converting the statement context information of the target query statement into a third vector includes: Obtain the previous query statement of the target query statement, the database load information during the execution of the target query statement, and the user identifier of the user who input the target query statement to obtain the statement context information of the target query statement; The statement context information is converted into a third vector.

6. The method according to claim 1, characterized in that, Also includes: The query features and / or performance metrics of the target query statement are converted into a fourth vector; The step of constructing the target feature vector of the target query statement based on the first vector, the second vector, and the third vector includes: The target feature vector of the target query statement is constructed based on the first vector, the second vector, the third vector, and the fourth vector.

7. A query statement detection device, characterized in that, include: The first vector conversion module is used to convert the abstract syntax tree corresponding to the target query statement into a first vector, wherein the target query statement is a slow query statement, and the slow query statement is a query statement whose execution time is greater than a preset threshold. The second vector conversion module is used to convert the target execution plan tree corresponding to the target query statement into a second vector. The third vector conversion module is used to convert the statement context information of the target query statement into a third vector; A vector construction module is used to construct a target feature vector of the target query statement based on the first vector, the second vector, and the third vector. The cause type determination module is used to input the target feature vector into a preset detection model and determine the slow cause type of the target query statement based on the output of the preset detection model.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the query statement detection method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the query statement detection method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the query statement detection method according to any one of claims 1-6.