Data processing method and device, electronic equipment, storage medium and program product
By integrating initial encoding and position encoding of the planning nodes in the database query plan tree, the problem of poor database query encoding effect is solved, and the learning effect of parameter adjustment model and parameter adjustment effect of database parameters is improved.
Patent Information
- Application Number
- CN202510168932.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, the encoding effect of database query is poor, which affects the learning effect of parameter adjustment model, resulting in poor database parameter adjustment effect.
By obtaining the query plan tree of the database, initial encoding and location encoding are performed for each planning node, and the two are fused to obtain the target encoding result, which is used to adjust the configuration parameters of the database.
It improves the coding effect of the database query plan, enhances the learning effect of the parameter adjustment model on the correlation relationship between database query, parameters and performance, thereby improving the parameter adjustment effect of the database parameters.
Smart Images

Figure CN120011345A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a data processing method, device, electronic device, storage medium, and program product. Background Art
[0002] In the process of adjusting the parameters of the database, the query and parameters of the database are usually sampled first. Then, the query, parameter and corresponding database performance information are combined to form training data. Then, the query is encoded so that the parameter adjustment model can learn the relationship between the database query, parameter and database performance to obtain the adjusted parameters. In the related art, due to the poor encoding effect of the database query, the learning effect of the parameter adjustment model is affected, resulting in poor database parameter adjustment effect. Summary of the invention
[0003] In view of this, the present disclosure provides a data processing method, device, electronic device, storage medium and program product to solve the coding problem of database query.
[0004] In a first aspect, the present disclosure provides a data processing method, the method comprising:
[0005] Obtaining a query plan tree of a database, wherein the query plan tree includes at least one plan node;
[0006] For each plan node in the query plan tree, encoding the node attribute of the plan node to obtain an initial code of the plan node;
[0007] Based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree, the plan node is encoded to obtain a position code of the plan node;
[0008] The initial coding and position coding of the plan node are merged to obtain the coding result of the plan node, so as to obtain the target coding result of the query plan tree, and the target coding result is used to adjust the configuration parameters of the database.
[0009] In a second aspect, the present disclosure provides a data processing device, the device comprising:
[0010] A data acquisition module, used to acquire a query plan tree of a database, wherein the query plan tree includes at least one plan node;
[0011] A first encoding module, used for encoding the node attributes of each plan node in the query plan tree to obtain an initial code of the plan node;
[0012] A second encoding module, configured to encode the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain a position code of the plan node;
[0013] The data fusion module is used to fuse the initial coding and position coding of the plan node to obtain the coding result of the plan node to obtain the target coding result of the query plan tree, and the target coding result is used to adjust the configuration parameters of the database.
[0014] In a third aspect, the present disclosure provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the data processing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the data processing method of the first aspect or any corresponding embodiment thereof.
[0016] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to enable a computer to execute the data processing method of the first aspect or any corresponding embodiment thereof.
[0017] The data processing method provided by the embodiment of the present disclosure, after encoding the node attributes of the plan node in the query plan tree to obtain the initial code of the plan node, further encodes the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain the position code of the plan node, so that the structural information of the query plan tree and the dependency relationships such as the connectivity between the plan nodes can be retained by using the position code. Then, the initial code and the position code of the plan node are merged to obtain the coding results of the plan node and the query plan tree, so that the coding effect of the query plan of the database can be effectively improved, thereby effectively improving the learning effect of the parameter adjustment model on the relationship between the query, parameters, and database performance of the database, and then improving the parameter adjustment effect of the database parameters.
[0018] The beneficial effects of the data processing device, the electronic device, the storage medium, and the program product correspond to the beneficial effects of the data processing method and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the specific embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0020] Figure 1 is a flowchart of a data processing method according to an embodiment of the present disclosure;
[0021] Figure 2 is a flowchart of another data processing method according to an embodiment of the present disclosure;
[0022] Figure 3 is a flowchart of another data processing method according to an embodiment of the present disclosure;
[0023] Figure 4 is a coding schematic diagram of a database query according to an embodiment of the present disclosure;
[0024] Figure 5 is a structural block diagram of a data processing device according to an embodiment of the present disclosure;
[0025] Figure 6 is a structural block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0027] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, scope of use, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0028] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0029] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0030] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0031] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0032] There are many adjustable configuration parameters (also known as system parameters) in the database. Database operation and maintenance personnel usually rely on their own experience to tune these parameters, which takes a long time. Therefore, in related technologies, a learning-based parameter tuning method is proposed to automatically tune parameters. This parameter tuning method can effectively reduce the time required for parameter tuning and improve the tuning quality.
[0033] In the learning-based parameter adjustment method, the query plan and parameters of the database are usually sampled first. Then, the query plan, parameters and corresponding database performance information are combined to form training data. Furthermore, the query plan is encoded so that the parameter adjustment model can learn the relationship between the database query plan, parameters and database performance to obtain the adjusted parameters. Among them, the encoding effect of the query plan will greatly affect the learning effect of the parameter adjustment model. If the encoding accuracy of the query plan is insufficient, it is difficult for the parameter adjustment model to accurately learn the relationship between the encoded query plan, parameters and database performance, thereby affecting the quality of the recommended parameters and resulting in poor database parameter adjustment. It can be seen that the encoding of the query plan is a key link in the entire parameter adjustment process.
[0034] Regarding the encoding of query plans, related technologies have proposed an encoding method based on query text. However, this encoding method has limitations. It is difficult to cover the information of the query execution process, which easily leads to the lack of fine-grained effective information in the encoding results. The query execution process includes many key factors, such as execution order, data flow, and processing of intermediate results. If these key factors are lacking, it will make it difficult for the encoding results to fully reflect the actual situation of the query, which may cause the subsequent parameter adjustment process to be inaccurate or invalid.
[0035] Therefore, some schemes for encoding query plans have been proposed in the related art. However, these schemes are difficult to accurately encode the node information in the query plan, such as query operation, table name, and cost information. In addition, since the query plan is a tree-shaped result, some existing schemes for encoding sequential structures are difficult to encode the tree-shaped query plan, thereby ignoring the positional relationship between the query plan nodes, resulting in poor encoding effect of the query plan.
[0036] In view of this, according to an embodiment of the present disclosure, a data processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0037] In this embodiment, a data processing method is provided, which can be used in a database management system. Figure 1 is a flow chart of a data processing method according to an embodiment of the present disclosure, such as Figure 1 As shown, the process includes the following steps:
[0038] Step S101, obtaining a query plan tree of a database, where the query plan tree includes at least one plan node.
[0039] The query plan of a database is a tree structure, which presents the logical steps and operation sequence of database query execution in a tree form. The plan node as the root node in the query plan tree represents the starting point of the entire query plan, and the plan node as the leaf node is associated with the table or data source in the database. Each plan node in the query plan tree corresponds to a specific operator, such as a scan operator, a join operator, an aggregation operator, etc. Among them, the scan operator is used to read data from the table, the join operator is responsible for associating data from different tables according to specified conditions, and the aggregation operator is used to summarize and calculate the data.
[0040] Through the query plan tree, the database system can clearly plan the query execution path and determine the execution order and method of each query operation according to the optimization strategy to achieve efficient data retrieval and processing, thereby meeting the query requirements while minimizing resource consumption and execution time, improving the overall performance of the database, and ensuring the accuracy and completeness of data processing.
[0041] Step S102, for each plan node in the query plan tree, encode the node attribute of the plan node to obtain the initial code of the plan node.
[0042] In the complex system of analytical queries, the node attributes of the plan nodes of the query plan tree are extremely critical to the processing and optimization of the entire query in the database.
[0043] The node attributes of the plan node include at least one of the operator type, the data source involved, the predicate, the cardinality, and the cost estimate. The operator type is used to clarify the operation category performed in the query execution process, such as data scanning, data filtering, etc. The data source includes at least one of a table and a column. The data source is used to locate the source and scope of the data. The table is used by the database system to determine the data reading path, and the column is used by the database system to determine the processing object. The predicate is the key basis for filtering data and is presented in the form of a specific triple. The triple includes a column element (such as a column name), a comparison operator, and a value element. The comparison operator is used to define the conditional rules for data filtering. The comparison operators include "greater than", "less than", "equal to", etc. The cardinality and cost estimate are used to provide a quantitative reference for the evaluation and optimization of the query plan from the perspective of data volume and execution overhead.
[0044] In practical applications, the encoding of different node attributes can be designed based on the analytical query engine of the database and the characteristics of the workload environment in which the engine is located, so as to encode different node attributes more effectively.
[0045] Step S103, encoding the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain the position code of the plan node.
[0046] It is worth noting that after successfully obtaining the initial encoding of each plan node, the self-attention mechanism needs to be used to encode the entire query plan. In this process, in order to effectively control and reduce the computational cost, only the attention scores between the connected plan nodes can be calculated, thereby reducing unnecessary computational overhead to a certain extent and improving the encoding efficiency. However, when encoding the query plan, how to properly retain the query plan's own organizational information and the intricate relationships between plan nodes is an urgent problem to be solved.
[0047] Since the standard self-attention mechanism is carefully designed and developed for sequential data, the data processing logic and structural characteristics of the self-attention mechanism determine that it is difficult to directly apply it to query plan scenarios with tree-like structure characteristics (i.e., query plan trees). Tree-like query plans have their unique hierarchical, branching, and complex association patterns between plan nodes, which are significantly different from sequential data in nature.
[0048] Therefore, in order to solve the problem of self-attention mechanism, the data processing method of the present disclosure introduces a coding scheme of hierarchical spectral position encoding (HSPE). The core purpose of hierarchical spectral position encoding is to cleverly incorporate the location information of the plan node into the entire coding system of the query plan tree comprehensively and accurately. Among them, hierarchical spectral position encoding combines hierarchical coding and spectral coding.
[0049] Specifically, the plan node is hierarchically encoded based on the hierarchical position of the plan node in the query plan tree to obtain a first encoding of the plan node. Based on the connectivity between the plan nodes in the query plan tree, the query plan tree is spectrally encoded to obtain a second encoding. The first encoding and the second encoding of the plan node are fused to obtain a position encoding of the plan node.
[0050] It is worth mentioning that this hierarchical spectral position coding method can organically combine the local neighborhood information of the query plan tree with the entire long-range dependency of the query plan tree, thereby effectively improving the accuracy, completeness and effectiveness of the plan nodes and even the query plan tree encoding, laying a solid and reliable foundation for a series of subsequent query plan optimization, execution, data processing and other tasks.
[0051] Step S104, the initial coding and position coding of the plan node are merged to obtain the coding result of the plan node, so as to obtain the target coding result of the query plan tree, and the target coding result is used to adjust the configuration parameters of the database.
[0052] Specifically, the initial coding and position coding of the plan node are concatenated to obtain the coding result of the plan node.
[0053] Furthermore, the initial coding and position coding of the planning node are fused by the following formula: n' i =n i ||p i ; where n' i is the encoding result of the i-th plan node, n i is the initial code of the i-th planning node, p i Encode the position of the i-th plan node.
[0054] After obtaining the target encoding result of the query plan tree, the target encoding result is used as the input of the self-attention module of the database. The self-attention mechanism is used in the self-attention module to process the target encoding result and obtain the node embedding of the plan node in the query plan tree. The length of the node embedding is equal to the number of plan nodes in the query plan tree. The query plan of the database is adjusted using the node embedding.
[0055] The data processing method provided in this embodiment, after encoding the node attributes of the plan node in the query plan tree to obtain the initial code of the plan node, further encodes the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain the position code of the plan node, so that the structural information of the query plan tree and the connectivity between the plan nodes and other dependencies can be retained by using the position code. Then, the initial code and the position code of the plan node are merged to obtain the coding results of the plan node and the query plan tree, so that the coding effect of the query plan of the database can be effectively improved, thereby effectively improving the learning effect of the parameter adjustment model on the relationship between the query, parameters, and database performance of the database, and then improving the parameter adjustment effect of the database parameters.
[0056] In this embodiment, another data processing method is provided, which can be used in a database management system. Figure 2 is a flow chart of another data processing method according to an embodiment of the present disclosure, such as Figure 2 As shown, the process includes the following steps:
[0057] Step S201, obtaining a query plan tree of the database, the query plan tree including at least one plan node. Please refer to the above step S101, which will not be described in detail here.
[0058] Step S202, for each plan node in the query plan tree, encode the node attribute of the plan node to obtain the initial code of the plan node.
[0059] Specifically, the above step S202 includes:
[0060] Step S2021, performing discrete feature encoding on the discrete attributes of the plan node to obtain the encoding result of the discrete attributes.
[0061] Among them, discrete attributes include attributes whose value range is relatively static and limited.
[0062] Optionally, the discrete attribute includes at least one of an operator type, a data source involved, a column element of a predicate, and a comparison operator of a predicate. In practical applications, the discrete attribute may be adjusted according to actual conditions.
[0063] Discrete feature coding is used to convert discrete elements into a coded form that is easy for computers to process.
[0064] Optionally, the discrete feature encoding is one-hot encoding. In addition, the discrete feature encoding can also be label encoding, binarization, and other encodings for discrete attributes, which are not limited here.
[0065] It is understandable that since the value ranges of discrete attributes such as operator types and data sources involved are relatively static and limited in a relatively stable and specific environment, the use of encoding methods such as one-hot encoding for these discrete attributes can concisely and efficiently convert these elements with limited categories of values into an encoding form that is easy for computers to process.
[0066] Among them, for the column elements of the predicate, discrete feature encoding is performed according to their roles and attributes in the query environment. For the comparison operators of the predicate, discrete feature encoding is performed according to their logical meanings.
[0067] Step S2022, normalize and encode the continuous attributes of the plan node to obtain the encoding result of the continuous attributes.
[0068] Among them, continuous attributes include attributes with a wide and continuous value range.
[0069] Optionally, the continuous attribute includes at least one of a cardinality, a cost estimate, and a value element of a predicate. In practical applications, the continuous attribute may be adjusted according to actual conditions.
[0070] Understandably, since the value range of continuous attributes is relatively wide and continuous, in order to facilitate the processing and comparison of continuous attributes in a unified coding system, a normalization algorithm is used to map the value range of continuous attributes to a preset value range, for example, [0,1]. Therefore, through normalized coding, these continuous attributes can work better with other coding elements (such as discrete attributes) in the subsequent calculation and analysis process.
[0071] Among them, the value element of the predicate is normalized and encoded according to its data type and value range.
[0072] Step S2023, the encoding results of the node attributes are merged to obtain the initial encoding of the planned node.
[0073] After completing the individual encoding of each element of the predicate, the encoding results of each element of the predicate are connected to obtain the encoding result of the predicate.
[0074] For each plan node, after obtaining the encoding results of each node element, the encoding results of each node attribute are connected according to the target connection order and rules to obtain a single vector as the initial encoding of the plan node. Among them, the initial encoding is used to represent the unique identification and encoding form of the plan node.
[0075] Step S203, based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree, encode the plan node to obtain the position code of the plan node. Please refer to the above step S103, which will not be described in detail here.
[0076] Step S204, the initial coding and position coding of the plan node are merged to obtain the coding result of the plan node, so as to obtain the target coding result of the query plan tree, and the target coding result is used to adjust the configuration parameters of the database. Please refer to the above step S104, which will not be described in detail here.
[0077] The data processing method provided in this embodiment performs discrete feature encoding on the discrete attributes of the plan nodes based on the value range characteristics of the node attributes, so that the discrete attributes can be succinctly and efficiently converted into a coding form that is easy for the computer to process, thereby improving the coding efficiency. The continuous attributes of the plan nodes are normalized and encoded, so that the continuous attributes and discrete attributes can be easily processed and compared in a unified coding system, so as to better coordinate the continuous attributes and discrete attributes for processing in the subsequent calculation and analysis process.
[0078] In some optional implementations, the above step S2023 includes:
[0079] Step a1, obtaining a target connection sequence, where the target connection sequence includes a connection sequence between multiple preset attributes, and the multiple preset attributes include node attributes.
[0080] It can be understood that the preset attributes include all node attributes involved in the planned node in the database query.
[0081] Step a2, determining the missing attributes of the plan node based on the attributes other than the node attributes of the plan node in the plurality of preset attributes.
[0082] Understandably, individual plan nodes do not contain specific information, such as predicates or join conditions, due to their relatively simple or special query logic. Therefore, it is necessary to find the missing attributes of the plan nodes through preset attributes to fill in the encoding results of the missing attributes, thereby ensuring the consistency of the dimensions represented by all plan nodes in the entire query plan tree.
[0083] Step a3, merging the encoding results of the node attributes based on the target connection order. During the fusion process, if there are missing attributes in the planned node, the encoding results of the missing attributes are filled with preset values to obtain the initial encoding of the planned node.
[0084] Optionally, the preset value is zero. In addition, the preset value may be adjusted according to actual conditions.
[0085] In practical applications, the encoding length of the encoding result of each node attribute can be preset, and the encoding result of the missing attribute can be filled according to the preset value and the encoding length to obtain the initial encoding of the planned node.
[0086] The data processing method provided in this embodiment uses the connection order between multiple preset attributes to fuse the encoding results of the node attributes for the plan node, so that the format of the initial encoding of each plan node can be ensured to be consistent. In the fusion process, the encoding results of the missing attributes of the plan node are filled with preset values to obtain the initial encoding of the plan node, so that the initial encoding dimensions of all plan nodes in the query plan tree can be ensured to be consistent. Furthermore, whether it is a complex plan node or a simple plan node, they can participate in a series of operation processes such as subsequent query plan processing, optimization, and data transmission in a unified dimensional form at the data structure level, effectively improving the stability, efficiency, and accuracy of the entire query processing system.
[0087] In this embodiment, another data processing method is provided, which can be used in a database management system. Figure 3 is a flow chart of another data processing method according to an embodiment of the present disclosure, such as Figure 3 As shown, the process includes the following steps:
[0088] Step S301, obtaining a query plan tree of the database, the query plan tree including at least one plan node. Please refer to the above step S101, which will not be described in detail here.
[0089] Step S302, for each plan node in the query plan tree, encode the node attribute of the plan node to obtain the initial code of the plan node. Please refer to the above step S202, which will not be described in detail here.
[0090] Step S303, encoding the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain the position code of the plan node.
[0091] Specifically, the above step S303 includes:
[0092] Step S3031, based on the hierarchical position of the plan node in the query plan tree, hierarchically encode the plan node to obtain a first code of the plan node.
[0093] Optionally, based on a breadth-first search (BFS) algorithm and the hierarchical position of the plan node, the plan node is hierarchically encoded to obtain a first encoding of the plan node.
[0094] In addition, you can also adaptively select algorithms for hierarchical coding, such as Depth-First Search (DFS) algorithm and Level Order Traversal (Level Order Traversal) algorithm, to perform hierarchical coding on the plan nodes to obtain the first coding of the plan nodes, without any limitation here.
[0095] Step S3032: Based on the connectivity between the plan nodes in the query plan tree, perform spectrum coding on the query plan tree to obtain a second code.
[0096] Optionally, based on the Laplace eigenvector of the query plan tree and the connectivity between plan nodes in the query plan tree, spectral encoding is performed on the query plan tree to obtain a second encoding.
[0097] Among them, the Laplace eigenvector can deeply explore and reflect the global properties of the query plan tree. The Laplace eigenvector not only focuses on the relationship between a single plan node or a local plan node, but also focuses on the macroscopic structural characteristics of the entire query plan tree. For example, the Laplace eigenvector can accurately capture the connectivity pattern between the large branches or main branches of the query plan tree. This connectivity pattern contains important global structural information such as the degree of connection between plan nodes, information transmission paths, and data flow. By analyzing the Laplace eigenvector, we can fully understand the structural characteristics and information interaction patterns of the entire query plan tree at the macro level.
[0098] In addition, breadth-first search, depth-first search, etc. can also be used to determine the connected components of the query plan tree, assign a unique code to each connected component, and assign relative codes to the plan nodes in the connected components to obtain a second code. The spectral coding method of the query plan tree can be adjusted according to actual conditions and is not limited here.
[0099] Step S3033, fusing the first code and the second code of the planning node to obtain the position code of the planning node.
[0100] Specifically, the first code and the second code of the plan node are spliced to obtain the position code of the plan node.
[0101] Furthermore, the first code and the second code of the plan node are fused by the following formula: Among them, p i is the position code of the i-th planning node, is the first code of the i-th plan node, p Lap For the second code.
[0102] Step S304, the initial coding and position coding of the plan node are merged to obtain the coding result of the plan node, so as to obtain the target coding result of the query plan tree, and the target coding result is used to adjust the configuration parameters of the database. Please refer to the above step S104, which will not be described in detail here.
[0103] The data processing method provided in this embodiment uses the hierarchical position of the plan node in the query plan tree to perform hierarchical coding to obtain the first code of the plan node. Therefore, the first code can be used to capture the hierarchical structural information of the query plan tree and accurately locate the hierarchical position of each plan node. In addition, the connectivity between the plan nodes in the query plan tree is used to perform spectrum coding to obtain the second code. Therefore, the second code can be used to capture the structural information and information interaction mode of the query plan tree at the macro level, so that the fused position code can comprehensively and accurately reflect the position information of the plan node.
[0104] In some optional implementations, the above step S2031 includes:
[0105] Step b1, obtaining the depth information of the plan node in the query plan tree, where the depth information is used to characterize the hierarchical position of the plan node in the query plan tree.
[0106] Specifically, starting from the plan node that is the root node in the query plan tree, the plan nodes in the query plan tree are traversed and visited in sequence according to the hierarchical order to obtain the depth information of the plan node in the query plan tree.
[0107] It can be understood that the above traversal access method can clearly capture the hierarchical structural characteristics of the query plan tree and accurately determine the hierarchical position of each plan node, that is, the distance between the plan node and the root node, thereby providing important structural basic information for subsequent encoding operations.
[0108] Step b2: hierarchically encode the plan nodes in the query plan tree based on the depth information to obtain a first code of the plan node.
[0109] Specifically, when the breadth-first search algorithm is used for encoding, the first encoding of the plan node represents the breadth-first search distance from the root node to the plan node, and the breadth-first search distance captures the hierarchical information of the plan node in the query plan tree.
[0110] The data processing method provided in this embodiment performs hierarchical encoding based on the depth information of the plan node in the query plan tree. Therefore, the first code obtained by encoding can provide the structural information of the plan node.
[0111] In some optional implementations, the above step S2032 includes:
[0112] Step c1, constructing a target matrix of the query plan tree, where the target matrix is used to characterize the connectivity between plan nodes in the query plan tree.
[0113] Among them, the Laplace matrix can indirectly or directly reflect the connectivity between plan nodes to capture the global structure of the query plan tree, and the computational complexity of the Laplace matrix is linearly related to the number of plan nodes. It is more efficient to construct the Laplace matrix of the query plan tree and calculate the eigenvalues and eigenvectors of the Laplace matrix. Therefore, the target matrix can be obtained by constructing the Laplace matrix of the query plan tree. In addition, other matrices that can be used to characterize the connectivity between plan nodes in the query plan tree can also be selected, which is not limited here.
[0114] Step c2, operate on the target matrix to obtain at least one eigenvector and an eigenvalue corresponding to the eigenvector, the eigenvector is used to characterize the connectivity between the planned nodes, and the eigenvalue is used to characterize the frequency of occurrence of the corresponding connectivity.
[0115] It should be noted that the eigenvalues corresponding to the eigenvectors are all non-negative numbers. The eigenvectors of the target matrix can refer to the relevant description of the Laplace eigenvectors mentioned above, and will not be elaborated here.
[0116] Step c3, screening the feature vector based on the feature value to obtain the target feature vector.
[0117] Specifically, the above step c3 includes: determining a non-zero and minimum target eigenvalue among all eigenvalues; and taking a eigenvector corresponding to the target eigenvalue as a target eigenvector.
[0118] Among them, the target feature vector captures the most important global structure in the query plan tree.
[0119] The target eigenvalue is also called algebraic connectivity. The target eigenvalue is used to reflect the connectivity of the plan nodes in the query plan tree. If the target eigenvalue is greater than zero, it means that the plan nodes of the query plan tree are connected. The larger the target eigenvalue, the better the connectivity of the query plan tree.
[0120] Step c4, fusing the target feature vector to obtain a second code.
[0121] Specifically, if the target eigenvalue is associated with k eigenvectors, such as v1, v2, ..., v k , then the spectrum coding of the second coding corresponding to the i-th planning node is represented by the following formula: in, is the spectrum encoding of the i-th planning node, v 1i is the value of the first target feature vector v1 at the i-th planning node, v 2iis the value of the second target feature vector v2 at the i-th planning node, v ki is the kth target feature vector v k The value at the i-th plan node.
[0122] It can be obtained that the position encoding of the i-th plan node is
[0123] The data processing method provided in this embodiment constructs a target matrix for characterizing the connectivity between plan nodes in a query plan tree, and calculates the eigenvectors and eigenvalues of the target matrix. Therefore, the eigenvectors can be used to reflect the connectivity between all plan nodes in the query plan tree, and the eigenvalues can be used to reflect the frequency of occurrence of these connectivity modes. Thus, the most important global structure in the query plan tree can be quickly determined using the target eigenvector corresponding to the non-zero and minimum target eigenvalue. Therefore, the second code obtained based on the fusion of the target eigenvector can be used to provide important global structural information such as the degree of connection between plan nodes in the query plan tree, the information transmission path, and the data flow direction, so as to fully understand the structural characteristics and information interaction mode of the entire query plan tree at the macro level.
[0124] As a specific example, Figure 4 As shown, in the data processing method disclosed in the present invention, the encoding method of database query mainly includes the following contents: obtaining the query plan tree of the database, encoding the node attributes of the plan node in the query plan tree, and obtaining the initial encoding of the plan node, such as n1, n2, ..., n m Perform hierarchical spectral position encoding on the query plan tree to obtain the position encoding of the plan node, such as p1, p2, ..., p m Then, for each plan node, the initial code and position code of the plan node are concatenated to obtain the coding result of the plan node, such as n'1, n'2, ..., n' m The encoding results of the plan nodes are processed using the self-attention mechanism to obtain the node embedding of the plan nodes to obtain the target encoding results of the query plan tree.
[0125] It is worth noting that in the data processing method disclosed in the present invention, a coding scheme for database query plans based on hierarchical spectral position coding is proposed, which can not only encode detailed information such as operator type, table name, cost (such as cardinality, cost estimate) of the plan nodes of the query plan, but also fully encode the structural information of the query plan, that is, the positional relationship between the plan nodes. Therefore, it can effectively improve the coding effect of database queries, so that the parameter adjustment model can better learn the relationship between database queries, parameters, and database performance, thereby improving the parameter adjustment effect of database parameters.
[0126] As a specific application example, a target program is installed on a database management system, and the target program is used to execute the data processing method of the present disclosure. The database management system can encode the query plan tree of the database through the execution of the target program, so as to adjust the configuration parameters of the database using the target encoding result of the query plan tree.
[0127] In the present embodiment, a data processing device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0128] This embodiment provides a data processing device, such as Figure 5 As shown, including:
[0129] The data acquisition module 501 is used to acquire a query plan tree of a database, where the query plan tree includes at least one plan node;
[0130] A first encoding module 502 is used to encode the node attributes of each plan node in the query plan tree to obtain an initial code of the plan node;
[0131] A second encoding module 503 is used to encode the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain a position code of the plan node;
[0132] The data fusion module 504 is used to fuse the initial coding and position coding of the plan node to obtain the coding result of the plan node to obtain the target coding result of the query plan tree. The target coding result is used to adjust the configuration parameters of the database.
[0133] In some optional implementations, the node attribute of the plan node includes at least one of a discrete attribute and a continuous attribute. The first encoding module 502 includes:
[0134] The first attribute encoding unit is used to perform discrete feature encoding on the discrete attributes of the plan node to obtain the encoding result of the discrete attributes;
[0135] The second attribute encoding unit is used to perform normalized encoding on the continuous attributes of the plan node to obtain an encoding result of the continuous attributes;
[0136] The first fusion unit is used to fuse the encoding results of the node attributes to obtain the initial encoding of the planned node.
[0137] In some optional implementations, the encoding fusion unit includes:
[0138] A sequence acquisition subunit, used to acquire a target connection sequence, the target connection sequence includes a connection sequence between a plurality of preset attributes, the plurality of preset attributes including a node attribute;
[0139] A data operation subunit, used for determining the missing attributes of the plan node based on the attributes other than the node attributes of the plan node among the plurality of preset attributes;
[0140] The coding fusion subunit is used to fuse the coding results of node attributes based on the target connection order. During the fusion process, if there are missing attributes in the planned node, the coding results of the missing attributes are filled with preset values to obtain the initial coding of the planned node.
[0141] In some optional implementations, the second encoding module 503 includes:
[0142] A hierarchical encoding unit, used to perform hierarchical encoding on the plan node based on the hierarchical position of the plan node in the query plan tree to obtain a first code of the plan node;
[0143] A spectrum encoding unit, configured to perform spectrum encoding on the query plan tree based on connectivity between plan nodes in the query plan tree to obtain a second code;
[0144] The second fusion unit is used to fuse the first code and the second code of the planning node to obtain the position code of the planning node.
[0145] In some optional implementations, the hierarchical coding unit includes:
[0146] A depth acquisition subunit is used to acquire the depth information of the plan node in the query plan tree, and the depth information is used to characterize the hierarchical position of the plan node in the query plan tree;
[0147] The hierarchical encoding subunit is used to hierarchically encode the plan nodes in the query plan tree based on the depth information to obtain the first code of the plan node.
[0148] In some optional implementations, the spectrum encoding unit includes:
[0149] A matrix construction subunit, used to construct a target matrix of the query plan tree, where the target matrix is used to characterize the connectivity between plan nodes in the query plan tree;
[0150] A matrix operation subunit, used to operate the target matrix to obtain at least one eigenvector and an eigenvalue corresponding to the eigenvector, wherein the eigenvector is used to characterize the connectivity between the plan nodes, and the eigenvalue is used to characterize the occurrence frequency of the corresponding connectivity;
[0151] A data screening subunit, used for screening the feature vectors based on the feature values to obtain the target feature vectors;
[0152] The vector fusion subunit is used to fuse the target feature vector to obtain the second code.
[0153] In some optional implementations, the data screening subunit is specifically used to: determine a non-zero and minimum target eigenvalue among all eigenvalues; and use a eigenvector corresponding to the target eigenvalue as a target eigenvector.
[0154] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0155] The data processing device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0156] The present disclosure also provides an electronic device having the above Figure 5 The data processing device shown.
[0157] See also Figure 6 , Figure 6 is a structural block diagram of an electronic device provided by an optional embodiment of the present disclosure, such as Figure 6 As shown, the electronic device includes: one or more processors 601, memory 602, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 601 is taken as an example.
[0158] The processor 601 may be a central processing unit, a network processor or a combination thereof. The processor 601 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0159] The memory 602 stores instructions executable by at least one processor 601 , so that the at least one processor 601 executes the method shown in the above embodiment.
[0160] The memory 602 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 602 may optionally include a memory remotely arranged relative to the processor 601, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0161] The memory 602 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 602 may also include a combination of the above types of memory.
[0162] The electronic device also includes an input device 603 and an output device 604. The processor 601, the memory 602, the input device 603 and the output device 604 may be connected via a bus or other means. Figure 6 The example of connecting through bus is taken in the following.
[0163] The input device 603 can receive input digital or character information, and generate key signal input related to the user settings and function control of the electronic device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator rod, one or more mouse buttons, a trackball, a joystick, etc. The output device 604 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0164] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0165] A part of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0166] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A data processing method, characterized in that: The method comprises: Obtaining a query plan tree of a database, wherein the query plan tree includes at least one plan node; For each plan node in the query plan tree, encoding the node attribute of the plan node to obtain an initial code of the plan node; Based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree, the plan node is encoded to obtain a position code of the plan node; The initial coding and position coding of the plan node are merged to obtain the coding result of the plan node, so as to obtain the target coding result of the query plan tree, and the target coding result is used to adjust the configuration parameters of the database.
2. The data processing method according to claim 1, characterized in that: The step of encoding the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain the position code of the plan node includes: Based on the hierarchical position of the plan node in the query plan tree, hierarchically encode the plan node to obtain a first code of the plan node; Based on the connectivity between the plan nodes in the query plan tree, the query plan tree is spectrally encoded to obtain a second code; The first code and the second code of the plan node are merged to obtain the position code of the plan node.
3. The data processing method according to claim 2, characterized in that: The step of hierarchically encoding the plan node based on the hierarchical position of the plan node in the query plan tree to obtain a first code of the plan node includes: Acquire depth information of the plan node in the query plan tree, where the depth information is used to characterize the hierarchical position of the plan node in the query plan tree; The plan nodes in the query plan tree are hierarchically encoded based on the depth information to obtain a first encoding of the plan node.
4. The data processing method according to claim 2, characterized in that: The step of performing spectrum coding on the query plan tree based on the connectivity between plan nodes in the query plan tree to obtain a second code comprises: Constructing a target matrix of the query plan tree, wherein the target matrix is used to characterize connectivity between plan nodes in the query plan tree; Performing operations on the target matrix to obtain at least one eigenvector and an eigenvalue corresponding to the eigenvector, wherein the eigenvector is used to characterize the connectivity between the nodes of the plan, and the eigenvalue is used to characterize the occurrence frequency of the corresponding connectivity; Screening the feature vector based on the feature value to obtain a target feature vector; The target feature vectors are fused to obtain the second code.
5. The data processing method according to claim 4, characterized in that: The step of screening the feature vector based on the feature value to obtain a target feature vector includes: Determine a non-zero and minimum target eigenvalue among all the eigenvalues; The eigenvector corresponding to the target eigenvalue is used as the target eigenvector.
6. The data processing method according to claim 1, characterized in that: The node attribute of the plan node includes at least one of a discrete attribute and a continuous attribute; the node attribute of the plan node is encoded to obtain the initial encoding of the plan node, including Performing discrete feature encoding on the discrete attributes of the plan node to obtain encoding results of the discrete attributes; Normalizing and encoding the continuous attributes of the plan node to obtain an encoding result of the continuous attributes; The encoding results of the node attributes are merged to obtain the initial encoding of the planned node.
7. The data processing method according to claim 6, characterized in that: The encoding results of the node attributes are merged to obtain the initial encoding of the planned node, including: Acquire a target connection sequence, where the target connection sequence includes a connection sequence between a plurality of preset attributes, where the plurality of preset attributes include the node attribute; Determine the missing attributes of the plan node based on the attributes other than the node attributes of the plan node among the plurality of preset attributes; The encoding results of the node attributes are fused based on the target connection order. During the fusion process, if there are missing attributes in the planned node, the encoding results of the missing attributes are filled with preset values to obtain the initial encoding of the planned node.
8. A data processing device, characterized in that: The device comprises: A data acquisition module, used to acquire a query plan tree of a database, wherein the query plan tree includes at least one plan node; A first encoding module, used for encoding the node attributes of each plan node in the query plan tree to obtain an initial code of the plan node; A second encoding module, configured to encode the plan node based on the hierarchical position of the plan node in the query plan tree and the connectivity between the plan nodes in the query plan tree to obtain a position code of the plan node; The data fusion module is used to fuse the initial coding and position coding of the plan node to obtain the coding result of the plan node to obtain the target coding result of the query plan tree, and the target coding result is used to adjust the configuration parameters of the database.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data processing method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data processing method according to any one of claims 1 to 7.
11. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the data processing method according to any one of claims 1 to 7.
Citation Information
Cited By
Database configuration parameter tuning method based on load characterization and large model exploration
CN120821715A