Access processing method and apparatus for distributed database
By constructing and merging execution plan features and node features, and combining index configuration features, predicting the execution cost, the problem of low routing accuracy in distributed databases is solved and processing efficiency is improved.
Patent Information
- Application Number
- PCT/CN2024/127396
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-05
- Filing Date
- 2024-10-25
- Publication Date
- 2025-06-12
AI Technical Summary
In distributed databases, when routing is performed based on the query cost of query requests, it is affected by network delay and cross-storage node access bandwidth, resulting in low accuracy of routing.
By constructing execution plan features and node features, combining these features to generate merged features, and combining index configuration features of the target data table of the access statement, predict the execution cost of each execution plan to determine the target execution plan.
It improves the accuracy of the execution cost prediction of the execution plan, and improves the timeliness and efficiency of the processing of access statements by distributed databases.
Smart Images

Figure CN2024127396_12062025_PF_FP_ABST
Abstract
Description
Access processing method and device for distributed database Technical Field
[0001] The present disclosure relates to the field of database technology, and in particular to a method and device for processing access to a distributed database. Background Art
[0002] Distributed databases provide distributed storage, computing and other functions through a service cluster composed of several storage nodes. Different data in the data tables of the distributed database can be allocated to different storage nodes for storage according to the partitioning strategy. Query requests for different data need to be routed to their corresponding storage nodes for processing as much as possible. In order to ensure the efficient operation of the distributed database, in the routing process of query requests, routing is often selected based on the query cost of the query request. However, due to the influence of network latency and access bandwidth across storage nodes, the accuracy of routing based on query cost is low.
[0003] Summary of the Invention
[0004] One or more embodiments of the present disclosure provide an access processing method for a distributed database, comprising: constructing execution plan features based on the execution logic, execution nodes, and data partitions of each execution plan of an access statement of the distributed database. Generating node features based on the data partitions and node communication bandwidth of the execution nodes of each execution plan. Merging the execution plan features and the node features according to the data partitions to obtain merged features. Predicting the execution costs of each execution plan based on the merged features and the index configuration features of the target data table of the access statement, so as to determine the target execution plan based on the execution cost and execute the access statement.
[0005] One or more embodiments of the present disclosure provide an access processing device for a distributed database, comprising: a plan feature construction module, configured to construct execution plan features based on the execution logic, execution nodes, and data partitions of each execution plan of the access statement of the distributed database. A node feature generation module, configured to generate node features based on the data partitions and node communication bandwidths of the execution nodes of the respective execution plans. A feature merging module, configured to merge the execution plan features and the node features according to the data partitions to obtain merged features. An execution cost prediction module, configured to predict the execution costs of the respective execution plans based on the merged features and the index configuration features of the target data table of the access statement, so as to determine the target execution plan based on the execution cost and execute the access statement.
[0006] One or more embodiments of the present disclosure provide an access processing device for a distributed database, comprising: a processor; and a memory configured to store computer-executable instructions, wherein the computer-executable instructions, when executed, cause the processor to: construct execution plan features based on the execution logic, execution nodes, and data partitions of each execution plan of the access statement of the distributed database. Generate node features based on the data partitions and node communication bandwidth of the execution nodes of each execution plan. Merge the execution plan features and the node features according to the data partitions to obtain merged features. Predict the execution costs of each execution plan based on the merged features and the index configuration features of the target data table of the access statement, so as to determine the target execution plan based on the execution cost and execute the access statement.
[0007] One or more embodiments of the present disclosure provide a storage medium for storing computer-executable instructions, which implement the following process when executed by a processor: construct execution plan features based on the execution logic, execution nodes, and data partitions of each execution plan of an access statement of a distributed database. Generate node features based on the data partitions and node communication bandwidth of the execution nodes of each execution plan. Merge the execution plan features and the node features according to the data partitions to obtain merged features. Predict the execution costs of each execution plan based on the merged features and the index configuration features of the target data table of the access statement, so as to determine the target execution plan based on the execution cost and execute the access statement. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in one or more embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0009] FIG1 is a schematic diagram of an implementation environment of a distributed database access processing method provided by one or more embodiments of the present disclosure.
[0010] FIG2 is a processing flow chart of a distributed database access processing method provided by one or more embodiments of the present disclosure.
[0011] FIG3 is a deployment architecture diagram of an OceanBase distributed database provided by one or more embodiments of the present disclosure.
[0012] FIG4 is a schematic diagram of executing a query plan provided by one or more embodiments of the present disclosure.
[0013] FIG5 is an architectural block diagram of a cost prediction model provided by one or more embodiments of the present disclosure.
[0014] FIG6 is a flow chart of a distributed database access processing method applied to a database query scenario provided by one or more embodiments of the present disclosure.
[0015] FIG7 is a schematic diagram of an embodiment of an access processing device for a distributed database provided by one or more embodiments of the present disclosure.
[0016] FIG8 is a schematic structural diagram of a distributed database access processing device provided by one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of the present disclosure, the technical solutions in one or more embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in one or more embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on one or more embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present disclosure.
[0018] The access processing method for a distributed database provided by one or more embodiments of the present disclosure is applicable to an implementation environment of a distributed database system. Referring to FIG1 , the implementation environment includes at least: a distributed database 101, a plan generation unit 102 for generating one or more execution plans for access statements, a plan scheduling unit 103 for selecting a target execution plan from the execution plan for scheduling based on the execution cost of the execution plan, and a cost prediction model 104 for predicting the execution cost of the execution plan; wherein the distributed database 101 is deployed with at least one processing cluster, each processing cluster is composed of one or more processing nodes, and each processing node contains multiple data partitions.
[0019] In this implementation environment, for access statements that access data to the distributed database 101, an execution plan for the access statement is generated by the plan generation unit 102, and the generated execution plan is sent to the plan scheduling unit 103 for scheduling processing of the execution plan. During the scheduling processing, the plan scheduling unit 103 needs to call the cost prediction model 104 to predict the execution cost of each execution plan, and determine the target execution plan based on the predicted execution cost to execute the access statement.
[0020] Specifically, in the execution plan prediction process of the access statement, on the one hand, the execution plan features are constructed according to the execution logic, execution nodes and data partitions of each execution plan, and on the other hand, the node features are generated based on the data partitions and node communication bandwidth of the execution nodes of each execution plan. Then, the execution plan features and the node features are merged according to the data partitions to obtain merged features. According to the merged features and the index configuration features of the target data table of the access statement, the execution cost of each execution plan is predicted. In this way, in the execution cost prediction process, the execution logic, execution nodes and data partitions as well as the node communication bandwidth can be combined to perform execution cost prediction, so that the prediction of the execution cost of the execution plan is more accurate, which helps to improve the timeliness and efficiency of processing access statements for data access to the distributed database 101.
[0021] One or more embodiments of a distributed database access processing method provided by the present disclosure are as follows.
[0022] 2 , the distributed database access processing method provided in this embodiment specifically includes steps S202 to S208 .
[0023] Step S202 : constructing execution plan features according to the execution logic, execution nodes, and data partitions of each execution plan of the access statement of the distributed database.
[0024] The access statement for the distributed database described in this embodiment is a statement for accessing data in the distributed database. The access statement can be a data operation statement for adding, deleting, modifying, or querying the distributed database, such as an SQL query statement for querying data in the OceanBase distributed database: "SELECT * FROM t1, t2 WHERE t1.c1 = t2.c1".
[0025] Optionally, the access statements of the distributed database are executed by at least one processing cluster of the distributed database; the processing cluster is composed of one or more processing nodes, each processing node being configured with at least one data partition corresponding to a data table of the distributed database. The processing nodes may be computing nodes or storage nodes, and the computing nodes or storage nodes may specifically be servers responsible for executing data operations.
[0026] For example, the deployment architecture diagram of the OceanBase distributed database shown in Figure 3 includes three server clusters: server cluster 1, server cluster 2, and server cluster 3. Each server cluster consists of three server nodes, and four data partitions are configured on each server node. These four data partitions correspond to corresponding data tables respectively; at the same time, a Paxos group is also set up.
[0027] During the specific execution process, an execution plan for the access statement is generated after the access statement of the distributed database is obtained. During the process of generating the execution plan, the execution processing of an access statement may involve multiple data operations. Different data operations can be executed and processed on different processing nodes. Since there are multiple processing nodes in the distributed database, the execution of the access statement may be executed and processed by different processing nodes. Accordingly, the execution plan generated for the access statement may also be different. This embodiment predicts the execution cost of each execution plan of the access statement, and determines which execution plan to use to execute the access statement based on the execution cost.
[0028] In an optional implementation provided by this embodiment, each execution plan of the access statement is generated in the following manner: the access statement is parsed to obtain a syntax tree structure, and the syntax tree is converted into an access tree structure; the access tree structure is optimized to obtain an execution tree structure, and the execution plans are generated according to the execution tree structure and the node topology structure of the execution node.
[0029] For example, in the OceanBase distributed database SQL query statement "SELECT * FROM t1, t2 WHERE t1.c1 = t2.c1", when generating the execution plan for this query statement, the SQL query statement is first parsed to obtain a parsing result with a tree-like hierarchical structure, namely a syntax tree structure. Then, the syntax tree structure is converted into a query tree structure for data query on the OceanBase distributed database, and the query is optimized using optimization rules to obtain an optimized query tree structure. Finally, based on the query characteristics, the topological structure information of the OceanBase distributed database, and the optimized query tree structure, multiple query plans for the SQL query statement are generated.
[0030] As shown in FIG4 , there is a schematic diagram of the execution of one of the multiple query plans. The execution of the query plan involves three server nodes: server node 1 , server node 2 , and server node 3 in server cluster 1 .
[0031] During the execution of the query plan, server node 1 sends data operations to server node 2, and server node 2 performs data queries on data table t1 of the OceanBase distributed database. In addition, server node 1 sends data operations to server node 3 in server cluster 1, and server node 3 performs data queries on data table t2 of the OceanBase distributed database. Server node 2 and server node 3 communicate data during the data query process, and finally return the queried data to server node 1.
[0032] During specific implementation, based on the execution plan generated for the access statement, execution plan features are constructed based on the execution logic, execution nodes, and data partitions of each execution plan of the access statement. This is used to obtain the execution plan features that carry the execution logic, execution nodes, and data partitions of the execution plan, providing a data basis for the calculation of the execution cost of the subsequent execution plan.
[0033] In an optional implementation provided by this embodiment, execution plan features are constructed based on the execution logic, execution nodes, and data partitions of each execution plan of the access statement of the distributed database, including: constructing graph structure data based on the execution logic, execution nodes, and data partitions of each execution plan; performing vector conversion on the graph structure data, and using the execution plan vector obtained by the conversion as the execution plan feature.
[0034] Among them, the execution node of the execution plan refers to the processing node in the distributed database that participates in the execution processing of the execution plan. For example, in the schematic diagram of a query plan of an SQL query statement shown in Figure 4, the execution nodes participating in the execution processing of the query plan are server node 1, server node 2 and server node 3 in server cluster 1.
[0035] The execution logic of an execution plan refers to the data operations that need to be performed on each execution node during the execution of the execution plan. For example, in the query plan diagram of a SQL query statement shown in Figure 4, the execution logic includes: server node 1 in server cluster 1 sends data operations to server node 2 and server node 3, performs data query on data table t1 of the OceanBase distributed database through server node 2 in server cluster 1, and performs data query on data table t2 of the OceanBase distributed database through server node 3 in server cluster 1.
[0036] The data partitions of the execution plan refer to the data partitions on the processing nodes corresponding to the data tables of the distributed database targeted by the access statement. For example, the data partitions corresponding to data table t1 of the OceanBase distributed database include: data partitions P7 and data partition P8 of server node 2 in server cluster 1; the data partitions corresponding to data table t2 of the OceanBase distributed database include: data partitions P11 and data partition P12 of server node 3 in server cluster 1; then the data partitions of the current execution plan include data partitions P7 and data partition P8 of server node 2 in server cluster 1, and data partitions P11 and data partition P12 of server node 3 in server cluster 1.
[0037] During the specific execution process, the execution plan data of the three dimensions of execution logic, execution node and data partition of each execution statement of the access statement is integrated through a graph structure to obtain graph structure data carrying the execution plan data of the three dimensions of execution logic, execution node and data partition. The graph structure data is then converted into vector form to obtain three-dimensional execution plan features in vector form.
[0038] In addition, in the process of obtaining three-dimensional execution plan features based on the execution plan data of the three dimensions of execution logic, execution node and data partition of each execution statement of the access statement, the execution plan data of the three dimensions of execution logic, execution node and data partition can also be encoded respectively, and converted into vector form according to the encoding results to obtain three-dimensional execution plan features in vector form.
[0039] Step S204: generating node characteristics based on the data partitions and node communication bandwidths of the execution nodes of the respective execution plans.
[0040] During specific implementation, during the execution of the execution plan of the access statement, the execution of an execution plan may involve multiple processing nodes. When the execution plan is executed by multiple processing nodes, the processing nodes may perform data communication or data transmission during the execution of the execution plan. The data communication or data transmission between the processing nodes requires a certain processing time. Taking into account the data transmission between the processing nodes of the distributed database, in order to more accurately predict the execution cost of the execution plan, by extracting the node communication bandwidth-related features of the execution nodes of the execution plan, the network overhead in the distributed database environment can be better considered, thereby helping to improve the accuracy and comprehensiveness of the execution cost prediction when the extracted node communication bandwidth-related features are used to predict the execution cost of the following execution plan.
[0041] In an optional implementation provided by this embodiment, node features are generated based on the data partitions and node communication bandwidths of the execution nodes of the respective execution plans, including: constructing graph structure data based on the data partitions and node communication bandwidths of the execution nodes; graph nodes in the graph structure data correspond to execution nodes, node attributes of the graph nodes correspond to the data partitions of the execution nodes, and connection weights between graph nodes correspond to the node communication bandwidth; performing vector conversion on the graph structure data, and using the node vectors obtained by the conversion as the node features.
[0042] Among them, the node communication bandwidth refers to the communication bandwidth for data communication or data transmission between the execution nodes included in the execution plan. The node communication bandwidth can be read from the pre-stored node configuration of the processing node; based on the node communication bandwidth and the number of data operations allocated to each execution node in the actual execution scenario, the transmission time for data communication or data transmission between the processing nodes can be calculated.
[0043] During the specific execution process, the processing node data of the two dimensions of data partition and node communication bandwidth of each execution statement of the access statement are integrated through a graph structure to obtain graph structure data carrying the processing node data of the two dimensions of data partition and node communication bandwidth, and the graph structure data is converted into vector form, and the node vector obtained by the conversion is used as the node feature of the execution plan; in addition, the data partition and node communication bandwidth of each execution statement can be encoded separately, and converted into vector form according to the encoding result, and the node vector obtained by the conversion is used as the node feature of the execution plan.
[0044] Step S206: Merge the execution plan feature and the node feature according to the data partition to obtain a merged feature.
[0045] The above-mentioned execution plan features are constructed according to the execution logic, execution nodes and data partitions of each execution plan of the access statement of the distributed database, and graph structure data of the execution plan data with three dimensions of execution logic, execution nodes and data partitions is obtained. The graph structure data is converted into vector form to obtain three-dimensional execution plan features in vector form, and node features are generated based on the data partitions and node communication bandwidth of the execution nodes of each execution plan to obtain processing node data with two dimensions of data partitions and node communication bandwidth.
[0046] On this basis, based on the data partitions carried by both the execution plan features and the node features, the two parts of data are merged. Specifically, the execution logic and execution nodes carried in the execution plan features are merged with the node communication bandwidth corresponding to the same data partition in the node features. Alternatively, according to the data partition, the node communication bandwidth carried in the node features is merged into the execution logic and execution nodes corresponding to the same data partition in the execution plan features, thereby obtaining execution plan data carrying four-dimensional data of execution logic, execution node, data partition and node communication bandwidth. This can better consider the network overhead in the distributed database environment, and thus can make more accurate execution cost predictions for the execution plan based on the network overhead in the distributed database environment.
[0047] In an optional implementation provided by this embodiment, the execution plan feature and the node feature are merged according to the data partition to obtain a merged feature, including: determining the vector elements corresponding to the same data partition in the execution plan vector and the node vector by parsing the execution plan vector and the node vector; merging the vector elements corresponding to the same data partition in the execution plan vector and the node vector, and using the merged vector as the merged feature.
[0048] During the specific execution process, the execution plan vector is a three-dimensional vector carrying execution plan data of three dimensions: execution logic, execution node, and data partition. The node vector is a two-dimensional vector carrying data partition and node communication bandwidth. For each node vector, based on the vector element of the data partition in the node vector, the execution plan vector containing the same data vector element is determined, and the other node communication element other than the data partition in the node vector is added as a new vector element to the determined execution plan vector, thereby obtaining a four-dimensional execution plan vector, that is, obtaining execution plan data carrying four-dimensional data of execution logic, execution node, data partition, and node communication bandwidth.
[0049] Step S208 , predicting the execution costs of the respective execution plans based on the merge feature and the index configuration feature of the target data table of the access statement, so as to determine a target execution plan based on the execution cost and execute the access statement.
[0050] The target data table of the access statement described in this embodiment refers to the data table corresponding to the data table identifier carried by the access statement in the distributed database. For example, the SQL query statement "SELECT*FROM t1,t2 WHERE t1.c1=t2.c1" of the OceanBase distributed database carries data table identifiers t1 and t2, so the target data tables of the SQL query statement are data tables t1 and t2 in the OceanBase distributed database.
[0051] The index configuration characteristics of the target data table refer to the characteristics of the index configuration information of the target data table carried. The index configuration information includes the index name, index type, table name and / or number of rows. The index configuration information of the OceanBase distributed database can be obtained by querying the OceanBase system view. The index configuration information can be configured according to the actual access requirements of the distributed database. Different index configuration information can be configured for different query scenarios.
[0052] This embodiment, based on the execution plan data carrying four-dimensional data of execution logic, execution nodes, data partitions and node communication bandwidth obtained by data merging as mentioned above, obtains merged features carrying execution logic, execution nodes, data partitions and node communication bandwidth. In the process of predicting the execution cost of the execution plan, the execution cost is predicted based on the merged features and the underlying logical configuration of the distributed database. Specifically, the execution cost of the execution plan of the access statement is predicted based on the merged features and the index configuration features of the target data table of the access statement, so as to improve the accuracy of the execution cost prediction and thus improve the access processing efficiency of the distributed database.
[0053] When predicting the execution cost of the execution plan of the access statement based on the merge features and the index configuration features of the target data table of the access statement, before predicting the execution cost of the execution plan of the access statement, the index configuration features required for the execution cost prediction can also be obtained from the access statement. At the same time, in order to improve the prediction efficiency of the execution cost, the index configuration features of each data table in the distributed database can be generated and stored in advance. When predicting the execution cost of the execution plan of the access statement, it is only necessary to read the index configuration features of the target data table of the corresponding access statement.
[0054] In an optional implementation provided by this embodiment, the index configuration characteristics of the target data table of the access statement are obtained in the following manner: according to the data table identifier carried by the access statement, the data table corresponding to the data table identifier in the distributed database is determined to be the target data table, and the pre-generated index configuration characteristics of the target data table are read.
[0055] In the process of generating index configuration features for each data table in a distributed database, in order to improve the comprehensiveness of the underlying logical configuration of the distributed database, so as to help improve the accuracy of execution cost prediction based on merge features and index configuration features, in the process of generating index configuration features for data tables in a distributed database, the local relationship between indexes can be learned by introducing an attention mechanism to improve the feature expression effect of the index configuration features. The following is an example of the process of determining the index configuration features of the target data table of the access statement in the distributed data. The process of determining the index configuration features of other data tables other than the target data table in the distributed database is similar, and this embodiment will not be repeated here.
[0056] In an optional implementation provided by this embodiment, the index configuration feature of the target data table is determined in the following manner: a configuration embedding vector is generated based on the index configuration information of the target data table, and a configuration embedding matrix is constructed based on the index configuration vector; self-attention calculation is performed on the configuration embedding matrix to obtain an attention sequence, and an index configuration vector is generated based on the attention sequence as the index configuration feature.
[0057] For example, the index configuration sequence Index={I1,I2,…,I k}, I1, I2, I k Represents the index configuration information of the data table, and uses the index configuration sequence Index to generate the index configuration embedding matrix E Index =(E1 T ,E2 T ,…,E k T ), where each index configuration I1 corresponds to an embedding vector E1 T ; Then the self-attention mechanism is used to embed the index configuration matrix E Index Self-attention calculation is performed to learn the internal correlation between the index configuration information of the data table. The self-attention mechanism allows the index configuration information at each position in the sequence to interact and transfer information with the index configuration information at other positions, thereby capturing the long-range dependency between the index configuration information at different positions. After the self-attention calculation, a vector sequence is obtained. Each vector in the vector sequence represents the index configuration information that integrates the comprehensive internal correlations of different positions, and the vector sequence is weighted and summed to obtain the index configuration vector. The index configuration vector is also the index configuration feature that carries the local relationship and long-range dependency between the index configuration information.
[0058] In specific implementation, in the process of predicting the execution cost of the execution plan of the access statement based on the merged features of the four-dimensional data carrying the execution logic, execution nodes, data partitions and node communication bandwidth, and the index configuration features of the target data table of the access statement, the execution cost prediction efficiency can be improved by data vectorization. In an optional implementation provided by this embodiment, the execution cost of each execution plan is predicted based on the merged features and the index configuration features of the target data table of the access statement, including: vector connection of the merged vector of the merged features and the index configuration vector of the index configuration features to obtain a prediction input vector; input the prediction input vector into the cost prediction algorithm to perform execution cost prediction to obtain the execution cost of each execution plan.
[0059] The determination of the merging vector of the merging feature and the index configuration vector of the index configuration feature may refer to the specific processing procedures of generating the merging vector and generating the index configuration vector provided above.
[0060] In actual applications, the execution cost of the execution plan of the access statement is predicted by combining the merge characteristics of the four-dimensional data carrying execution logic, execution nodes, data partitions and node communication bandwidth, as well as the index configuration characteristics of the target data table of the access statement. After obtaining the execution cost of each access plan of the access statement, the target execution plan for the current execution of the access statement can be determined in each access plan based on the execution cost of each access plan, and the execution processing of the access statement can be performed according to the determined target execution plan, so as to improve the processing and response efficiency of the distributed database to the access statement.
[0061] Specifically, in an optional implementation provided by this embodiment, a target execution plan is determined based on the execution cost and the access statement is executed, including: determining the execution cost that meets the execution conditions among the execution costs of each execution plan as the target execution cost; executing the access statement according to the execution plan corresponding to the target execution cost, and obtaining the execution result of the access statement.
[0062] The execution cost that satisfies the execution condition may be an execution plan whose execution cost is less than the execution costs of other execution plans, such as the first execution plan after sorting the execution plans in ascending order of execution cost.
[0063] In this embodiment, the execution cost of the execution plan of the access statement is predicted after the execution plan of the access statement is generated and before the access statement is scheduled according to the execution plan. The target execution plan for executing the access statement is determined by calculating the execution cost of each execution plan of the access statement, so that the access statement can be scheduled according to the target execution plan and dispatched to the corresponding processing node to execute the corresponding data operation of the access statement. To achieve automated scheduling of access statements, the execution cost of each execution plan of the access statement can be predicted using a cost prediction model, thereby improving the processing efficiency of access statements through automated scheduling.
[0064] Optionally, the access processing method of the distributed database provided in this embodiment is executed by a cost prediction model; the cost prediction model includes: a plan feature extraction unit, a node feature extraction unit, a feature merging unit, an index feature extraction unit and a cost prediction unit; wherein, the step of constructing execution plan features according to the execution logic, execution nodes and data partitions of each execution plan of the access statement of the distributed database is executed by the plan feature extraction unit; the step of generating node features based on the data partitions and node communication bandwidth of the execution nodes of each execution plan is executed by the node feature extraction unit; the step of merging the execution plan features and the node features according to the data partitions to obtain merged features is executed by the feature merging unit; the index configuration features of the target data table are obtained by the index feature extraction unit; the operation of predicting the execution cost of each execution plan based on the merged features and the index configuration features of the target data table of the access statement is executed by the cost prediction unit.
[0065] For example, the following model function is established for the cost prediction model:
[0066] C=LearnedCost(qry,svr,idx)
[0067] Among them, qry represents the execution plan data of each query plan of the query statement, svr represents the data partition and node communication bandwidth of the processing node of each query plan, idx represents the index configuration information of the target data table in the OceanBase distributed database of the data table identifier carried by the query statement, and C represents the query cost of each query plan of the query statement.
[0068] As shown in Figure 5, the cost prediction model obtained based on the training of the model function specifically includes: a plan feature extraction layer, a node feature extraction layer, a feature merging layer, an index feature extraction layer and a cost prediction layer; wherein, the plan feature extraction layer is used to construct query plan features based on the execution logic of each query plan of the input query statement, the execution node and the execution plan data composed of the data partition; the node feature extraction layer is used to generate node features based on the data partition and the node communication bandwidth of the execution node of each query plan; the feature merging layer is used to merge the query plan features and the node features according to the data partition to obtain the merged features; the index feature extraction layer is used to extract the index configuration features of the target data table corresponding to the data table identifier carried by the query statement in the pre-generated and stored index configuration set of the OceanBase distributed database; the cost prediction layer is used to predict the query costs of each query plan of the query statement based on the merged features and the index configuration features of the target data table of the query statement, and output the predicted query costs of each query plan of each query statement.
[0069] The following takes the application of a distributed database access processing method provided by this embodiment in a database query scenario as an example, and combines Figure 6 to further illustrate the distributed database access processing method provided by this embodiment. Referring to Figure 6, the distributed database access processing method applied to the database query scenario includes steps S602 to S618.
[0070] Step S602: construct graph structure data according to the execution logic, execution nodes and data partitions of multiple query plans of query statements of the distributed database.
[0071] Step S604: perform vector conversion on the graph structure data to obtain a query plan vector.
[0072] Step S606 : constructing graph structure data based on the data partitions and node communication bandwidths of the execution nodes of the multiple query plans.
[0073] Step S608: Perform vector conversion on the graph structure data to obtain node vectors.
[0074] Step S610: Determine, by analysis, the vector elements in the query plan vector and the node vector that correspond to the same data partition.
[0075] Step S612: Merge the vector elements corresponding to the same data partition in the query plan vector and the node vector to obtain a merged vector.
[0076] Step S614: determining the target data table corresponding to the data table identifier carried by the query statement in the distributed database, and reading the pre-generated and stored index configuration vector of the target data table.
[0077] Step S616: perform vector concatenation on the merged vector and the index configuration vector to obtain a prediction input vector.
[0078] Step S618: Input the predicted input vector into the cost prediction algorithm to perform execution cost prediction to obtain the execution cost of each query plan.
[0079] It should be pointed out that the query processing method of the distributed database shown in Figure 6 can be applied to the cost prediction model, which includes a plan feature extraction layer, a node feature extraction layer, a feature merging layer, an index feature extraction layer and a cost prediction layer; accordingly, the above steps S602 to S604 can be executed by the plan feature extraction layer, the above steps S606 to S608 can be executed by the plan feature extraction layer, the above steps S610 to S612 can be executed by the feature merging layer, the above step S614 can be executed by the index feature extraction layer, and the above steps S616 to S618 can be executed by the cost prediction layer.
[0080] An embodiment of an access processing device for a distributed database provided by the present disclosure is as follows.
[0081] In the above embodiment, a distributed database access processing method is provided, and correspondingly, a distributed database access processing device is also provided, which will be described below with reference to the accompanying drawings.
[0082] Since the device embodiment corresponds to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the corresponding description of the method embodiment provided above. The device embodiment described below is only illustrative.
[0083] This embodiment provides an access processing device for a distributed database, which includes: a plan feature construction module 702, configured to construct execution plan features based on the execution logic, execution nodes and data partitions of each execution plan of the access statement of the distributed database; a node feature generation module 704, configured to generate node features based on the data partitions and node communication bandwidths of the execution nodes of the respective execution plans; a feature merging module 706, configured to merge the execution plan features and the node features according to the data partitions to obtain merged features; an execution cost prediction module 708, configured to predict the execution costs of the respective execution plans based on the merged features and the index configuration features of the target data table of the access statement, so as to determine the target execution plan based on the execution costs and execute the access statement.
[0084] An embodiment of an access processing device for a distributed database provided by the present disclosure is as follows: Corresponding to the access processing method for a distributed database described above, based on the same technical concept, one or more embodiments of the present disclosure also provide an access processing device for a distributed database, which is used to execute the access processing method for a distributed database provided above. Figure 8 is a structural schematic diagram of an access processing device for a distributed database provided by one or more embodiments of the present disclosure.
[0085] As shown in FIG8 , distributed database access processing devices can vary significantly due to different configurations or performance. This embodiment provides a distributed database access processing device comprising: one or more processors 801 and memory 802. Memory 802 can store one or more applications or data. Memory 802 can be either transient or persistent storage. Applications stored in memory 802 can include one or more modules (not shown), each of which can include a series of computer-executable instructions for the distributed database access processing device. Furthermore, processor 801 can be configured to communicate with memory 802, executing the series of computer-executable instructions in memory 802 on the distributed database access processing device. The distributed database access processing device can also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, one or more keyboards 806, and the like.
[0086] In a specific embodiment, an access processing device for a distributed database includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for the access processing device for the distributed database, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: constructing execution plan features based on the execution logic, execution nodes and data partitions of each execution plan of the access statement of the distributed database; generating node features based on the data partitions and node communication bandwidth of the execution nodes of the each execution plan; merging the execution plan features and the node features according to the data partitions to obtain merged features; predicting the execution costs of the each execution plan based on the merged features and the index configuration features of the target data table of the access statement, so as to determine the target execution plan based on the execution cost and execute the access statement.
[0087] An embodiment of a storage medium provided by the present disclosure is as follows: Corresponding to the above-described method for processing access to a distributed database, based on the same technical concept, one or more embodiments of the present disclosure further provide a storage medium.
[0088] The storage medium provided in this embodiment is used to store computer-executable instructions, which implement the following process when executed by a processor: constructing execution plan features based on the execution logic, execution nodes and data partitions of each execution plan of an access statement of a distributed database; generating node features based on the data partitions and node communication bandwidth of the execution nodes of each execution plan; merging the execution plan features and the node features according to the data partitions to obtain merged features; predicting the execution cost of each execution plan based on the merged features and the index configuration features of the target data table of the access statement, so as to determine the target execution plan based on the execution cost and execute the access statement.
[0089] It should be noted that the embodiment of a storage medium in the present disclosure and the embodiment of a distributed database access processing method in the present disclosure are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned corresponding method, and the repeated parts will not be repeated.
[0090] The various embodiments in the present disclosure are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. For example, the device embodiment, equipment embodiment and storage medium embodiment are similar to the method embodiment, so the description is relatively simple. To read the relevant content in the device embodiment, equipment embodiment and storage medium embodiment, please refer to the partial description of the method embodiment.
[0091] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0092] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0093] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0094] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0095] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of the present disclosure, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0096] Those skilled in the art will appreciate that one or more embodiments of the present disclosure may be provided as a method, system, or computer program product. Therefore, one or more embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0097] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0098] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0100] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0101] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0102] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0103] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising at least one ..." does not exclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.
[0104] One or more embodiments of the present disclosure may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0105] The foregoing is an embodiment of the present disclosure and is not intended to limit the present disclosure. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure are intended to be included within the scope of the claims of the present disclosure.
Claims
1. A distributed database access processing method, comprising: Construct execution plan features based on the execution logic, execution nodes, and data partitions of each execution plan of the access statement of the distributed database; Generate node characteristics based on data partitions and node communication bandwidths of execution nodes of each execution plan; Merging the execution plan feature and the node feature according to the data partition to obtain a merged feature; The execution costs of the respective execution plans are predicted according to the merge feature and the index configuration feature of the target data table of the access statement, so as to determine the target execution plan according to the execution cost and execute the access statement.
2. The distributed database access processing method according to claim 1, wherein: The execution plan features are constructed according to the execution logic, execution nodes and data partitions of each execution plan of the access statement of the distributed database, including: Constructing graph structure data according to the execution logic, execution nodes and data partitions of each execution plan; The graph structure data is vectorized, and an execution plan vector obtained by the transformation is used as the execution plan feature.
3. The distributed database access processing method according to claim 2, wherein: The generating node characteristics based on the data partition and node communication bandwidth of the execution nodes of each execution plan includes: Graph structure data is constructed based on the data partitions of the execution nodes and the node communication bandwidth; the graph nodes in the graph structure data correspond to the execution nodes, the node attributes of the graph nodes correspond to the data partitions of the execution nodes, and the connection weights between the graph nodes correspond to the node communication bandwidth; The graph structure data is subjected to vector conversion, and the node vector obtained by the conversion is used as the node feature.
4. The distributed database access processing method according to claim 3, wherein: The step of merging the execution plan feature and the node feature according to the data partition to obtain the merged feature includes: Determine vector elements in the execution plan vector and the node vector corresponding to the same data partition by parsing the execution plan vector and the node vector; The vector elements corresponding to the same data partition in the execution plan vector and the node vector are merged, and the merged merged vector is used as the merged feature.
5. The distributed database access processing method according to claim 1, wherein: Each execution plan of the access statement is generated in the following way: Performing grammatical analysis on the access statement to obtain a grammatical tree structure, and converting the grammatical tree into an access tree structure; The access tree structure is optimized to obtain an execution tree structure, and the execution plans are generated according to the execution tree structure and the node topology structure of the execution nodes.
6. The distributed database access processing method according to claim 1, wherein: The access statement of the distributed database is executed by at least one processing cluster of the distributed database; The processing cluster is composed of one or more processing nodes, and each processing node is configured with at least one data partition corresponding to a data table of the distributed database.
7. The distributed database access processing method according to claim 1, wherein: Before the step of predicting the execution cost of each execution plan according to the merge feature and the index configuration feature of the target data table of the access statement to determine the target execution plan according to the execution cost and execute the access statement is executed, the step further includes: According to the data table identifier carried by the access statement, it is determined that the data table corresponding to the data table identifier in the distributed database is the target data table, and the pre-generated index configuration feature of the target data table is read.
8. The distributed database access processing method according to claim 7, wherein: The index configuration characteristics of the target data table are determined in the following manner: Generate a configuration embedding vector based on the index configuration information of the target data table, and construct a configuration embedding matrix based on the index configuration vector; A self-attention calculation is performed on the configuration embedding matrix to obtain an attention sequence, and an index configuration vector is generated based on the attention sequence as the index configuration feature.
9. The distributed database access processing method according to claim 1, wherein: The predicting the execution cost of each execution plan according to the merge feature and the index configuration feature of the target data table of the access statement includes: Performing vector connection on the merged vector of the merged feature and the index configuration vector of the index configuration feature to obtain a prediction input vector; The predicted input vector is input into a cost prediction algorithm to perform execution cost prediction, so as to obtain the execution cost of each execution plan.
10. The distributed database access processing method according to claim 1, wherein: The step of determining a target execution plan according to the execution cost and executing the access statement includes: Determining the execution cost that satisfies the execution condition from among the execution costs of each execution plan as the target execution cost; The access statement is executed according to the execution plan corresponding to the target execution cost, and the execution result of the access statement is obtained.
11. The distributed database access processing method according to claim 1, wherein: The distributed database access processing method is executed through a cost prediction model; The cost prediction model includes: a plan feature extraction unit, a node feature extraction unit, a feature merging unit, an index feature extraction unit and a cost prediction unit; The step of constructing execution plan features according to the execution logic, execution nodes and data partitions of each execution plan of the access statement of the distributed database is performed by the plan feature extraction unit; the step of generating node features based on the data partitions and node communication bandwidth of the execution nodes of each execution plan is performed by the node feature extraction unit; the step of merging the execution plan features and the node features according to the data partitions to obtain the merged The feature step is performed by the feature merging unit; the index configuration feature of the target data table is obtained by the index feature extraction unit; the operation of predicting the execution cost of each execution plan based on the merged feature and the index configuration feature of the target data table of the access statement is performed by the cost prediction unit.
12. A distributed database access processing device, comprising: A plan feature building module is configured to build execution plan features according to the execution logic, execution nodes and data partitions of each execution plan of the access statement of the distributed database; A node feature generation module, configured to generate node features based on data partitions and node communication bandwidths of execution nodes of each execution plan; A feature merging module is configured to merge the execution plan feature and the node feature according to the data partition to obtain a merged feature; The execution cost prediction module is configured to predict the execution cost of each execution plan according to the merge feature and the index configuration feature of the target data table of the access statement, so as to determine the target execution plan according to the execution cost and execute the access statement.
13. A distributed database access processing device, comprising: processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: Construct execution plan features based on the execution logic, execution nodes, and data partitions of each execution plan of the access statement of the distributed database; Generate node characteristics based on data partitions and node communication bandwidths of execution nodes of each execution plan; Merging the execution plan feature and the node feature according to the data partition to obtain a merged feature; The execution costs of the respective execution plans are predicted according to the merge feature and the index configuration feature of the target data table of the access statement, so as to determine the target execution plan according to the execution cost and execute the access statement.
14. A storage medium for storing computer executable instructions, which when executed by a processor implement: Construct execution plan features based on the execution logic, execution nodes, and data partitions of each execution plan of the access statement of the distributed database; Generate node characteristics based on data partitions and node communication bandwidths of execution nodes of each execution plan; Merging the execution plan feature and the node feature according to the data partition to obtain a merged feature; The execution costs of the respective execution plans are predicted according to the merge feature and the index configuration feature of the target data table of the access statement, so as to determine the target execution plan according to the execution cost and execute the access statement.
Citation Information
Patent Citations
SQL query method and device for distributed database
CN113934763A
Deep learning cost estimation system, method and equipment for cloud side-end collaborative query
CN114911823A
Execution cost evaluation method and device and electronic equipment
CN116431448A
Query optimization system based on cost estimation
CN116521719A
Access processing method and device for distributed database
CN117762986A
Cited By
A graph database query method and system
CN122654373A