Database distributed parallel execution method and device, equipment and storage medium

By optimizing the collaborative work of distributed optimizers and actuators, using load prediction models and adaptive scanning mechanisms, the computational load imbalance caused by data skew in traditional databases is solved, and the performance and resource utilization of the database system are improved.

CN120371896APending Publication Date: 2025-07-25SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510519017.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Data skew caused by static resource allocation in traditional databases in cloud environments leads to unbalanced computing load and query performance bottlenecks.

Method used

The initial distributed optimizer is optimized through network transmission technology and resource pooling storage characteristics, and the target distributed optimizer is generated, and a distributed execution plan is generated by combining cost model and adaptive scanning mechanism. The load prediction model is used to allocate query tasks to optimize the workload of the query executor.

Benefits of technology

It improves the performance, scalability and resource utilization of distributed database systems, effectively responds to the challenges of large-scale data queries, and prevents unbalanced computing load.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371896A_ABST
    Figure CN120371896A_ABST
Patent Text Reader

Abstract

The invention discloses a database distributed parallel execution method and device, equipment and a storage medium, and relates to the technical field of databases, and the method comprises the steps: carrying out the extension and optimization of an initial distributed optimizer based on a network transmission technology and resource pooling storage characteristics, and obtaining a target distributed optimizer; generating a node scanning plan based on the query request by using a target distributed optimizer, evaluating the node scanning plan by using a cost model to obtain a target node scanning plan, and generating a distributed execution plan based on the target node scanning plan and a preset adaptive scanning mechanism; and predicting the load condition of each node by using a target load prediction model, and allocating a query task in the distributed execution plan to a target node by using a query coordinator in the distributed executor based on the obtained predicted load value and the physical position of each node, and executing the query task by using the query executor on the target node to obtain an execution result. Therefore, the unbalanced calculation load caused by data skew is prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of databases, and particularly relates to a method, device, equipment and storage medium for distributed parallel execution of databases. Background Art

[0002] With the wide popularization of cloud computing, cloud native technology has gradually become the standard for cloud service enterprises to build and manage applications. Due to the limitations of static resource allocation and expansion ability, traditional databases are difficult to meet the dynamic requirements of the cloud environment. Resource pooling technology realizes efficient sharing through unified management of computing and storage resources and combined with dynamic scheduling, significantly improving resource utilization rate and reducing costs. However, in traditional distributed query execution, each query executor is usually assigned a fixed data range for scanning. This static allocation method is prone to data skew and cause query performance bottlenecks in the case of uneven data distribution or changing query load.

[0003] As can be seen from the above, how to prevent the uneven computing load caused by data skew is an urgent problem to be solved at present. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and storage medium for distributed parallel execution of databases, which can prevent the uneven computing load caused by data skew. The specific solutions are as follows:

[0005] In the first aspect, the present application provides a method for distributed parallel execution of databases, including:

[0006] Expanding and optimizing an initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer;

[0007] Generating a number of node scan plans based on a query request and using the target distributed optimizer, evaluating each node scan plan using a cost model to obtain a target node scan plan, and then generating a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism;

[0008] Predicting the load conditions of each node using a target load prediction model to obtain predicted load values, distributing the query tasks in the distributed execution plan to target nodes based on the predicted load values and the physical locations of each node and using a query coordinator in a distributed executor, and executing the query tasks using a query executor on the target node to obtain corresponding execution results.

[0009] Optionally, the expanding and optimizing an initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer includes:

[0010] Evaluate the network transmission capacity of the initial distributed optimizer, and determine the target transmission strategy based on the result of the capacity evaluation;

[0011] Analyze the data distribution, storage characteristics, and access patterns in the database by using the resource pooling storage characteristics to obtain corresponding analysis results;

[0012] Optimize the initial distributed optimizer based on the analysis results and the target transmission strategy to obtain the target distributed optimizer;

[0013] Package the target distributed optimizer into a dynamic library, and load the dynamic library into the kernel of the database.

[0014] Optionally, generating several node scan plans based on the query request and using the target distributed optimizer, and evaluating each of the node scan plans by using a cost model to obtain a target node scan plan, including:

[0015] Obtain a query request for the database, and generate node scan plans including different scan nodes, scan ranges, and scan strategies based on the query statement in the query request and by using the target distributed optimizer;

[0016] Generate respective cost values representing the resources and performance consumption required for executing each of the node scan plans based on each of the node scan plans and by using the cost model;

[0017] Compare each of the cost values, and determine the node scan plan with the lowest cost value as the target node scan plan based on the comparison result.

[0018] Optionally, generating a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism, including:

[0019] Monitor the load and data distribution conditions of each of the nodes based on the preset adaptive scan mechanism, and adjust the scan strategy in the target node scan plan based on the monitoring result to obtain an adjusted scan strategy;

[0020] Generate a corresponding distributed execution plan based on the target node scan plan including the adjusted scan strategy and by using the target distributed optimizer.

[0021] Optionally, generating a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism, including:

[0022] Determine a preset scan operator as the basic operator of the distributed execution plan to read target data corresponding to the target node scan plan from a preset repository by using the basic operator;

[0023] Introduce a pre-designed computing operator in the distributed execution plan to perform data processing operations on the target data using the pre-designed computing operator to obtain processed data;

[0024] Generate a distributed execution plan based on the processed data, the target node scan plan, and a preset adaptive scan mechanism.

[0025] Optionally, predicting the load conditions of each node using a target load prediction model to obtain predicted load values includes:

[0026] Obtain the historical load data of each node and generate corresponding time series data based on the historical load data;

[0027] Construct an initial load prediction model based on the autoregressive moving average model and the long short-term memory network, and train the initial load prediction model using the time series data to obtain a target load prediction model;

[0028] Predict the load conditions of each node using the target load prediction model to obtain the predicted load values of the nodes in a preset time period.

[0029] Optionally, based on the predicted load values and the physical locations of each node and using the query coordinator in the distributed executor, allocate the query tasks in the distributed execution plan to target nodes, and use the query executor on the target nodes to execute the query tasks to obtain corresponding execution results, including:

[0030] Use the query coordinator in the distributed executor to analyze the query tasks in the distributed execution plan, and determine the task requirements of the query tasks based on the analysis results;

[0031] Allocate the query tasks to target nodes that meet the preset load conditions and preset distance conditions based on the predicted load values and the physical locations of each node and using remote procedure calls;

[0032] Use the query executor on the target nodes to execute the query tasks, and during the execution of the query tasks, monitor the load conditions of the target nodes and the task execution progress corresponding to the query tasks based on the query coordinator;

[0033] If the monitoring results indicate that the load conditions of the target nodes reach the target load threshold or the query tasks are executed abnormally, jump to the step of using the query coordinator in the distributed executor to analyze the query tasks in the distributed execution plan and determine the task requirements of the query tasks based on the analysis results.

[0034] In a second aspect, the present application provides a database distributed parallel execution device, including:

[0035] An initial optimizer optimization module, configured to expand and optimize an initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer;

[0036] An execution plan generation module, configured to generate a plurality of node scan plans based on a query request and by using the target distributed optimizer, evaluate each of the node scan plans by using a cost model to obtain a target node scan plan, and then generate a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism;

[0037] A query task execution module, configured to predict the load conditions of each node by using a target load prediction model to obtain predicted load values, and based on the predicted load values and the physical locations of the nodes, and by using a query coordinator in a distributed executor, allocate the query tasks in the distributed execution plan to target nodes, and execute the query tasks by using query executors on the target nodes to obtain corresponding execution results.

[0038] In a third aspect, the present application provides an electronic device, including:

[0039] A memory, configured to store a computer program;

[0040] A processor, configured to execute the computer program to implement the foregoing database distributed parallel execution method.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, wherein when the computer program is executed by a processor, the foregoing database distributed parallel execution method is implemented.

[0042] The present application expands and optimizes an initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer; generates a plurality of node scan plans based on a query request and by using the target distributed optimizer, evaluates each of the node scan plans by using a cost model to obtain a target node scan plan, and then generates a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism; predicts the load conditions of each node by using a target load prediction model to obtain predicted load values, and based on the predicted load values and the physical locations of the nodes, and by using a query coordinator in a distributed executor, allocates the query tasks in the distributed execution plan to target nodes, and executes the query tasks by using query executors on the target nodes to obtain corresponding execution results.

[0043] As can be seen from the above, in this application, the initial distributed optimizer is optimized through network transmission technology and resource pooling storage characteristics to obtain the target distributed optimizer, and the target distributed optimizer is used to generate a distributed execution plan adapted to the distributed environment. The distributed executor is responsible for the query task allocation and execution in the distributed execution plan. Through the collaborative work of the query coordinator and the query executor, and by introducing the target load prediction model and the physical location of the nodes, efficient task allocation and query processing are achieved. In this way, the query coordinator and the query executor based on the optimized target distributed optimizer and the distributed executor perform query task allocation and execution, greatly improving the performance, scalability, and resource utilization rate of the distributed database system, and being able to effectively cope with the challenges brought by large-scale data queries. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0045] Figure 1 It is a flowchart of a database distributed parallel execution method disclosed in this application;

[0046] Figure 2 It is a schematic diagram of generating a distributed execution plan provided by this application;

[0047] Figure 3 It is a schematic diagram of distributed parallel query provided by this application;

[0048] Figure 4 It is a schematic diagram of query task allocation provided by this application;

[0049] Figure 5 It is a schematic diagram of the structure of a database distributed parallel execution device disclosed in this application;

[0050] Figure 6 It is a structural diagram of an electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0052] Currently, in traditional distributed query execution, each query executor is usually assigned a fixed data range for scanning. This static allocation method is prone to data skew in the case of uneven data distribution or changing query loads. To address this, the present application provides a database distributed parallel execution method that distributes and executes query tasks based on an optimized target distributed optimizer, a query coordinator, and query executors, greatly enhancing the performance, scalability, and resource utilization of the distributed database system and effectively coping with the challenges brought by large-scale data queries.

[0053] See Figure 1 As shown, an embodiment of the present invention discloses a database distributed parallel execution method, including:

[0054] Step S11: Expand and optimize the initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer.

[0055] In this embodiment, the network transmission ability of the initial distributed optimizer is evaluated, and a target transmission strategy is determined based on the ability evaluation result to make its transmission ability more adaptable to the distributed environment. Then, the data distribution, storage characteristics, and access patterns in the database are analyzed using the resource pooling storage characteristics, and the initial distributed optimizer is optimized based on the analysis results and the target transmission strategy to obtain a target distributed optimizer. For the convenience of database use, the target distributed optimizer can be encapsulated into a dynamic library and the dynamic library can be loaded into the kernel of the database.

[0056] Specifically, expanding and optimizing the initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer includes: evaluating the network transmission ability of the initial distributed optimizer and determining a target transmission strategy based on the ability evaluation result; analyzing the data distribution, storage characteristics, and access patterns in the database using the resource pooling storage characteristics to obtain corresponding analysis results; optimizing the initial distributed optimizer based on the analysis results and the target transmission strategy to obtain a target distributed optimizer; encapsulating the target distributed optimizer into a dynamic library and loading the dynamic library into the kernel of the database.

[0057] Step S12: Generate a number of node scan plans based on the query request and using the target distributed optimizer, evaluate each node scan plan using a cost model to obtain a target node scan plan, and then generate a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism.

[0058] In this embodiment, after obtaining the target distributed optimizer, a query request for the database is acquired. Based on the query statement in the query request to understand information such as which tables and fields need to be accessed for the query, and then the target distributed optimizer is used to generate a node scan plan including different scan nodes, scan ranges, and scan strategies. A cost model for estimating the resources and performance consumption required to execute the node scan plan is utilized to generate respective cost values, and the respective cost values are compared. Based on the comparison result, the node scan plan with the lowest cost value is determined as the target node scan plan; the target node scan plan is the scan plan with the lowest execution cost and the highest efficiency.

[0059] Specifically, generating a number of node scan plans based on the query request and using the target distributed optimizer, and evaluating each of the node scan plans using the cost model to obtain the target node scan plan includes: acquiring a query request for the database, generating a node scan plan including different scan nodes, scan ranges, and scan strategies based on the query statement in the query request and using the target distributed optimizer; generating respective cost values representing the resources and performance consumption required to execute each of the node scan plans based on each of the node scan plans and using the cost model; comparing the respective cost values, and based on the comparison result, determining the node scan plan with the lowest cost value as the target node scan plan.

[0060] It can be understood that a preset adaptive scan mechanism is used to monitor the load and data distribution of each of the nodes, and based on parameters such as the computing resources, memory, and network bandwidth of each of the nodes, the scan strategy in the target node scan plan is adjusted to obtain an adjusted scan strategy. Then, based on the target node scan plan including the adjusted scan strategy and using the target distributed optimizer, a corresponding distributed execution plan is generated. Specifically, generating the distributed execution plan based on the target node scan plan and the preset adaptive scan mechanism includes: monitoring the load and data distribution of each of the nodes based on the preset adaptive scan mechanism, and adjusting the scan strategy in the target node scan plan based on the monitoring result to obtain an adjusted scan strategy; generating a corresponding distributed execution plan based on the target node scan plan including the adjusted scan strategy and using the target distributed optimizer.

[0061] In this embodiment, a preset scan operator is determined as the basic operator of the distributed execution plan. The preset scan operator includes various types such as full table scan, index scan, and partition scan. The target data corresponding to the target node scan plan is read from a preset repository by using the basic operator, and then a preset calculation operator is introduced into the distributed execution plan. The preset calculation operator is used to perform data processing operations on the target data to obtain processed data. A distributed execution plan is generated based on the processed data, the target node scan plan, and a preset adaptive scan mechanism. Specifically, generating the distributed execution plan based on the target node scan plan and the preset adaptive scan mechanism includes: determining a preset scan operator as the basic operator of the distributed execution plan to read the target data corresponding to the target node scan plan from a preset repository by using the basic operator; introducing a preset calculation operator into the distributed execution plan to perform data processing operations on the target data by using the preset calculation operator to obtain processed data; and generating a distributed execution plan based on the processed data, the target node scan plan, and the preset adaptive scan mechanism.

[0062] Figure 2 FIG. 134 is a schematic diagram of generating a distributed execution plan provided in this embodiment. First, a user terminal inputs a SQL (Structured Query Language) query statement, and then the SQL query statement is parsed syntactically, and the parsed query statement obtained is input to an optimizer so that the optimizer optimizes the parsed query statement to obtain an execution plan. Specifically, a data dictionary containing database structure information is used to help the optimizer generate a more efficient execution plan. Then, the SQL query statement is converted into DXL (Data Exchange Language), and the DXL language is queried by using database metadata information in a metadata adapter to obtain a query result, and the query result is converted into a specific execution plan. The target distributed optimizer is used to perform global optimization in a distributed database environment to ensure that the query task is efficiently executed on multiple nodes.

[0063] Step S13: Use a target load prediction model to predict the load conditions of each node to obtain predicted load values. Based on the predicted load values and the physical locations of the nodes and by using a query coordinator in a distributed executor, the query tasks in the distributed execution plan are assigned to target nodes, and the query executor on the target nodes is used to execute the query tasks to obtain corresponding execution results.

[0064] In this embodiment, an initial load prediction model is constructed based on the autoregressive integrated moving average model (ARIMA) and combined with the long short-term memory network (LSTM). The initial load prediction model is trained using time series data containing the historical load data of each node to obtain a target load prediction model, and the target load prediction model is used to predict the predicted load values of each node in a preset time period. Specifically, predicting the load conditions of each node using the target load prediction model to obtain predicted load values includes: obtaining the historical load data of each node and generating corresponding time series data based on the historical load data; constructing an initial load prediction model based on the autoregressive integrated moving average model and the long short-term memory network, and training the initial load prediction model using the time series data to obtain a target load prediction model; using the target load prediction model to predict the load conditions of each node to obtain the predicted load values of the nodes in a preset time period.

[0065] It can be understood that after obtaining the predicted load values, the query coordinator in the distributed executor is used to analyze the query tasks in the distributed execution plan to determine the task requirements of the query tasks based on the analysis results, such as the required computing resources, memory, network bandwidth, etc. Then, based on the predicted load values and the physical locations of the nodes, the query tasks are assigned to target nodes that meet the preset load conditions and preset distance conditions, that is, target nodes with lighter loads and shorter distances, to ensure that the target nodes have sufficient resources to process the query tasks and minimize network latency and improve query efficiency. Then, the query executor on the target node is used to execute the query tasks, and during the execution of the query tasks, the query coordinator monitors the load conditions of the target nodes and the task execution progress corresponding to the query tasks. If the monitoring results indicate that the load conditions of the target nodes reach the target load threshold or the query tasks are executed abnormally, it jumps to the step of using the query coordinator in the distributed executor to analyze the query tasks in the distributed execution plan and determining the task requirements of the query tasks based on the analysis results, that is, if the maximum load threshold of the target node is reached or an abnormality occurs during the execution of the query tasks, the analysis and allocation of the query tasks are performed again. In a specific implementation, if the predicted load value of a certain node is predicted to increase significantly in the future period, the number of query tasks allocated to this node is correspondingly reduced, and the query tasks are reasonably allocated to nodes with lower predicted load values.

[0066] Figure 3A schematic diagram of a distributed parallel query provided in this embodiment. A client sends query requests to different databases. The client can be an application program, a user interface, or other systems that need to access the database. After receiving the query request from the client, the database passes the query request to a distributed plugin to utilize the query coordinator in the distributed plugin and schedule a suitable query executor to process the query request based on the query request and the status of the current system. The query executor executes specific query tasks according to the scheduling of the query coordinator, and after the query tasks are completed, returns the execution results of the query executor to the corresponding database so that the database can return the execution results to the corresponding client.

[0067] Specifically, based on the predicted load value and the physical locations of the nodes, and by using the query coordinator in the distributed executor, the query tasks in the distributed execution plan are assigned to target nodes, and the query executor on the target node is used to execute the query tasks to obtain corresponding execution results, including: analyzing the query tasks in the distributed execution plan by using the query coordinator in the distributed executor, and determining the task requirements of the query tasks based on the analysis results; based on the predicted load value and the physical locations of the nodes, and by using remote procedure call, assigning the query tasks to target nodes that meet the preset load conditions and preset distance conditions; using the query executor on the target node to execute the query tasks, and during the execution of the query tasks, monitoring the load situation of the target node and the task execution progress corresponding to the query tasks based on the query coordinator; if the monitoring results indicate that the load situation of the target node reaches the target load threshold or the query task execution is abnormal, then jump to the step of analyzing the query tasks in the distributed execution plan by using the query coordinator in the distributed executor and determining the task requirements of the query tasks based on the analysis results.

[0068] Furthermore, Figure 4A schematic diagram of query task allocation provided in this embodiment. When receiving the query request, the location information of the data storage node corresponding to the query request is determined by using the coordination node in the query coordinator. Then, the physical distances between each node and the data storage node are determined by using GIS (Geographic Information System) and the coordination node. A target weighted function is constructed based on the load conditions of each node and by using the coordination node. The resource conditions of each node are evaluated by using the target weighted function and the coordination node, and the real-time load conditions of each node to be selected are monitored to avoid allocating the query task to a node with too high a load. If multiple nodes meet the preset allocation conditions, the node to be selected with the closest physical distance is taken as the target node. If there are multiple nodes to be selected with equal physical distances, the node to be selected with the highest efficiency is determined as the target node based on the historical task completion efficiency of the node.

[0069] As can be seen from the above, the present application optimizes the initial distributed optimizer through network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer, and uses the target distributed optimizer to generate a distributed execution plan adapted to the distributed environment. The distributed executor is responsible for the query task allocation and execution in the distributed execution plan. Through the collaborative work of the query coordinator and the query executor, and by introducing the target load prediction model and the physical location of the node, efficient task allocation and query processing are realized. In this way, the query coordinator and the query executor based on the optimized target distributed optimizer and the distributed executor perform the query task allocation and execution, greatly improving the performance, scalability and resource utilization rate of the distributed database system, and being able to effectively cope with the challenges brought by large-scale data queries.

[0070] Correspondingly, as shown in Figure 5 the present application also provides a database distributed parallel execution device, including:

[0071] An initial optimizer optimization module 11, configured to expand and optimize the initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer;

[0072] An execution plan generation module 12, configured to generate a plurality of node scan plans based on the query request and by using the target distributed optimizer, evaluate each node scan plan by using a cost model to obtain a target node scan plan, and then generate a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism;

[0073] The query task execution module 13 is used to predict the load conditions of each node by using the target load prediction model to obtain predicted load values, and based on the predicted load values and the physical locations of the nodes, and by using the query coordinator in the distributed executor, allocate the query tasks in the distributed execution plan to the target nodes, and use the query executors on the target nodes to execute the query tasks to obtain corresponding execution results.

[0074] As can be seen from the above, in this application, the initial distributed optimizer is optimized through network transmission technology and resource pooling storage characteristics to obtain the target distributed optimizer, and the target distributed optimizer is used to generate a distributed execution plan adapted to the distributed environment. The distributed executor is responsible for the allocation and execution of query tasks in the distributed execution plan. Through the collaborative work of the query coordinator and the query executor, and by introducing the target load prediction model and the physical locations of the nodes, efficient task allocation and query processing are realized. In this way, the allocation and execution of query tasks are carried out based on the query coordinator and the query executor of the optimized target distributed optimizer and the distributed executor, which greatly improves the performance, scalability and resource utilization rate of the distributed database system and can effectively cope with the challenges brought by large-scale data queries.

[0075] In some specific embodiments, the initial optimizer optimization module 11 may specifically include:

[0076] The transmission capacity evaluation unit is used to evaluate the network transmission capacity of the initial distributed optimizer and determine the target transmission strategy based on the capacity evaluation result;

[0077] The data analysis unit is used to analyze the data distribution, storage characteristics and access patterns in the database by using the resource pooling storage characteristics to obtain corresponding analysis results;

[0078] The optimizer optimization unit is used to optimize the initial distributed optimizer based on the analysis results and the target transmission strategy to obtain the target distributed optimizer;

[0079] The optimizer encapsulation unit is used to encapsulate the target distributed optimizer into a dynamic library and load the dynamic library into the kernel of the database.

[0080] In some specific embodiments, the execution plan generation module 12 may specifically include:

[0081] The scan plan generation unit is used to obtain a query request for the database, and based on the query statement in the query request and by using the target distributed optimizer, generate a node scan plan including different scan nodes, scan ranges and scan strategies;

[0082] A cost value generation unit, configured to generate respective cost values representing the resources and performance consumption required for executing each of the node scan plans based on each of the node scan plans and by using a cost model;

[0083] A cost value comparison unit, configured to compare each of the cost values, and determine, based on the comparison result, the node scan plan with the lowest cost value as the target node scan plan.

[0084] In some specific embodiments, the execution plan generation module 12 may specifically include:

[0085] A scan strategy adjustment unit, configured to monitor the load and data distribution conditions of each of the nodes based on the preset adaptive scan mechanism, and adjust the scan strategy in the target node scan plan based on the monitoring result to obtain an adjusted scan strategy;

[0086] A first plan generation unit, configured to generate a corresponding distributed execution plan based on the target node scan plan including the adjusted scan strategy and by using the target distributed optimizer.

[0087] In some specific embodiments, the execution plan generation module 12 may specifically include:

[0088] A target data reading unit, configured to determine a preset scan operator as a basic operator of the distributed execution plan, so as to read target data corresponding to the target node scan plan from a preset repository by using the basic operator;

[0089] A data processing unit, configured to introduce a preset calculation operator into the distributed execution plan, so as to perform a data processing operation on the target data by using the preset calculation operator to obtain processed data;

[0090] A second plan generation unit, configured to generate a distributed execution plan based on the processed data, the target node scan plan, and the preset adaptive scan mechanism.

[0091] In some specific embodiments, the query task execution module 13 may specifically include:

[0092] A time series data determination unit, configured to obtain historical load data of each node, and generate corresponding time series data based on the historical load data;

[0093] An initial prediction model construction unit, configured to construct an initial load prediction model based on the autoregressive moving average model and the long short-term memory network, and train the initial load prediction model by using the time series data to obtain a target load prediction model;

[0094] A load condition prediction unit, configured to predict the load conditions of each node by using the target load prediction model, so as to obtain the predicted load values of the nodes in a preset time period.

[0095] In some specific embodiments, the query task execution module 13 may specifically include:

[0096] A query task analysis unit, configured to analyze the query tasks in the distributed execution plan by using the query coordinator in the distributed executor, and determine the task requirements of the query tasks based on the analysis results;

[0097] A query task allocation unit, configured to allocate the query tasks to target nodes that meet the preset load conditions and preset distance conditions by using remote procedure call based on the predicted load values and the physical locations of the nodes;

[0098] An execution progress monitoring unit, configured to execute the query tasks by using the query executor on the target nodes, and monitor the load conditions of the target nodes and the task execution progress corresponding to the query tasks based on the query coordinator during the execution of the query tasks;

[0099] An execution exception unit, configured to, if the monitoring result indicates that the load condition of the target node reaches the target load threshold or the query task execution is abnormal, jump to the step of analyzing the query tasks in the distributed execution plan by using the query coordinator in the distributed executor and determining the task requirements of the query tasks based on the analysis results.

[0100] Furthermore, an electronic device is also disclosed in an embodiment of the present application. Figure 6 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Wherein, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the database distributed parallel execution method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0101] In this embodiment, the power supply 23 is used to provide operating voltages for the various hardware devices on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and specific limitations thereof are not provided herein; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type thereof can be selected according to specific application requirements, and specific limitations are not provided herein.

[0102] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc., and the resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0103] Among them, the operating system 221 is used to manage and control the various hardware devices and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the database distributed parallel execution method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.

[0104] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the database distributed parallel execution method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.

[0105] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0106] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0107] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0108] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0109] The technical solutions provided in this application have been introduced in detail above. Specific examples are used herein to illustrate the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for distributed parallel execution of a database, characterized in that, Including: Expanding and optimizing the initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer; Based on a query request and using the target distributed optimizer to generate several node scan plans, evaluating each node scan plan using a cost model to obtain a target node scan plan, and then generating a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism; Using a target load prediction model to predict the load conditions of each node to obtain predicted load values, based on the predicted load values and the physical locations of each node and using a query coordinator in the distributed executor, allocating the query tasks in the distributed execution plan to target nodes, and using a query executor on the target nodes to execute the query tasks to obtain corresponding execution results.

2. The database distributed parallel execution method according to claim 1, characterized in that The expanding and optimizing the initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer includes: Evaluating the network transmission ability of the initial distributed optimizer and determining a target transmission strategy based on the ability evaluation result; Analyzing the data distribution, storage characteristics, and access patterns in the database using the resource pooling storage characteristics to obtain corresponding analysis results; Optimizing the initial distributed optimizer based on the analysis results and the target transmission strategy to obtain a target distributed optimizer; Encapsulating the target distributed optimizer into a dynamic library and loading the dynamic library into the kernel of the database.

3. The database distributed parallel execution method according to claim 1, wherein The generating several node scan plans based on a query request and using the target distributed optimizer, and evaluating each node scan plan using a cost model to obtain a target node scan plan includes: Obtaining a query request for the database, and generating node scan plans including different scan nodes, scan ranges, and scan strategies based on the query statement in the query request and using the target distributed optimizer; Generating respective cost values representing the resources and performance consumption required for executing each node scan plan based on each node scan plan and using a cost model; Comparing each cost value and determining the node scan plan with the lowest cost value as the target node scan plan based on the comparison result.

4. The database distributed parallel execution method according to claim 3, wherein The generating a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism includes: Monitoring the load and data distribution conditions of each node based on the preset adaptive scan mechanism, and adjusting the scan strategy in the target node scan plan based on the monitoring result to obtain an adjusted scan strategy; Generating a corresponding distributed execution plan based on the target node scan plan including the adjusted scan strategy and using the target distributed optimizer.

5. The database distributed parallel execution method according to claim 1, wherein The generating a distributed execution plan based on the target node scan plan and a preset adaptive scan mechanism includes: Determining a preset scan operator as the basic operator of the distributed execution plan to use the basic operator to read target data corresponding to the target node scan plan from a preset repository; Introduce a pre-designed computing operator in the distributed execution plan to perform data processing operations on the target data using the pre-designed computing operator to obtain processed data; Generate a distributed execution plan based on the processed data, the target node scan plan, and a preset adaptive scanning mechanism.

6. The database distributed parallel execution method according to claim 1, characterized in that, The predicting the load conditions of each node using the target load prediction model to obtain predicted load values includes: Obtain the historical load data of each node and generate corresponding time series data based on the historical load data; Construct an initial load prediction model based on the autoregressive moving average model and the long short-term memory network, and train the initial load prediction model using the time series data to obtain the target load prediction model; Use the target load prediction model to predict the load conditions of each node to obtain the predicted load values of the nodes in a preset time period.

7. The database distributed parallel execution method according to any one of claims 1 to 6, characterized in that, The distributing the query tasks in the distributed execution plan to target nodes based on the predicted load values and the physical locations of the nodes and using the query coordinator in the distributed executor, and executing the query tasks using the query executor on the target nodes to obtain corresponding execution results includes: Analyze the query tasks in the distributed execution plan using the query coordinator in the distributed executor, and determine the task requirements of the query tasks based on the analysis results; Allocate the query tasks to target nodes that meet the preset load conditions and preset distance conditions based on the predicted load values and the physical locations of the nodes using remote procedure call; Execute the query tasks using the query executor on the target nodes, and during the execution of the query tasks, monitor the load conditions of the target nodes and the task execution progress corresponding to the query tasks based on the query coordinator; If the monitoring results indicate that the load conditions of the target nodes reach the target load threshold or the query tasks are executed abnormally, jump to the step of analyzing the query tasks in the distributed execution plan using the query coordinator in the distributed executor and determining the task requirements of the query tasks based on the analysis results.

8. A database distributed parallel execution device, characterized in that, Includes: An initial optimizer optimization module for expanding and optimizing the initial distributed optimizer based on network transmission technology and resource pooling storage characteristics to obtain a target distributed optimizer; An execution plan generation module for generating several node scan plans based on a query request and using the target distributed optimizer, evaluating each node scan plan using a cost model to obtain a target node scan plan, and then generating a distributed execution plan based on the target node scan plan and a preset adaptive scanning mechanism; A query task execution module, which is used to predict the load conditions of each node by using a target load prediction model to obtain predicted load values, and based on the predicted load values and the physical locations of the nodes, use a query coordinator in a distributed executor to allocate the query tasks in the distributed execution plan to target nodes, and use a query executor on the target nodes to execute the query tasks to obtain corresponding execution results.

9. An electronic device, characterized in that, It includes: A memory, which is used to store computer programs; A processor, which is used to execute the computer programs to implement the database distributed parallel execution method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing computer programs, wherein when the computer programs are executed by a processor, the database distributed parallel execution method according to any one of claims 1 to 7 is implemented.